- GPT Image 2.5 ブログ
- GPT Image 2.5 vs GPT Image 2:同じプロンプト、3つのモデル、見えてくる違い
GPT Image 2.5 vs GPT Image 2:同じプロンプト、3つのモデル、見えてくる違い
GPT Image 2.5にはFlareとSunburstの2つのモデルがあり、このページでは両方を同じプロンプトでGPT Image 2と並べて比較しています。以下の画像はすべてこのサイトで生成したもので、合計101枚、各行の上にプロンプトを記載しているので自分でも同じように生成できます。キャプションは各画像に写っているものと計測結果を説明するものであり、モデルを採点するものではありません。サンプルの前に、他の公開テストが報告している内容をまとめているので、その結果とここで見るものを比較できます。
GPT Image 2.5 vs GPT Image 2 のひと目比較
| GPT Image 2.5 | GPT Image 2 | |
|---|---|---|
| モデル数 | 2つ:FlareとSunburst | 1つ |
| OpenAIの位置付け | Flare:速度重視、品質はGPT Image 2と同等。Sunburst:品質重視、GPT Image 2より上 | ベースライン |
| 今回の4K生成時間 | 約32 s(FlareとSunburstとも) | 約66 s |
| 今回の1K生成時間 | 平均72〜73 s | 平均75 s |
| 参照画像 | 最大16枚まで入力可能 | 複数入力可能 |
| 透過PNG | 対応、実際のアルファチャンネルあり | 対応 |
他のGPT Image 2.5 vs GPT Image 2テストが報告している内容
これらは2026年9月16日時点で見つけることができた、計測結果や比較画像を含む公開テストです。それぞれが報告した内容を一覧にしており、リンクは元記事につながっています。
| 報告された結果 | 情報源 | 根拠 |
|---|---|---|
| GPT Image 2.5の方が高速 | Tosea、Hacker Newsの開発者、@levelsio | 同じトークン予算で比較:Flare 19.7 s、Sunburst 27.7 s、GPT Image 2 37.3 s(41回のAPI呼び出し)。4K:28〜34 s 対 92 s。あるデベロッパーは約104 sが35〜40 sに短縮したと報告。 |
| GPT Image 2.5の方が参照写真により忠実 | @levelsio、Hacker News、PixVerseが引用するRedditのUIスレッド | 「GPT Image 2は参照画像をより文字通りに使い、写真に貼り付けるような形になるのに対し、GPT Image 2.5は実際に参照として使っている」。参照ベースのUIプロンプトは明らかな改善と評されている。 |
| GPT Image 2.5の方が連続編集でのドリフトが少ない | Tosea、Renoiseが引用する4回編集のデスクテスト | 3ターンのピクセルドリフト:GPT Image 2が11.4%、Sunburstが9.9%、Flareが9.2%(それぞれ1チェーン)。4回の編集を通じて、未編集領域の構造的類似度は0.87〜0.99。 |
| GPT Image 2.5の方が質感が豊かで小さな文字も保持 | ImagineArt、Pixmax | 6つのプロンプトをそれぞれ1回生成:「2.5は紙、金属、布地の質感をより多く再現し、小さな文字もより一貫して保持する」。Pixmaxは素材表現と仕上がりの美しさでSunburstを最高評価とした。 |
| GPT Image 2.5は編集時にチャートのデータを保持 | Simon Willison、Renoiseが引用 | 軸、グリッドライン、データ曲線をそのままに、既存の折れ線グラフにアライグマを追加。 |
| GPT Image 2.5はアリーナ投票でGPT Image 2を上回る | Artificial Analysis | テキストから画像へのElo:Flare(最大)1189、Sunburst(最大)1183、GPT Image 2(高)1173、信頼区間は±9〜±11。 |
| 文字の正確さは同等 | Tosea、Pixmax、Picsart | 図表を含む15枚のスライドで、3モデルすべてが見出し・箇条書き・数字をすべて正しく描画。タイムテーブルの文字も3モデルとも正確に再現。 |
| GPT Image 2の方が空間的精度と再現性で高評価 | Pixmax | プロンプトへの忠実性、空間的一貫性、生成間の一貫性で5つ星、GPT Image 2.5の2モデルは4〜4.5。 |
| 違いを感じないユーザーもいる | PixVerseが引用するr/ChatGPTのローンチスレッド | 生成が速くなった、表情がよくなったと感じた人もいれば、Images 2.0との違いが分かりにくいという声もあった。 |
以下の私たちのタスクは、速度、参照写真、連続編集、質感と小さな文字、チャート編集、密度の高いテキスト、低照度という同じ観点をカバーするように作成しました。
GPT Image 2.5 vs GPT Image 2 の速度比較
| モデル | 4K 1回目 | 4K 2回目 | 1K平均(10回) | 1Kの範囲 |
|---|---|---|---|---|
| GPT Image 2.5 Flare | 32 s | 160 s | 72 s | 67–80 s |
| GPT Image 2.5 Sunburst | 33 s | 33 s | 73 s | 65–96 s |
| GPT Image 2 | 68 s | 65 s | 75 s | 68–90 s |
時間はこのサイト上での実測値で、リクエストから結果が出るまで(キューイングを含む)を計測しています。各モデルの1回目の実行、4Kでの同じポスタープロンプト:
A gig poster for an indie jazz night, two-color risograph print style in deep navy and warm orange, heavy paper grain. Large hand-set headline at the top reading "MIDNIGHT BRASS". Below it, in smaller type, exactly these lines: "Live Jazz Quartet", "Friday 24 October, 9 PM", "The Copper Room, 18 Harbor Street", "Tickets $15 at the door". A stylized trumpet silhouette fills the lower half. All text must be spelled exactly as written, clean and legible, no extra words.



参照画像を 3 枚使った場合(下の合成)、どのモデルも時間が長くなり、6 回の生成で 95〜172 秒でした。
GPT Image 2.5 vs GPT Image 2 の参照写真比較
スタジオ写真の男性を生成し、各モデルに同じ人物を新しいシーン・新しい表情で配置するよう指示しました。
以下の2行で使用する参照写真。

The same man from the reference photo, now laughing, three-quarter view, photographed outdoors at golden hour on a hiking trail with pine trees behind him, wearing a dark green fleece. Keep his face, red curly hair, round tortoiseshell glasses and freckles exactly as in the reference. 85mm portrait, shallow depth of field, photorealistic.



3枚の参照画像を同時に使用:同じ男性、ジャケット、通りを1枚の写真に組み合わせます。
以下の行で使用する入力:画像1は人物、画像2はジャケット、画像3は通り。



Show the man from image 1 wearing the jacket from image 2, standing on the street from image 3 in front of the bakery window, looking at the camera. Keep his face, hair, glasses and freckles identical to image 1. Keep the jacket's mustard corduroy, four brown buttons and green pine-tree embroidery identical to image 2. Keep the street, bakery window, blue door and string lights identical to image 3. Photorealistic, night, wet cobblestones.



GPT Image 2.5 vs GPT Image 2 のローカル編集比較
生成したアパートメント写真の中でアームチェアを1つだけ差し替えました。パーセンテージは、チェア以外の領域で元画像と比べて一定のしきい値を超えて変化したピクセルの割合です。
以下の編集で使用する元の写真。

Replace only the green velvet armchair with a tan leather mid-century lounge chair on a walnut base. Keep everything else in the photo exactly the same: the window, curtains, side table, floor lamp, bookshelf, rug, lighting, camera angle and framing.



チャートにキャラクターを追加し、データはそのままにするよう指示しました。パーセンテージは、アライグマの領域以外で変化したピクセルの割合です。
以下の編集で使用する元のチャート。

Add a small cartoon raccoon in a white lab coat standing on the September peak of the line and pointing at it. Keep the title, every axis label, every gridline, the legend and the data line exactly as they are. Do not change any numbers or move the line.



アパートメント写真に対して5回連続で編集を実施し、各ステップは前のステップの出力を引き継ぎました:チェア、ランプシェード、壁の色、観葉植物、ラグ。掲載画像は5番目のステップの結果です。
ステップ1:アームチェアを交換。ステップ2:マットブラックのランプシェード。ステップ3:本棚の背後の壁をセージグリーンに。ステップ4:隅にフィドルリーフフィグを追加。ステップ5:ラグをダークネイビーの平織りに交換。各ステップの最後には「他はすべてそのまま」と指示。



3つのチェーンすべてで、「本棚の背後の壁だけ」という指示が、ドア脇の隣接する壁の色も変えてしまいました。
GPT Image 2.5 vs GPT Image 2 の質感とディテール比較
5種類の素材と手書きラベルによる2Kの静物写真。
Overhead still life on a dark walnut desk, hard side light: a sheet of crumpled kraft paper, a brushed brass pocket compass with visible machining marks, a folded piece of undyed linen with loose weave, a red wax seal with a stamped anchor, and a small handwritten paper label that reads "No. 27 — Harbor Street" in black ink. Every material should show its texture: paper fibres, brass grain, linen threads, wax gloss. Photorealistic, 100mm macro.



E-commerce hero photo of a matte charcoal ceramic pour-over coffee dripper standing on a small slab of raw slate, a thin ribbon of steam rising, soft diffused studio light from the upper left, pale warm grey seamless background, shallow depth of field. The word "NORD" is debossed in small clean sans-serif letters near the base of the dripper. Photorealistic, 100mm macro lens, product photography for a store listing.



Editorial portrait of a woman in her late thirties working at a potter's wheel in a small ceramics studio, clay on her hands and forearms, linen apron, hair tied back, looking at the wheel with concentration. Soft north-facing window light from the left, shelves of unglazed bowls blurred in the background. Shot on 85mm at f/2, natural skin texture, no retouching look, realistic photograph.



夜の通りを、影は中立的な色味に保つよう指示して生成。各画像について、最も暗い5分の1のピクセルを対象に、赤マイナス青のバランス(プラスが暖色寄り)とノイズレベルを計測しました。
Night photograph of a quiet residential street after rain, lit only by two sodium street lamps and one lit window, deep shadows, neutral white balance so the shadows stay grey rather than yellow or green, clean low-noise rendering, a parked bicycle against a brick wall in the foreground, 35mm, f/2, photorealistic.



GPT Image 2.5 vs GPT Image 2 のテキストとレイアウト比較
A gig poster for an indie jazz night, two-color risograph print style in deep navy and warm orange, heavy paper grain. Large hand-set headline at the top reading "MIDNIGHT BRASS". Below it, in smaller type, exactly these lines: "Live Jazz Quartet", "Friday 24 October, 9 PM", "The Copper Room, 18 Harbor Street", "Tickets $15 at the door". A stylized trumpet silhouette fills the lower half. All text must be spelled exactly as written, clean and legible, no extra words.



A clean flat-design infographic poster titled "How Sourdough Rises" at the top. Below the title, a 2x2 grid of four numbered steps, each with a simple line icon, a bold heading and one caption sentence, exactly as follows. 1. "Mix" — "Flour, water and starter come together into a shaggy dough." 2. "Rest" — "Autolyse for 45 minutes so gluten begins to form." 3. "Fold" — "Four sets of stretch and folds over two hours build strength." 4. "Bake" — "Bake at 250°C in a covered pot for 20 minutes, then uncovered for 25." A footer line at the bottom reads "Total time: about 24 hours". Cream background, terracotta and olive green accents, all text spelled exactly as written and fully legible.



A tablet dashboard screen for a bicycle repair shop's booking system, modern flat UI, light theme. Left sidebar with the shop name "Spoke & Chain" and menu items "Bookings", "Customers", "Inventory", "Reports". Top row of three stat cards: "Today 8 bookings", "Open repairs 5", "Revenue this week $2,340". Below, a table titled "Today's appointments" with exactly five rows showing time, customer name and job: "09:00 Maya Lindqvist Brake bleed", "10:30 Tom Okafor Tubeless setup", "11:15 Priya Nair Full service", "13:00 Jonas Weber Wheel true", "15:30 Aiko Tanaka Chain replace". All text crisp and spelled exactly as written.



GPT Image 2.5 vs GPT Image 2 のスタイル比較
Mid-century gouache illustration in the style of a 1960s travel magazine: a bustling night market street in Taipei, strings of paper lanterns overhead, steam rising from food stalls, crowds of small stylized figures, scooters parked along the curb. Flat shapes, visible brush texture, a limited palette of five colors: ink black, cream, vermilion, teal and mustard yellow. No text or lettering anywhere in the image.



GPT Image 2.5 vs GPT Image 2 の透過PNG比較
背景を透過に設定して生成しました。各ファイルのアルファチャンネルを読み取り、パーセンテージは部分的に透明なピクセル(ソフトエッジや影がある部分)の割合を示します。
A single red enamel camping mug with a white rim, isolated on a transparent background, studio product shot, PNG with alpha.



GPT Image 2.5 vs GPT Image 2 のテスト方法
- 日付: 2026年9月16日、このサイトのジェネレーターにて。
- プロンプト: このテストのために作成し、上記に全文を掲載。OpenAIのプロンプトガイドからのものは含まれていません。
- 実行回数: 各モデル・各プロンプトにつき2回(透過PNGと5ステップの編集チェーンは各モデル1回)。合計101枚の画像を生成。2件のリクエストがエラーとなり再試行しました。エラーはデータに残しています。
- 設定: すべての行でモデル間のアスペクト比と解像度を統一。参照画像と編集用の入力画像はまずここで生成し、その後変更せずに3モデルすべてで再利用しました。
- 計測: 編集領域外のピクセル変化は、各出力を同じサイズにリサイズした上で入力画像と比較しています。シャドウバランスとノイズは各夜間画像の最も暗い5分の1で算出。アルファの割合はPNGファイルから読み取っています。
- 限界: サンプル数は少なく、パイプラインは1つ、画像を記述するオペレーターも1人です。この数値はこれらの実行結果を表すものであり、確率を示すものではありません。
よくある質問
GPT Image 2.5はGPT Image 2より高速ですか? このサイトでの4Kでは、今回の実行でGPT Image 2.5の両モデルともGPT Image 2の約半分の時間で完了しました。1Kでは、3モデルとも65–96 sの同じ範囲に収まりました。同じ設定での公開APIテストでは、Flareが19.7 s、Sunburstが27.7 s、GPT Image 2が37.3 sと報告されています。
FlareとSunburst、どちらを選ぶべきですか? OpenAIはFlareを、品質はGPT Image 2と同等の速度重視プロファイルとして位置付け、Sunburstを品質重視プロファイルとして位置付けています。上記のプロンプトでの出力を並べて掲載しているので、直接比較できます。
透過背景は本物ですか? はい。3モデルすべてがアルファチャンネル付きのPNGファイルを返しました。上記のパーセンテージはそのチャンネルを読み取った結果です。
GPT Image 2用のプロンプトはGPT Image 2.5でも使えますか? このページのすべてのプロンプトは、変更なしで3モデルすべてに使用しました。
更新履歴
- 2026年9月16日: 初版。101枚の画像、3モデル、13のプロンプト、および公開テストのまとめ。
