Google’s Gemini 3.7 Flash targets coding and agents with a 50% introductory price cut
Summary
<p>Google is <a href="https://blog.google/innovation-and-ai/models-and-research/gemini-models/introducing-gemini-3-7-flash/">rolling out Gemini 3.7 Flash</a>, a new version of its workhorse AI model that puts coding, agentic workflows and knowledge work at the center of the upgrade — while temporarily cutting API prices in half.</p><p>The release arrives just <a href="https://venturebeat.com/technology/googles-gemini-3-6-flash-model-cuts-ai-agent-token-costs-by-up-to-65-on-long-horizon-engineering-tasks-and-3-5-pro-is-on-the-way">three weeks after the release of Gemini 3.6 Flash</a>, an unusually short turnaround that Google attributes to developer feedback and algorithmic improvements. </p><p>For enterprise developers, the more consequential story may be the combination of those intelligence gains with lower inference costs: through the end of 2026, <a href="https://ai.google.dev/gemini-api/docs/pricing">Gemini 3.7 Flash costs $0.75 per million input tokens and $3.75 per million output tokens</a>.</p><p>Starting Jan. 1, 2027, pricing rises to $1.50 per million input tokens and $7.50 per million output tokens. That means the current discount is temporary, but it gives teams deploying high-volume coding and business agents several months to evaluate whether Google's claimed reductions in retries and manual oversight translate into lower total operating costs.</p><p>The launch also underscores Google's rapid iteration on its Flash line while its next flagship Pro model remains absent. Google did not provide a release date for Gemini 3.5 Pro with Thursday's announcement, Reuters reported, despite the model having previously been described as undergoing partner testing. Axios similarly noted that 3.7 Flash arrives before the anticipated Pro release.</p><h2><b>A three-week upgrade focused on getting work done</b></h2><p>Google describes Gemini 3.7 Flash as its "most intelligent workhorse model yet for coding and agents." The company says the model is better at adapting when it encounters roadblocks, clarifying intent when necessary and following instructions with greater fidelity.</p><p>Those improvements matter beyond benchmark scores. In an enterprise coding agent, a model that makes fewer unnecessary changes, recovers from errors and executes multi-step plans more reliably can reduce the number of human interventions needed to complete a task. The same principle applies to business agents operating across documents and applications, where an incorrect tool call or poorly interpreted instruction can derail an otherwise useful workflow.</p><p>Google says 3.7 Flash "thinks more diligently," applying more effort to multi-step planning and tool calls. Its stated goal is more disciplined execution with fewer retries and less manual supervision.</p><p>That represents an interesting evolution from Gemini 3.6 Flash. Google's developer documentation described 3.6 as reducing reasoning steps, conversational turns and tool calls compared with earlier models while attempting to limit execution-loop spiraling. With 3.7, the emphasis shifts toward putting sufficient effort into planning while improving the quality of execution — potentially a more useful optimization than simply minimizing the number of steps an agent takes.</p><p>Google DeepMind said in a post accompanying the release that 3.7 Flash shows gains in debugging and issue resolution, generates more functional web layouts and applications with fewer prompts, and improves reasoning and accuracy on real-world business workflows.</p><h2><b>Coding gains are substantial, but not universal</b></h2><p>Google's benchmarks show a large generational improvement in several software engineering tests.</p><p>On FrontierCode 1.1 Main, which measures production code quality, Gemini 3.7 Flash scores 43.6%, up from 34.4% for Gemini 3.6 Flash. That also narrowly exceeds the 42.7% Google reports for Claude Sonnet 5 and 41.3% for GPT-5.6 Terra.</p><p>On DeepSWE v1.1, a long-horizon software engineering evaluation, 3.7 Flash reaches 65.3%, compared with 49.0% for its predecessor. GPT-5.6 Terra remains ahead at 69.6% in Google's table.</p><p>Web development shows another notable gain. Gemini 3.7 Flash receives an Elo score of 1588 on Code Arena, versus 1538 for 3.6 Flash, 1541 for Claude Sonnet 5 and 1523 for GPT-5.6 Terra. Google says the new model can produce more functional layouts and feature-complete applications in fewer prompts while more closely following reference screenshots, images and design systems.</p><p>The broader benchmark table is more mixed, which is important for enterprises evaluating the model against particular workloads rather than looking for a single "best" model.</p><p>Gemini 3.7 Flash scores 85.8% on Terminal-bench 2.1, compared with 87.4% for GPT-5.6 Terra. Terra also leads Google's comparisons on Terminal-bench 3.0 and OSWorld-2.0. Claude Sonnet 5 leads the Agent's Last Exam multimodal desktop and operating-system tasks with a 33.3% pass rate, versus 26.3% for Gemini 3.7 Flash.</p><p>In other words, Google's own results do not show 3.7 Flash universally displacing higher-priced competitors. They instead suggest a model that has become substantially more competitive in coding and agent workloads while occupying a lower price tier.</p><h2><b>Enterprise workflows may be the more important test</b></h2><p>The gains extend beyond software development.</p><p>On AutomationBench, which Google describes as measuring enterprise workflow automation, Gemini 3.7 Flash scores 30.4%, up sharply from 17.0% for 3.6 Flash. Google's table lists Claude Sonnet 5 at 10.7% and GPT-5.6 Terra at 23.6%.</p><p>The model also reaches 34.0% on GDP.PDF, an evaluation of complex PDF comprehension, compared with 22.0% for 3.6 Flash, 28.0% for Claude Sonnet 5 and 24.7% for GPT-5.6 Terra.</p><p>That combination is relevant for enterprise agents because many practical deployments require more than generating text or code. An agent may need to interpret a long report, identify relevant information, decide which tool to invoke, update another system and produce a document for a human reviewer. Reliability across that chain can matter more than performance on an isolated reasoning benchmark.</p><p>Google is putting that thesis into practice with Gemini Spark. Google AI Pro and Ultra subscribers can use 3.7 Flash in Spark, the company's personal AI agent. Google says the upgrade improves Spark's knowledge work and tool use across Google Workspace applications, including workflows that consolidate files, draft emails and update status documents.</p><p>For enterprises, 3.7 Flash is also available through the Gemini Enterprise Agent Platform and Gemini Enterprise app.</p><h2><b>Price becomes part of the model competition</b></h2><p>Gemini 3.7 Flash's introductory pricing is a notable bid to embed the model into enterprise workflows.</p><p>Until Dec. 31, developers pay $0.75 per million input tokens and $3.75 per million output tokens. Context caching costs $0.075 per million tokens during the introductory period. Google says standard prices will double on Jan. 1, 2027, to $1.50 for input and $7.50 for output, with context caching rising to $0.15.</p><p>For comparison, Gemini 3.6 Flash's standard API pricing is $1.50 per million input tokens and $7.50 per million output tokens. Google's benchmark table lists Claude Sonnet 5 at $2 and $10, respectively, while GPT-5.6 Terra is listed at $2 and $12.</p><p>The economics become more pronounced for autonomous agents because a single user request can produce a long sequence of model calls, reasoning tokens and tool interactions. A model that costs less per token but requires substantially more retries may not ultimately be cheaper. </p><p>Conversely, Google's combination of lower introductory token pricing and claimed improvements in first-pass accuracy could materially change the cost of running high-volume coding or document-processing agents if those gains carry over to production.</p><p>That is the metric enterprise teams will ultimately need to test: not price per million tokens in isolation, but cost per successfully completed task.</p><h2><b>Google’s AI shake-up raises the stakes for Gemini</b></h2><p>Gemini 3.7 Flash arrives amid a broader debate over whether Google is losing ground at the AI frontier. The company has not released Gemini 3.5 Pro, despite saying in May that the flagship model would arrive the following month. </p><p>By July, Google said it remained in partner testing and would become broadly available when ready; Thursday’s announcement offered no further timetable. Google’s latest released general-purpose Pro model therefore remains <a href="https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-1-pro/">Gemini 3.1 Pro</a>, introduced in February. </p><p><a href="https://www.reuters.com/business/google-gemini-launch-delayed-tech-falls-short-internal-goals-bloomberg-news-2026-07-16/">Reuters reported in July that Gemini 3.5 Pro missed its original target</a> after falling short of internal goals, particularly in coding, even as Google began training what it calls its most ambitious model yet, Gemini 4.</p><p>The delay coincides with a <a href="https://blog.google/company-news/inside-google/message-ceo/next-chapter-ai-momentum/">major overhaul of Google’s AI leadership announced last week.</a> </p><p>Google DeepMind co-founder and Nobel Prize Winner Demis Hassabis has relinquished day-to-day control of the company's famed DeepMind AI division to become its chair and, simultaneously, to take on the role of Alphabet’s chief scientist.</p><p>Meanwhile, former DeepMind CTO Koray Kavukcuoglu now runs the unit as a senior vice president reporting directly to CEO Sundar Pichai. </p><p>Kavukcuoglu controls Gemini model development, frontier research, the Gemini app and developer teams—effectively consolidating the full Gemini chain under a more product-focused operator. </p><p>Chief scientist Jeff Dean, Gemini co-lead Oriol Vinyals, Quoc Le and Sanjay Ghemawat left to establish the <a href="https://www.wired.com/story/jeff-dean-google-discovery-loop-startup/">research startup Discovery Loop. </a></p><p>Those exits followed Gemini co-lead Noam Shazeer’s move to OpenAI and Nobel Prize-winning AlphaFold scientist John Jumper’s departure for Anthropic. <a href="https://www.reuters.com/world/inside-google-executive-moves-that-led-its-big-ai-reshuffle-2026-08-12/">Reuters reported</a> that internal disagreements, constrained compute allocation and Google’s bureaucracy contributed to slower releases and weaknesses in coding.</p><p>Outside interpretations range from organizational repair to a more fundamental retreat. </p><p><a href="https://newsletter.semianalysis.com/p/gemini-is-cooked-but-gcp-is-cooking">SemiAnalysis has argued</a> that Google is increasingly prioritizing the highly profitable business of supplying cloud infrastructure to AI companies—including Gemini competitors—over keeping its own models at the absolute frontier. That analysis also claimed Google had effectively canceled 3.5 Pro, although Google has not confirmed that and continues to describe the model as delayed. </p><p><a href="https://www.theverge.com/podcast/979370/google-deepmind-ai-race-lose-jeff-dean-demis-hassabis">The Verge offered a more measured assessment</a>: the departures and model delays are serious, but Google retains enormous advantages through Search, Workspace, Android, Cloud, custom AI chips and consumer distribution. Google says the Gemini app has surpassed 950 million monthly users, giving it a reach that does not depend entirely on owning the highest-scoring model.</p><p>Current benchmarks similarly depict a company behind the overall leaders but still firmly competitive. Artificial Analysis places Claude Opus 5 at 63 on its overall model <a href="https://artificialanalysis.ai/evaluations/artificial-analysis-intelligence-index">Intelligence Index</a>, while Google reports a score of 56 for Gemini 3.7 Flash—an improvement from 52 for 3.6 Flash but not a return to the top. </p><p>Arena’s early human-preference results are more favorable, provisionally ranking 3.7 Flash ninth overall and eighth for web development. </p><p>The resulting picture is not that Google has abandoned advanced AI, but that it has become stronger at rapidly shipping efficient Flash models while struggling to deliver the premium flagship required to reclaim broad leadership. Gemini 4 will now serve as the clearest test of whether the leadership reorganization fixes that execution gap.</p><h2><b>Available now across Google's developer stack</b></h2><p>Developers can access Gemini 3.7 Flash through the Gemini API in Google AI Studio and Android Studio, as well as Google's Antigravity environment. Enterprises can deploy it through Gemini Enterprise Agent Platform and Gemini Enterprise, while consumers with Google AI Pro or Ultra subscriptions can access the model through Spark in supported countries.</p><p>Google is also shipping updated safeguards covering chemical, biological, radiological and nuclear risks and cyber-offense misuse, according to the company.</p><p>The unusually fast jump from Gemini 3.6 Flash to 3.7 Flash points toward a model development cycle in which algorithmic improvements can reach production products without waiting for a new flagship generation. Ars Technica also highlighted the three-week interval between the two releases, while Google says the techniques behind the update will inform future models.</p><p>For developers, that faster cadence creates its own operational question. Models can improve quickly, but production teams still have to benchmark new releases against their own repositories, prompts, tool schemas and failure modes before changing a deployment.</p><p>Gemini 3.7 Flash gives those teams a particularly strong incentive to run that evaluation. Google's own numbers show major improvements in production coding, web development, document comprehension and workflow automation without claiming leadership everywhere. At its introductory price, Google is effectively betting that developers will value a model that is competitive enough with more expensive systems while being cheap enough to run repeatedly inside agents.</p><p>Whether that advantage survives the return to full pricing in January will depend less on leaderboard positions than on how reliably 3.7 Flash completes real work.</p>