OpenAI halves GPT‑5.6 Sol price as Moonshot AI rolls out cheaper
Photo by panumas nikhomkhai on Pexels
OpenAI announced a 50 % price cut for its GPT‑5.6 Sol model, and Moonshot AI responded with two new releases that undercut the incumbent on both cost and capability. The moves tighten a pricing battle that could reshape where developers run their most demanding workloads.
The price reduction was posted on Hacker News under the headline “GPT‑5.6 Sol Pricing Cut by 50%”1. The same forum also highlighted Moonshot AI’s Kimi K2.5, a model that “beats GPT‑5 on reasoning” and introduces an “Agent Swarm” orchestration layer23456. A separate Hacker News thread announced Kimi K3, a 2.8‑trillion‑parameter open‑source model now available through Telnyx’s inference API78.
OpenAI’s half‑price gamble
OpenAI’s decision to halve the cost of GPT‑5.6 Sol is a direct response to mounting pressure from cheaper alternatives. The cut brings the per‑token fee down to an undisclosed level, but the headline alone signals that the company is willing to sacrifice margin to keep volume high. For developers who already pay per‑token, a 50 % reduction can turn a marginally profitable use case into a break‑even or profitable one.
The move also forces competitors to justify higher fees with tangible performance gains. OpenAI has not published new benchmark numbers alongside the price change, leaving the community to wonder whether the cut reflects a genuine efficiency gain or a defensive tactic. Either way, the pricing shift will likely accelerate migration experiments, especially among startups that can’t afford runaway API bills.
Kimi K2.5’s agent swarm and price advantage
Moonshot AI’s Kimi K2.5 arrives with a claim of beating GPT‑5 on reasoning while costing five times less24. According to the Hacker News post, the model can orchestrate up to 100 parallel agents, allowing a single request to spawn dozens of sub‑tasks23456. The post gives the example of asking the model to “analyze 50 competitors” and having it launch 50 research agents simultaneously2.
Pricing for K2.5 is laid out in concrete terms: direct usage costs $0.60 per 1 M input tokens23, but a Gold Plan priced at $30 delivers $90 worth of compute. That works out to an effective $0.20 per 1 M tokens, a rate the post describes as “unparalleled in the industry”. The plan also includes a “Credit Multiplier” that inflates API credits, making the model attractive for heavy‑weight workloads that would otherwise generate costly token loops.
The model’s OpenAI‑compatible endpoint means migration is a simple key swap. Users can upload a screenshot and receive pixel‑perfect React / Tailwind code, a feature highlighted in the announcement2345. The post emphasizes that the switch from OpenAI to RouterLab + K2.5 is “trivial”, positioning the offering as a low‑friction, high‑value alternative for teams already familiar with OpenAI’s API shape.
K3 pushes open‑source into the trillion‑parameter arena
Moonshot AI’s flagship Kimi K3 is billed as the world’s first open‑source model in the 3‑trillion‑parameter class78. The model packs 2.8 trillion parameters, runs on a 1 M‑token context window, and includes native vision capabilities78. Its architecture builds on “Kimi Delta Attention” and “Attention Residuals”, technical details that differentiate it from earlier open‑source releases8.
Benchmark claims place K3 on par with closed‑source frontier models from Anthropic and OpenAI for coding, reasoning, and agentic knowledge work78. The post frames the achievement as evidence that “the model side of the AI competition is solving itself”, suggesting that open‑source projects can now match proprietary offerings without sacrificing performance7.
K3 is hosted on Telnyx‑owned GPU infrastructure and accessed via an OpenAI‑compatible API78. The availability through Telnyx’s Inference API means developers can tap into the model without building their own hardware stack, echoing the broader industry trend of abstracting compute behind managed services.
Industry implications and what to watch
The simultaneous price cut from OpenAI and the launch of two Moonshot AI models create a three‑way tension: cost, capability, and infrastructure control. OpenAI’s lower fees may retain price‑sensitive customers, but K2.5’s agent swarm and K3’s trillion‑parameter scale offer functional advantages that could lure workloads requiring parallelism or vision.
Developers will now weigh token economics against features like parallel agent orchestration and native image handling. The ease of swapping API keys—thanks to OpenAI compatibility—lowers the barrier for rapid experimentation, meaning we may see a wave of proof‑of‑concept projects that benchmark K2.5 and K3 against GPT‑5.6 Sol in real‑world pipelines.
What to watch next is OpenAI’s response beyond pricing. Will the company release a new model tier, adjust token limits, or introduce its own parallel‑agent framework? On the Moonshot side, the adoption metrics for K2.5’s Gold Plan and the volume of K3 requests through Telnyx will indicate whether the market is ready to shift from proprietary to open‑source giants. The next quarter’s usage reports from RouterLab and Telnyx will be the barometer for this emerging pricing‑performance equilibrium.
What to watch: OpenAI’s upcoming token‑pricing updates, Moonshot AI’s adoption numbers for K2.5’s Gold Plan, and Telnyx’s reported traffic to the K3 inference endpoint. These data points will reveal whether cost alone can tip the balance or if the functional edge of agent swarms and trillion‑parameter vision models will drive the next wave of AI infrastructure choices.
Footnotes
Related Articles
Googlebook challenges Apple’s retail play as Mac Studio prices
Google’s $899 AI‑native laptop and Apple’s store‑first philosophy clash amid rising component costs and AI‑driven hardware bets.
Amiga Unix resurfaces on modern hardware
A community revives the 1990 Amiga Unix port with AI‑assisted tools, new drivers, and a modern package manager for classic 68k machines.
Laya runs offline on M4, ChatGPT tracks ads
New offline inference benchmarks, cross‑site data collection, and rescue services raise fresh questions for AI developers.