Thinking Machines has released Inkling-Small, a 276-billion-parameter multimodal reasoning model under an Apache 2.0 license, two weeks after introducing its first open-source model, Inkling.
The company said Inkling-Small comes within one point of Inkling on the Artificial Analysis Intelligence Index, despite Inkling having 975 billion parameters. Artificial Analysis assigned Inkling-Small a score of 40 and Inkling a score of 41.
Thinking Machines said Inkling-Small accepts text, image and audio inputs, produces text, and supports a context window of up to 1 million tokens. The company has released the full weights on Hugging Face and added fine-tuning support through its Tinker API.
At launch, Thinking Machines is offering a limited-time 50% discount on API pricing for the standard 64K-context Inkling-Small model. The discounted rates are $0.58 per million prefill tokens, $1.44 per million sampled tokens, $1.73 per million training tokens, and $0.116 per million cached prefill tokens.
Artificial Analysis reported that no open-weight model at Inkling-Small’s size or smaller scored higher on its index. Thinking Machines said Inkling-Small uses 12 billion active parameters per token, compared with 41 billion for Inkling.
On benchmark tests reported by Thinking Machines, Inkling-Small scored 80.2% on SWE-bench Verified versus 77.6% for Inkling, and 64.7% on Terminal Bench 2.1 versus 63.8% for the larger model. The company also said Inkling-Small scored higher on SciCode, Humanity’s Last Exam, GPQA Diamond and CritPt.
Thinking Machines said Inkling still leads on factual knowledge and some agentic tasks. Inkling-Small scored 15.5% on τ³-Banking compared with 23.7% for Inkling, and the company said its AA Omniscience score was negative.
According to Thinking Machines’ model card, Inkling-Small is a sparse Mixture-of-Experts model with a 42-layer decoder that routes each token to six of 256 specialized experts, alongside two shared experts. The architecture allows the model to activate 12 billion parameters at a time from a total of 276 billion.
Thinking Machines said the model processes images, audio and text in a shared representation rather than through separate external systems. The company also said developers can adjust reasoning effort at test time to trade off quality, latency and cost.
The standard BF16 checkpoint requires at least 600 GB of aggregate GPU memory, according to Thinking Machines. The company listed 4 NVIDIA B300 GPUs or 8 NVIDIA H200 GPUs as supported configurations.
Thinking Machines said a quantized NVFP4 checkpoint lowers the memory requirement to about 180 GB of aggregate VRAM. That version can run in W4A4 mode on a single NVIDIA B300 or in W4A16 mode on two H200 GPUs, the company said.
Thinking Machines researcher Horace He said in a post on X that releasing Inkling-Small “felt much more routine” than Inkling because the team reused the earlier pipeline with a smaller model. Thinking Machines said Inkling-Small benefited from an improved pre-training data mix, changes to its machine-learning recipe, on-policy distillation using Inkling as a teacher, and two weeks of continued agentic coding reinforcement learning.
