Thinking Machines, the AI startup founded by former OpenAI executive Mira Murati, has expanded its model lineup with the introduction of Inkling-Small. This release arrives just two weeks after the debut of its initial open-source language model, Inkling. Despite being roughly one-fourth the size of the original 975-billion-parameter flagship, the new model manages to maintain near-identical performance levels while surpassing its predecessor in specific technical benchmarks.
According to VentureBeat, Inkling-Small is a 276-billion-parameter multimodal reasoning engine capable of processing text, image, and audio inputs. It utilizes an Apache 2.0 license and supports a context window of up to one million tokens. While the model remains too large to operate on standard consumer hardware, its 12-billion active parameters—a significant reduction from the 41-billion active parameters in the flagship—drastically lower the barrier for enterprise deployment. This optimization makes the model a more viable option for organizations that require powerful reasoning capabilities without the massive GPU overhead associated with the original Inkling.
Benchmark data underscores the efficiency of this new release. The Artificial Analysis Intelligence Index granted the smaller model a score of 40, narrowly trailing the flagship’s 41, while simultaneously outperforming the larger model on specialized evaluations like SWE-bench Verified and Terminal Bench 2.1. While the larger Inkling maintains a slight edge in factual knowledge and certain agentic tasks, the efficiency gains in coding and multimodal performance make Inkling-Small a compelling alternative for developers. Thinking Machines has published the model weights on Hugging Face and is offering promotional pricing for its Tinker API to encourage enterprise adoption.
Reader Discussion & Insights