Gemini 3.5 Flash vs Inkling Small 256K

rates, context and capabilities side by side on the impossibl api, and what the same workload costs on each.

the models at a glance: base rates, context window and modalities
specificationGemini 3.5 Flash

google/gemini-3.5-flash

Inkling Small 256K

thinkingmachines/inkling-small-256k

input · usd / 1m tok$1.50$1.16
output · usd / 1m tok$9.00$2.89
context window1mtokens262.14ktokens
modalities · in → out

takes text, image, audio and video in, gives text out

takes text and image in, gives text out

pricing

base api rates by model
base api rates · usd per 1m tokensGemini 3.5 Flash

google/gemini-3.5-flash

Inkling Small 256K

thinkingmachines/inkling-small-256k

input$1.50$1.16
output$9.00$2.89
cache read$0.15$0.232

the estimate below prices one workload at these rates on every model.

specifications

published specifications by model
as publishedGemini 3.5 Flash

google/gemini-3.5-flash

Inkling Small 256K

thinkingmachines/inkling-small-256k

creatorGoogleThinking Machines
model typeText generationText generation
context window1,000,000tokens262,144tokens
input

takes text, image, audio and video in, gives text out

takes text and image in, gives text out

outputtexttext
tool callingnot verifiednot verified
releasedmay 19, 2026not published
endpoint/v1/chat/completions/v1/chat/completions

unknown means the public catalog does not establish that fact. a model's documented capabilities can differ from what a particular api supports. see the model documentation for integration details.

benchmarks

benchmark scores by model, primary tested configuration
primary configurationGemini 3.5 Flash

google/gemini-3.5-flash

high

Inkling Small 256K

thinkingmachines/inkling-small-256k

no benchmarks yet

intelligence index · gateway rank33.0#31 of 90not measured
coding index70.1#23 of 60not measured
agentic index27.3#30 of 49not measured
omniscience21.2#15 of 86not measured
gpqa92.2%#22 of 88
not measured
humanity’s last exam42.7%#17 of 89
not measured
terminal-bench v2.178.7%#23 of 57
not measured
tau295.3%
not measured
mmmu-pro84.3%
not measured
harvey lab82.1%
not measured
ifbench76.3%
not measured
aa-lcr73.3%
not measured
scicode53.9%
not measured
apex-agents47.1%
not measured
analystagent45.0%
not measured
automationbench partial42.1%
not measured
terminal-bench hard40.9%
not measured
itbench40.3%
not measured
tau banking32.2%
not measured
gdp.pdf all-pass19.8%
not measured
multilingual lcr18.3%
not measured
critpt13.1%
not measured
terminal-bench v4.06.6%
not measured

Gemini 3.5 Flash source ↗

each column describes the tested configuration named under the model; performance and costs can differ on impossibl. a score one model was not measured on is left blank rather than scored zero. indices are ranked against every model on the gateway with one. last checked 2026-09-11.

estimate your cost

one token workload, priced on every model at the published rates, in usd. change the numbers to match yours.

token workload

total input, including cached tokens

include billable reasoning tokens

the same token usage in each request

a subset of total input, never additional tokens

Gemini 3.5 Flash

google/gemini-3.5-flash

estimated usage cost

$6.00

1,000 requests × (1,000 uncached input × $1.50 + 500 output × $9.00) / 1,000,000 = $6.00

base rates; no context pricing tiers are published.

Inkling Small 256K

thinkingmachines/inkling-small-256k

estimated usage cost

$2.605

1,000 requests × (1,000 uncached input × $1.16 + 500 output × $2.89) / 1,000,000 = $2.605

base rates; no context pricing tiers are published.

estimates use published impossibl api usage rates. configured billing adjustments, funding fees, taxes, and custom workspace rates are excluded and may change the final amount.

image and audio usage, cache creation, and additional tool charges are excluded. reasoning tokens count as billable output, even when they are not visible in the response.

summary

Gemini 3.5 Flash and Inkling Small 256K are available through the Impossibl API. they share the /v1/chat/completions endpoint; the model identifier selects the model for a request.

Gemini 3.5 Flash is created by Google. It has a 1,000,000-token context window. Base API pricing (USD): $1.50 per million input tokens; $9.00 per million output tokens.

Inkling Small 256K is created by Thinking Machines. It has a 262,144-token context window. Base API pricing (USD): $1.16 per million input tokens; $2.89 per million output tokens.

Gemini 3.5 FlashInkling Small 256K

this page compares Gemini 3.5 Flash vs Inkling Small 256K as listed by the impossibl public api. prices and availability come from that catalog, refreshed here every five minutes. this is an impossibl offering, not a comparison of independent hosting-provider quotes.

view the source catalog ↗comparison as markdown ↗report a correction →