DatatalkwithAnkit
Ankit Dongare on the National Mall, with the US Capitol behind him
Washington, DC

Ankit Dongare, AI engineer

I measure what language models cost to run.

I benchmark LLM inference: GPUs, serving configs, throughput, and price. When I run a benchmark, I publish the configs with it so you can rerun it and check my numbers.

Will your model fit?

Pick a model size and precision to see how many GPUs it takes to hold, and what that costs per hour on the cheapest published neocloud rates.

Model size, in billions of parameters
Precision

NVIDIA H200141 GB per GPU
AMD MI300X192 GB per GPU

Rough estimate. Rates are the low end of published prices, Aug 2026: H200 $2.60/hr, MI300X $1.71/hr. How I got these numbers

Writing

GPU comparisons, inference benchmarks, and notes on the AI market. Posts go up when the work is done.

Latest post

H200 vs MI300X: what you're actually paying for

AMD sells memory, NVIDIA sells interconnect and a software stack that already works. Which is the better buy depends on whether your problem is fitting the model or making many GPUs act like one.

H200
141 GB
MI300X
192 GB

Memory per GPU

Coming soon

Quantization in practice: quality vs. cost on consumer GPUs

What happens to output quality and tokens per second going from FP16 to INT8 and INT4, measured on real prompts.

In progress
Coming soon

How batch size changes vLLM throughput

A single-GPU sweep of max_num_seqs, measuring tokens per second against p50 and p95 latency.

In progress
Coming soon

KV cache is the real memory budget

Why context length, not parameter count, usually decides how many users one GPU can hold.

In progress

See everything I'm writing

Projects

Things I've built and run. More on GitHub.

My product

Aurapal

A career platform I design, build, and run end to end. It's where my engineering, automation, and product work come together.

Visit aurapal.org
  • Built soloFrom design and code to deployment and support.
  • Automation-firstRepetitive work is scripted, so the time goes into the product.
  • LiveRunning in production at aurapal.org.

Automated YouTube pipeline

Generates, renders, captions, and publishes videos with almost no manual steps.

Python, FFmpeg, Whisper, YouTube Data API
View on GitHub

LLM inference benchmarks

An ongoing series on how batch size, quantization, and context length change throughput, latency, and cost.

vLLM, PyTorch, Python
Read the series

EWG, Elite White Glove

A conversion-focused website for a white-glove service business, from design to deployment.

HTML, CSS, Vercel
View source

HNH Handyman

A booking-focused web app for a handyman service, with a typed, maintainable codebase.

TypeScript
View source

About

I started as a data analyst turning messy tables into answers, and kept following the harder questions until they led to building AI systems full time.

Today I work on the practical side of large language models: how they're served, how fast they run, what they cost, and how configuration choices change all three. Most of my week goes into benchmarking models and tuning serving stacks like vLLM.

That data background still shapes how I work. I trust measurements over opinions, and if I do something twice, I write a script for it.

Day to day: Python, PyTorch, vLLM, Hugging Face, Docker, and AWS. The full list is on the uses page.

Ankit Dongare outdoors with a dog Ankit Dongare with the New York City skyline behind him

Stay in touch

Get new posts by email, or send me a note about a benchmark, a collaboration, or Aurapal.

The newsletter

New benchmarks and write-ups, sent when they're published. No schedule and no filler. Unsubscribe any time.

Signups come straight to me while the list is small.

Send a message

I usually reply within two days. Or email dongare.ankit29@gmail.com.