BLOG ARTICLE • 26 SEPTEMBER 2026 • SUGGESTED READING TIME: 12–15 MINUTES
Inference Infrastructure Questionnaire
Inference Infrastructure Sizing Questionnaire
A plain-language guide for customers
About this guide
Inference is the stage where a trained AI model is put to work: answering questions, summarising documents, writing code or processing requests for real users. Sizing inference infrastructure is fundamentally a service-capacity problem. It must account for memory residency, prefill, decode, concurrency, queueing, latency and availability, not merely whether the model fits into GPU memory.
The fifteen questions below help us understand how your AI service will be used, so we can recommend infrastructure that responds quickly, stays available and is cost-effective. Each question is followed by a short explanation in plain terms. Approximate answers and “not sure yet” are both useful starting points.
1. What business application or inference use case will the platform serve?
Why it matters: Chat, RAG, coding, agents, vision, speech, classification and batch processing have very different request patterns.
What to capture: Interactive versus batch use, application type, criticality and expected user experience.
Sizing implication: Sets the latency, concurrency and throughput profile that the infrastructure must satisfy.
In simple terms: Different applications place very different demands on the system. A chatbot has people waiting for each reply, so speed matters most. A batch job, such as summarising a million documents overnight, cares about total volume rather than the speed of any single request. AI agents often make many model calls to complete one task.
Tell us what the application does, who uses it, how business-critical it is, and what experience users expect.
2. Which model or models will be served?
Why it matters: Parameter count, architecture and model family determine weight memory, compute intensity and runtime support.
What to capture: Model names, dense/MoE, parameter count, custom fine-tuned variants, number of simultaneously resident models.
Sizing implication: Establishes baseline accelerator memory and whether model sharding is required.
In simple terms: A model’s weights (its learned parameters) must stay loaded in GPU memory the whole time it is serving requests. A 70-billion-parameter model needs about 140 GB at standard precision, which is more than a single 80 GB GPU can hold, so it must be split (sharded) across several GPUs.
For Mixture of Experts (MoE) models, all experts must be loaded even though only some are used per word. If you plan to serve several models or fine-tuned versions at the same time, each one needs its own memory.
3. What precision or quantisation will be used for inference?
Why it matters: FP16/BF16, FP8, INT8 and INT4 can dramatically change memory use and throughput.
What to capture: Permitted precision, quality tolerance, quantisation method and validation requirements.
Sizing implication: Changes how many model replicas fit per accelerator and the achievable tokens/sec.
In simple terms: Quantisation stores the model’s numbers with fewer bits, making it smaller and faster. A 70B model needs about 140 GB at 16-bit, about 70 GB at 8-bit and roughly 35 to 40 GB at 4-bit.
Smaller models need fewer GPUs and respond faster, but heavy quantisation can slightly reduce answer quality. Let us know how much quality trade-off is acceptable and whether you will test quantised models on your own tasks before approving them.
4. What is the maximum and average context length?
Why it matters: Long contexts increase KV-cache memory and prefill compute even when the model weights are unchanged.
What to capture: Average, P95 and maximum context tokens, conversation history policy and RAG payload size.
Sizing implication: Directly affects memory headroom, concurrency and prefill capacity.
In simple terms: The context is all the text the model considers for one request: its instructions, the conversation so far, any documents retrieved for it, and the answer it is writing. Text is measured in tokens, roughly three-quarters of a word each.
For every active request, the model keeps a working memory of this context called the KV cache. With long contexts and many users at once, the KV cache can need more GPU memory than the model itself. P95 means the length that 95% of requests stay within, which is more useful for sizing than the average alone.
5. What is the average and peak number of input tokens per request?
Why it matters: Input processing is the prefill phase and can dominate latency for long prompts or document-heavy RAG.
What to capture: System prompt, user prompt, retrieved context, history and P95/peak input tokens.
Sizing implication: Sizes prompt-processing capacity and helps distinguish prefill-bound from decode-bound workloads.
In simple terms: Before the model writes anything, it must read the whole input. This reading stage is called prefill. A short question is read almost instantly, but a 20-page document takes noticeably longer.
The input includes more than what the user types: hidden instructions (the system prompt), previous messages and any documents retrieved by a search (RAG, or Retrieval-Augmented Generation). Please estimate all of these together.
6. What is the average and peak number of output tokens per request?
Why it matters: Long generated responses occupy decode capacity for longer and reduce the number of concurrent requests a server can sustain.
What to capture: Average, P95 and maximum output length by use case.
Sizing implication: Drives decode throughput, session duration and total token capacity.
In simple terms: The model writes its answer one token at a time, a stage called decode. Longer answers keep the GPU busy for longer, which means fewer users can be served at the same moment.
A classification task that answers in a few words and a report generator that writes 2,000 words place very different loads on the system, even with the same model.
7. How many total users and how many simultaneous active users are expected?
Why it matters: Registered users are not a sizing metric; concurrent active requests are.
What to capture: Named users, daily active users, simultaneous users and concurrency at normal/P95/peak periods.
Sizing implication: Determines replica count, accelerator count and queueing headroom.
In simple terms: What matters for sizing is how many requests are being processed at the same moment, not how many people have access. For example, 10,000 employees may have access, 2,000 may use it on a given day, and perhaps 100 to 200 have a request in progress at the busiest moment.
The number of simultaneous requests decides how many copies of the model (replicas) and GPUs are needed.
8. What request rate must be supported?
Why it matters: Concurrency alone is insufficient because request arrival rate and session duration determine queue behaviour.
What to capture: Requests per second/minute, average and burst rate, batch jobs and scheduled peaks.
Sizing implication: Defines aggregate service capacity and autoscaling/reservation requirements.
In simple terms: Think of a supermarket checkout: the queue length depends on how many customers arrive per minute and how long each takes to serve. The same applies here.
Tell us how many requests arrive per second or minute on average and during bursts, such as Monday mornings, month-end processing or scheduled batch jobs.
9. What Time to First Token target is required?
Why it matters: Interactive applications are highly sensitive to the delay before the first response appears.
What to capture: Average/P95 TTFT SLA by application and geography.
Sizing implication: Influences prefill capacity, batching policy, network placement and required headroom.
In simple terms: Time to First Token (TTFT) is the wait between a user sending a request and the first words of the answer appearing. For chat applications, a delay of more than a second or two starts to feel slow.
TTFT depends on input length, how busy the system is, and how far users are from the data centre. Targets for different applications or regions can differ.
10. What output generation speed is acceptable?
Why it matters: Tokens per second determines how quickly users receive the completed response after generation begins.
What to capture: Minimum tokens/sec per session and aggregate tokens/sec target.
Sizing implication: Sizes decode capacity and enables comparison between GPU servers and specialised inference platforms.
In simple terms: Once the answer starts appearing, generation speed is measured in tokens per second. People read at roughly 5 to 8 tokens per second, so around 20 to 30 or more feels smooth for chat. Coding tools and agents, where software rather than a person reads the output, often need much higher speeds.
We also need the total tokens per second across all users, which sets the overall capacity of the platform.
11. What end-to-end response-time SLA is required?
Why it matters: Model serving is only one component; retrieval, reranking, networking and application logic also consume latency budget.
What to capture: P50/P95/P99 completion-time targets and acceptable queueing delay.
Sizing implication: Determines overall headroom, architecture and whether dedicated capacity is required.
In simple terms: An SLA (Service Level Agreement) is a promised level of performance. End-to-end time is everything the user waits for: searching documents, network travel, application processing and the model itself. If the whole answer must arrive within 5 seconds and search takes 1 second, the model has only 4 seconds left.
P50 is the typical experience, while P95 and P99 describe the slowest 5% and 1% of requests. Strict P99 targets need extra spare capacity.
12. Does the application include embeddings, reranking, vision, speech or other models?
Why it matters: Production AI applications often contain several inference stages, each needing separate compute and memory.
What to capture: Embedding model, reranker, OCR/vision encoder, speech models, guardrails and moderation services.
Sizing implication: Prevents sizing only the LLM while ignoring adjacent AI workloads.
In simple terms: Real applications usually use several AI models together. Embedding models turn documents into a searchable form. Rerankers pick the most relevant search results. OCR and vision models read scanned pages and images. Speech models convert between voice and text. Guardrails and moderation check requests and answers for unsafe content.
Each of these needs its own computing resources. They are usually smaller than the main model, but leaving them out leads to an undersized system.
13. What availability, redundancy and maintenance behaviour are required?
Why it matters: A system that must survive node or accelerator failure cannot be sized at 100% normal utilisation.
What to capture: Target availability, N+1/N+N policy, maintenance windows, failover, multi-site requirements.
Sizing implication: Adds replicas, spare capacity and potentially multi-zone/site architecture.
In simple terms: Availability is how much downtime is acceptable. 99.9% availability allows about 8.8 hours of downtime a year; 99.99% allows only about 53 minutes. Higher availability needs spare capacity.
N+1 means one spare server beyond what normal load needs; N+N means a complete duplicate. Tell us whether the system must keep running during maintenance or the loss of a whole site.
14. Where must inference occur and what data/privacy restrictions apply?
Why it matters: The answer determines whether the customer can use an API, cloud GPU, dedicated capacity or only local infrastructure.
What to capture: On-prem/cloud/API preference, data residency, prompt retention policy, private connectivity, encryption and Internet dependency.
Sizing implication: Defines the eligible deployment models before comparing performance and cost.
In simple terms: Options range from calling a provider’s API over the Internet, to renting cloud GPUs, to dedicated hosted capacity, to running on your own hardware. Sensitive data may require that information stays within a country, that prompts are not stored by a provider, that connections are private, or that the system works without the Internet.
These rules decide which options are allowed before cost and performance are compared, so it helps to know them early.
15. What is the expected daily/monthly token volume and utilisation profile?
Why it matters: Economics change dramatically between bursty low-volume use and continuous high-volume inference.
What to capture: Input/output tokens per day/month, hourly demand curve, business-hour peaks, growth forecast and seasonality.
Sizing implication: Allows a true TCO comparison using cost per useful inference rather than only server price or GPU-hour rate.
In simple terms: Total volume and how it varies over time largely decide the most economical option. Low or irregular usage is often cheaper on pay-per-use APIs or cloud; steady, high usage is often cheaper on dedicated or owned hardware.
TCO (Total Cost of Ownership) includes power, cooling, support and staff as well as hardware. The fairest comparison is the cost per useful answer, rather than the price of a server or a GPU-hour alone. Please share expected growth and any seasonal peaks.
Quick glossary
| Term | Meaning |
| Inference | Using a trained model to respond to requests. |
| Token | A small piece of text, roughly three-quarters of a word. |
| Context length | All the text the model considers for one request, including its answer. |
| Prefill | The stage where the model reads the input before answering. |
| Decode | The stage where the model writes its answer, one token at a time. |
| KV cache | The model’s working memory for each active request; grows with context length. |
| Quantisation | Storing model numbers with fewer bits to save memory and increase speed. |
| Replica | A complete, independent copy of a model serving requests. |
| Concurrency | The number of requests being processed at the same moment. |
| TTFT | Time to First Token: the wait before the answer starts appearing. |
| P50 / P95 / P99 | The response time that 50%, 95% or 99% of requests stay within. |
| RAG | Retrieval-Augmented Generation: searching documents and giving them to the model with the question. |
| Embedding / reranker | Supporting models that make documents searchable and rank the best results. |
| SLA | Service Level Agreement: a promised level of performance or availability. |
| N+1 / N+N | One spare unit beyond need, or a full duplicate, for resilience. |
| TCO | Total Cost of Ownership: hardware plus power, cooling, space, support and staff. |