BLOG ARTICLE • 24 SEPTEMBER 2026 • SUGGESTED READING TIME: 14–18 MINUTES
Training Infrastructure Questionnaire
A plain-language guide for customers
About this guide
Sizing the right infrastructure for AI training depends on many factors beyond the model itself. The fifteen questions below help us understand your workload so we can recommend hardware, networking and storage that match your needs without overspending. Each question is followed by a short explanation in plain terms. You do not need to answer every question in full; approximate answers and “not sure yet” are both useful starting points.
These questions are intended for full training, continued pre-training and substantial fine-tuning workloads. Smaller PEFT/LoRA projects may require far less infrastructure, but the same discovery logic applies.
1. What model are you planning to train or fine-tune, and how many parameters does it have?
Why it matters: Parameter count drives accelerator memory, compute volume and the feasibility of single-node versus distributed training.
What to capture: Model family, dense or MoE (Mixture of Experts) architecture, total parameters, active parameters per token, custom architecture or standard framework.
Sizing implication: Establishes the base compute and memory scale and whether model parallelism is required.
In simple terms: The model is the AI system you want to train, and its parameters are the numbers that store what it has learned. A “70B” model has about 70 billion of them. The more parameters, the more GPU memory and computing power you need.
Models come in two main designs. A dense model uses all of its parameters for every word it processes. A Mixture of Experts (MoE) model is split into smaller sections called experts, and only a few are used for each word. For MoE models we need two numbers: total parameters (which decide how much memory is needed) and active parameters (which decide how much computing work is done).
Example answer: “Llama 3.1 70B, dense, standard Hugging Face code.”
2. Are you training from scratch, continuing pre-training, or fine-tuning an existing model?
Why it matters: These workloads can differ by orders of magnitude in compute requirement and duration.
What to capture: Training method, LoRA/QLoRA/PEFT versus full-parameter tuning, number of training phases and expected experiments.
Sizing implication: Determines whether a workstation/server is sufficient or whether a cluster/cloud burst model is needed.
In simple terms: Training from scratch means building a model’s knowledge from nothing. It is by far the most expensive option, often needing hundreds or thousands of GPUs for weeks. Continued pre-training means feeding an existing model large volumes of new material, for example a whole industry’s documents. Fine-tuning means teaching an existing model a specific task using a smaller dataset.
Fine-tuning can be done in two ways. Full-parameter tuning updates the whole model. LoRA, QLoRA and other PEFT (Parameter-Efficient Fine-Tuning) methods train only small add-on pieces, which often fits on a single server. Tell us which approach you plan and how many rounds of training or experiments you expect.
3. What numerical precision and optimisation method will be used?
Why it matters: BF16, FP16, FP8 and mixed-precision approaches affect memory footprint, performance and accelerator compatibility.
What to capture: Training precision, optimizer, gradient checkpointing, quantisation-aware training, framework requirements.
Sizing implication: Changes GPU memory requirement, achievable batch size and accelerator selection.
In simple terms: Precision is how many digits each number in the model is stored with. Lower precision formats such as BF16 or FP8 use less memory and run faster, but not every GPU supports every format.
The optimizer is the method used to adjust the model during training. Popular optimizers such as Adam keep extra working data for every parameter, which adds significantly to memory. Gradient checkpointing is a technique that saves memory at the cost of some extra computation. If you are unsure of these details, sharing your training code or framework settings is enough for us to work them out.
4. What is the dataset size and composition?
Why it matters: Training capacity is not sized by model alone; total tokens, images, audio or video determine the amount of work that must pass through the system.
What to capture: Total training tokens or samples, file formats, average sample size, number of epochs, data growth.
Sizing implication: Determines total training FLOPs, storage capacity, preprocessing and I/O requirement.
In simple terms: The model size tells us how heavy each step is; the dataset size tells us how many steps there are. Text is measured in tokens, which are small pieces of words (roughly three-quarters of a word each). An epoch is one full pass through the whole dataset.
Dataset size also affects how much storage you need and how quickly data must be read. Let us know how large the data is today, what format it is in, and how fast it is growing.
5. What training completion time is acceptable?
Why it matters: The same model can be trained on fewer accelerators over a longer period or on more accelerators to reduce time-to-result.
What to capture: Target wall-clock time per training run, deadline, experiment cadence and acceptable queue time.
Sizing implication: Converts compute requirement into the number of accelerators/nodes required.
In simple terms: There is a direct trade-off between time and hardware. A job that needs about 3,000 GPU-hours would take around 15 days on 8 GPUs, or about 4 days on 32 GPUs. Telling us your deadline, and how often you expect to retrain, lets us choose the right number of GPUs rather than over- or under-buying.
6. What global batch size, sequence length and micro-batch size are expected?
Why it matters: These directly influence activation memory, throughput and communication patterns.
What to capture: Sequence length, tokens/sample, batch strategy, gradient accumulation and maximum acceptable batch.
Sizing implication: Defines memory pressure and helps calculate per-GPU batch feasibility.
In simple terms: Sequence length is how much text the model reads in one go, for example 4,096 or 32,000 tokens. Batch size is how many examples the model processes before it updates itself. The micro-batch is the smaller share each GPU handles at one time; gradient accumulation adds several micro-batches together to form the full batch.
Longer sequences and larger micro-batches need more GPU memory. These values are usually set in your training configuration file, so sharing that file is often the easiest answer.
7. How many training experiments will run concurrently?
Why it matters: A cluster sized for one job may become unusable if multiple data scientists need simultaneous experiments.
What to capture: Number of teams, concurrent jobs, development versus production training, expected scheduling policy.
Sizing implication: Determines capacity headroom and whether a scheduler/multi-tenant cluster is required.
In simple terms: If several people or teams will train models at the same time, the system must be large enough for all of them, or they will wait in a queue. Shared clusters usually need a scheduler, such as Slurm or Kubernetes, to divide GPUs fairly between users and priorities.
8. What framework and software ecosystem are mandatory?
Why it matters: CUDA, ROCm, PyTorch, TensorFlow, JAX, DeepSpeed, Megatron-LM and vendor libraries can constrain accelerator choice.
What to capture: Framework versions, custom kernels, dependencies, containers and existing codebase.
Sizing implication: Prevents selecting hardware that is technically capable but incompatible with the customer software stack.
In simple terms: Frameworks are the software tools your team uses to write and run training code. Some software runs only on NVIDIA GPUs (through CUDA), while some also supports AMD GPUs (through ROCm). Choosing hardware your existing code cannot run on would be costly, so please list the tools, versions and any custom code you depend on.
9. How will the model be parallelised across accelerators?
Why it matters: Large models may require data, tensor, pipeline, expert or sequence parallelism, each with different communication behaviour.
What to capture: Parallelism strategy, current distributed-training code, preferred orchestration framework.
Sizing implication: Determines the importance of GPU-to-GPU fabric, node count and scale-out network.
In simple terms: When a model or job is too large for one GPU, the work is split across many. In data parallelism, each GPU holds a copy of the model and works on different data. In tensor and pipeline parallelism, the model itself is divided between GPUs. Expert parallelism places different experts of an MoE model on different GPUs, and sequence parallelism splits very long inputs.
The more a job is split, the more the GPUs need to talk to each other, which affects the network design. If your team already has distributed training code, it tells us which approach is used.
10. What interconnect performance is required between GPUs and between nodes?
Why it matters: Distributed training frequently becomes communication-bound if accelerators cannot exchange gradients/activations fast enough.
What to capture: NVLink/NVSwitch needs, Ethernet or InfiniBand, link speed, RDMA, topology and oversubscription tolerance.
Sizing implication: Drives node architecture, NIC count, switch design and cluster fabric cost.
In simple terms: The interconnect is the set of high-speed links GPUs use to share results. Inside a server, GPUs connect through technologies such as NVLink. Between servers (nodes), they connect over InfiniBand or high-speed Ethernet, often using RDMA for direct memory transfers.
If these links are too slow, expensive GPUs spend time waiting instead of working. Oversubscription means several links share a smaller uplink; it lowers cost but can slow large jobs.
11. What sustained storage throughput must feed the training cluster?
Why it matters: Expensive accelerators can sit idle if datasets and checkpoints cannot be delivered fast enough.
What to capture: Dataset read rate, checkpoint size/frequency, local NVMe cache, shared filesystem/object storage, metadata intensity.
Sizing implication: Sizes NVMe, parallel storage, object storage and east-west network capacity.
In simple terms: Throughput is how fast storage can deliver data to the GPUs, and how fast it can save results. Training also writes checkpoints, which are saved copies of the model’s progress; for large models these can be hundreds of gigabytes or more each time.
Datasets with millions of small files put extra load on storage, which is what “metadata intensity” refers to. Knowing your data and checkpoint patterns lets us size fast local drives (NVMe) and shared storage correctly.
12. What CPU and system-memory preprocessing is required?
Why it matters: Tokenisation, augmentation, decoding and data loaders can become CPU or RAM bottlenecks before the GPU is fully utilised.
What to capture: CPU threads/processes, preprocessing libraries, RAM working set, NUMA sensitivity, host-to-device transfer pattern.
Sizing implication: Determines CPU socket/core count, RAM capacity and PCIe lane requirements.
In simple terms: Before data reaches the GPUs, the server’s regular processors (CPUs) prepare it: splitting text into tokens, decoding images or video, and applying changes such as cropping. If the CPUs or memory (RAM) cannot keep up, the GPUs sit idle.
This is especially important for image, audio and video workloads. Telling us what preprocessing happens during training helps us choose the right CPUs, memory and internal connections (PCIe).
13. What checkpointing, resilience and restart requirements exist?
Why it matters: Long training jobs need protection from hardware failure; checkpoint frequency affects storage and network load.
What to capture: Maximum acceptable lost compute, checkpoint interval, checkpoint size, replication and restart process.
Sizing implication: Adds storage capacity/throughput and may require redundant nodes/fabric.
In simple terms: On large clusters running for days or weeks, hardware failures are expected rather than rare. Checkpoints let a job restart from its last save instead of from the beginning.
The key question is how much work you can afford to lose. Saving every hour means at most an hour lost, but more frequent saves add load on storage and network. Some customers also keep spare nodes ready so jobs can resume quickly.
14. Where must the training run, and are there data-sovereignty or security constraints?
Why it matters: Sensitive datasets, IP or regulated data can rule out certain cloud or multi-tenant options.
What to capture: On-premises/colo/cloud preferences, residency, encryption, isolation, Internet restrictions and model/data export policy.
Sizing implication: Determines deployment model and security architecture before hardware sizing is finalised.
In simple terms: Training can run in your own data centre (on-premises), in hardware you place in a third-party facility (colocation), or in the cloud. Data residency rules may require data to stay within a particular country, and some organisations cannot share hardware with other tenants or connect training systems to the Internet.
These requirements can narrow the options before sizing begins, so it helps to know them early, along with any rules on moving trained models or data elsewhere.
15. What is the expected utilisation and three-year training roadmap?
Why it matters: A large on-premises cluster is economical only when it will be used; sporadic jobs may be better suited to cloud or hybrid capacity.
What to capture: Runs/month, hours/run, annual growth, future model sizes, budget horizon and internal operations capability.
Sizing implication: Enables TCO comparison among purchase, cloud GPU, reserved capacity and hybrid architectures.
In simple terms: Buying GPUs is cost-effective only if they are kept busy. If training happens a few times a year, renting cloud GPUs may be cheaper; if it runs constantly, owning hardware usually wins. Many organisations use a mix (hybrid).
TCO (Total Cost of Ownership) includes not just the purchase price but also power, cooling, space, support and the staff needed to run the system. Sharing how often you train today, and how you expect models and usage to grow over three years, lets us compare these options fairly.
Quick glossary
| Term | Meaning |
| Accelerator / GPU | A specialised processor that performs the heavy calculations in AI training. |
| Parameter | One of the numbers inside a model that stores what it has learned. |
| Token | A small piece of text, roughly three-quarters of a word, used to measure text data. |
| Dense model | A model that uses all its parameters for every word it processes. |
| MoE (Mixture of Experts) | A model split into experts, of which only a few are used for each word. |
| Node | A single server, typically containing 4 or 8 GPUs. |
| Cluster | Many nodes connected by a high-speed network to work as one system. |
| FLOPs | Floating-point operations; the unit used to measure total computing work. |
| Fine-tuning | Adapting an existing model to a specific task or domain. |
| LoRA / QLoRA / PEFT | Lightweight fine-tuning methods that train small add-ons instead of the full model. |
| Checkpoint | A saved copy of the model’s progress, used to resume training after a failure. |
| TCO | Total Cost of Ownership: purchase plus power, cooling, space, support and staff. |