Recommended vLLM runtime configuration
Use the recommended vLLM startup parameters when deploying supported foundation models for IBM Bob self-hosted.
Use the recommended vLLM startup parameters when deploying supported foundation models for IBM Bob self-hosted.
The recommended configurations focus on:
- Speculative decoding to reduce latency.
- Prefix caching and chunked prefill to improve long-context performance.
- Model-specific acceleration backends for optimal throughput.
- Efficient memory utilization to support large context windows and higher concurrency.
The vLLM startup arguments in this section are based on internal validation testing conducted with NVIDIA H200 GPUs and are provided as recommended configuration values. These settings are not hardcoded requirements and might not be suitable for all deployment scenarios. Evaluate and adjust configuration values to accommodate the specific hardware platform, model variant, performance objectives, and workload characteristics of your environment before deploying in production.
The models listed on this page support text-based inputs only and do not support image processing or multimodal image analysis. Do not upload screenshots, architecture diagrams, application diagrams, scanned documents, or other image-based content as model input. If information is contained in an image, convert it to text before submitting it to the model.
Nvidia Nemotron 3 Ultra (FP8, NV-FP4)
Nemotron 3 Ultra 550B is a hybrid state-space model (SSM) and attention-based model. The FP8 and NV-FP4 variants use the same vLLM configuration. Only the model location and served model name differ between variants.
For optimal performance, enable speculative decoding and Mamba-specific optimizations.
FP8 configuration
args:
- --tensor-parallel-size=8
- --enable-expert-parallel
- --max-model-len=262144
- --max-num-seqs=32
- --max-num-batched-tokens=32768
- --gpu-memory-utilization=0.90
- --enable-chunked-prefill
- --enable-prefix-caching
- --mamba-cache-mode=align
- --mamba-ssm-cache-dtype=float16
- --mamba-backend=flashinfer
- --enable-mamba-cache-stochastic-rounding
- --mamba-cache-philox-rounds=5
- '--speculative-config={"method": "nemotron_h_mtp", "num_speculative_tokens": 5}'
- '--model-loader-extra-config={"enable_multithread_load": true, "num_threads": 64}'
- --reasoning-parser=nemotron_v3
- --tool-call-parser=qwen3_coder
- --enable-auto-tool-choice
- --enable-prompt-tokens-details
- --trust-remote-codeNV-FP4 configuration
args:
- --tensor-parallel-size=8
- --enable-expert-parallel
- --max-model-len=262144
- --max-num-seqs=32
- --max-num-batched-tokens=32768
- --gpu-memory-utilization=0.90
- --enable-chunked-prefill
- --enable-prefix-caching
- --mamba-cache-mode=align
- --mamba-ssm-cache-dtype=float16
- --mamba-backend=flashinfer
- --enable-mamba-cache-stochastic-rounding
- --mamba-cache-philox-rounds=5
- '--speculative-config={"method": "nemotron_h_mtp", "num_speculative_tokens": 5}'
- '--model-loader-extra-config={"enable_multithread_load": true, "num_threads": 64}'
- --reasoning-parser=nemotron_v3
- --tool-call-parser=qwen3_coder
- --enable-auto-tool-choice
- --enable-prompt-tokens-details
- --trust-remote-codeWhy these settings are used
- Speculative decoding: The
nemotron_h_mtpspeculative decoding method generates draft tokens and verifies them in a single pass. This configuration can improve response latency and reduce time to first token (TTFT) under concurrent workloads. - Mamba backend: The
flashinferbackend is recommended for Nemotron hybrid SSM layers. Combined with stochastic rounding, it helps provide stable throughput and predictable memory consumption. - Chunked prefill and prefix caching: Chunked prefill processes large prompts in smaller segments, while prefix caching reuses previously computed prompt prefixes. These options help improve performance for long-context workloads.
- Multi-threaded model loading: Using 64 loading threads can reduce model startup times when model weights are stored on persistent storage. This setting does not affect inference performance after the model is loaded.
Laguna S 2.1 (FP8)
Laguna S 2.1 is a mixture-of-experts (MoE) model that uses dFlash speculative decoding. The following configuration was validated with FP8-quantized weights on eight GPUs.
Configuration
args:
- --trust-remote-code
- --tensor-parallel-size=8
- --tool-call-parser=poolside_v1
- --enable-auto-tool-choice
- --reasoning-parser=poolside_v1
- --enable-prompt-tokens-details
- --max-num-batched-tokens=16384
- --enable-prefix-caching
- '--default-chat-template-kwargs={"preserve_thinking": true, "enable_thinking": true}'
- '--speculative-config={"num_speculative_tokens": 7, "method": "dflash", "model": "/mnt/models/dflash-fp8"}'
- --moe-backend=tritonWhy these settings are used
- dFlash speculative decoding: The dFlash draft model generates seven speculative tokens for each decoding step. The primary model validates these tokens in a single pass, reducing latency and improving throughput.
- Triton MoE backend: The
tritonbackend is optimized for mixture-of-experts workloads and can provide more stable performance under concurrent load. - Prefix caching: Prefix caching reduces redundant computation by reusing common prompt prefixes across multiple requests.
Mistral Medium 3.5 (FP8)
Mistral Medium 3.5 is a dense model that uses Eagle speculative decoding.
Configuration
args:
- --tensor-parallel-size=8
- --tool-call-parser=mistral
- --enable-auto-tool-choice
- --reasoning-parser=mistral
- --max-num-batched-tokens=32768
- --max-num-seqs=32
- --kv-cache-dtype=fp8
- --max-model-len=234800
- --gpu-memory-utilization=0.92
- '--speculative-config={"model": "/mnt/models/eagle", "num_speculative_tokens": 3, "method": "eagle"}'Why these settings are used
- FP8 KV-cache compression: The
--kv-cache-dtype=fp8setting reduces KV-cache memory requirements, enabling support for large context windows without modifying model weights. - Eagle speculative decoding: The Eagle draft model predicts three tokens ahead of the primary model. The primary model verifies the generated tokens in a single pass, helping reduce time to first token and response latency.
- GPU memory utilization: A GPU memory utilization target of 92% maximizes available GPU resources while maintaining stable operation for the tested workload.
Performance tuning
You might need to adjust the following parameters to meet your performance and capacity requirements:
| Parameter | Description |
|---|---|
--tensor-parallel-size | Scales inference across multiple GPUs. |
--max-model-len | Controls the maximum supported context length. |
--max-num-seqs | Increases or decreases concurrent request capacity. |
--max-num-batched-tokens | Controls batching behavior. |
--gpu-memory-utilization | Balances memory usage and stability. |
--speculative-config | Optimizes latency and throughput. |
--enable-prefix-caching | Improves performance for repeated prompts. |
--enable-chunked-prefill | Improves processing of large prompts. |
After deployment, monitor latency, throughput, GPU utilization, and memory consumption, and adjust configuration values as required for your workload.
Configuring the Model Gateway
Learn how to create and manage a Model Gateway configuration file, including provider settings, model definitions, credential management, TLS configuration, and deployment examples for supported model providers.
Validation before installation
Validate model connectivity, credentials, certificates, and configuration settings before installation to identify and resolve Model Gateway issues before deploying Bob self-hosted.