vLLM/Recipes
Moonshot AI

moonshotai/Kimi-K2.5

Open-source native multimodal agentic MoE model with vision-language understanding, tool calling, and thinking modes

Multimodal agentic MoE model with DeepSeek-V3 backbone and MLA attention

moe1T / 32B262,144 ctxvLLM 0.19.1+multimodaltext
Guide

Overview

Kimi K2.5 is an open-source, native multimodal agentic model built through continual pretraining on approximately 15 trillion mixed visual and text tokens atop Kimi-K2-Base. It seamlessly integrates vision and language understanding with advanced agentic capabilities, instant and thinking modes, as well as conversational and agentic paradigms.

Prerequisites

  • vLLM version: >= 0.15.0 (speculative decoding with Eagle3 requires >= 0.18.0)
  • Hardware (BF16): 8x H200 GPUs (verified), or equivalent aggregate VRAM (~640 GB)
  • Hardware (NVFP4): 4x Blackwell GPUs (e.g. GB200)
  • AMD support: 8x MI300X / MI325X / MI355X with ROCm 7.2.1 and Python 3.12

Install vLLM

Pip (NVIDIA):

uv venv
source .venv/bin/activate
uv pip install vllm --torch-backend auto

Pip (AMD ROCm):

uv venv --python 3.12
source .venv/bin/activate
uv pip install vllm --extra-index-url https://wheels.vllm.ai/rocm

Docker (NVIDIA):

docker pull vllm/vllm-openai:latest

AMD MI300X/MI325X

On 8x MI300X or MI325X (gfx942), use the standard W4A16 MoE path with AITER and INT4 QuickReduce.

export VLLM_ROCM_USE_AITER=1
export VLLM_ROCM_QUICK_REDUCE_QUANTIZATION=INT4

vllm serve moonshotai/Kimi-K2.5 \
  --host 0.0.0.0 \
  --port 8000 \
  --trust-remote-code \
  --tensor-parallel-size 8 \
  --tool-call-parser kimi_k2 \
  --enable-auto-tool-choice \
  --reasoning-parser kimi_k2 \
  --mm-encoder-tp-mode data

AMD MI350X/MI355X

On 8x MI350X or MI355X (gfx950), add --moe-backend flydsl to use the optimized FlyDSL W4A16 MoE kernel. Keep LoRA disabled for this path.

export VLLM_ROCM_USE_AITER=1
export VLLM_ROCM_QUICK_REDUCE_QUANTIZATION=INT4

vllm serve moonshotai/Kimi-K2.5 \
  --tensor-parallel-size 8 \
  --trust-remote-code \
  --mm-encoder-tp-mode data \
  --moe-backend flydsl \
  --compilation-config '{"pass_config": {"fuse_allreduce_rms": false}}'

Notes:

  • The FlyDSL INT4 MoE path does not support expert parallelism; do not add --enable-expert-parallel.
  • Keep --compilation-config '{"pass_config": {"fuse_allreduce_rms": false}}'; it is required for this FlyDSL path on MI350X / MI355X.
  • vLLM has tuned MI350X/MI355X FlyDSL configs for this Kimi shape at TP=8 and TP=4.
  • Keep vLLM's default block size unless you are tuning long-context throughput; --block-size 64 is safe to try.

AMD MI355X (MXFP4 checkpoint)

The MXFP4 variant serves the AMD Quark checkpoint amd/Kimi-K2.5-MXFP4 — a different quantization from the packed-INT4 moonshotai/Kimi-K2.5 FlyDSL path above. It runs the AITER MXFP4 MoE path and adds an fp8 KV cache.

export VLLM_ROCM_USE_AITER=1
export VLLM_ROCM_QUICK_REDUCE_QUANTIZATION=INT4

vllm serve amd/Kimi-K2.5-MXFP4 \
  --tensor-parallel-size 8 \
  --trust-remote-code \
  --mm-encoder-tp-mode data \
  --kv-cache-dtype fp8 \
  --gpu-memory-utilization 0.90

Notes:

  • fp8 KV cache is the main MI355X win: it doubles resident KV capacity and halves KV read bandwidth. Measured +27% throughput at 8k/1k input with no accuracy change (GSM8K 0.9697, unchanged from bf16 KV).
  • AITER RMSNorm stays on. A local TP4 GSM8K check on the current nightly measured 0.9697.
  • Old firmware only: if rocm-smi --showfw reports a MEC firmware version below 177, add export HSA_NO_SCRATCH_RECLAIM=1 — older firmware can't reclaim RCCL scratch and vLLM crashes without it. Current MI355X firmware doesn't need it.
  • For maximum throughput on fixed-length benchmark 8k/1k or 1k/1k workloads, see the benchmark reproduction section below. These are throughput-sweep tunings; leave vLLM's defaults for general serving.

MXFP4 benchmark reproduction (InferenceX MI355X sweep)

The command above is the default serving configuration. The SemiAnalysis InferenceX MI355X benchmark lane for this checkpoint is a separate high-concurrency benchmark configuration — not the recipe default. To reproduce that sweep, start from the default command above and apply these deltas:

  • Use --tensor-parallel-size 4 (the sweep runs TP4; the default recipe keeps TP8 for KV-cache and multimodal-encoder headroom).
  • Add --kv-cache-dtype fp8.
  • Use the benchmark block size and scheduler knobs: --block-size 16, --max-num-batched-tokens 16384, --max-num-seqs 512, --async-scheduling, and --no-enable-prefix-caching.
  • Run on a current vllm/vllm-openai-rocm:nightly that contains the required AITER Kimi MXFP4 backend.
  • Export the AITER and INT4 quantized all-reduce env vars before launch:
# Kernel selection and benchmark runtime knobs.
export VLLM_ROCM_USE_AITER=1
export VLLM_ROCM_QUICK_REDUCE_QUANTIZATION=INT4
export VLLM_ROCM_USE_AITER_FUSION_SHARED_EXPERTS=1
export VLLM_ROCM_USE_SKINNY_GEMM=0
export VLLM_ROCM_USE_AITER_RMSNORM=0
export AITER_MXFP4_INTERMEDIATE=1
export AITER_BYPASS_TUNE_CONFIG=0
export AITER_MOE_SORT_BACKEND=auto
export OMP_NUM_THREADS=1

The full reproduction serve command is then:

vllm serve amd/Kimi-K2.5-MXFP4 \
  --port "${PORT:-8000}" \
  --tensor-parallel-size 4 \
  --gpu-memory-utilization 0.85 \
  --max-model-len "$MAX_MODEL_LEN" \
  --kv-cache-dtype fp8 \
  --block-size 16 \
  --max-num-batched-tokens 16384 \
  --max-num-seqs 512 \
  --async-scheduling \
  --trust-remote-code \
  --no-enable-prefix-caching \
  --mm-encoder-tp-mode data

The benchmark script enforces MAX_MODEL_LEN >= 9472 so the 1k/1k configuration matches the locally validated server. The env vars and scheduler knobs are benchmark-specific throughput settings, not default serving choices.

Tuned AITER MXFP4 MoE (MI355X)

For higher MoE throughput on MI355X, run the AITER MXFP4 (RadeonFlow) intermediate GEMM path with fused shared experts. This requires a vllm/vllm-openai-rocm:nightly built after vLLM #48683 (AITER v0.1.16.post5, which includes the ROCm/aiter#3832 gfx950 MXFP4 MoE backend). Export the AITER MoE kernel-selection vars and add the batching/scheduling flags below. AITER_MXFP4_INTERMEDIATE=1 is what selects the RadeonFlow MXFP4 intermediate path — without it (and VLLM_ROCM_USE_AITER=1) the MoE oracle falls back to a different kernel. Runs at TP4 or TP8; for TP < 8 disable AITER RMSNorm.

export VLLM_ROCM_USE_AITER=1
export VLLM_ROCM_QUICK_REDUCE_QUANTIZATION=INT4
export VLLM_ROCM_USE_AITER_FUSION_SHARED_EXPERTS=1
export VLLM_ROCM_USE_SKINNY_GEMM=0
export AITER_MXFP4_INTERMEDIATE=1
export AITER_BYPASS_TUNE_CONFIG=0
export AITER_MOE_SORT_BACKEND=auto
export OMP_NUM_THREADS=1
# TP < 8 only:
export VLLM_ROCM_USE_AITER_RMSNORM=0

vllm serve amd/Kimi-K2.5-MXFP4 \
  --tensor-parallel-size 4 \
  --trust-remote-code \
  --gpu-memory-utilization 0.90 \
  --kv-cache-dtype fp8 \
  --block-size 16 \
  --async-scheduling \
  --no-enable-prefix-caching \
  --mm-encoder-tp-mode data

These AITER MXFP4 MoE env vars are kernel-selection choices (same MXFP4 math, faster kernel), not precision reductions.

For fixed-length throughput sweeping, additionally cap --max-model-len to the working sequence length (e.g. 9472 for an 8k/1k workload) and set --max-num-batched-tokens 16384 --max-num-seqs 512. Leave these off for general serving.

Client Usage

Once the vLLM server is running, consume it via the OpenAI-compatible API:

import time
from openai import OpenAI

client = OpenAI(
    api_key="EMPTY",
    base_url="http://localhost:8000/v1",
    timeout=3600
)

messages = [
    {
        "role": "user",
        "content": [
            {
                "type": "image_url",
                "image_url": {
                    "url": "https://ofasys-multimodal-wlcb-3-toshanghai.oss-accelerate.aliyuncs.com/wpf272043/keepme/image/receipt.png"
                }
            },
            {
                "type": "text",
                "text": "Read all the text in the image."
            }
        ]
    }
]

start = time.time()
response = client.chat.completions.create(
    model="moonshotai/Kimi-K2.5",
    messages=messages,
    max_tokens=2048
)
print(f"Response costs: {time.time() - start:.2f}s")
print(f"Generated text: {response.choices[0].message.content}")

Troubleshooting

  • OOM errors: Lower --gpu-memory-utilization or adjust TP/EP to match your GPU count.
  • Vision encoder performance: Use --mm-encoder-tp-mode data to run the vision encoder in data-parallel mode. The encoder is small, so TP adds communication overhead with little gain.
  • Unique multimodal inputs: Pass --mm-processor-cache-gb 0 to avoid caching overhead. For repeated inputs, --mm-processor-cache-type shm uses host shared memory for better performance at high TP settings.
  • MoE kernel tuning: Use the benchmark_moe script from vLLM to tune Triton kernels for your specific hardware.
  • Async scheduling: Enabled by default for better throughput. Disable if you encounter issues, and file a bug report to vLLM.

References