llama-benchy

Meeko1 Qwen/Qwen3.8-27B-FP8 Meeko2 Qwen/Qwen3.8-27B-FP8 Setup/execute vllm: Meeko1 Qwen/Qwen3.8-27B-FP8 Warm up model vllm running, connect to container get it run it vllM: Meeko2 Qwen/Qwen3.8-27B-FP8 Start Warm up model vllm running, connect to container get it run it Results

FAILED: vllm/vllm-openai:qwen38-flash-next

vLLM: Command BUGGER: NotImplementedError: Qwen4Exp N-gram PLE embedding requires pipeline_parallel_size=1 because non-first pipeline ranks do not receive the raw input_ids it needs. Please run with PP=1. (Worker_TP5 pid=242) ERROR 09-23 18:07:39 [multiproc_executor.py:944] NotImplementedError: Qwen4Exp QSA requires a BF16 main KV cache RuntimeError: PLE offload worker failed during startup: ValueError(“There is no module or parameter named […]