https://github.com/theroyallab/tabbyAPI
Meeko1 Qwen/Qwen3.8-27B-FP8 Meeko2 Qwen/Qwen3.8-27B-FP8 Setup/execute vllm: Meeko1 Qwen/Qwen3.8-27B-FP8 Warm up model vllm running, connect to container get it run it vllM: Meeko2 Qwen/Qwen3.8-27B-FP8 Start Warm up model vllm running, connect to container get it run it Results
Supported sources:
vLLM: Command BUGGER: NotImplementedError: Qwen4Exp N-gram PLE embedding requires pipeline_parallel_size=1 because non-first pipeline ranks do not receive the raw input_ids it needs. Please run with PP=1. (Worker_TP5 pid=242) ERROR 09-23 18:07:39 [multiproc_executor.py:944] NotImplementedError: Qwen4Exp QSA requires a BF16 main KV cache RuntimeError: PLE offload worker failed during startup: ValueError(“There is no module or parameter named […]
Investigate and for SGLang? Meeko1: single Table Gemma4 Meeko1 + Meeko2: TheDrummer/Artemis-31B-v1.2 Meeko2 Qwen/Qwen3.8-27B-FP8 Gemma4
Command? Ran it and it produced Final command: Warmup Test single Test Multi NCCL changes Test single Log:
unsloth/Qwen3.8-27B-NVFP4 Command Startup log
Meeko 2 almost fully built
Replace thermal Pads
docker-date drive For your current example: Then six months from now you can have: without confusing who produced it with how it’s quantized.