•Wafer achieved 2626 tok/s/node for GLM5.2 LLM inference on AMD MI355X, demonstrating over 2x lower cost per token compared to NVIDIA Blackwell B300/B200 GPUs.
•Key optimizations included MXFP4 quantization with AMD Quark and utilizing the sglang inference engine, along with custom fixes for speculative decoding within the ROCm stack.
•This benchmark highlights AMD's growing viability for cost-effective AI inference, challenging NVIDIA's market dominance and emphasizing the importance of software optimization for hardware efficiency...
•Wafer achieved 2626 tok/s/node for GLM5.2 LLM inference on AMD MI355X, demonstrating over 2x lower cost per token compared to NVIDIA Blackwell B300/B200 GPUs.
•Key optimizations included MXFP4 quantization with AMD Quark and utilizing the sglang inference engine, along with custom fixes for speculative decoding within the ROCm stack.
•This benchmark highlights AMD's growing viability for cost-effective AI inference, challenging NVIDIA's market dominance and emphasizing the importance of software optimization for hardware efficiency...