Benchmarking DFlash on a 30B Model: Why Tokens per Second Can Mislead A field guide for ML engineers, LLM DevOps, and system architects deploying open-weights 30B-scale models with block-diffusion speculative decoding. 1. Introduction & Background 1.1 What DFlash actual... Benchmarks & Architecture Guides