Hardware

MLCommons Debuts MLPerf Client v2.0 AI PC Benchmark

MLCommons has launched MLPerf Client v2.0, expanding its PC benchmark suite with new tests for agentic workflows and image generation to better evaluate local hardware capabilities.

ML Commons2 days agoHardware
Image: ML Commons

MLCommons has officially released MLPerf Client v2.0, a major update to its open-source benchmarking suite designed to measure local AI performance on personal computers. The updated tool allows hardware manufacturers, software developers, and enterprise buyers to evaluate how laptops, desktops, and workstations handle modern AI workloads. By reporting both responsiveness and throughput, the benchmark provides a standardized way to gauge the efficiency of consumer-grade hardware.

The v2.0 release introduces two major evaluation categories to reflect the shifting landscape of consumer AI applications. A new Image Generation category debuts with Flux.2 klein 4B included as an experimental test, allowing users to measure local text-to-image performance. Additionally, the benchmark now features an Agentic AI category that tests Software Engineering (SWE) Agent and Data Analyst Agent scenarios. This category tracks end-to-end performance, breaking down the time spent on large language model inference versus external tool execution.

Existing large language model inference tests have also received significant upgrades. MLPerf Client v2.0 replaces the older Phi 3.5 mini instruct model with the newer Phi 4 Mini Instruct for its required workloads. It also introduces Qwen 3 8B as an experimental test. To better simulate real-world productivity demands, the base testing suite now includes an Intermediate Summarization task that processes input prompts of approximately 4K tokens.

Developed through a broad industry collaboration involving AMD, Intel, Microsoft, Nvidia, and major PC original equipment manufacturers, the benchmark aims to establish a cross-platform standard. For AI practitioners and hardware developers, these updates offer a more realistic framework for optimizing local models and validating hardware claims. The benchmarking software is open-source and freely available for Windows, macOS, and Linux via the MLCommons GitHub repository.

This is our own summary of reporting by ML Commons

More in Hardware