MLCommons Releases New MLPerf Storage v3.0 Benchmark Results
Benchmark suite now covers the full range of AI workloads for storage systems
Press Release Disclaimer: This is a press release distributed through the XPR Media network. It has not been independently verified by our newsroom.

![]()
SAN FRANCISCO, Sept. 01, 2026 (GLOBE NEWSWIRE) — Today, MLCommons® announced the results of its industry-standard MLPerf® Storage v3.0 benchmark suite, which measures the performance of storage systems for machine learning (ML) workloads in an architecture-neutral, representative, and reproducible manner. Version 3.0 expands the tests in the suite to represent the breadth of storage workloads that AI systems can generate, and adds support for an S3 object storage access layer alongside the existing POSIX layer.
Two new benchmark tests
Version 3.0 adds a new KV Cache test, which measures storage performance for LLM inference cache read/write operations. KV caching is a widely adopted technique to increase performance in transformer-based AI inference applications, particularly autoregressive ones such as LLMs that iteratively access the same key-value vectors.
It also adds a new Vector Database (VDB) test that measures storage performance for vector indexing and querying workloads. Vector databases are used to store high-dimensional data that exceed the limits of traditional database architectures, and they are frequently used by AI applications to store unstructured data such as text and media.
“These new additions to the benchmark suite round out the test collection, covering a larger range of AI inference workloads that drive storage needs,” said Brian Belgodere, MLPerf Storage working group co-chair. “On the training side, we have training and checkpointing tests, and on the inference side we now have KV cache and vector database tests. Including tests that decompose monolithic AI systems and focus on specific storage uses and patterns, such as checkpointing, KV caching and vector databases, gives stakeholders a much clearer idea of how to engineer and provision AI systems to minimize storage performance bottlenecks.”
New S3 object storage access layer
In addition, Version 3.0 adds new support for an S3 object storage access layer, alongside and as an alternative to the existing POSIX-compliant layer. Adding this capability in v3.0 enables performance comparisons across a wider range of hosted storage options for the same workload. The S3 layer is supported for training, checkpointing, and some VDB tests in the version 3.0 benchmark suite. Approximately one-sixth of the total submissions in this round utilized the S3 storage access layer.
“Further broadening our support for a diverse set of storage systems, including S3, in the MLPerf Storage benchmark suite gives organizations that are provisioning AI systems even greater ability to select and combine technologies to meet the specific technical requirements of their application,” said Curtis Anderson, MLPerf Storage working group co-chair. “As the scale of AI contexts reaches into the trillions, we expect object-based storage systems to emerge as a viable – and possibly preferred – alternative to filesystem-based storage. By enabling S3 support now, we are ensuring that stakeholders will have the performance information they need to make smart decisions.”
11 new organizations submitting to Version 3.0
Nineteen organizations submitted to this round of the benchmark. “We would particularly like to welcome the eleven first-time submitters: Azure, Everpure, HolmesAI, Nebius, NewFW, NVIDIA, OpenLake, Suzhou Zishan Longlin, TuringData, XSKY, and ZettaLane,” said David Kanter, Head of MLPerf. “The MLPerf Storage benchmark has been embraced widely by the storage community, with submitters representing both cloud-based and on-premises solution providers, as well as organizations developing storage systems and devices. That broad level of engagement shows that the MLPerf Storage benchmark is providing critical information for stakeholders involved in all aspects of provisioning AI systems.”
The performance results provide unique and important insights into the state of the storage industry – including power efficiency. On-premises submissions for the checkpointing write test achieved a median rate of 14 GB/second per watt, with a maximum of 201. Similarly, submissions for the UNet3D read test achieved a median of 34 GB/second per watt, with a maximum of 277. “There is a wide range of power efficiencies represented in the results,” said Anderson, “which is critical performance information for stakeholders buying storage solutions for AI applications. It also shows that there is ample room for further improvement, and we encourage all suppliers to optimize for that metric.”
The MLPerf Storage benchmark was created through a collaborative engineering process over five years by 35 leading storage solution providers and academic research groups. The open-source and peer-reviewed benchmark suite offers a level playing field for competition, driving innovation, performance, and energy efficiency across the industry. It also provides critical technical information for customers who are procuring and tuning AI training and inference systems.
We invite stakeholders to join the MLPerf Storage working group and help us continue to evolve the benchmark suite.
View the Results
To view the results for MLPerf Storage v3.0, please visit the Storage benchmark results. For more information on the results, please visit the Supplemental Discussion.
About MLCommons
MLCommons is the world’s leader in AI benchmarking. An open engineering consortium supported by over 125 members and affiliates, MLCommons has a proven record of bringing together academia, industry, and civil society to measure and improve AI. The foundation for MLCommons began with the MLPerf benchmarks in 2018, which rapidly scaled into a set of industry metrics for measuring machine learning performance and promoting transparency in machine learning techniques. Since then, MLCommons has continued to use collective engineering to build the benchmarks and metrics required for better AI – ultimately helping to evaluate and improve the accuracy, safety, speed, and efficiency of AI technologies.
For additional information on MLCommons and details on becoming a member, please visit MLCommons.org or email participation@mlcommons.org.


