ZeroQ – A Prospective Solution to Scalable AI Direct NVMe® SSD Access
ZeroQ addresses existing SSD bottlenecks by introducing a highly scalable, energy-efficient "green superhighway" that enables more than 10,000 AI cores to directly access NVMe® SSDs.
The Industry Challenge
Over the past several years, AI and storage experts have explored multiple approaches to enable large numbers of AI cores to access NVMe® SSDs directly. However, existing solutions have not fully met industry expectations for scalability, efficiency and cost optimization.
Traditional NVMe architecture is based on a queue-pair model, consisting of Submission Queues (SQ) and Completion Queues (CQ), typically assigned per host-side CPU core. This approach works effectively for conventional CPU environments because it minimizes inter-core coordination overhead. However, the model becomes impractical for large-scale AI deployments.
As AI workloads continue to expand, the number of processing cores is increasing rapidly. NVIDIA has already announced support for architectures exceeding 16,000 cores, with significantly larger deployments anticipated in the future.
Challenges With Current Architecture
![]() |
Current SSD access methods face several critical limitations when scaled for AI. Delivering sufficient data access and throughput for tens of thousands of AI cores remains a major challenge.
- The current four-step AI-to-NVMe command path (steps 1 through 4, above) through the CPU introduces performance bottlenecks for AI workloads
- The traditional two-step CPU-to-NVMe command path (steps 2 and 3, above) is only efficient in one-to-one core-to-queue-pair access models
- The option of the NVMe ASIC providing a queue pair per AI core is impractical
- ASIC complexity increases substantially
- Silicon gate requirements grow substantially
- Power consumption increases
- Overall implementation costs become prohibitive
- Delivering fair and efficient resource allocation across a very large number of cores is difficult. Submission queues are processed in arbitrary order and that order is unlikely to match the order commands were issued by the AI cores. Response times would be inconsistent
As AI systems scale to larger core counts, these challenges become increasingly difficult to address. This situation demands a new approach.
The ZeroQ Solution
Microchip’s AI architecture team is developing a new approach called “ZeroQ,” named for its ability to bypass traditional SQ/CQ mechanisms. ZeroQ creates a highly efficient direct-access pathway ("green superhighway") between AI cores and NVMe SSDs.
Beyond concept development, we demonstrated a working prototype at FMS 2026, validating ZeroQ operation across more than 10,000 AI cores. The technology is now positioned for broader industry evaluation as a scalable solution for next-generation AI storage architectures.
![]() |
ZeroQ Delivers Advantages That Solve AI Scalability Challenges
- Enables direct NVMe SSD access for more than 10,000 AI cores
- Achieves scalability through FIFO structures in SSD memory rather than expensive silicon gate expansion associated with traditional SQ/CQ per AI core implementations
- Maintains a small and fixed silicon footprint even at 100,000-core scale
- Enables effective queue depths of up to 100,000, well beyond conventional benchmarks
- Eliminates the cost and power limitations associated with SQ/CQ scaling
- Preserves command ordering through the PCIe® fabric and SSD FIFO, achieving more consistent response times
- Removes the need for locks and synchronization between host-side cores
- Reduces per-core resource overhead, including large SQ/CQ memory allocations
- Maintains compatibility with PCIe, NVMe and SR-IOV standards
How ZeroQ Works
ZeroQ uses a streamlined architecture designed for massive scalability:
- All AI cores share a single ZeroQ Command List Register for each NVMe SSD
- Each core maintains its own command lists
- AI cores issue PCIe memory-write operations containing the 64-bit host address of their command list to the ZeroQ List Register FIFO
- Hardware automation transfers command list addresses into the SSD-resident FIFO and initiates processing
- Each command list may contain multiple commands, enabling more efficient processing compared to traditional single-command SQ/CQ operations
- Command lists and their associated commands can be executed in parallel by the SSD
Key Benefits of ZeroQ
ZeroQ provides several advantages over conventional SQ/CQ architectures:
Host Driver Operations The host driver performs several key functions:
| About the ZeroQ Demo Host Driver Shown at Flash Memory Summit (FMS)
|
ZeroQ Process Within the SSD
- Hardware automation accepts incoming Command List addresses and appends to the FIFO
- SSD firmware arbitrates between ZeroQ traffic and traditional SQ/CQ traffic
- The SSD retrieves ZeroQ Command Lists from host memory
- Available SSD resources process commands in parallel
- Multiple ZeroQ FIFO entries may be processed concurrently
- Completion notifications are returned using PCIe memory-write operations similar to MSI/MSI-X mechanisms
Demonstrated Results
Prototype testing has demonstrated the scalability and efficiency of ZeroQ:
- Successfully completed 10 million READ operations across 10,000 threads
- Delivered fair access across all participating cores
- Enabled a shared-access architecture using a single ZeroQ List Register
- Demonstrated potential improvements in performance, latency and power efficiency
Industry Collaboration
We welcome industry feedback and participation in ongoing discussions around ZeroQ as the ecosystem explores scalable AI storage solutions.
This QR code allows you to get your email feedback to the right team at Microchip.
![]() |
Or click here to send an email.
Q&A
- Is ZeroQ compatible with PCIe, NVMe, SR-IOV and MSI-X?
Yes. ZeroQ is designed as an extension to existing NVMe host-driver and SSD implementations. It operates as a parallel path alongside traditional SQ/CQ mechanisms and remains compatible with PCIe, SR-IOV and MSI/MSI-X technologies. - What is the current development status?
It is a work in progress. We wanted to share our ZeroQ idea and progress with the industry and solicit feedback. There’s plenty to do for all interested parties. - Will ZeroQ be proposed as an industry standard?
Microchip actively participates in standards organizations and industry initiatives including NVMe and OCP. We intend to continue engaging with the industry regarding ZeroQ evolution, adoption and potential standardization opportunities. - Does ZeroQ require separate FIFOs for SR-IOV Physical Functions (PFs) and Virtual Functions (VFs)?
The current prototype does not implement SR-IOV functionality. Future SR-IOV support is under evaluation, and each VF would require its own ZeroQ Command List Register. - How is arbitration handled between SQ/CQ traffic and ZeroQ traffic within the SSD?
Arbitration mechanisms are implementation-specific and would be determined by SSD vendors based on design priorities and customer requirements. - Does ZeroQ impact SSD security?
ZeroQ relies on the existing security framework provided by the host driver, PCIe infrastructure and NVMe architecture. - How can ZeroQ improve SSD performance?
ZeroQ allows hosts to submit complete command lists instead of individual commands, improving PCIe and interrupt efficiency while enabling more optimized SSD processing. Performance gains will vary depending on workload characteristics and implementation details. Early evaluations suggest that ZeroQ may improve performance across many AI and high-throughput workloads.


