Shared mixture of experts
The shared MoE architecture addresses the issue of duplicated experts in MoE models by deduplicating and parallel processing, enhancing GPU memory usage and processing efficiency.
WO2026109962A1PCT designated stage Publication Date: 2026-05-28INTERNATIONAL BUSINESS MACHINE CORPORATION +2
View PDF 0 Cites 0 Cited by
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- INTERNATIONAL BUSINESS MACHINE CORPORATION
- Filing Date
- 2025-10-30
- Publication Date
- 2026-05-28
AI Technical Summary
Technical Problem
Existing machine learning models with Mixture of Experts (MoE) architectures suffer from excessive GPU memory consumption and increased serving costs due to duplicated experts across model instances, leading to low GPU utilization and inefficient resource allocation.
Method used
Implement a shared MoE architecture that deduplicates identical experts across model instances and batches requests to these experts in parallel, using a fused gate mechanism to optimize memory usage and improve processing efficiency.
Benefits of technology
Reduces memory footprint and improves processing efficiency by sharing experts, allowing for higher throughput and better utilization of GPU resources.
✦ Generated by Eureka AI based on patent content.
Smart Images

Figure IB2025061063_28052026_PF_FP_ABST
Abstract
A method, computer program product, and computer system for using a shared mixture of experts (MoE) architecture to implement inference with respect to a machine learning model. A gate mechanism receives, from N clients, N requests and an identification of N MoE models. N is at least 2. The gate mechanism selects, from N sets of experts, E experts to process the N requests. The N sets of experts collectively include at least one duplicative expert that is common to more than one set of the N sets of experts. Each duplicative expert has been deduplicated by being stored only once in the one or more graphic processing units (GPUs). The N requests are routed to the E experts. The E experts are executed to generate N respective responses to the N requests. Each response of the N responses is transmitted to the respective client of the N clients.
Need to check novelty before this filing date? Find Prior Art