Shared mixture of experts

The shared MoE architecture addresses the issue of duplicated experts in MoE models by deduplicating and parallel processing, enhancing GPU memory usage and processing efficiency.

WO2026109962A1PCT designated stage Publication Date: 2026-05-28INTERNATIONAL BUSINESS MACHINE CORPORATION +2
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
INTERNATIONAL BUSINESS MACHINE CORPORATION
Filing Date
2025-10-30
Publication Date
2026-05-28

AI Technical Summary

Technical Problem

Existing machine learning models with Mixture of Experts (MoE) architectures suffer from excessive GPU memory consumption and increased serving costs due to duplicated experts across model instances, leading to low GPU utilization and inefficient resource allocation.

Method used

Implement a shared MoE architecture that deduplicates identical experts across model instances and batches requests to these experts in parallel, using a fused gate mechanism to optimize memory usage and improve processing efficiency.

Benefits of technology

Reduces memory footprint and improves processing efficiency by sharing experts, allowing for higher throughput and better utilization of GPU resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IB2025061063_28052026_PF_FP_ABST
    Figure IB2025061063_28052026_PF_FP_ABST
Patent Text Reader

Abstract

A method, computer program product, and computer system for using a shared mixture of experts (MoE) architecture to implement inference with respect to a machine learning model. A gate mechanism receives, from N clients, N requests and an identification of N MoE models. N is at least 2. The gate mechanism selects, from N sets of experts, E experts to process the N requests. The N sets of experts collectively include at least one duplicative expert that is common to more than one set of the N sets of experts. Each duplicative expert has been deduplicated by being stored only once in the one or more graphic processing units (GPUs). The N requests are routed to the E experts. The E experts are executed to generate N respective responses to the N requests. Each response of the N responses is transmitted to the respective client of the N clients.
Need to check novelty before this filing date? Find Prior Art