Wireless Network ML Operation Distribution for Device-Limited Inference

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Generative artificial intelligence models are computationally expensive and impractical for deployment on devices with limited computing resources due to their large size and memory bandwidth requirements, leading to inefficiencies in generating responses.

Innovation Solution

Distribute machine learning model operations across entities in a wireless communications network by deploying differently sized generative models on devices with varying computing capabilities, leveraging their unique resources for inferencing and training, and using control signaling to coordinate operations across these entities.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If generative artificial intelligence models are deployed on devices with limited computing resources, then the model size and memory bandwidth requirements increase, but the computing capabilities of the device decrease

Engineering Contradiction:
Improvemodel deployment capabilityVSAvoidmemory bandwidth requirements
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

The patent divides the large language model into multiple smaller sub-models that can be distributed across different entities in the wireless network. Each entity (user equipment or network entity) hosts a subset of sub-models, allowing the system to leverage collective computing resources while reducing individual device memory bandwidth requirements.

Inventive Principle:
Principle #1Segmentation

2Measurement precision

If the number of parameters within the large language model increases, then the computational power and processing capability improve, but the computational expense and resource consumption increase

Engineering Contradiction:
Improvemodel accuracyVSAvoidcomputational resource expense
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The model is segmented into sub-models distributed across multiple entities. This allows the system to achieve high computational power collectively while each individual entity consumes fewer computational resources, balancing model accuracy with resource efficiency.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Multiple smaller sub-models hosted on different entities are combined through coordinated execution to deliver the computational power equivalent to a single large model, while distributing the energy and resource consumption across the network.

Inventive Principle:
Principle #5Merging (Combining)

3Device complexity

If machine learning model operations are centralized on a single entity, then the system complexity is reduced, but the computing resource bottleneck increases

Engineering Contradiction:
Improvesystem architecture complexityVSAvoidinferencing throughput
Core Design Contradiction:
Device complexityVSProductivity

Solution Approach 1:

The patent segments the model execution across multiple entities, with each entity handling a portion of the inferencing workload. This distribution increases overall productivity and throughput while maintaining manageable complexity through coordinated control signaling.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20250247718A1Distributing machine learning model operations across entities in a wireless communications network
Publication Date: 2025.07.31 QUALCOMM INC
  • US20250247718A1 patent drawing
  • US20250247718A1 patent drawing
  • US20250247718A1 patent drawing

AI summary

Certain aspects of the present disclosure provide techniques and apparatus for distributing machine learning model operations across entities in a wireless communications network. The method generally includes receiving, at the entity, an input prompt for processing using a machine learning model including a plurality of sub-models including a first sub-model configured to execute on a user equipment and one or more second sub-models configured to execute on network entities in the wireless communications network. Execution of one or more operations for the machine learning model based on the input prompt is coordinated via transmission of control signaling by the entity to one or more of the user equipment or network entities in the wireless communications network. Generally, the operations use a set of sub-models from the plurality of sub-models. A result responsive to the input prompt is generated based on the one or more operations, and the result is output.