Wireless Network ML Operation Distribution for Device-Limited Inference
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Generative artificial intelligence models are computationally expensive and impractical for deployment on devices with limited computing resources due to their large size and memory bandwidth requirements, leading to inefficiencies in generating responses.
Innovation Solution
Distribute machine learning model operations across entities in a wireless communications network by deploying differently sized generative models on devices with varying computing capabilities, leveraging their unique resources for inferencing and training, and using control signaling to coordinate operations across these entities.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If generative artificial intelligence models are deployed on devices with limited computing resources, then the model size and memory bandwidth requirements increase, but the computing capabilities of the device decrease
Solution Approach 1:
The patent divides the large language model into multiple smaller sub-models that can be distributed across different entities in the wireless network. Each entity (user equipment or network entity) hosts a subset of sub-models, allowing the system to leverage collective computing resources while reducing individual device memory bandwidth requirements.
2Measurement precision
If the number of parameters within the large language model increases, then the computational power and processing capability improve, but the computational expense and resource consumption increase
Solution Approach 1:
The model is segmented into sub-models distributed across multiple entities. This allows the system to achieve high computational power collectively while each individual entity consumes fewer computational resources, balancing model accuracy with resource efficiency.
Solution Approach 2:
Multiple smaller sub-models hosted on different entities are combined through coordinated execution to deliver the computational power equivalent to a single large model, while distributing the energy and resource consumption across the network.
3Device complexity
If machine learning model operations are centralized on a single entity, then the system complexity is reduced, but the computing resource bottleneck increases
Solution Approach 1:
The patent segments the model execution across multiple entities, with each entity handling a portion of the inferencing workload. This distribution increases overall productivity and throughput while maintaining manageable complexity through coordinated control signaling.
Data Source
AI summary
Certain aspects of the present disclosure provide techniques and apparatus for distributing machine learning model operations across entities in a wireless communications network. The method generally includes receiving, at the entity, an input prompt for processing using a machine learning model including a plurality of sub-models including a first sub-model configured to execute on a user equipment and one or more second sub-models configured to execute on network entities in the wireless communications network. Execution of one or more operations for the machine learning model based on the input prompt is coordinated via transmission of control signaling by the entity to one or more of the user equipment or network entities in the wireless communications network. Generally, the operations use a set of sub-models from the plurality of sub-models. A result responsive to the input prompt is generated based on the one or more operations, and the result is output.


