Distributed Model Inference Across Control RAN Nodes
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing model inference methods in wireless communication networks face challenges such as data security issues, increased network load, delayed feedback, and insufficient computing power due to centralized processing by the OAM, especially in scenarios with high mobility and limited resources.
Innovation Solution
Distribute model inference tasks across multiple gNB-CUs based on their AI processing capacity, segmenting models, and enabling collaborative inference to balance computing power and ensure timely feedback to the terminal.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If model inference is centralized at OAM, then model inference accuracy is maintained, but network load increases and feedback delay occurs
Solution Approach 1:
The patent segments the centralized model inference function into distributed inference nodes at gNB-CU and gNB-DU levels. The model is divided into multiple segments that can be processed in parallel across different network elements, reducing the time required for complete inference while maintaining accuracy through coordinated processing.
Solution Approach 2:
The patent introduces a new dimensional approach by enabling model inference at multiple hierarchical levels (gNB-CU and gNB-DU) rather than solely at the OAM level. This multi-dimensional inference architecture allows simultaneous processing paths, reducing feedback delay while preserving inference quality through cross-validation.
2Device complexity
If model inference is centralized at OAM, then model processing is simplified, but computing power becomes insufficient
Solution Approach 1:
The patent divides the computing workload by segmenting the model inference function across multiple distributed nodes (gNB-CU, gNB-DU). Each node processes specific model segments, distributing the computational burden and enabling the system to handle more complex models without overwhelming a single processing unit.
Solution Approach 2:
The patent combines the computing resources of multiple network elements (gNB-CU, gNB-DU, and OAM) into a distributed computing architecture. This merging of computational capabilities across hierarchical levels provides sufficient total computing power while maintaining coordinated processing through standardized interfaces.
3Reliability
If model inference is centralized at OAM, then data security is compromised less, but network load increases
Solution Approach 1:
The patent segments data processing tasks so that only necessary model segments and corresponding data portions are transmitted between distributed inference nodes. This reduces the volume of data traffic compared to centralized processing, where all data would need to be transmitted to and from the OAM, thereby reducing network load while maintaining security through localized processing.
4Loss of time
If model inference is distributed across gNB-CU and gNB-DU, then feedback delay is reduced, but model inference complexity increases
Solution Approach 1:
The patent implements a dynamic model segmentation and distribution mechanism that adapts to different operational scenarios. The system can dynamically adjust which nodes perform inference and how models are segmented, allowing flexible coordination that reduces feedback delay while managing complexity through adaptive rather than static configurations.
Solution Approach 2:
The patent incorporates feedback mechanisms where distributed inference nodes exchange intermediate results and coordination information. This feedback loop enables synchronized processing across gNB-CU and gNB-DU, managing the complexity of distributed inference through structured communication protocols that coordinate model segment processing and aggregate results efficiently.
Data Source
AI summary
A method for model inference, applicable for an operation administration and maintenance (OAM) entity, includes determining a first model corresponding to model subscription request information in response to receiving the model subscription request information sent by a control radio access network (RAN) device; and obtaining a first number of model segmentation blocks by segmenting the first model, and distributing the first number of model segmentation blocks to a first number of control RAN devices.


