Split Inference for Communication Overhead and Privacy
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The existing communication technologies face challenges in reducing communication overheads and minimizing the risk of user privacy leakage when deep learning models with massive parameters perform inference tasks, as they require transmitting large amounts of data to central servers, which consumes significant resources and poses privacy risks.
Innovation Solution
A split inference method where a model is divided into submodels and distributed across multiple communication apparatuses, allowing each apparatus to perform inference locally and reduce the need for data transmission to central servers, thereby minimizing data leakage and communication overheads.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If the communication apparatus sends massive original data to the central server for inference, then the model can complete complex tasks with better performance, but communication overheads are high and user privacy leakage risk increases
Solution Approach 1:
The patent divides the central server into multiple distributed communication apparatuses, each holding a submodel. The original data is processed locally at each apparatus using its submodel, eliminating the need to transmit massive original data to a central server. This segmentation resolves the contradiction by maintaining inference performance through distributed computation while dramatically reducing communication overheads.
Solution Approach 2:
The patent introduces an intermediary mechanism where each communication apparatus processes data locally using its submodel before any potential aggregation or transmission. This intermediary local processing step eliminates the need for direct transmission of massive original data to a central server, reducing communication overheads while preserving inference capability.
2Reliability
If the communication apparatus sends massive original data to the central server for inference, then the model can complete complex tasks with better performance, but the risk of user privacy leakage increases
Solution Approach 1:
The patent segments the inference system across multiple distributed communication apparatuses, each processing data locally. This segmentation ensures that original data never leaves the local apparatus, eliminating the privacy leakage risk associated with transmitting data to a central server while maintaining inference performance through distributed computation.
Solution Approach 2:
Each communication apparatus performs self-service by processing inference tasks locally using its own submodel. This self-service capability eliminates the need to send original data to a central server, thereby removing the privacy leakage risk while maintaining the ability to complete complex inference tasks.
3Device complexity
If a single central server performs inference, then the system is simple to implement, but communication resources are consumed excessively
Solution Approach 1:
The patent segments the inference system into multiple distributed communication apparatuses, each with its own submodel. This segmentation reduces communication resources by enabling local processing, while the distributed architecture maintains system simplicity through standardized communication protocols and automated data aggregation.
4Extent of automation
If data is transmitted to central server for inference, then centralized processing is achieved, but communication overheads increase
Solution Approach 1:
The patent segments the centralized processing function across multiple distributed communication apparatuses. Each apparatus performs automated inference locally using its submodel, eliminating the need to transmit data to a central server. This segmentation maintains automation while dramatically reducing communication overheads.
Data Source
AI summary
Embodiments of this application relate to the field of communication technologies, and provide a split inference method and an apparatus, to reduce communication overheads and also reduce a risk of leaking original data to a central server. The method includes: A first communication apparatus receives, from a second communication apparatus, first indication information indicating a previous-hop communication apparatus and a next-hop communication apparatus of the first communication apparatus, where the previous-hop communication apparatus and the next-hop communication apparatus each include at least one terminal device. The first communication apparatus receives data from the previous-hop communication apparatus, performs inference on the data based on a submodel corresponding to the first communication apparatus, to obtain an inference result, and sends the inference result to the next-hop communication apparatus.


