Distributed AI Model Splitting for Mobile Resource Constraints

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The challenge is deploying high-accuracy AI processing models in mobile devices with limited resources due to high internal memory overhead and computing workloads, which existing model compression technologies cannot adequately address, resulting in reduced recognition accuracy and poor user experience.

Innovation Solution

The solution involves splitting an AI processing model into two sub-models, where a first sub-model with M neural network layers is deployed on the mobile device and a second sub-model with K neural network layers is deployed on a server, allowing for distributed processing to reduce computational load and memory requirements, while maintaining recognition accuracy through model compression and compensation techniques.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If a large neural network is used to achieve high-accuracy recognition, then recognition accuracy is improved, but internal memory overhead and computing workload increase significantly

Engineering Contradiction:
Improverecognition accuracyVSAvoidinternal memory overhead and computing workload
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent divides the large neural network into multiple sub-models distributed across different devices. The first sub-model is deployed on the mobile device while the second sub-model is deployed on the server, allowing the system to achieve high recognition accuracy without requiring the entire large neural network to be stored and executed on the resource-constrained mobile device.

Inventive Principle:
Principle #1Segmentation

2Device complexity

If model compression technology is applied to reduce computing workload, then device resource requirements are reduced, but model compression capability is limited and cannot adequately deploy high-accuracy models

Engineering Contradiction:
Improvecomputing workload and memory requirementsVSAvoidrecognition accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent transitions from a single-device deployment model to a distributed multi-device model. By moving part of the neural network computation from the mobile device to the server, the system overcomes the limitations of model compression technology and enables deployment of high-accuracy models that would otherwise be too large for mobile devices.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Speed

If the entire AI processing model is deployed on the mobile device, then processing speed is improved, but the device's limited resources cannot support the model

Engineering Contradiction:
Improveprocessing speedVSAvoidresource requirements
Core Design Contradiction:
SpeedVSDevice complexity

Solution Approach 1:

The patent segments the AI processing model into two sub-models executed on different devices. The first sub-model runs on the mobile device for local processing, while the second sub-model runs on the server for additional processing, achieving a balance between processing speed and resource constraints.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS11689607B2Data processing method and apparatus, storage medium, and electronic device
Publication Date: 2023.06.27 TENCENT TECHNOLOGY (SHENZHEN) CO LTD
  • US11689607B2 patent drawing
  • US11689607B2 patent drawing
  • US11689607B2 patent drawing

AI summary

The disclosure provides a data processing method and apparatus, a storage medium and an electronic device. The method includes: performing first processing on data by using a first sub-model corresponding to an artificial intelligence (AI) processing model, to obtain an intermediate processing result, the AI processing model corresponding to the first sub-model and a second sub-model, the first sub-model being generated according to M neural network layers in the AI processing model, the second sub-model being generated according to K neural network layers in the AI processing model, M and K being positive integers greater than or equal to 1; transmitting the intermediate processing result to a first server; and receiving a target processing result from the first server, the target processing result being based on a result of second processing on the intermediate processing result by using the second sub-model.