Adaptive Deep Learning Inference for Deterministic Latency

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The latency of mobile edge computing-based deep learning data analysis services is affected by changes in wireless data transmission latency, making it challenging to provide time-sensitive services like autonomous driving and XR applications with deterministic latency.

Innovation Solution

An adaptive deep learning inference system that adjusts deep learning model inference based on varying wireless network latency, ensuring end-to-end data processing service latency by selecting an appropriate deep learning model inference computation method.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If deep learning inference service is provided in mobile edge computing environment, then data processing capability is improved, but service latency becomes variable due to wireless network latency changes

Engineering Contradiction:
Improvedata processing capabilityVSAvoidservice latency determinism
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The system dynamically adjusts deep learning inference parameters based on real-time wireless network latency measurements. The edge computing server monitors network conditions and adapts inference computation depth, model complexity, or execution timing to maintain deterministic end-to-end latency while preserving data processing capability.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The invention changes inference parameters such as computation depth, model precision, or batch size based on measured network latency. When network latency is high, the system reduces inference computation depth or uses quantized models to compensate, ensuring total service latency remains within deterministic bounds.

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If deep learning model inference computation is increased to improve accuracy, then data analysis performance is improved, but end-to-end service latency increases

Engineering Contradiction:
Improvedata analysis accuracyVSAvoidend-to-end service latency
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs partial inference computations based on network conditions rather than always executing full-depth models. When network latency allows, more comprehensive inference is performed for higher accuracy; when latency is constrained, reduced-computation models are used to meet timing requirements.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The edge computing server performs preliminary measurements of wireless network latency and pre-determines appropriate inference computation depth before data processing begins. This allows the system to prepare inference configurations in advance and execute them efficiently without last-minute delays.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12218801B2Adaptive deep learning inference apparatus and method in mobile edge computing
Publication Date: 2025.02.04 ELECTRONICS & TELECOMM RES INST
  • US12218801B2 patent drawing
  • US12218801B2 patent drawing
  • US12218801B2 patent drawing

AI summary

Disclosed is an adaptive deep learning inference system that adapts to changing network latency and executes deep learning model inference to ensure end-to-end data processing service latency when providing a deep learning inference service in a mobile edge computing (MEC) environment. An apparatus and method for providing a deep learning inference service performed in an MEC environment including a terminal device, a wireless access network, and an edge computing server are provided. The apparatus and method provide deep learning inference data having deterministic latency, which is fixed service latency, by adjusting service latency required to provide a deep learning inference result according to a change in latency of the wireless access network when at least one terminal device senses data and requests a deep learning inference service.