Distributed AI Inference for Hardware-Limited Terminal Devices

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Terminal devices with insufficient processing capabilities struggle to execute AI model inference tasks due to hardware limitations or incompatible AI processing platforms.

Innovation Solution

A method where a first device, such as a server or processor outside the wireless cellular system, assists a second device, like a mobile phone, in completing AI model inference tasks by providing direct or indirect support, including joint inference, capability information sharing, and model transmission, leveraging a third device like a network node to facilitate AI model inference.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If terminal devices use high-end hardware capabilities for AI inference, then inference performance is improved, but device cost and complexity increase

Engineering Contradiction:
Improveinference performanceVSAvoidhardware capabilities
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent introduces a network device as an intermediary that provides AI inference capabilities to terminal devices. The network device hosts AI models and performs inference computations, while terminal devices only need to send inference requests and receive results, eliminating the need for complex local AI hardware

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent enables terminal devices to access AI inference capabilities by copying or replicating the network device's inference service through virtualization. Multiple terminal devices can share the same AI model instances hosted on the network device, achieving resource utilization without duplicating hardware

Inventive Principle:
Principle #26Copying

2Adaptability or versatility

If terminal devices are equipped with AI processing platforms, then inference capability is improved, but compatibility requirements and system complexity increase

Engineering Contradiction:
Improveinference capabilityVSAvoidsoftware platforms
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The network device acts as an intermediary that handles AI model management, deployment, and execution. Terminal devices interact with the network device through standardized APIs, eliminating the need for complex local AI software platforms and ensuring compatibility across different terminal types

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The network device provides a universal AI inference platform that can serve multiple terminal devices with different capabilities. The network device handles diverse AI models and inference tasks, allowing terminal devices to access AI capabilities without needing device-specific AI software

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Speed

If AI inference is performed locally on terminal devices, then response time is improved, but processing power requirements increase

Engineering Contradiction:
Improveinference speedVSAvoidprocessing capabilities
Core Design Contradiction:
SpeedVSPower

Solution Approach 1:

The patent segments the AI inference system into two parts: the network device handles computationally intensive model execution, while terminal devices handle lightweight tasks like data preprocessing and result processing. This segmentation allows fast inference without requiring full processing power at the terminal

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20260065092A1Ai model reasoning method
Publication Date: 2026.03.05 BEIJING XIAOMI MOBILE SOFTWARE CO LTD
  • US20260065092A1 patent drawing
  • US20260065092A1 patent drawing
  • US20260065092A1 patent drawing

AI summary

Disclosed in the embodiments of the present application are a reasoning method and apparatus, which can be applied to wireless artificial intelligence (AI) systems. The method comprises: in the solution, a third device sending an AI model reasoning task to a second device; and when the second device does not have a condition for independent reasoning, in response to receiving an AI model reasoning request, which is sent by means of the second device, the first device assisting the second device with completing the AI model reasoning task. Therefore, the second device can be able to indirectly perform reasoning in response to a requirement for providing or using an AI model reasoning result, thereby benefiting from wireless AI.