Edge AI Model Segmentation for Lower-Cost LLM Processing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The computational demands and costs associated with implementing large language models (LLMs) in distributed computing networks, particularly in edge networks, are unsustainable due to the high resource requirements and scalability issues.

Innovation Solution

Deploying smaller, less accurate machine learning models on CPUs or GPUs at the edge network to preprocess requests, with the option to forward them to a larger LLM on GPUs for further processing, thereby reducing computational load and costs.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If large language models are deployed in edge networks, then AI capabilities and accuracy are improved, but computational resource requirements and costs increase significantly

Engineering Contradiction:
ImproveAI model accuracyVSAvoidcomputational resource consumption
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The patent segments the AI processing workload into two parts: a smaller preliminary model deployed at the edge network for initial processing, and a larger main model deployed in the data center for final processing. This segmentation allows the edge to handle routine tasks with low resource consumption while maintaining access to high-accuracy models when needed.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary preliminary model that acts as a gateway between the edge network and the main large language model. This intermediary handles preprocessing and filtering tasks, reducing the burden on both the edge resources and the main model, thereby optimizing the overall system efficiency.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If larger and more capable language models are deployed, then model capabilities and accuracy are improved, but computational costs and resource requirements increase

Engineering Contradiction:
Improvemodel accuracyVSAvoidcomputational infrastructure complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent divides the computational infrastructure into two segments: edge servers with smaller models for local processing and data center servers with larger models for complex tasks. This segmentation reduces the complexity requirement at each individual node while maintaining overall system capability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent employs a preliminary model that performs partial processing of requests before they reach the main model. This partial action filters out many routine queries, allowing the main large model to focus only on complex tasks, thereby reducing the effective computational burden.

Inventive Principle:
Principle #16Partial or excessive action

3Productivity

If smaller machine learning models are used at the edge, then computational cost is reduced, but processing accuracy decreases

Engineering Contradiction:
Improveprocessing efficiencyVSAvoidmodel accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent introduces a two-stage processing system where the smaller edge model acts as an intermediary that handles preliminary processing. When the edge model's confidence is low or the task is complex, it forwards the request to the more accurate main model in the data center, thus combining the benefits of both small and large models.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system implements a feedback mechanism where the preliminary model's output is evaluated, and low-confidence or complex predictions are fed back to the main model for reprocessing. This feedback loop ensures that accuracy is maintained for critical predictions while using the more efficient small model for routine tasks.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS20250254539A1Artificial Intelligence (AI) on an Edge Network
Publication Date: 2025.08.07 AKAMAI TECHNOLOGIES INC
  • US20250254539A1 patent drawing
  • US20250254539A1 patent drawing
  • US20250254539A1 patent drawing

AI summary

This disclosure provides for Artificial Intelligence (AI) support in a distributed computing environment. First machine learning models are configured in a first network located between requesting clients, and a second network, which hosts a second machine learning model, such as a Large Language Model (LLM). Significant processing efficiencies are obtained by provisioning these ML models on the respective networking components. Preferably, and as between a first machine learning model and the second machine learning model, the first machine learning model provides inferencing at a lower cost but with less accuracy. In response to receipt of a request by a first machine learning model, a response is generated. The response is forwarded onward to the second machine learning model for additional handling. The first machine learning model executes primarily on Central Processing Units (CPUs), and the second machine learning model executes primarily on Graphics Processing Units (GPUs).