A large model task layering task offloading and resource allocation method for an industrial internet

By employing a hierarchical multi-agent deep reinforcement learning approach, the problem of improper resource scheduling for large language models in mobile edge computing was solved, improving memory security and service stability, reducing task latency, and optimizing resource utilization.

CN122431839APending Publication Date: 2026-07-21HENAN UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HENAN UNIVERSITY
Filing Date
2026-06-04
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing mobile edge computing offloading methods fail to effectively distinguish between computationally intensive and memory-intensive stages in the Large Language Model (LLM) inference process, leading to improper resource scheduling, memory overflow, and inference delays. Furthermore, traditional reinforcement learning algorithms struggle to simultaneously optimize the hybrid decision-making of computational and memory resources.

Method used

A hierarchical multi-agent deep reinforcement learning approach is adopted to establish a cloud-edge-device collaborative architecture. The LLM inference task is modeled as a phased resource consumption model. Through a hierarchical Markov decision process and a composite reward function, the discrete offloading decision and continuous resource allocation are decoupled. The resource allocation is optimized by using a two-layer collaborative architecture of bottom-layer distributed offloading and upper-layer centralized allocation, combined with GRU memory mechanism and dynamic masking mechanism.

Benefits of technology

It significantly improves memory security and service stability during large model inference, reduces task queuing latency, ensures service continuity and resource utilization, and achieves rapid convergence and efficient decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122431839A_ABST
    Figure CN122431839A_ABST
Patent Text Reader

Abstract

This invention proposes a hierarchical task offloading and resource allocation method for large-scale model tasks in the Industrial Internet, belonging to the technical fields of Industrial Internet and mobile edge computing. The invention includes: establishing a cloud-edge-device collaborative reasoning architecture for large-scale models in Industrial Internet scenarios, and modeling the large language model reasoning task as a phased resource consumption model including a pre-filling stage and a decoding stage; modeling the large-scale model task offloading and resource allocation problem as a hierarchical Markov decision process, constructing a composite reward function; and utilizing the proposed hierarchical multi-agent large-scale model task offloading and resource allocation algorithm based on deep reinforcement learning, adopting a two-layer collaborative architecture of distributed offloading at the bottom layer and centralized allocation at the top layer to achieve task offloading and resource allocation. This invention significantly improves memory security and service stability during the large-scale model reasoning process; greatly reduces the action space dimension, and achieves rapid convergence and efficient decision-making in large-scale scenarios.
Need to check novelty before this filing date? Find Prior Art