A large model task layering task offloading and resource allocation method for an industrial internet
By employing a hierarchical multi-agent deep reinforcement learning approach, the problem of improper resource scheduling for large language models in mobile edge computing was solved, improving memory security and service stability, reducing task latency, and optimizing resource utilization.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HENAN UNIVERSITY
- Filing Date
- 2026-06-04
- Publication Date
- 2026-07-21
AI Technical Summary
Existing mobile edge computing offloading methods fail to effectively distinguish between computationally intensive and memory-intensive stages in the Large Language Model (LLM) inference process, leading to improper resource scheduling, memory overflow, and inference delays. Furthermore, traditional reinforcement learning algorithms struggle to simultaneously optimize the hybrid decision-making of computational and memory resources.
A hierarchical multi-agent deep reinforcement learning approach is adopted to establish a cloud-edge-device collaborative architecture. The LLM inference task is modeled as a phased resource consumption model. Through a hierarchical Markov decision process and a composite reward function, the discrete offloading decision and continuous resource allocation are decoupled. The resource allocation is optimized by using a two-layer collaborative architecture of bottom-layer distributed offloading and upper-layer centralized allocation, combined with GRU memory mechanism and dynamic masking mechanism.
It significantly improves memory security and service stability during large model inference, reduces task queuing latency, ensures service continuity and resource utilization, and achieves rapid convergence and efficient decision-making.
Smart Images

Figure CN122431839A_ABST