D2D auxiliary MEC computing resource allocation method based on bidirectional unloading mechanism
By introducing a D2D assisted MEC computing resource allocation method based on a two-way offload mechanism in the industrial Internet, a multi-agent deep reinforcement learning framework is used to make task offload and resource allocation decisions, solving the problem of resource waste in the existing technology and low decision efficiency in dynamic environments, and achieving efficient and stable computing resource utilization.
Patent Information
- Application Number
- CN202510203805.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-05-06
AI Technical Summary
In the existing industrial Internet, task offloading and resource allocation cannot fully utilize device computing resources, resulting in waste of resources, and traditional model-driven algorithms perform poorly in dynamic environments.
A D2D assisted MEC computing resource allocation method based on a bidirectional offload mechanism is proposed. By establishing a D2D assisted MEC computing offloading architecture, the task offloading and resource allocation problems are modeled as distributed partially observable Markov decision-making process based on graph structure, and the multi-agent deep reinforcement learning framework that integrates Transformer and task shielding mechanism are used for solving.
In the industrial Internet scenario where resources are limited, we can improve task offload efficiency, make full use of computing resources, significantly improve the security and stability of the system, and achieve efficient allocation and utilization of resources under the delay constraints of computing-intensive and delay-sensitive tasks.
Smart Images

Figure CN119946725A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of resource allocation, and in particular to a D2D-assisted MEC computing resource allocation method based on a bidirectional offloading mechanism. A lightweight offloading and resource allocation method with a task selection mechanism is designed. Background Art
[0002] With the advancement of digital transformation, the Industrial Internet has become an important bridge connecting the physical and digital worlds. However, the rapid development of information technology and the surge in the number of industrial equipment have led to an exponential growth in computing tasks in the Industrial Internet. Automated Guided Vehicles (AGVs) and industrial video surveillance tasks usually require high computing power and real-time performance, while the computing resources of existing industrial robots can no longer meet these requirements. Mobile Edge Computing (MEC), as an emerging technology, has attracted widespread attention because of its ability to provide low latency and high bandwidth.
[0003] Mobile edge computing deploys computing resources at the edge of the network, allowing tasks to access resources with lower latency. However, the constant access of industrial robots and the limited computing resources in mobile edge computing make it difficult to offload tasks in a timely manner. Therefore, the introduction of device-to-device (D2D) communication can assist mobile edge computing servers in processing tasks and utilize nearby idle robots for computing, thereby reducing the computing pressure on edge servers.
[0004] D2D (Device to Device) communication directly transmits data between devices, optimizes task offloading, reduces latency and improves task completion rate. However, existing research usually statically divides industrial robots into "idle class" and "demand class", so that idle robots can only accept task offloading in one direction, simplifying the dynamic adjustment of offloading demand. In actual industrial Internet scenarios, tasks are random and sudden, and robots may have offloading demands at different time points. This static division cannot fully utilize device computing resources, resulting in resource waste. Therefore, there is an urgent need for a D2D-assisted MEC two-way task offloading and resource allocation solution that is more suitable for actual scenarios.
[0005] Due to the non-stationarity of industrial scenarios and the limitations of information acquisition, traditional model-driven algorithms perform poorly in dynamic environments. In contrast, data-driven reinforcement learning algorithms are more suitable for dynamic task offloading by interactively approximating the optimal solution. With the development of deep reinforcement learning, algorithms that use centralized training and decentralized execution architectures can effectively cope with the non-stationary state of equipment and implement adaptive offloading strategies. However, when the number of robots increases, the increase in the global information dimension affects the algorithm performance and prolongs the training and reasoning time. To meet these challenges, a dynamic task offloading and resource allocation method is urgently needed.
[0006] The invention patent with application number 202311801292.4 discloses a D2D user resource allocation method based on multi-agent reinforcement learning, including the following steps: constructing a heterogeneous wireless network model in which D2D communication and cellular network share spectrum, and discretely dividing channel resource blocks and transmission power in the heterogeneous wireless network model; constructing a resource allocation optimization problem model based on the heterogeneous wireless network model with the goal of maximizing the total network throughput and ensuring user QoS; according to the resource allocation optimization problem model, a multi-agent reinforcement learning model is constructed with a D2D communication pair as an agent; the multi-agent reinforcement learning model is trained according to the MA-POCA algorithm; and the respective observation values of each D2D communication pair in the subsequent time slot are extracted, and the resource allocation scheme of each D2D communication pair is obtained by inputting the trained multi-agent reinforcement learning model. The above invention guarantees user QoS and reduces link interference in the network while maximizing system throughput. However, in the heterogeneous wireless network model, the discrete division of channel resource blocks and transmission power in the above invention may limit the flexibility of the system. Summary of the invention
[0007] In view of the technical problem that task offloading and resource allocation in the existing industrial Internet cannot fully utilize device computing resources, resulting in resource waste, the present invention proposes a D2D-assisted MEC computing resource allocation method based on a two-way offloading mechanism, which integrates a dynamic task shielding mechanism, improves task offloading efficiency, and is lightweight.
[0008] In order to achieve the above object, the technical solution of the present invention is implemented as follows: a D2D-assisted MEC computing resource allocation method based on a bidirectional offloading mechanism, the steps of which are as follows:
[0009] Step 1: Establish a D2D-assisted MEC computing offloading architecture in resource-constrained industrial Internet scenarios;
[0010] Step 2: Model the task offloading and resource allocation problem of the D2D-assisted MEC computing offloading architecture as a distributed partially observable Markov decision process based on a graph structure;
[0011] Step 3: A multi-agent deep reinforcement learning framework based on the fusion of Transformer and task shielding mechanism is used to solve the distributed partially observable Markov decision process based on the graph structure, and each industrial robot performs distributed adaptive task offloading and resource allocation decisions.
[0012] Preferably, the D2D-assisted MEC computing offloading architecture includes 1 industrial base station, 1 MEC server and K industrial robots. The industrial robot exchanges data with other heterogeneous devices in the scene through D2D communication. The industrial robot exchanges data with the MEC server through a cellular link. The MEC server is connected to the industrial base station through a wired network. The tasks generated by the industrial robot are transmitted to the industrial base station through the cellular link, and are thus offloaded to the MEC server for processing. The industrial robot set is represented as Industrial Robots Generate a task at time slot t express, d k (t) is the task size of robot k, e k It is the time slot mark of the task. is the maximum waiting delay of robot k’s task, It is a collection of tasks generated by industrial robots; the industrial robot buffer is used to store tasks that need to be calculated.
[0013] Preferably, the graph-structure-based distributed partially observable Markov decision process includes modeling of a graph-structure observation space and an action space of an industrial robot;
[0014] The modeling method of the graph structure observation space is:
[0015] At time slot t, industrial robot k observes other industrial robots l, The characteristics are:
[0016] ent k,l (t) = [ACK l (t),d l (t),c l (t),T l ,b l (t)],
[0017] Among them, ACK l (t)∈{-2,-1,0,1,2} represents the offloading signal of the industrial robot l, 0, 1, 2 represent the task offloading locally, offloading to D2D or offloading to MEC server, respectively, -1 and -2 represent the offloading failure caused by conflict during D2D offloading or MEC offloading, respectively, d l (t), c l (t), T l are the size of the first task in the buffer of industrial robot l, the CPU cycles required for calculation and the remaining delay, and b l (t) is the length of the buffer zone of the industrial robot l; when l = k, ent k,k (t) = [ACKk (t),d k (t),c k (t),T k ,b k (t)] represents the characteristics of the industrial robot k observed by itself;
[0018] At time slot t, the graph observation space o of industrial robot k k (t) = [ent k,1 (t),ent k,2 (t),...,ent k,K (t)] is the matrix of K*z, and the overall observation space O(t)={o1(t),o2(t),...,o K (t)}; where z is the total number of observable features of the industrial robot;
[0019] The modeling method of the action space of the industrial robot is:
[0020] At time slot t, industrial robot k observes and selects action a k (t) is:
[0021] a k (t) = [x k (t),β k (t)];
[0022] Among them, the industrial robot selects the action and saves it as the access signal x k (t)∈(0,K-1) and the channel access signal β k (t)∈(0,C), according to the access signal x k (t) value to obtain the service robot information of task unloading generated by industrial robot k, C is the number of channels, and according to the channel access signal β k The value of (t) obtains the channel information of the task offloaded to the MEC server.
[0023] Preferably, the reward function R(t) of the industrial robot k in time slot t is expressed as:
[0024] R(t)=R d (t)+R o (t)+R p (t);
[0025] Among them, the computational delay reward R of industrial robot k is d (t) is
[0026]
[0027] in, is the local processing delay of the task generated by industrial robot k in time slot t, is the D2D offloading delay of industrial robot k in time slot t, Offload the delay of the MEC server for industrial robot k in time slot t;
[0028] The reward R for successful or failed task processing on the MEC server o (t) is:
[0029]
[0030] Timeout penalty R for task offloading p (t) is
[0031]
[0032] After each task offloading or resource allocation decision, the industrial robot, as an intelligent agent, obtains rewards through the reward function R(t) according to the selected action, thereby continuously adjusting the strategy, optimizing the effects of offloading and resource allocation, and achieving the goal of efficient task processing under limited resources.
[0033] Preferably, the multi-agent deep reinforcement learning framework based on the fusion of Transformer and task shielding mechanism includes an Agent network deployed on the industrial robot and a Mixing network deployed on the MEC server.
[0034] The Agent network consists of an input layer, a gated recurrent unit, and an output layer I. Each industrial robot k independently observes the graph space o k (t) Generate the Q value, a core indicator in reinforcement learning for evaluating the quality of an action of an industrial robot in a certain state;
[0035] The core architecture of the Mixing network consists of an Embedder layer, a Transformer module and an output layer II. The output layer II includes a parameter generation network and an inference network. The parameter generation network receives the output results of the Transformer module and generates neuron weights and biases in the inference network. The inference network receives the Q values of all industrial robots and assigns the neuron weights and biases generated by the parameter generation network to the network itself, thereby inferring the global Q value.
[0036] Preferably, the graph observation space o k (t) The multi-layer perceptron entering the input layer performs feature extraction. The features extracted by the multi-layer perceptron and the hidden state h of the previous time step are continuously updated and transmitted by the gated recurrent unit according to the previous input and current timing information. k (t-1) is passed to the gated recurrent unit, which outputs the hidden state h at the current moment k(t), and the hidden state h k (t) It is passed to the next time step. The output of the gated recurrent unit is processed by the multi-layer perceptron of the output layer to further extract features and calculate the Q value, which is then passed to the Mixing network of the MEC server.
[0037] The input of the Embedder layer of the Mixing network is the global observation state s(t) composed of observations of each industrial robot. The global observation state after one-hot encoding is sent to the Transformer module for feature learning. The parameter generation network generates neuron weights w1 and w2. The inference network generates a global Q value Q through the neuron weights w1 and w2 and the Q values of all industrial robots. tot (τ,a), where τ represents the hidden state h at the current time step k (t) and the observation space O(t) of all industrial robots, a represents the action selected by all industrial robots at the current time step; the optimization goal of the Agent network is to maximize its own Q value, and the optimization goal of the Mixing network is to minimize the global Q value Q tot The mean square error between (τ,a) and the target Q value used to guide the Agent network to learn the optimal task offloading and resource allocation strategy. The Mixing network is constrained by Ensure the consistency of the two optimization goals and learn the optimal network resource allocation strategy.
[0038] Preferably, the Transformer module comprises a Linear layer, a Multi-Head Attention layer, a first Add&Norm layer, a Feed Forward layer and a second Add&Norm layer connected in sequence; the Multi-Head Attention layer comprises a plurality of parallel attention mechanisms; each attention mechanism is connected to a Linear layer;
[0039] The feature representation after the Embedder layer processing and the hidden state h of the previous time step k (t-1) is converted into a vector representation of fixed dimension through the Linear layer, and the sequence information is introduced through position encoding; the attention mechanism of the Multi-Head Attention layer allows the input vector of each position to pay attention to all positions in the sequence; the first Add&Norm layer and the second Add&Norm layer ensure the smooth flow of information and improve the training stability; the Feed Forward layer further enriches the feature representation through nonlinear transformation.
[0040] Preferably, the task shielding mechanism implements the steps of bidirectional task unloading as follows:
[0041] (1) The industrial robot is modeled as an intelligent agent that generates tasks with a fixed probability in each time slot and collects interaction data with other devices through D2D communication, including observations, actions, and rewards;
[0042] (2) The interaction data generated by the industrial robot in different time slots are constructed into a data frame in pandas format, and the deep learning-based DataWig model is used to predict the possible tasks of the industrial robot in the next time slot based on the existing features;
[0043] (3) The categorical target is categorized and encoded using a one-hot vector, and features are extracted through the Embedder layer; the remaining continuous numerical data is normalized and processed by a neural network, and concatenated with the embedding vector to generate complete input features;
[0044] (4) Use the classification cross entropy loss function and regression mean square error loss function to jointly train the classification target and continuous numerical data respectively;
[0045] (5) Use the trained neural network model to determine the task category of the industrial robot. The neural network model includes an embedding layer and a fully connected layer. The embedding layer processes the classification features after classification encoding and the processed continuous numerical data point features. The output layer uses the SoftMax function to generate the probability distribution of the task category and selects the category with the highest probability as the prediction result. Make task offloading decisions based on the prediction results and system status.
[0046] Preferably, a method for implementing joint task offloading and resource allocation comprises the following steps:
[0047] 1) Initialize the parameters of the neural network model and the Transformer module in the Mixing network, including initializing weights, biases, the number of attention heads, and other related hyperparameters;
[0048] 2) Build a graph structure based on the D2D communication range between industrial robots, and integrate the features of other industrial robots in the communication range into the graph observation space of the current industrial robot. For industrial robots beyond the observation range, the features are filled with 0 vectors.
[0049] 3) Observation results o of industrial robot k through D2D communication graph structure k (t) = [ent k,1 (t),ent k,2 (t),...,ent k,l (t)] to understand the specific situation of all current industrial robots generating tasks; at each time step, the industrial robot observes the space o based on the current graph k(t), the global observed state s(t) and the hidden state h(t), and select the action a(t) through the learned strategy;
[0050] 4) The environment evaluates the action a(t) performed by the industrial robot. The industrial robot receives the reward r(t) given by the environment and updates its strategy based on the reward. The environment transfers from the graph observation space O(t) and the global observation state s(t) to the graph observation space O(t+1) and the global observation state s(t+1) of the next time slot, and stores the industrial robot's global observation state s(t), the overall observation space O(t), the action a(t), the reward, and the next time slot global observation state s(t+1) and the overall observation space O(t+1) in the experience replay pool.
[0051] 5) When the amount of experience data in the experience replay pool reaches a certain amount, batch data is sampled from the experience replay pool, and the Transformer module is used to extract features from the batch data and update the hidden state h(t), and the parameters of the neural network model are updated based on the sampled experience data;
[0052] 6) Repeat iterative steps 1)-5) until the maximum number of simulation training rounds is reached. After the training is completed, each industrial robot performs distributed adaptive task offloading and resource allocation decisions to maximize the use of computing resources in the scene while meeting the delay constraints of computationally intensive tasks and delay-sensitive tasks.
[0053] Preferably, the classification target includes an offload signal ACK l (t)∈{-2,-1,0,1,2} and action a k (t) = [x k (t),β k (t)];
[0054] The execution results of the task offloading decision can be used as feedback data for retraining the neural network model;
[0055] After performing task offloading and resource allocation, rewards are obtained based on environmental feedback. Through the reward mechanism, the strategy is gradually adjusted and the decision-making process is continuously optimized to maximize the reward function R(t).
[0056] Compared with the prior art, the beneficial effects of the present invention are as follows: for industrial Internet scenarios with limited resources, a D2D (Device to Device)-assisted MEC (Mobile Edge Computing) computing offloading architecture is established; the D2D-assisted MEC task offloading and resource allocation problem is modeled as a distributed partially observable Markov decision process (POMDP) based on a graph structure; based on a multi-agent deep reinforcement learning framework that integrates Transformer and task shielding mechanism, a lightweight two-way task offloading and resource allocation algorithm is designed to deal with the complex decision-making and large state space of introducing D2D-assisted MEC two-way task offloading when resources are insufficient. The present invention builds a framework that integrates Transformer and multi-agent reinforcement learning to cope with scenarios where the complexity of offloading decisions increases as the number of agents increases, and a dynamic task shielding mechanism is designed to flexibly classify industrial robots in the scene, improve task offloading efficiency, and provide an efficient solution for task offloading and resource allocation in the industrial Internet. The beneficial effects of the present invention are specifically manifested in:
[0057] 1. Aiming at the challenge of insufficient computing resources in industrial Internet scenarios, the present invention proposes a joint task offloading and resource allocation algorithm based on data-driven multi-agent reinforcement learning, which gets rid of the dependence on precise system modeling, can support industrial robots to make autonomous decisions on task offloading solutions, and significantly improve the security and stability of the system.
[0058] 2. The present invention can fully tap and utilize the computing resources in the scene to cope with computing-intensive and industry-sensitive tasks, and at the same time achieve efficient allocation and utilization of resources while meeting the task latency requirements.
[0059] 3. The present invention has the characteristics of a lightweight model. Even when the number of intelligent agents increases, resulting in increased decision-making complexity, the model scale remains unchanged and various performance indicators still perform well. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0061] Figure 1 This is a schematic diagram of the present invention.
[0062] Figure 2 This is a model schematic diagram of the bidirectional task offloading resource allocation algorithm based on multi-agent deep reinforcement learning of the present invention.
[0063] Figure 3 The present invention is a flowchart of a method for industrial Internet task offloading and resource allocation under resource constraints.
[0064] Figure 4 1 is a comparison chart of the global rewards of the present invention and the existing algorithm on multiple industrial robots, where (a) is 10 and (b) is 20.
[0065] Figure 5 The figure is a comparison chart of the average processing delay of the present invention and the existing algorithm on multiple industrial robots, where (a) is 10 and (b) is 20.
[0066] Figure 6 The figure is a comparison chart of the average unloading success rate of the present invention and the existing algorithm on multiple industrial robots, where (a) is 10 and (b) is 20. DETAILED DESCRIPTION
[0067] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0068] like Figure 3 As shown, the present invention proposes a D2D-assisted MEC computing resource allocation method based on a bidirectional offloading mechanism, comprising the following steps:
[0069] Step 1: In resource-constrained industrial Internet scenarios, establish a D2D-assisted MEC computing offloading architecture.
[0070] like Figure 1 As shown in the figure, in order to make full use of computing resources and improve task processing efficiency and system stability, a D2D-assisted MEC computing offloading architecture is established. The architecture includes 1 industrial base station, 1 mobile edge server (MEC server) and K industrial robots. The industrial robot exchanges data with other heterogeneous devices in the scene through D2D communication, and exchanges data through cellular links and mobile edge services, prompting devices without sufficient computing resources to offload tasks to nearby idle devices or mobile edge servers for auxiliary computing. The mobile edge server is connected to the industrial base station through a wired network. The tasks generated by the industrial robot can be transmitted to the industrial base station through the cellular link, and then offloaded to the mobile edge server for processing. Industrial robots include distributed intelligent devices such as robotic arms and welding robots. The set of industrial robots is represented as Smart devices are connected to industrial base stations via wireless networks, including industrial robots Generate a task at time slot t express, d k (t) is the task size of robot k, e k It is the time slot mark of the task. is the maximum waiting delay of robot k’s task, It is a collection of tasks generated by industrial robots. The industrial robot buffer is used to store tasks that need to be calculated. In each time slot, the industrial robots are dynamically divided into task robots that need to offload tasks and service robots that provide computing resources according to the storage of tasks in the industrial robot buffer. The decision of task offloading becomes complicated and the complexity increases. The complexity of task offloading is reflected in the different roles played by industrial robots in different time slots. Industrial robots play different roles in different time slots. Each industrial robot may be a task robot that needs to offload tasks in the current time slot, but it may be transformed into a service robot that provides its own computing resources for auxiliary offloading in the next time slot. Due to the characteristics of the tasks and the changes in the number of tasks in the buffer, the industrial robot must analyze its own status in real time and make decisions. This dynamic role switching and task characteristics make the industrial robot face higher complexity when making offloading decisions, resulting in the continuous expansion of the state space of task offloading decisions, which increases the difficulty of model convergence.
[0071] Step 2: Considering the high requirements of computationally intensive and time-sensitive tasks in the scenario on heterogeneous device resources, the task offloading and resource allocation problem of D2D-assisted MEC is modeled as a distributed partially observable Markov decision process based on a graph structure.
[0072] How industrial robots make real-time task offloading decisions in a dynamic environment based on limited resources, partially observable device states and task requirements can be summarized as the task offloading and resource allocation problem of D2D-assisted MEC. Considering the high demand for heterogeneous device resources for computationally intensive tasks and time-sensitive tasks in the scenario and the fact that limited computing resources may cause long waiting delays when tasks are offloaded to mobile edge servers, this paper proposes a collaborative offloading strategy to make full use of the computing resources of idle D2D devices. The observation space information of the graph structure is dynamically acquired by the industrial robot through D2D communication to achieve more efficient task scheduling and resource allocation, while reducing offloading delays.
[0073] The distributed partially observable Markov decision process based on graph structure includes the modeling of graph structure observation space and device action space.
[0074] First, each of the K industrial robots is set as an observable entity with z features. The observations of the industrial robot include its own features and the features of K-1 other entities obtained through D2D communication.
[0075] At time slot t, industrial robot k observes other industrial robots l, The feature definition is as follows:
[0076] ent k,l (t) = [ACK l (t),d l (t),c l (t),T l ,b l (t)],
[0077] Among them, ACK l (t)∈{-2,-1,0,1,2} represents the offloading signal of the industrial robot l, where 0, 1, and 2 represent the task offloading locally, offloading to D2D, or offloading to the MEC server, respectively; -1 and -2 represent the offloading failure caused by the conflict during D2D offloading or MEC offloading, respectively; d l (t), c l (t), T l They are the size of the first task in the buffer of the industrial robot l, the CPU cycles required for calculation, and the remaining delay. These three characteristics are important factors in task offloading calculation, which can help determine whether the task can be completed on time or whether it has timed out, thus affecting the decision and scheduling of task offloading. l (t) is the length of the buffer zone of the industrial robot l. When l = k, ent k,k (t) = [ACK k (t),d k (t),c k (t),T k ,b k (t)], which represents the characteristics of the industrial robot k observed by itself.
[0078] At time slot t, the graph observation space o of industrial robot k k (t) = [ent k,1 (t),ent k,2 (t),...,ent k,K (t)] is a K*z matrix, and the overall observation space of the system can be expressed as O(t)={o1(t),o2(t),...,o K (t)}.
[0079] At time slot t, industrial robot k observes and selects action a k (t) is defined as follows:
[0080] a k (t) = [x k (t),β k (t)];
[0081] Among them, the industrial robot selects the action and saves it as the access signal x k (t)∈(0,K-1) and the channel access signal β k (t)∈(0,C), according to x k The value of (t) can be used to obtain the service robot information generated by the industrial robot k to offload the task, and the channel access signal β k (t)∈(0,C), C is the number of channels, according to β k The value of (t) can obtain the channel information of the task offloaded to the MEC server. The state space of the graph structure is conducive to the designed algorithm to better extract device feature information.
[0082] The longer the task unloading takes, the smaller the reward is, and the shorter the time is, the greater the reward is. Taking into account the comprehensive consideration, the composite reward function in the model is designed to combine delayed reward, unloading success reward and timeout penalty, including the following steps:
[0083] At time slot t, the computation delay reward R of industrial robot k is d (t) is expressed as
[0084]
[0085] in, is the local processing delay of the task generated by industrial robot k in time slot t, is the D2D offloading delay of industrial robot k in time slot t, Offload the delay of the MEC server for industrial robot k in time slot t.
[0086] The reward R for successful or failed task processing on the MEC server o (t) Yes
[0087]
[0088] Timeout penalty R for task offloading p (t) is expressed as
[0089]
[0090] In summary, by fully combining the above rewards and penalties, the reward function R(t) of the industrial robot k in time slot t can be expressed as
[0091] R(t)=R d (t)+R o (t)+R p (t).
[0092] Step 3: Based on the multi-agent deep reinforcement learning framework that integrates Transformer and task shielding mechanism, the distributed partially observable Markov decision process based on the graph structure is solved. Each industrial robot can make adaptive task offloading and resource allocation decisions in a distributed manner, and judge the quality of the offloading decision from three indicators: global reward, average task success rate, and average task offloading delay. In the algorithm, the reward function is closely integrated with the state and action of the agent. After each task offloading or resource allocation decision is executed, the agent will obtain rewards through the reward function R(t) according to its selected action, thereby continuously adjusting the strategy, optimizing the effect of offloading and resource allocation, and finally achieving the goal of efficient task processing under limited resources.
[0093] Based on a multi-agent deep reinforcement learning framework that integrates Transformer and task shielding mechanism, a task offloading and resource allocation algorithm is designed to solve the problem of complex offloading decisions and difficult model convergence caused by introducing D2D-assisted MEC when resources are insufficient. The overall network architecture of the algorithm is as follows: Figure 2 shown.
[0094] First, we build an agent network deployed on an industrial robot. The agent network consists of an input layer (MLP), a gated recurrent unit (GRU), and an output layer (MLP). Each industrial robot independently observes the space o k (t) Generate the Q value, the core indicator for evaluating the quality of an industrial robot’s action in a certain state in reinforcement learning.
[0095] Graph observation space o k (t) First, the multi-layer perceptron (MLP) network in the input layer is used for feature extraction. This step can map the original observation information and action information to a higher-dimensional or more suitable feature space. The features extracted by MLP and the hidden state h of the previous time step are continuously updated and transmitted by GRU based on the previous input and current timing information. k (t-1) is passed to GRU (Gate Recurrent Unit) together, which can effectively capture the previous and next dependencies in the time series.
[0096] GRU outputs the hidden state h at the current moment k (t), and the hidden state h k (t) is passed to the next time step, and the output of the GRU is processed by the second MLP to further extract features and calculate the Q value, which is then passed to the Transformer-Mixing network.
[0097] Secondly, build a Mixing network deployed on the MEC server. The core architecture of the Mixing network consists of an Embedder layer, a Transformer module, and an output layer. The output layer consists of two deep neural networks, namely the parameter generation network and the inference network. The parameter generation network receives the output results of the Transformer module and generates the neuron weights and biases in the inference network. The inference network receives the Q values of all industrial robots and assigns the neuron weights and biases generated by the parameter generation network to the network itself, thereby inferring the global Q value.
[0098] The input of the Embedder layer of the Mixing network is the global observation state s(t) composed of observations of each industrial robot. The global observation state s(t) after one-hot encoding is sent to the Transformer module for feature learning. The parameter generation network generates neuron weights w1 and w2. Finally, the inference network generates a global Q value Q through neuron weights w1 and w2 and the Q values of all industrial robots. tot (τ,a), where τ represents the hidden state h at the current time step k (t) and the observation space O(t) of all industrial robots, a represents the action selected by all industrial robots at the current time step. Global Q value Q tot (τ,a) can reflect the performance of the entire system. By integrating the Q value of each industrial robot and performing feature learning, it helps optimize global resource allocation and task scheduling. The optimization goal of the Agent network is to maximize its own Q value, and the optimization goal of the Mixing network is to minimize the global Q value Q tot The mean square error between (τ,a) and the target Q value used to guide the Agent network to learn the optimal task offloading and resource allocation strategy. The Mixing network is constrained by Ensure the consistency of the two optimization goals and learn the optimal network resource allocation strategy.
[0099] Finally, a task shielding mechanism is designed to implement the two-way task offloading method. The task shielding mechanism design includes the following steps:
[0100] (1) The industrial robot is modeled as an intelligent agent that generates tasks with a fixed probability in each time slot and collects interaction data with other devices through D2D communication, including observation, action, and reward data.
[0101] (2) The interaction data generated by the robot at different time slots is constructed into a data frame in pandas format, making the data analysis and modeling process more efficient, and using the deep learning-based DataWig model to predict the tasks that the industrial robot may generate in the next time slot based on the existing features. The prediction results will help the robot better understand the task mode, improve its responsiveness to future tasks, and enable it to make more intelligent decisions in a dynamic environment.
[0102] (3) Among them, the unloading signal ACK l (t)∈{-2,-1,0,1,2} and action a k (t) = [x k (t),β k (t)] uses one-hot vectors for classification encoding and extracts features through the embedding layer; the remaining numerical data is processed by a normalized neural network and concatenated with the embedding vector to generate complete input features, which can significantly improve the expressiveness of the model. The embedding layer can learn the latent semantics of categorical features, improve the model's ability to process categorical data, and reduce feature dimensions, avoid dimensionality disasters, and accelerate model convergence. Normalization ensures that features are on the same scale, improving the stability and convergence speed of the model. The concatenated input features can effectively combine categorical and numerical features, enhance feature interactions, and improve the model's ability to learn complex patterns. Neural networks can make full use of all types of input features, making the model more intelligent and stable, and adapting to complex task offloading and decision-making needs.
[0103] (4) The classification cross entropy loss function and the regression mean square error loss function are used to jointly train the classification targets (such as ACK signals and device actions) and the continuous numerical targets (such as observation values and reward values) respectively, so as to improve the model's adaptability to classification and regression tasks.
[0104] (5) To use the trained neural network model to determine the task category of the industrial robot, a model including an embedding layer and a fully connected layer needs to be constructed. The input layer processes classification features (such as one-hot encoding of ACK signals and device actions) and numerical features (such as normalization of observations and rewards) through the embedding layer. The output layer uses SoftMax to generate the probability distribution of task categories and selects the category with the highest probability as the prediction result. Task offloading decisions are made based on the predicted category and system status to improve resource utilization efficiency and accuracy. In addition, execution feedback data can be used for model retraining to ensure its continuous optimization in a dynamic environment, thereby improving the intelligence and adaptability of the robot.
[0105] like Figure 3 As shown, a joint task offloading and resource allocation algorithm mainly includes the following steps:
[0106] (1) Initialize the parameters of the algorithm's neural network model and the Transformer module in the Mixing network, including initializing weights, biases, the number of attention heads, and other related hyperparameters. The Transformer module is nested in the Mixing network and is mainly used to process the complex interactions between agents and generate joint feature representations. The structure of the Transformer module consists of multiple core modules, which can efficiently capture the complex dependencies in sequence data. Input the feature representation processed by the Embedder layer and the hidden state h of the previous time step k (t-1) is converted into a vector representation of fixed dimension through the Linear layer, and the sequence information is introduced through position encoding. The Multi-Head Attention mechanism allows the input vector of each position to weight the attention of all positions in the sequence, thereby capturing global dependencies and enhancing the expressiveness through parallel multiple independent attention calculations. The Add&Norm layer ensures smooth information flow and improves training stability. The Feed Forward layer further enriches the feature representation through nonlinear transformation and improves the expressiveness of the model. These modules work together to enable the Transformer module to adapt to a variety of task scenarios, while improving expressiveness and speeding up training.
[0107] (2) A graph structure is constructed based on the D2D communication range between industrial robots, and the features of other industrial robots within the communication range are integrated into the observation space of the current industrial robot. For devices beyond the observation range, their features are filled with 0 vectors.
[0108] (3) Observation results o of industrial robot k through the D2D communication graph structure k (t) = [ent k,1 (t),ent k,2 (t),...,ent k,l (t)] can understand the specific situation of all current industrial robots generating tasks. The industrial robot generating the task does not provide computing resources to other industrial robots. This is called the task shielding mechanism. At each time step, the industrial robot observes the space o based on the current graph. k(t), global observation state s(t) and hidden state h(t), the present invention uses the bidirectional task offloading and resource allocation (Bidirectional Task Offloading and Resource Allocation, BTORT) algorithm to learn the given strategy π. The system initially uses a simple strategy, performs task offloading and resource allocation, and obtains rewards based on environmental feedback; through the reward mechanism, the strategy is gradually adjusted, and the decision-making process is continuously optimized to maximize long-term rewards. This process is repeated through iterations to eventually find the optimal task offloading and resource allocation strategy. Then execute actions a(t) (such as performing task offloading, adjusting computing resources, etc.) to interact with the industrial wireless network environment.
[0109] (4) After executing action a(t), the environment will evaluate the action performed by the industrial robot. The industrial robot will receive the reward r(t) given by the environment and update the strategy based on the reward. The environment moves from the graph observation space o k (t), the global observation state s(t) is transferred to the next moment graph observation space o k (t+1), the global observation state s(t+1), and then the complete observation, global observation state, action, reward and the next time slot observation and state [s(t), O(t), R(t), a(t), s(t+1), O(t+1)] of the industrial robot are stored in the experience replay pool, which will be used for subsequent training.
[0110] (5) When the amount of experience data in the experience replay pool reaches a certain amount (for example, greater than the size of the small batch sampling), start sampling small batch data from the experience replay pool, use the Transformer module to extract features from the batch data and update the hidden state h(t), and update the neural network parameters based on the sampled experience data. The industrial robot interacts with the environment, collects experience and trains the neural network.
[0111] (6) Repeat iterative steps (1)-(5) until the maximum number of simulation training rounds is reached. After the training is completed, each industrial robot can perform adaptive task offloading and resource allocation decisions in a distributed manner, thereby maximizing the utilization of computing resources in the scene while satisfying the delay constraints of computing-intensive tasks and delay-sensitive tasks.
[0112] The BTORT algorithm of the present invention is used to solve the problem of task offloading and resource allocation. Therefore, the cumulative rewards, average offloading delay of tasks and average completion rate of tasks of the two contrasting algorithms, QMIX algorithm and IQL algorithm, are compared. QMIX is the abbreviation of Q-value Mixing, which is commonly used for joint strategy learning in multi-agent systems. The core idea of the QMIX algorithm is to fuse the local Q value of each agent to form a global Q value, so that the multi-agent system can make decisions on the basis of collaboration. IQL is the abbreviation of Implicit Q-learning, which is an algorithm for Q learning without an explicit objective function, and is usually used to process complex tasks and large-scale state spaces. The IQL algorithm makes decisions by implicitly learning the state value function, avoiding the explicit calculation and storage of all state-action pairs in traditional Q learning.
[0113] Figure 4 It is clearly shown that the BDTORA algorithm of the present invention outperforms the benchmark algorithms, such as the QMIX algorithm and the IQL algorithm, in terms of cumulative rewards. The Transformer module introduced in the BDTORA algorithm of the present invention can dynamically adjust the attention weights to adapt to the changing network environment. This capability ensures that the strategy can be quickly optimized when dealing with fluctuating network conditions and task requirements, thereby obtaining higher cumulative rewards.
[0114] like Figure 5 As shown in the figure, the BDTORA algorithm of the present invention performs well in reducing the average processing delay, and its value is significantly lower than the baseline algorithm. The ability of the Transformer module to capture long-range dependencies plays a key role in the task offloading problem, making the prediction of future states and potential rewards more accurate. This prediction accuracy helps to optimize the task offloading strategy, thereby effectively reducing processing delays.
[0115] Figure 6 It is shown that the BDTORA algorithm of the present invention exceeds the benchmark algorithms such as QMIX and IQL in terms of task offloading success rate. The use of the self-attention mechanism in the BDTORA algorithm of the present invention makes it possible to selectively focus on features that are closely related to task offloading. This precise focus on relevant features enhances the decision-making process, thereby improving the success rate of task offloading.
[0116] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. A D2D-assisted MEC computing resource allocation method based on a bidirectional offloading mechanism, characterized in that: The steps are as follows: Step 1: Establish a D2D-assisted MEC computing offloading architecture in resource-constrained industrial Internet scenarios; Step 2: Model the task offloading and resource allocation problem of the D2D-assisted MEC computing offloading architecture as a distributed partially observable Markov decision process based on a graph structure; Step 3: A multi-agent deep reinforcement learning framework based on the fusion of Transformer and task shielding mechanism is used to solve the distributed partially observable Markov decision process based on the graph structure, and each industrial robot performs distributed adaptive task offloading and resource allocation decisions.
2. The D2D-assisted MEC computing resource allocation method based on a bidirectional offloading mechanism according to claim 1 is characterized in that: The D2D-assisted MEC computing offloading architecture includes 1 industrial base station, 1 MEC server and K industrial robots. The industrial robot exchanges data with other heterogeneous devices in the scene through D2D communication. The industrial robot exchanges data with the MEC server through a cellular link. The MEC server is connected to the industrial base station through a wired network. The tasks generated by the industrial robot are transmitted to the industrial base station through a cellular link, and then offloaded to the MEC server for processing. The industrial robot set is represented as Industrial Robots Generate a task at time slot t express, d k (t) is the task size of robot k, e k It is the time slot mark of the task. is the maximum waiting delay of robot k’s task, It is a collection of tasks generated by industrial robots; the industrial robot buffer is used to store tasks that need to be calculated.
3. The D2D-assisted MEC computing resource allocation method based on a bidirectional offloading mechanism according to claim 1 or 2, characterized in that: The graph-based distributed partially observable Markov decision process includes modeling of a graph-structured observation space and an action space of an industrial robot; The modeling method of the graph structure observation space is: At time slot t, industrial robot k observes other industrial robots The characteristics are: ent k,l (t)=[ACK l (t),d l (t),c l (t),T l ,b l (t)], Among them, ACK l (t)∈{-2,-1,0,1,2} represents the offloading signal of the industrial robot l, 0, 1, 2 represent the task offloading locally, offloading to D2D or offloading to MEC server, respectively, -1 and -2 represent the offloading failure caused by conflict during D2D offloading or MEC offloading, respectively, d l (t), c l (t), T l are the size of the first task in the buffer of industrial robot l, the CPU cycles required for calculation and the remaining delay, and b l (t) is the length of the buffer zone of the industrial robot l; when l = k, ent k,k (t) = [ACK k (t),d k (t),c k (t),T k ,b k (t)] represents the characteristics of the industrial robot k observed by itself; At time slot t, the graph observation space o of industrial robot k k (t) = [ent k,1 (t),ent k,2 (t),...,ent k,K (t)] is the matrix of K*z, and the overall observation space O(t)={o1(t),o2(t),...,o K (t)}; where z is the total number of observable features of the industrial robot; The modeling method of the action space of the industrial robot is: At time slot t, industrial robot k observes and selects action a k (t) is: a k (t)=[x k (t),β k (t)]; Among them, the industrial robot selects the action and saves it as the access signal x k (t)∈(0,K-1) and the channel access signal β k (t)∈(0,C), according to the access signal x k (t) value to obtain the service robot information of task unloading generated by industrial robot k, C is the number of channels, and according to the channel access signal β k The value of (t) obtains the channel information of the task offloaded to the MEC server.
4. The D2D-assisted MEC computing resource allocation method based on a bidirectional offloading mechanism according to claim 3 is characterized in that: The reward function R(t) of the industrial robot k in time slot t is expressed as: R(t)=R d (t)+R o (t)+R p (t); Among them, the computational delay reward R of industrial robot k is d (t) is in, is the local processing delay of the task generated by industrial robot k in time slot t, is the D2D offloading delay of industrial robot k in time slot t, Offload the delay of the MEC server for industrial robot k in time slot t; The reward R for successful or failed task processing on the MEC server o (t) is: Timeout penalty R for task offloading p (t) is After each task offloading or resource allocation decision, the industrial robot, as an intelligent agent, obtains rewards through the reward function R(t) according to the selected action, thereby continuously adjusting the strategy, optimizing the effects of offloading and resource allocation, and achieving the goal of efficient task processing under limited resources.
5. The D2D-assisted MEC computing resource allocation method based on a bidirectional offloading mechanism according to any one of claims 1, 2, and 4, characterized in that: The multi-agent deep reinforcement learning framework based on the fusion of Transformer and task shielding mechanism includes an Agent network deployed on an industrial robot and a Mixing network deployed on a MEC server. The Agent network consists of an input layer, a gated recurrent unit, and an output layer I. Each industrial robot k independently observes the graph space o k (t) Generate the Q value, a core indicator in reinforcement learning for evaluating the quality of an action of an industrial robot in a certain state; The core architecture of the Mixing network consists of an Embedder layer, a Transformer module and an output layer II. The output layer II includes a parameter generation network and an inference network. The parameter generation network receives the output results of the Transformer module and generates neuron weights and biases in the inference network. The inference network receives the Q values of all industrial robots and assigns the neuron weights and biases generated by the parameter generation network to the network itself, thereby inferring the global Q value.
6. The D2D-assisted MEC computing resource allocation method based on a bidirectional offloading mechanism according to claim 5 is characterized in that: The graph observation space o k (t) The multi-layer perceptron entering the input layer performs feature extraction. The features extracted by the multi-layer perceptron and the hidden state h of the previous time step are continuously updated and transmitted by the gated recurrent unit according to the previous input and current timing information. k (t-1) is passed to the gated recurrent unit, which outputs the hidden state h at the current moment k (t), and the hidden state h k (t) It is passed to the next time step. The output of the gated recurrent unit is processed by the multi-layer perceptron of the output layer to further extract features and calculate the Q value, which is then passed to the Mixing network of the MEC server. The input of the Embedder layer of the Mixing network is the global observation state s(t) composed of observations of each industrial robot. The global observation state after one-hot encoding is sent to the Transformer module for feature learning. The parameter generation network generates neuron weights w1 and w2. The inference network generates a global Q value Q through the neuron weights w1 and w2 and the Q values of all industrial robots. tot (τ,a), where τ represents the hidden state h at the current time step k (t) and the observation space O(t) of all industrial robots, a represents the action selected by all industrial robots at the current time step; the optimization goal of the Agent network is to maximize its own Q value, and the optimization goal of the Mixing network is to minimize the global Q value Q tot The mean square error between (τ,a) and the target Q value used to guide the Agent network to learn the optimal task offloading and resource allocation strategy. The Mixing network is constrained by Ensure the consistency of the two optimization goals and learn the optimal network resource allocation strategy.
7. The D2D-assisted MEC computing resource allocation method based on a bidirectional offloading mechanism according to claim 5 is characterized in that: The Transformer module includes a Linear layer, a Multi-Head Attention layer, a first Add&Norm layer, a Feed Forward layer, and a second Add&Norm layer connected in sequence; the Multi-Head Attention layer includes multiple parallel attention mechanisms; each attention mechanism is connected to a Linear layer; The feature representation after the Embedder layer processing and the hidden state h of the previous time step k (t-1) is converted into a vector representation of fixed dimension through the Linear layer, and the sequence information is introduced through position encoding; the attention mechanism of the Multi-Head Attention layer allows the input vector of each position to pay attention to all positions in the sequence; the first Add&Norm layer and the second Add&Norm layer ensure the smooth flow of information and improve the training stability; the Feed Forward layer further enriches the feature representation through nonlinear transformation.
8. The D2D-assisted MEC computing resource allocation method based on a bidirectional offloading mechanism according to claim 6 or 7, characterized in that: The steps of implementing bidirectional task unloading by the task shielding mechanism are as follows: (1) The industrial robot is modeled as an intelligent agent that generates tasks with a fixed probability in each time slot and collects interaction data with other devices through D2D communication, including observations, actions, and rewards; (2) The interaction data generated by the industrial robot in different time slots are constructed into a data frame in pandas format, and the deep learning-based DataWig model is used to predict the tasks that the industrial robot may generate in the next time slot based on the existing features; (3) The categorical target is categorized and encoded using a one-hot vector, and features are extracted through the Embedder layer; the remaining continuous numerical data is normalized and processed by a neural network, and concatenated with the embedding vector to generate complete input features; (4) Use the classification cross entropy loss function and regression mean square error loss function to jointly train the classification target and continuous numerical data respectively; (5) Use the trained neural network model to determine the task category of the industrial robot. The neural network model includes an embedding layer and a fully connected layer. The embedding layer processes the classification features after classification encoding and the processed continuous numerical data point features. The output layer uses the SoftMax function to generate the probability distribution of the task category and selects the category with the highest probability as the prediction result. Make task offloading decisions based on the prediction results and system status.
9. The D2D-assisted MEC computing resource allocation method based on a bidirectional offloading mechanism according to claim 8 is characterized in that: A method for implementing joint task offloading and resource allocation comprises the following steps: 1) Initialize the parameters of the neural network model and the Transformer module in the Mixing network, including initializing weights, biases, the number of attention heads, and other related hyperparameters; 2) Build a graph structure based on the D2D communication range between industrial robots, and integrate the features of other industrial robots in the communication range into the graph observation space of the current industrial robot. For industrial robots beyond the observation range, the features are filled with 0 vectors; 3) Observation results o of industrial robot k through D2D communication graph structure k (t) = [ent k,1 (t),ent k,2 (t),...,ent k,l (t)] to understand the specific situation of all current industrial robots generating tasks; at each time step, the industrial robot observes the space o based on the current graph k (t), the global observed state s(t) and the hidden state h(t), and select the action a(t) through the learned strategy; 4) The environment evaluates the action a(t) performed by the industrial robot. The industrial robot receives the reward r(t) given by the environment and updates its strategy based on the reward. The environment transfers from the graph observation space O(t) and the global observation state s(t) to the graph observation space O(t+1) and the global observation state s(t+1) of the next time slot, and stores the industrial robot's global observation state s(t), the overall observation space O(t), the action a(t), the reward, and the next time slot global observation state s(t+1) and the overall observation space O(t+1) in the experience replay pool. 5) When the experience data in the experience replay pool reaches a certain amount, batch data is sampled from the experience replay pool, and the Transformer module is used to extract features from the batch data and update the hidden state h(t), and the parameters of the neural network model are updated based on the sampled experience data; 6) Repeat iterative steps 1)-5) until the maximum number of simulation training rounds is reached. After the training is completed, each industrial robot performs distributed adaptive task offloading and resource allocation decisions to maximize the use of computing resources in the scene while meeting the delay constraints of computing-intensive tasks and delay-sensitive tasks.
10. The D2D-assisted MEC computing resource allocation method based on a bidirectional offloading mechanism according to claim 9 is characterized in that: The classification target includes an offload signal ACK l (t)∈{-2,-1,0,1,2} and action a k (t) = [x k (t),β k (t)]; The execution results of the task offloading decision can be used as feedback data for retraining the neural network model; After performing task offloading and resource allocation, rewards are obtained based on environmental feedback. Through the reward mechanism, the strategy is gradually adjusted and the decision-making process is continuously optimized to maximize the reward function R(t).
Citation Information
Patent Citations
D2D user resource allocation method based on multi-agent reinforcement learning
CN118118908A