Multi-Agent Collaboration Method for MEC Systems Integrating AoI and Intrinsic Motivation

By introducing a multi-agent deep reinforcement learning framework and local information self-identification module in the MEC system, the problems of limited resources and heterogeneous AoI in edge devices are solved, and collaborative learning between edge devices and cloud centers are realized, resource management is optimized, delay is reduced, and data freshness and system efficiency are improved.

CN119788702BActive Publication Date: 2025-07-08HARBIN INST OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411987177.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-07-08
Estimated Expiration
2044-12-31

AI Technical Summary

Technical Problem

In MEC systems, edge device resources are limited and data tasks change dynamically, resulting in preemption of communication resources and computing resources, resulting in increased task delay and energy consumption, and heterogeneous AoI and resource heterogeneity bring about inefficient collaboration.

Method used

Using a multi-agent collaboration method that integrates AoI and intrinsic incentives, we use the multi-agent deep reinforcement learning (MADRL) framework, combined with heterogeneous AoI threshold optimization mechanism and local information self-identification module (LISIM), optimize resource management and scheduling strategies to achieve collaborative learning at the edge and cloud center.

Benefits of technology

Effectively reduce the information age of data, improve the collaboration efficiency between devices, optimize resource utilization, and ensure data freshness and system performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119788702B_ABST
    Figure CN119788702B_ABST
Patent Text Reader

Abstract

A multi-agent cooperation method for MEC systems integrating AoI and intrinsic motivation belongs to the technical field of cloud computing and edge computing. The method is as follows: generation of heterogeneous resources at the edge side and edge movement decision-making; communication link modeling; collection of heterogeneous resources and AoI threshold limitation; heterogeneous age-sensitive optimization problem; formulation of external reward function and intrinsic reward function by MDP; construction and functional analysis of heterogeneous multi-agent Actor-Critic; intrinsic reward based on local information recognition. The present invention extends computing power to the network edge through MEC, improves the real-time performance of data processing and reduces latency. The concept of AoI is introduced to measure data freshness, and through optimizing resource management and scheduling strategies, efficient utilization of resources and meeting dynamic data requirements are achieved. A multi-agent deep reinforcement learning strategy is proposed to ensure data freshness and improve the cooperation efficiency between devices.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a multi - agent cooperation method for a MEC system integrating AoI and intrinsic motivation, belonging to the technical fields of cloud computing and edge computing. Background Art

[0002] The data sources in manufacturing are highly diverse, including various data from devices, sensors, production lines, product tracking systems, etc., and involving different formats, structures, frequencies, and real - time requirements. These heterogeneous data need effective governance and integration to provide valuable insights for enterprises. Edge computing (EC) realizes lower - latency and higher - efficiency real - time data processing by migrating data - processing capabilities from the cloud to the edge of the network. Mobile edge computing (MEC) further optimizes this process, especially for applications with mobility requirements. MEC pushes computing, storage, and network capabilities closer to the data source, thus solving the latency and bandwidth pressure caused by excessive data volume in the cloud computing center. In manufacturing, MEC can effectively handle the massive heterogeneous data generated by devices, sensors, and systems, especially in industrial environments that require real - time response. For example, in an intelligent factory, tasks such as equipment monitoring, production line optimization, and fault warning rely on low - latency and high - reliability computing capabilities, and MEC provides an ideal solution.

[0003] The MEC system usually adopts a three - layer structure (as Figure 1 shown), which helps to provide efficient edge - computing services. The bottom layer is the data source, referring to the devices closest to the user equipment connected to the network, usually including sensors, smartphones, and smart home devices. The middle layer is the edge device, which can be a smart router, a smart base station, a drone (UAV), and a vehicle with an in - vehicle processor, playing a relaying role. The edge device can collect data from edge sensors, assist in performing some computing tasks, and then relay them to the cloud center. The top - layer cloud data center plays a central control role in the edge - computing solution. It integrates and analyzes the data from edge nodes and assembles tasks according to actual service requirements for processing at the edge. However, the resources of edge servers are not infinite. If the task distribution is uneven, it will cause the preemption of communication resources and edge - server computing resources, resulting in increased latency and energy consumption for task execution.

[0004] This study focuses on the real - time and low - latency performance of image and video transmission and processing in the MEC system. In critical applications such as urban surveillance, data - processing latency directly affects the quality of service. The age of information (AoI) is introduced to measure data freshness. In extreme cases, an AoI threshold is set. For example, data exceeding the threshold loses its value rapidly in disaster response. Therefore, controlling AoI below the threshold and reducing the total AoI are the core objectives.

[0005] In the MEC system, due to the limited computing resources of mobile edge devices, the dynamic changes in data tasks, and the communication bandwidth limitations, it is crucial to design effective scheduling strategies to optimize resource utilization and meet data requirements. This involves resource management, edge decision-making, and data offloading. The globally optimal strategy requires cooperation and coordination among edge devices, but resource heterogeneity poses challenges. Different devices and cloud centers have different capabilities, resources, goals, and sensing capabilities, and devices in different geographical locations generate data of different importance.

[0006] To address these challenges, the MEC system is regarded as a multi-agent system, where edge devices such as drones and vehicle-mounted processors act as agents with intelligence and autonomy. Multi-agent deep reinforcement learning (MADRL) provides a new idea for solving sequential decision-making problems. Value-based methods such as Q-learning and DQN make decisions by predicting action values, while policy-based methods such as DDPG are suitable for tasks with a large or continuous control space. Federated learning (FL) combined with MADRL allows agents to collaborate and learn while protecting privacy, optimizing system performance.

[0007] Given the above background and challenges, the present invention proposes a new MADRL strategy for the heterogeneous AoI and resource collaboration problems of different edge devices in the MEC environment to ensure data freshness and the collaboration efficiency among devices, and attempts to solve the problem that multi-agents rely on global information under heterogeneous AoI and resource conditions. The Internet of Things technology and sensor networks are fully utilized to achieve real-time monitoring and data collection of various devices and sensors. At the same time, through the data processing capabilities of edge computing nodes, preliminary analysis and preprocessing of data can be realized, reducing the data processing pressure on the cloud. Summary of the Invention

[0008] To solve the problems existing in the background technology, the present invention provides a multi-agent collaboration method for an MEC system that integrates AoI and intrinsic incentives.

[0009] To achieve the above object, the present invention adopts the following technical solution: A multi-agent collaboration method for an MEC system that integrates AoI and intrinsic incentives, the method comprising the following steps:

[0010] S1: Generation of end-side heterogeneous resources and edge movement decision-making;

[0011] S101: Generation of end-side device data sources;

[0012] S10101: End-side devices collect computing task data generated by intelligent devices of distributed sensors in the Internet of Things to generate data sources;

[0013] S10102: Due to differences in geographical environment and hardware configuration, the data sources form heterogeneous resource data to generate data packets in different forms;

[0014] S10103: Assign weights to each data source S n respectively to represent the importance of different positions, and assume that the data sources are independently generated;

[0015] S10104: Define the data size of the data packet as d(t), the data generation time of the data packet as w(t), and the data source index as idx(t).

[0016] S102: The edge device E k collects data packets from the data sources during movement and unloads the data to the cloud data center after local data processing;

[0017] S103: Establish a mobility model for the edge device:

[0018] pos k (t + 1) = pos k (t) + move k (t)(1)

[0019] In Equation (1):

[0020] pos k (t) = [x k (t), y k (t), h k (t)] represents the position of the edge device E k at the t-th time slot, x represents the horizontal distance difference, y represents the vertical distance difference, and h represents the height difference;

[0021] |move k (t)|2 ≤ r k move represents the movement in each time slot;

[0022] r k move represents the movement radius of the edge device E k ;

[0023] |·|2 represents the 2-norm of the vector;

[0024] S104: Define movement, data collection, and local execution as the core decision-making contents in the multi-agent deep reinforcement learning framework.

[0025] S2: Communication link modeling;

[0026] S201: Identify that the transmission links between agents in the three-layer system environment include source-edge, edge-edge, and edge-cloud;

[0027] ​S202: Since only the status, observations, and learning parameters are shared among edge devices without involving the transmission of data and tasks, the transmission cost of edge-edge communication is thus ignored;

[0028] S203: Model the transmission processes of source-edge and edge-cloud;

[0029] S20301: Considering the path loss of line-of-sight and non-line-of-sight, introduce the air-ground channel;

[0030]

[0031] In Equation (2):

[0032] f represents the carrier frequency;

[0033] c represents the speed of light;

[0034] represents the distance between the edge device and the ground entity;

[0035] η ξ represents the path loss in line-of-sight and non-line-of-sight cases, ξ = {0, 1};

[0036] S20302: Obtain the average air-ground path loss of the communication channel between the edge device and the data source or between the edge device and the cloud center C at the t-th time slot as:

[0037]

[0038] In Equation (3):

[0039] p1(t) represents the probability of non-line-of-sight;

[0040] represents the probability of line-of-sight; where: both a and b represent parameters related to the environment;

[0041] represents the angle between the edge-ground link and the horizontal plane.

[0042] S204: Considering the frequency division mode with a total bandwidth of W, calculate the transmission rate between the edge device and the data source or between the edge device and the cloud center C:

[0043]

[0044] In Equation (4):

[0045] b k,n (t) represents the allocated bandwidth ratio;

[0046] N0 represents the noise power spectral density;

[0047] Indicates the power for which the transmission is satisfied.

[0048] S3: Heterogeneous resource collection and AoI threshold limitation;

[0049] S301: When the edge device collects heterogeneous resource data, the edge device runs near the data source to collect all the data packets in the data source buffer and occupies a data block in the edge device's data collection buffer in the buffer.

[0050] S302: Use the local processor to preprocess the collected data buffer for preprocessing;

[0051] S303: Establish a time model for the cumulative execution of edge processing of the data packets collected by the edge device from the data source in time slot t:

[0052]

[0053] In Equation (5):

[0054] represents the data rate of the edge device for preprocessing the preset task in time slot t, which is related to the CPU cycle frequency;

[0055] S304: Assume that in each time slot, the edge device allocates its edge computing resources to a data block in the buffer. Then, the edge execution decision for the t-th time slot is represented by a one-hot vector as follows:

[0056]

[0057] In Equation (6):

[0058] exe i (t) ∈ {0, 1} represents the CPU allocation flag for each data block in the edge device's data collection buffer;

[0059] S305: Edge device local execution decision:

[0060]

[0061] S306: Cache the data in the executed data buffer and wait to be offloaded to the cloud center.

[0062] S307: Assume that in each time slot, the edge device decides to offload a data packet in the data buffer. Then, the offloading schedule is represented by a one-hot vector as follows:

[0063]

[0064] In Equation (8):

[0065] off i (t) ∈ {0, 1} represents the offloading decision for each data block;

[0066] S308: Edge device local offloading scheduling:

[0067]

[0068] S309: To quantify heterogeneity, HA-MAAC-Trans introduces weights for data sources to highlight the importance of different locations and corrects the packet size by weights:

[0069]

[0070] In Equation (10):

[0071] d(t) represents the data size of the data packet;

[0072] represents the data source weight;

[0073] S3010: Establish a data value depreciation model:

[0074]

[0075] In Equation (11):

[0076] d n,0 (t) represents the size of the initial data packet collected from the data source;

[0077] λ represents the penalty factor of the depreciation model, which is responsible for adjusting the value reduction speed after exceeding the threshold;

[0078] χ represents the AoI threshold;

[0079] AoI(t) represents the information delay of each data packet at time t;

[0080] S3011: Determine the data value based on the information delay of the data packet.

[0081] S3012: The edge device determines the subsequent travel route according to the depreciated data value.

[0082] S4: Heterogeneous age-sensitive optimization problem;

[0083] S401: Define the age of the data source at the t-th time slot as the difference Δ n (t) between the current time and the generation time of the latest data at the receiver:

[0084]

[0085] In Equation (12):

[0086] Indicates the generation time of the latest data packet of the data source received by the cloud center;

[0087] S402: Define the total number of data sources as N s , then a vector with the dimension of the total number of data sources can be used to record the age of each data source;

[0088] S403: Take the system model as a constraint condition to obtain an NP-hard optimization problem:

[0089]

[0090] C2: pos k (t + 1) = pos k (t) + move k (t)

[0091] C3: |move k (t)|2 ≤ r k move

[0092]

[0093]

[0094] S404: Discuss the solution based on MADRL.

[0095] S5: MDP formulates an external reward function and an intrinsic reward function;

[0096] S501: MDP is represented by a quadruple as follows:

[0097] M{S,A,R,P}(13)

[0098] In Equation (13)

[0099] S represents the state;

[0100] A represents the action;

[0101] R represents the reward;

[0102] P represents the transition strategy;

[0103] S502: The edge agent represents the quadruple local state M K {S K ,A K ,R K ,P K};

[0104] S503: The central agent represents the quadruple global state M C {SC , A C , R C , P C};

[0105] S504: The actions of the edge agent consist of movement, execution decision, and offloading scheduling:

[0106] {a k (t)} = {[move k (t), exe k (t), off k (t)]} (14)

[0107] S505: The action a c (t) of the central agent is to allocate the bandwidth ratio b(t);

[0108] S506: The synergistic effect of external reward and intrinsic reward.

[0109] S50601: The agents in HA-MAAC-Trans cooperate to minimize the average age of the data source as the goal;

[0110] S50602: Define the amount of data received by the cloud center as S(t), which is used to evaluate the data processing efficiency of the system at the current moment and motivate the system to process the backlogged data;

[0111] S50603: Describe the external reward of each agent at the t-th time slot as the weighted average of the data volume and the average age of the data source:

[0112]

[0113] In formula (15):

[0114] Both α and β are weight coefficients;

[0115] S50604: Calculate the intrinsic reward

[0116]

[0117] In formula (16):

[0118] P μ represents the personalized classifier, which is used to perform personalized classification on the local information of each edge agent;

[0119] P μ (k|o k ) represents the probability suitable for the edge device within the current observation range generated by the personalized classifier;

[0120] ω LISIM represents the adjustment weight;

[0121] S50605: HA-MAAC-Trans defines the global reward function, combines the long-term reward with attenuation, and studies the global optimality of the system. The long-term reward with attenuation is as follows:

[0122]

[0123] In Equation (17):

[0124] γ∈[0,1] represents the reward / punishment attenuation;

[0125] S6: Heterogeneous multi-agent Actor-Critic construction and functional analysis;

[0126] S601: Construct a heterogeneous multi-agent Actor-Critic, including a central Actor network, a central Critic network, an edge Actor network, and an edge Critic network; both the edge Actor network and the edge Critic network contain a neural network enhanced by Transformer to improve the ability to enhance spatial data features;

[0127] S602: Actor network A k (s k (t); θ k ), takes the state of each agent as input and outputs the current action a k (t);

[0128] S603: For each agent, design a Critic network C k (s k (t), a k (t); φ k ), combines the value evaluation module, and inputs the current state and action to estimate the state-action value function Q k (s k (t), a k (t));

[0129] S604: Design the edge Actor network;

[0130] S60401: Construct a multi-input / output neural network, integrate various edge observation information, and learn and process diverse operations;

[0131] S6040101: The inputs of the multi-input / output neural network include local observations from data sources, edge buffer states, and offloading channel states;

[0132] S6040102: Format the local observation data of the data source into the feature map of , and extract the edge observation information for the features of the feature map;

[0133] S60402: Build a CNN network to extract the key spatial features of the area with large data packets and high AoI;

[0134] S60403: Convert the feature map into a sequence and input it into the Transformer encoder, and use the multi-head self-attention mechanism of the Transformer network to capture long-range dependencies;

[0135] S60404: Combine the feature map with the CNN features and output the trajectory and action decisions of the device;

[0136] S60405: Use an MLP to process the execution buffer, completion buffer, and bandwidth allocation, and extract the edge status.

[0137] S605: Design a central Actor network;

[0138] The central Actor network takes the status information of the edge device as input. The Actor network outputs a one-sum vector representing the bandwidth allocation ratio, and uses an MLP to combine the multi-device status to optimize the communication scheduling between the center and the edge and allocate the bandwidth ratio for the edge-center communication;

[0139] S606: Design an edge Critic network;

[0140] The input of the edge Critic network includes the status information s k (t) of the edge device and the action policy a k (t) of the edge device output by the edge Actor network;

[0141] S60601: The edge observation information is preliminarily processed through a convolutional layer to extract local features;

[0142] S60602: Input the local features into the Transformer network to capture the temporal characteristics of the observation information, so as to enhance the understanding of the dynamic changes in the edge node environment;

[0143] S60603: Use multiple fully connected layers and ReLU activation functions to process the temporal characteristics, layer by layer to mine the complex features of the status information, so that the edge Critic network can more accurately judge the scenario in the current state and extract the edge status;

[0144] S60604: Process different types of edge actions through independent FC-ReLU branches respectively to ensure the effective extraction of different action features;

[0145] S60605: The processed action, state, and observation information are connected and integrated in the network, and finally the estimation of the action value function is output to evaluate the value of each action in the current state;

[0146] S607: Design the central Critic network;

[0147] The central Critic network inputs the edge device state information and the central bandwidth control, and obtains the estimation of the value function through the fully connected Softmax layer

[0148] S7: Intrinsic reward based on local information recognition.

[0149] S701: HA-MAAC-Trans introduces a local information self-recognition module to explore the personalized characteristics of edge devices through a self-supervised individual classifier;

[0150] S702: The local information self-recognition module, as the intrinsic reward function of HA-MAAC-Trans, aims to train a global probability classifier p with μ as the parameter μ , such that the input is the observation o k , and the output is the probability p that the observation belongs to each edge device μ (·|o k ), so as to identify an edge device from the given different observations o k for modeling the independent actions of each edge device; define the intrinsic reward p μ (k|o k ) to represent the possibility of accurately predicting / identifying a certain edge device from the observation. It can be seen that the total reward of the edge device is expressed as the combination of the external reward and the intrinsic reward:

[0151]

[0152] In Equation (18):

[0153] ∑ k p μ (k|o k ) = 1;

[0154] ω LISIM represents the adjustment weight used to measure the importance between the intrinsic reward and the extrinsic reward;

[0155] represents the extrinsic reward from the environment.

[0156] Compared with the prior art, the beneficial effects of the present invention are:

[0157] 1. The present invention extends computing power to the network edge through Mobile Edge Computing (MEC) to improve the real-time performance of data processing and reduce latency. In particular, the present invention focuses on the low-latency requirements of real-time data processing, introduces the concept of Age of Information (AoI) to measure data freshness, and strives to maintain the AoI below an acceptable threshold. In addition, in the face of the challenges of limited resources of edge devices and dynamic changes in system load, the present invention optimizes resource management and scheduling strategies to achieve efficient utilization of resources and meet dynamic data requirements. Finally, the present invention proposes a new Multi-Agent Deep Reinforcement Learning (MADRL) strategy to solve the heterogeneous AoI and resource collaboration problems in the MEC environment, ensuring data freshness and improving the collaboration efficiency between devices.

[0158] 2. The present invention proposes the HA-MAAC-Trans framework. Aiming at the heterogeneous AoI problem existing in edge devices, the data offloading is modeled as a non-convex and NP-hard optimization problem. HA-MAAC-Trans constructs a multi-agent reinforcement learning framework by integrating a heterogeneous AoI threshold optimization mechanism, realizing collaborative learning between the edge and the central model, which not only guarantees data freshness but also maximizes the cumulative reward of the system. Aiming at the problem of agent collaboration efficiency in multi-agent reinforcement learning, HA-MAAC-Trans introduces a Local Information Self-Identification Module (LISIM). The LISIM module models the cooperative behavior of heterogeneous agents by realizing the self-identification of each agent's state, thus overcoming the limitation of traditional multi-agent reinforcement learning's dependence on global information. Brief Description of the Drawings

[0159] Figure 1 is a schematic diagram of a three-layer mobile edge computing system;

[0160] Figure 2 is a schematic diagram of a reinforcement learning collaboration framework based on HA-MAAC-Trans;

[0161] Figure 3 is a schematic diagram of the neural network model of each Actor-Critic agent, where: (a) is the edge Actor network, (b) is the central Actor network, (c) is the edge Critic network, and (d) is the central Critic network;

[0162] Figure 4 is a schematic diagram of the influence of different AOI thresholds and depreciation factors on the amount of data received by the cloud center, where: (a) is the AoI threshold and (b) is the depreciation factor;

[0163] Figure 5 is a schematic diagram of the comparison of the worst AoI of different Actor-Critc networks;

[0164] Figure 6 It is a comparison schematic diagram of Actor-Critc networks designed with different network structures. Among them: (a) is a comparison schematic diagram of the average age of different network structures, (b) is a comparison schematic diagram of the amount of data received by different network structures, and (c) is a comparison schematic diagram of the received packet count of different network structures;

[0165] Figure 7 It is a comparison schematic diagram of different weighting parameters. Among them: (a) is a comparison schematic diagram of the average age of each parameter. α represents the weight of the number of tasks completed by the performance index system; β represents the weight of the total sample average age of the performance index; (b) is a comparison schematic diagram of the amount of data received by each parameter cloud center. The curve with α = 0.5 and β = 0.5 indicates that the finally received amount of data is second only to the highest value;

[0166] Figure 8 It is a comparison schematic diagram of the trajectories of edge devices obtained by different RL-based MEC cooperation algorithms. Among them: the red trajectory represents the 1st edge device, the green trajectory represents the 2nd edge device, the blue trajectory represents the 3rd edge device, and the magenta trajectory represents the 4th edge device; the data points enclosed by the dotted circle are the data points that the edge device may not be able to collect; (a) is DDPG, (b) is MADDPG, (c) is H-MAAC, and (d) is HA-MAAC-Trans;

[0167] Figure 9 It is a comparison schematic diagram of maps of 200×200 of four edge devices and RL-based MEC cooperation algorithms of 30 data sources. Among them, (a) is a schematic diagram of the change in the average age of the MEC system, (b) is a schematic diagram of the change in the worst (maximum average value) AoI in the data source, (c) is a schematic diagram of the amount of data received by the cloud center, and (d) is a schematic diagram of the count of data packets received by the cloud center. Specific implementation manners

[0168] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0169] A multi-agent cooperation method for an MEC system integrating AoI and intrinsic motivation, the method includes the following steps:

[0170] S1: Generation of end-side heterogeneous resources and edge movement decision-making;

[0171] S101: Generation of end-side device data sources;

[0172] S10101: The edge device collects the computing task data generated by intelligent devices of distributed sensors in the Internet of Things to generate a data source;

[0173] S10102: Due to differences in geographical environment and hardware configuration, the data source forms heterogeneous resource data to generate data packets in different forms;

[0174] S10103: Assign weights to each data source S n respectively to represent the importance of different locations and assume that the data sources are independently generated;

[0175] S10104: Define the data size of the data packet as d(t), the data generation time of the data packet as w(t), and the data source index as idx(t).

[0176] S102: In the MEC system, the edge device E k has mobility and computing capabilities, can collect data packets from data sources during movement, and unload the data to the cloud data center after local data processing;

[0177] S103: Establish a mobility model for the edge device:

[0178] pos k (t + 1)= pos k (t)+ move k (t)(1)

[0179] In formula (1):

[0180] pos k (t)=[x k (t), y k (t), h k (t)] represents the position of the edge device E k at the t-th time slot, x represents the horizontal distance difference, y represents the vertical distance difference, and h represents the height difference;

[0181] | move k (t)|2 ≤ r k move represents the movement in each time slot;

[0182] r k move represents the movement radius of the edge device E k ;

[0183] |·|2 represents the 2-norm of the vector;

[0184] S104: Define three main actions of the edge device - movement, data collection, and local execution - as the core decision-making content in the multi-agent deep reinforcement learning framework.​

[0185] S2: Communication link modeling;

[0186] S201: Identify that the transmission links between agents in the three-layer system environment include source-edge, edge-edge, and edge-cloud;

[0187] S202: Since only the state, observation, and learning parameters are shared between edge devices, and there is no transmission of data and tasks. Therefore, the transmission cost of edge-edge communication is ignored; this feature simplifies the communication model of the system and reduces the computational complexity.

[0188] S203: Model the transmission processes of source-edge and edge-cloud;

[0189] S20301: Considering the path loss of line-of-sight (LoS) and non-line-of-sight (NLoS), introduce the air-to-ground (A2G) channel;

[0190]

[0191] In Equation (2):

[0192] f represents the carrier frequency;

[0193] c represents the speed of light;

[0194] represents the distance between the edge device and the ground entity (a selected data source or cloud center);

[0195] η ξ represents the path loss in the line-of-sight and non-line-of-sight cases, ξ = {0, 1};

[0196] S20302: Obtain the average air-to-ground path loss of the communication channel between the edge device and the data source or between the edge device and the cloud center C in the t-th time slot as:

[0197]

[0198] In Equation (3):

[0199] p1(t) represents the probability of non-line-of-sight;

[0200] represents the probability of line-of-sight, which can approximately express the dynamic change of the path loss model; where: both a and b represent parameters related to the environment; represents the angle between the edge-ground link and the horizontal plane.

[0201] S204: Considering the frequency division mode with a total bandwidth of W, calculate the transmission rate between the edge device and the data source or between the edge device and the cloud center C:

[0202]

[0203] In Equation (4):

[0204] b k,n (t) represents the allocated bandwidth ratio;

[0205] N0 represents the noise power spectral density;

[0206] represents the power satisfied by the transmission.

[0207] This process takes into account factors such as the signal-to-noise ratio (SNR) of the channel, transmission power, and noise power spectral density, providing a theoretical basis for the communication performance evaluation of the source-edge and edge-cloud links. This modeling method effectively characterizes the characteristics of the transmission process and optimizes the communication and computing collaboration capabilities of the system.

[0208] S3: Heterogeneous resource collection and AoI threshold limitation;

[0209] S301: When the edge device collects heterogeneous resource data, the edge device runs near the data source to collect all the data packets in the data source buffer and occupies a data block in the edge device's data collection buffer (collected but unprocessed data)

[0210] S302: Use the local processor to preprocess the collected data buffer for preprocessing;

[0211] S303: Due to the determinism of the data packet formatting and edge preprocessing algorithm, establish a time model for the cumulative edge processing of the data packets collected by the edge device from the data source in time slot t through the data size and edge computing rate:

[0212]

[0213] In Equation (5):

[0214] represents the data rate at which the edge device preprocesses the preset task in time slot t and is related to the CPU cycle frequency;

[0215] S304: The unprocessed data packets are stored in the buffer of the edge device, waiting to be allocated computing resources. After processing, they are cached in the executed buffer, and finally, it is decided whether to offload them to the cloud center to ensure the reasonable allocation and processing of data blocks within each time slot; Assume that in each time slot, the edge device allocates its edge computing resources to a data block in the buffer, then the edge execution decision for the t-th time slot is represented by a one-hot vector as follows:

[0216]

[0217] In Equation (6):

[0218] exe i (t) ∈ {0, 1} represents the CPU allocation flag for each data block in the data collection buffer of the edge device, and it obviously satisfies the condition;

[0219] S305: Edge device local execution decision:

[0220]

[0221] S306: Cache the data in the executed data buffer and wait to be offloaded to the cloud center.

[0222] S307: Assume that in each time slot, the edge device decides to offload a data packet in the data buffer, then the offloading schedule is represented by a one-hot vector as follows:

[0223]

[0224] In Equation (8):

[0225] off i (t) ∈ {0, 1} represents the offloading decision for each data block, and they also satisfy the condition;

[0226] S308: Edge device local offloading schedule:

[0227]

[0228] S309: In the actual scenario, due to the heterogeneity of end-side devices, different types of devices (such as drones, cameras, Internet of Things sensors) will generate different amounts and importance of data due to differences in geographical location and computing power. This heterogeneity leads to differences in data processing time and the age of information (AoI). Therefore, different data source acquisition models are needed to describe the data acquisition process. To quantify the heterogeneity, HA-MAAC-Trans introduces weights for data sources, and the weights are affected by factors such as economic activities, population density, and traffic conditions to highlight the importance of different locations. The location weights will be affected by factors such as economic activities, traffic conditions, population density, and natural disaster risks. For example, areas with rich economic activities may be more important than areas with relatively less economic activities because they have greater demand for business and service industries. After investigation, it is found that the most important locations in a region account for 5% of the total, and the relatively important locations account for 25% of the total. And the data packet size is corrected by weights:

[0229]

[0230] In Equation (10):

[0231] d(t) represents the data size of the data packet;

[0232] represents the data source weight;

[0233] S3010: The Age of Information (AoI) of data is an important objective for system optimization. However, due to the mobility of edge devices, it is difficult to implement strict hard constraints. Therefore, HA-MAAC-Trans proposes a soft constraint mechanism. When the AoI value exceeds the set threshold, a penalty factor is introduced through the devaluation model to reduce the effective value of the data. This model dynamically adjusts the data value based on the information delay, thereby guiding the edge device to optimize the travel route and resource allocation to achieve a balance between data processing and transmission efficiency. Establish the data value devaluation model:

[0234]

[0235] In Equation (11):

[0236] d n,0 (t) represents the size of the initial data packet collected from the data source;

[0237] λ represents the penalty factor of the devaluation model, which is responsible for adjusting the value reduction speed after exceeding the threshold;

[0238] χ represents the AoI threshold;

[0239] AoI(t) represents the information delay of each data packet at time t;

[0240] S3011: Determine the data value based on the information delay of the data packet.

[0241] S3012: The edge device determines the subsequent travel route according to the devalued data value.

[0242] S4: Heterogeneous age-sensitive optimization problem;

[0243] S401: In the case of age sensitivity, the freshness of the data should be highly emphasized. Drawing on the concept of the Age of Information (AoI), in the system introduced above, the age of the data source at the t-th time slot is defined as the difference Δ n (t):

[0244]

[0245] In Equation (12):

[0246] represents the generation time of the latest data packet of the data source received by the cloud center;

[0247] Unlike the latency of each data packet, the difference Δ n (t) is a duration metric for each data source, indicating how frequently the data of the data source is collected, edge-executed, and offloaded. This is particularly important for such time-sensitive MEC systems because it directly affects the validity of the data and the timeliness of decision-making.

[0248] S402: Define the total number of data sources as N s , then a vector of the total number of data sources dimension can be used to record the age of each data source;

[0249] S403: To maintain data timeliness, the optimization goal of the system is to effectively allocate heterogeneous data resources of edge devices and minimize the average age of the data resources collected by all edge devices. Therefore, taking the system model as a constraint condition, an NP-hard optimization problem is obtained:

[0250]

[0251] C2: pos k (t + 1) = pos k (t) + move k (t)

[0252] C3: |move k (t)|2 ≤ r k move

[0253]

[0254]

[0255] S404: Before discussing the optimization problem, the present invention reexamines the system model and the problem from three dimensions. First, due to the significant hardware differences between edge devices and the optimization goal being divided into two stages, edge and cloud, the constraint conditions are heterogeneous and cannot be directly solved, requiring an iterative algorithm. Second, the optimization problem p1 is an instantaneous optimization, the environmental state is random, the strategy changes dynamically, and the agent's decision-making can be either competitive or cooperative. Finally, the optimization problem p1 is NP-hard and difficult to solve by traditional methods, so it is transformed into an MDP to explore the MADRL solution.

[0256] MADRL can dynamically handle the random states in the system and optimize the behavior decisions of each agent, thus effectively meeting the complexity and real-time requirements of the system. This method provides a feasible idea for solving the resource allocation problem in age-sensitive MEC systems.

[0257] S5: The MDP formulates an external reward function and an internal reward function;

[0258] S501: HA-MAAC-Trans realizes the interaction and cooperation of heterogeneous multi-agents by constructing a Markov decision process (MDP) model. The MDP is represented by a quadruple as follows and extended to a multi-agent version to adapt to edge devices and cloud centers in the mobile edge computing (MEC) scenario.

[0259] M{S,A,R,P}(13)

[0260] In formula (13)

[0261] S represents the state; the state of the edge agent includes local observations such as buffer state and allocated offloading bandwidth, while the state of the central agent integrates the states of all edge devices.

[0262] A represents the action;

[0263] R represents the reward;

[0264] P represents the transition policy;

[0265] S502: The edge agent represents the quadruple local state M K {S K ,A K ,R K ,P K};

[0266] S503: The central agent represents the quadruple global state M C {S C ,A C ,R C ,P C};

[0267] The HA-MAAC-Trans framework adopts a heterogeneous multi-agent Actor-Critic method to model the state input and output of the edge and the center, and realizes collaborative learning. The construction and learning process of the framework is as Figure 2 shown.

[0268] S504: The edge agent is responsible for device movement, data collection, local execution, and task offloading. The actions of the edge agent consist of movement, execution decision, and offloading scheduling:

[0269] {a k (t)} = {[move k (t),exe k (t),off k (t)]} (14)

[0270] S505: Coordinate the offloading resource allocation of edge devices to the cloud center agent by allocating the bandwidth ratio vector, that is: the action a of the central agent c(t) is the allocated bandwidth ratio b(t);

[0271] S506: HA-MAAC-Trans introduces a composite reward mechanism including external and intrinsic rewards to optimize heterogeneous AoI and resource allocation. The synergy between external rewards and intrinsic rewards.

[0272] S50601: The agents in HA-MAAC-Trans collaborate to minimize the average age of the data source as the goal to ensure the maximum freshness of the data;

[0273] S50602: Define the amount of data received by the cloud center as S(t), which is used to evaluate the data processing efficiency of the system at the current moment, and motivate the system to process the backlogged data quickly and effectively;

[0274] S50603: The external reward comprehensively considers the average age of the data source and the amount of data received by the cloud center, and realizes the balance between data freshness and system performance through a weighting strategy. The external reward of each agent at the t-th time slot is described as the weighting of the data volume and the average age of the data source:

[0275]

[0276] In formula (15):

[0277] Both α and β are weight coefficients. HA-MAAC-Trans can optimize the overall performance and resource utilization of the system while maintaining data freshness by designing the weights.

[0278] S50604: Calculate the intrinsic reward

[0279]

[0280] In formula (16):

[0281] P μ represents the personalized classifier, which is used to perform personalized classification on the local information of each edge agent;

[0282] P μ (k|o k ) represents the probability suitable for the edge device within the current observation range generated by the personalized classifier;

[0283] ω LISIM represents the adjustment weight;

[0284] S50605: HA-MAAC-Trans defines the global reward function, combines the long-term reward with decay, and studies the global optimality of the system. The long-term reward with decay is as follows:

[0285]

[0286] In formula (17):

[0287] γ ∈ [0, 1] represents the reward / punishment decay;

[0288] This multi-agent reinforcement learning framework not only optimizes the collaboration efficiency between the edge and the center, but also improves the overall performance and stability of the system through the composite reward mechanism.

[0289] S6: Construction and functional analysis of heterogeneous multi-agent Actor-Critic;

[0290] S601: Construct a heterogeneous multi-agent Actor-Critic, including a central Actor network, a central Critic network, an edge Actor network, and an edge Critic network; both the edge Actor network and the edge Critic network include a neural network enhanced by Transformer to improve the ability to enhance (abstract and model) spatial data features;

[0291] S602: Actor network A k (s k (t); θ k ), taking the state of each agent as input and outputting the current action a k (t); since the input states and output actions of each agent are multi-dimensional, the Actor network needs to be designed according to the specific data structure, as shown in Figure 3 (a) and Figure 3 (b).

[0292] S603: For each agent, design a Critic network C k (s k (t), a k (t); φ k ), combined with the value evaluation module, input the current state and action to estimate the state-action value function Q k (s k (t), a k (t)); as shown in Figure 3 (c) and Figure 3 (d).

[0293] S604: Design the edge Actor network;

[0294] S60401: Construct a multi-input and multi-output neural network, integrate various edge observation information, and learn and process diverse operations;

[0295] S6040101: The inputs of the multi-input and multi-output neural network include local observations from data sources, edge buffer status, offloading channel status, and other information;

[0296] S6040102: Format the local observation data of the data source into feature map, and extract edge observation information for the features of the feature map;

[0297] S60402: Construct a CNN network to extract key spatial features of areas with large data packets and high AoI;

[0298] S60403: Convert the feature map into a sequence input to the Transformer encoder, and use the multi-head self-attention mechanism of the Transformer network to capture long-range dependencies;

[0299] S60404: Combine the feature map with CNN features and output the trajectory and action decisions of the device;

[0300] S60405: Use an MLP to process information such as the execution buffer, completion buffer, and bandwidth allocation, and extract the edge status.

[0301] The CNN extracts key feature maps through convolutional and pooling layers. To enhance the adaptability of the model, the present invention introduces a Transformer network, uses its multi-head self-attention mechanism to capture global context information, and enhances feature expression through a feed-forward neural network. The Transformer effectively models long-range dependencies and is crucial for key information coordination and resource optimization in edge computing scenarios. The present invention serializes the CNN feature map and inputs it into the Transformer encoder (intermediate layer dimension 2, number of heads 2), and after processing, adjusts it back to the image shape and splices it with the original feature map to generate a more representative feature map. Project the observation map onto the mobile map to determine the trajectory action. On this basis, the extraction of the edge status is mainly achieved through an MLP (multi-layer perceptron). Input information such as the execution buffer, completion buffer, and allocated bandwidth is formatted into lists and scalar inputs, and the output is the execution and offloading scheduling decisions of the edge device. Finally, the concatenated feature map is further processed through a fully connected layer, and the observation map is projected onto the mobile map to determine the trajectory and action of the edge device, as shown in Figure 3 (a).

[0302] S605: Design a central Actor network;

[0303] The central Actor network takes the status information of edge devices as input. The status information includes the execution and completion data resource buffers of each edge device and its location in the environment. Based on these inputs, the Actor network outputs a one-sum vector representing the bandwidth allocation ratio, indicating the bandwidth allocation ratio for each edge device. Using an MLP to combine multi-device status, optimize the communication scheduling between the center and the edge, and allocate the bandwidth ratio for edge-center communication, as Figure 3 (b) shows;

[0304] S606: Design the edge Critic network;

[0305] The input of the edge Critic network includes the status information s k (t) (location information, buffer status, allocated offloading bandwidth) and the action policy a k (t) (movement, execution, and offloading) of the edge devices output by the edge Actor network;

[0306] S60601: The edge observation information is preliminarily processed through a convolutional layer (including convolution and average pooling) to extract local features;

[0307] S60602: Input the local features into a Transformer network to capture the temporal characteristics of the observation information, thereby enhancing the understanding of the dynamic changes in the edge node environment;

[0308] S60603: Use multiple fully connected layers and ReLU activation functions to process the temporal characteristics, layer by layer mining the complex features of the status information, enabling the edge Critic network to more accurately judge the scenario in the current state and extract the edge status;

[0309] S60604: Process different types of edge actions (movement actions, execution actions, and offloading actions) through independent FC-ReLU branches respectively to ensure the effective extraction of different action features;

[0310] S60605: The processed actions, status, and observation information are connected and integrated in the network, and finally the estimation of the action value function is output to evaluate the value of each action in the current state, as Figure 3 (c) shows;

[0311] S607: Design the central Critic network;

[0312] The central Critic network inputs the edge device status information (the execution / completion buffer and location of each edge device) and the central bandwidth control, and obtains the estimation of the value function through a fully connected Softmax layer

[0313] The central Critic network is used to calculate and evaluate the state value of the entire system, providing value evaluations for different actions to the central Actor, thereby helping the central Actor network to be more inclined to select high-value actions and optimize the decision-making effect. By integrating the state information, locations of each edge node, and related actions controlled by the center, it uniformly evaluates the value of the entire system in the current state. The design of the central Critic network is relatively lightweight, only using fully connected layers to process the location information of centralized control, thus achieving global estimation while ensuring a relatively low computational complexity. This design enables the central Critic network to effectively integrate information from each edge node, providing accurate value feedback to the central Actor network, and further promoting the coordination of the system at the global level. Through this global value evaluation, the system can ensure the consistency of the overall goal (such as minimizing the average age of information), as Figure 3 (d) shows.

[0314] S7: Intrinsic reward based on local information recognition.

[0315] S701: Due to the uneven distribution of data sources within the task area and each edge device can only observe its own local information, this leads to limitations in the cooperation efficiency among agents in the MADRL framework. The lack of local information makes it difficult for agents to comprehensively understand the global state, thus affecting the optimization of global decisions and further reducing the effect of multi-agent cooperation. Therefore, considering the continuous movement of devices in the MEC environment to achieve a reasonable division of labor in space and avoid resource waste caused by multiple devices arriving at the same area. In addition, considering that the task needs to handle the problems of heterogeneous data generation and the management of the age of information (AoI), HA-MAAC-Trans introduces a local information self-identification module (LISIM), which explores the personalized features of edge devices through a self-supervised individual classifier. Through sufficient individual exploration, devices can more effectively complete data collection tasks, thereby improving the overall performance when facing heterogeneous data and information delay problems;

[0316] S702: The local information self-identification module, as the intrinsic reward function of HA-MAAC-Trans, aims to train a global probability classifier p with parameter μ μ , such that the input is the observation o k , and the output is the probability p μ (·|o k ) that the observation belongs to each edge device, so as to accurately identify an edge device from the given different observations o k for modeling the independent actions of each edge device; define the intrinsic reward p μ (k|o k) represents the possibility of accurately predicting / identifying a certain edge device from observations. Thus, the total reward of the edge device is expressed as a combination of external rewards and intrinsic rewards:

[0317]

[0318] In Equation (18):

[0319] ∑ k p μ (k|o k ) = 1;

[0320] ω LISIM represents the adjustment weight, which is used to measure the importance between intrinsic rewards and extrinsic rewards;

[0321] represents the extrinsic reward from the environment;

[0322] The intrinsic reward helps the edge device achieve individual recognition through partial environmental observations, promotes division of labor, enhances policy differences, and improves classifier training. LISIM improves the cumulative external reward, strengthens the exploration enthusiasm, optimizes the agent task performance in data collection by discovering new data sources and trajectory patterns, and balances exploration and task completion through personalized recognition ability, thereby improving system efficiency.

[0323] Example 1

[0324] The present invention implements the HA-MAAC-Trans learning framework, and constructs a hybrid network of CNN-MLPs of an edge Transformer-enhanced Actor network and an MLP-based network as the central network through TensorFlow. To compare the performance of the proposed framework with other RL methods, the present invention selects two Actor-Critic-based algorithms, DDPG (centralized) and the popular MADDPG (multi-agent), which also use double neural networks to learn the hybrid strategy. In addition, an algorithm similar to joint collaboration through local information sharing, such as EdgeFed H-MAAC, is used as a baseline. For fairness, the present invention sets the same random seed for the MEC environment for all methods, and also fixes the random seed of the training process. Therefore, the results are general and easy to reproduce.

[0325] Figure 9 And Table 3 shows the comparison results of the proposed algorithm and the baseline on a 200×200 map with randomly initialized edge device and data source positions in an environment of 4 edge devices and 30 data sources.

[0326] Table 3 Comparison of AoI metric data

[0327]

[0328] Average age during online interaction As Figure 9 shown in (a), it can be found that Hybrid DDPG results in the highest, and the curve does not tend to be stable until 5K epochs. However, under Hybrid MADDPG, EdgeFed H-MAAC, and HA-MAAC-Trans, after sufficient iterations, the average age remains at a relatively low value. Specifically, the statistics are listed in Table 3. HA-MAAC-Trans not only achieves the lowest average age but also the lowest variance, which means that the HA-MAAC-Trans method is superior to DDPG and MADDPG in terms of both system penalty and learning stability. In addition, it is superior to the ordinary EdgeFed H-MAAC method in terms of convergence speed. This reflects the superiority of the Transformer and the intrinsic incentive mechanism of LISIM in terms of the delay of the MEC system. The right part of Table 3 is the comparison of the peak AoI. Obviously, the proposed HA-MAAC-Trans algorithm obtains the lowest peak PAoI and the lowest average This means that HA-MAAC-Trans can greatly reduce the peak value of AoI through the AoI threshold and improve the efficiency of data processing in MEC. The minimum variance of the peak also indicates that the updates of all data sources are frequent and fair. Figure 9 (b) shows the worst AoI (the AoI of the sample point with the largest average AoI) among the four methods. HA-MAAC-Trans also performs the best. It can be found that under centralized DDPG, some data sources are ignored for a long time. The explanation for this is that centralized collaborative algorithms require large neural network models with complex structures to extract the relationships between excessive global input states and the local policies of each individual agent, which also leads to difficulties in training.

[0329] In addition, Figure 9 (c) and Figure 9 (d) give the number and size of the aggregated data packets received at the cloud center (i.e., the last hop of the MEC system). It can be seen from the figure that the HA-MAAC-Trans collaborative algorithm improves the data utility of the MEC system by completing more data processing in the same time.

[0330] Embodiment 2

[0331] In S5, the present invention introduces a composite reward mechanism to simultaneously optimize multiple key performance indicators of the system, including the number of tasks completed by the system and the total sample average age.

[0332] In this experiment, first, based on different configurations of α and β weight coefficients, the proportion of different performance metrics in the reward mechanism was changed, and the average age of the system at the 5000th time slot was experimentally evaluated, as Figure 7 shown in (a).

[0333] From the box plot data, the medians of α = 0.1, β = 0.9 and α = 0.2, β = 0.8 in the experimental group are the lowest, at 11.5, indicating that the system can maintain a low average age, but the outliers are high, suggesting that there may be a situation of slow information update, affecting performance stability. The medians of another group α = 0.5, β = 0.5 are similar, only 0.5 higher, but the interquartile range is the lowest, the 75th percentile is 22, and the 25th percentile is 6, showing a high data concentration, small performance fluctuations, and the ability to maintain a stable low average age level in various situations, and low outliers, avoiding a significant performance decline even in extreme cases. Therefore, this group of configurations performs best in terms of average age performance.

[0334] Figure 7 (b) shows the change trend of the cumulative data volume received by the cloud center during the time slots from 4995 to 5000 under different weight coefficient configurations of α and β. The cumulative data volume received is a key indicator to measure the data transmission performance of the system. A higher received packet count means that the system can transmit data more efficiently and maintain better throughput. By analyzing the curves under different configurations, the optimal configuration under this experimental condition is found.

[0335] The figure shows that the received packet counts of all weight combinations increase linearly as the time slots progress. The two curves of α = 1, β = 0 and α = 0, β = 1 are located at the highest and lowest positions respectively, proving that α and β can affect the change of the reward. α = 1, β = 0 indicates that the reward only focuses on the number of tasks completed, so the data volume received by the cloud center is obviously the largest. And α = 0, β = 1 only focuses on the average age of the samples, and it is expected that the data volume received by the cloud center is the smallest. It is worth noting that the sub-optimal configuration groups are α = 0.6, β = 0.4, α = 0.3, β = 0.7 and α = 0.5, β = 0.5. They are second only to α = 1, β = 0 where the reward mechanism is completely for the number of tasks completed, and the data volume received by the cloud center is at a relatively high level. The cloud center can receive data packets at a higher rate, and the data transmission efficiency and throughput of the system are the highest.

[0336] From the results of the above two experiments, it can be seen that the most noteworthy set of parameter weights is α = 0.5 and β = 0.5. This set of weights not only performs optimally in terms of the average age performance but also achieves a level second only to α = 1 and β = 0 in terms of the performance metric of the data volume received at the cloud center. Thus, when the parameter weights are configured as α = 0.5 and β = 0.5, the MEC system not only focuses on the number of tasks completed to achieve high-speed data reception but also pays equal attention to the average age of the samples, ensuring that the system can quickly respond to data changes and avoid the accumulation of information delays.

[0337] In summary, the experimental results show that the parameter weights of α = 0.5 and β = 0.5 can achieve the optimal optimization of the system performance. Compared with the configurations of α = 1, β = 0 and α = 0, β = 1 that only focus on a single performance metric, the weight allocation that comprehensively considers α and β can more effectively improve the overall performance of the system.

[0338] Example 3

[0339] MEC cooperation algorithms such as DDPG and MADDPG often lead to similar access areas because edge devices start from the same location, resulting in resource waste and insufficient access to data sources. These algorithms do not consider device differences and have convergent behaviors, making it impossible to effectively cooperate to access more data sources. LISIM processes individual differences, avoids centralized access, encourages devices to explore diverse areas, reduces the AoI, and increases the data collection volume. The present invention selects several common cooperation algorithms and verifies the superiority of the LISIM algorithm by comparing the movement trajectories of edge devices on the data source graph under different algorithms, as Figure 8 shown.

[0340] It can be seen from the experimental results that sharing the same neural network parameters or using a centralized network is not a good choice (DDPG, MADDPG). The number of round trips of the paths obtained by the two algorithms is too large, and they are limited to obtaining rewards only from neighboring data sources. The H-MAAC algorithm can alleviate the path repetition problem to a certain extent by introducing the ε-exploration method. Edge devices can actively explore unvisited data sources and avoid path similarity between devices by learning the parameters of other devices. However, since H-MAAC does not consider the internal and external characteristics of the rewards between devices, there is still a problem of path repetition. In contrast, the HA-MAAC-Trans algorithm proposed by the present invention introduces the LISIM mechanism, effectively differentiates the influence of physically adjacent homogeneous neighbors on the task path, and avoids the problem of path repetition. As Figure 8As shown, the path task coverage using LISIM is close to 100%, and the trajectory of each agent is circular, successfully solving the problem of spatial division of labor. For the path without using LISIM (such as H-MAAC), there are some unexplored areas because the data source distribution is uneven and the observable range of each edge device is limited. Spatial division of labor must be achieved through continuous movement.

[0341] LISIM introduces personality into edge devices, conducts intrinsic incentives, enables them to realize their responsibilities of moving to remote areas, improves the coverage of task data sources, and ultimately improves efficiency.

[0342] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and the present invention can be implemented in other forms without departing from the spirit or basic characteristics of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent conditions of the claims are intended to be included in the present invention. Any reference signs in the claims should not be regarded as limiting the claims involved.

[0343] In addition, it should be understood that although this specification is described according to embodiments, not every embodiment only contains an independent technical solution. This narrative way of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.

Claims

1. A multi-agent cooperation method for MEC systems that integrates AoI and intrinsic motivation, characterized by: The method includes the following steps: S1: Edge-side heterogeneous resource generation and edge movement decision-making; The S1 includes the following steps: S101: Edge-side device data source generation; The S101 includes the following steps: S10101: The edge-side device collects the computing task data generated by intelligent devices of distributed sensors in the Internet of Things to generate a data source; S10102: Due to differences in geographical environment and hardware configuration, the data source forms heterogeneous resource data to generate data packets in different forms; S10103: Assign weights to each data source S n respectively to represent the importance of different positions and assume that the data sources are generated independently; S10104: Define the data size of the data packet as d(t), the data generation time of the data packet as w(t), and the data source index as idx(t); S102: Edge device E k Collects data packets from data sources while in motion, processes the data locally, and unloads the data to the cloud data center; S103: Establish a mobility model for edge devices: pos k (t + 1)= pos k (t)+ move k (t)(1) In Equation (1): pos k (t) = [x k (t), y k (t), h k (t)] represents the position of the edge device E k at the t-th time slot, where x represents the horizontal distance difference, y represents the vertical distance difference, and h represents the height difference; |move k (t)|2≤r k move represents the movement in each time slot; r k move represents the moving radius of the edge device E k ; |·|2 represents the 2-norm of a vector; S104: Define movement, data collection, and local execution as the core decision-making contents in the multi-agent deep reinforcement learning framework; S2: Communication link modeling; The S2 includes the following steps: S201: Clarify that the transmission links between agents in the three-layer system environment include source-edge, edge-edge, and edge-cloud; S202: Since only states, observations, and learning parameters are shared between edge devices, and data and tasks are not involved in transmission, the transmission cost of edge-edge communication is ignored; S203: Model the transmission processes of source-edge and edge-cloud; The S203 includes the following steps: S20301: Considering line-of-sight and non-line-of-sight path losses, introduce an air-ground channel; In Equation (2): f represents the carrier frequency; c represents the speed of light; Indicated as the distance between the edge device and the ground entity; η ξ represents the path loss in line-of-sight and non-line-of-sight scenarios, ξ = {0, 1}; S20302: Obtain the average air-ground path loss of the communication channel between the edge device and the data source or between the edge device and the cloud center C in the t-th time slot as: In Equation (3): p1(t) represents the probability of non-line-of-sight; represents the probability of the line-of-sight distance; where: both a and b represent parameters related to the environment; represents the angle between the edge link and the horizontal plane; S204: Considering a frequency division mode with a total bandwidth of W, calculate the transmission rate between the edge device and the data source or between the edge device and the cloud center C: In Equation (4): b k,n (t) represents the allocated bandwidth ratio; N0 represents the noise power spectral density; Indicates the power at which transmission is satisfied; S3: Heterogeneous resource collection and AoI threshold limitation; The S3 includes the following steps: S301: When the edge device collects heterogeneous resource data, the edge device runs near the data source to collect all the data packets in the data source buffer and occupies a data block in the edge device data collection buffer. in it S302: Preprocess the collected data buffer using the local processor ; S303: Establish a time model for the cumulative edge processing time of data packets collected by edge devices from data sources in time slot t: In Equation (5): Indicates the data rate at which the edge device preprocesses the preset task in time slot t and is related to the CPU cycle frequency; S304: Assume that in each time slot, the edge device allocates its edge computing resources to a data block in the buffer, and the edge execution decision in the t-th time slot is represented by a one-hot vector as follows: In Equation (6): exe i (t) ∈ {0, 1} represents the CPU allocation flag for each data block in the data collection buffer of the edge device; S305: Edge device local execution decision; S306: Cache the data in the executed data buffer and wait to be offloaded to the cloud center; S307: Assume that in each time slot, the edge device decides to offload a data packet in the data buffer, and the offloading schedule is represented by a one-hot vector as follows: In Equation (8): off i (t) ∈ {0, 1} represents the offloading decision for each data block; S308: Edge device local offloading schedule; S309: To quantify heterogeneity, HA-MAAC-Trans introduces weights for data sources to highlight the importance of different locations and modifies the data packet size through weights: In Equation (10): d(t) represents the data size of the data packet; Indicates the data source weight; S3010: Establish a data value depreciation model: In Equation (11): d n,0 (t) represents the size of the initial data packet collected from the data source; λ represents the penalty factor of the depreciation model, which is responsible for regulating the value reduction speed after exceeding the threshold; χ represents the AoI threshold; AoI(t) represents the information delay of each data packet at time t; S3011: Determine the data value based on the information delay of the data packet; S3012: The edge device determines the subsequent travel route according to the depreciated data value; S4: Heterogeneous age-sensitive optimization problem; The said S4 includes the following steps: S401: Define the age of the data source at the t-th time slot as the difference Δ between the current time and the generation time of the latest data at the receiver n (t): In formula (12): Indicates the generation time of the latest data packet of the data source received by the cloud center; S402: Define the total number of data sources as N s , and then a vector with the dimension of the total number of data sources can be used to record the age of each data source; S403: Take the system model as a constraint condition to obtain an NP-hard optimization problem: S404: Discuss the solution based on MADRL; S5: MDP formulates an external reward function and an intrinsic reward function; The said S5 includes the following steps: S501: MDP is represented by a quadruple as follows: M{S,A,R,P} (13) In formula (13) S represents the state; A represents the action; R represents the reward; P represents the transfer strategy; S502: Edge agent represents the local state M of the quadruple K {S K , A K , R K , P K}; S503: The central agent represents the four - tuple global state M C {S C , A C , R C , P C}; S504: The actions of the edge agent consist of movement, execution decision-making, and offloading scheduling; {a k (t)} = {[move k (t), exe k (t), off k (t)]} (14) S505: Action a of the central agent c (t) is the allocated bandwidth ratio b(t); S506: The synergistic effect of external reward and intrinsic reward; S50601: Agent Collaboration in HA-MAAC-Trans to Minimize the Average Age of Data Sources as the goal; S50602: Define the amount of data received by the cloud center as S(t), which is used to evaluate the data processing efficiency of the system at the current moment and encourage the system to process the backlogged data; S50603: Describe the external reward of each agent at the $t$-th time slot as a weighted combination of the data volume and the average age of the data source: In formula (15): Both α and β are weight coefficients; S50604: Calculate Intrinsic Rewards In formula (16): P μ represents a personalized classifier for performing personalized classification on the local information of each edge intelligent agent; P μ (k|o k ) represents the probability suitable for the edge device within the current observation range generated by the personalized classifier; ω LISIM represents an adjustment weight; S50605: HA-MAAC-Trans defines a global reward function, combines the long-term reward with attenuation, and studies the global optimality of the system. The long-term reward with attenuation is as follows: In formula (17): γ∈[0,1] represents the reward / punishment attenuation; S6: Construction and functional analysis of heterogeneous multi-agent Actor-Critic; The said S6 includes the following steps: S601: Construct a heterogeneous multi-agent Actor-Critic, including a central Actor network, a central Critic network, an edge Actor network, and an edge Critic network; both the edge Actor network and the edge Critic network contain a Transformer-enhanced neural network to improve the ability to enhance spatial data features; S602: Actor network A k (s k (t); θ k ), taking the states of each agent as input and outputting the current action a k (t); S603: For each agent, design a Critic network C based on the Actor network framework k (s k (t), a k (t); φ k ), combined with the value evaluation module, input the current state and action to estimate the state-action value function Q k (s k (t), a k (t)); S604: Design the edge Actor network; S60401: Construct a multi-input-output neural network, integrate various edge observation information, and learn and process diverse operations; S6040101: The inputs of the multi-input-output neural network include local observations from data sources, edge buffer status, and offloading channel status; S6040102: Format the local observation data of the data source into feature mapping graph, and extract edge observation information for the features of the feature mapping graph; S60402: Construct a CNN network to extract key spatial features of areas with large data packets and high AoI; S60403: Convert the feature map into a sequence and input it into the Transformer encoder, and use the multi-head self-attention mechanism of the Transformer network to capture long-range dependencies; S60404: Combine the feature map with the CNN features and output the trajectory and action decisions of the device; S60405: Use MLP to process the execution buffer, completion buffer, and bandwidth allocation, and extract the edge status; S605: Design the central Actor network; The central Actor network takes the status information of edge devices as input. The Actor network outputs a one-sum vector representing the bandwidth allocation ratio. It uses an MLP to combine multi-device status, optimize the communication scheduling between the center and the edge, and allocate the bandwidth ratio for edge-center communication; S606: Design the edge Critic network; The input of the Edge Critic network includes the state information s k (t) of the edge device and the action policy a k (t) output by the Edge Actor network; S60601: The edge observation information is preliminarily processed through a convolutional layer to extract local features; S60602: The local features are input into the Transformer network to capture the temporal characteristics of the observation information, thereby enhancing the understanding of the dynamic changes in the edge node environment; S60603: Use multiple fully connected layers and ReLU activation functions to process the temporal characteristics, layer by layer exploring the complex features of the status information, enabling the edge Critic network to more accurately judge the scenario in the current state and extract the edge state; S60604: Process different types of edge actions separately through independent FC-ReLU branches to ensure the effective extraction of different action features; S60605: The processed action, state, and observation information are connected and integrated in the network, and finally the estimation of the action value function is output to evaluate the value of each action in the current state; S607: Design the central Critic network; The central Critic network inputs the edge device status information and central bandwidth control, and obtains the estimation of the value function through a fully connected Softmax layer S7: Intrinsic reward based on local information recognition; The said S7 includes the following steps: S701: HA-MAAC-Trans introduces a local information self-recognition module to explore the personalized features of edge devices through a self-supervised individual classifier; S702: The local information self-identification module serves as the intrinsic reward function of HA-MAAC-Trans. The goal is to train a global probability classifier p parameterized by μ μ , such that for an input observation o k , the output is the probability p μ (·|o k ) that the observation belongs to each edge device. Thus, an edge device can be identified from the given different observations o k for modeling the independent actions of each edge device; Define the intrinsic reward p μ (k|o k ) to represent the likelihood of accurately predicting / identifying a certain edge device from the observation. It can be seen that the total reward of the edge device is expressed as a combination of the external reward and the intrinsic reward: In formula (18): ∑ k p μ (k|o k ) = 1; ω LISIM represents the adjustment weight, which is used to measure the importance between intrinsic rewards and extrinsic rewards; Represents an external reward from the environment.

Citation Information

Patent Citations

  • Unmanned group perception method for guaranteeing element universe application information quality

    CN117221829A

  • Satellite-ground collaborative edge network resource allocation method based on depth deterministic strategy gradient

    CN119109504A