Digital twin data synchronization optimization method fusing multi-agent and heterogeneous graph reinforcement learning

By integrating multi-agent and heterogeneous graph reinforcement learning, a Markov decision model is constructed to optimize data synchronization and processing decisions in cloud-edge-device collaborative networks. This solves the problem of correlation between discrete and continuous variables, reduces data synchronization deviation and processing latency, and improves the timeliness of digital twins and the business processing capabilities of collaborative networks.

CN121486985APending Publication Date: 2026-02-06FUZHOU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511620226.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-06
Publication Date
2026-02-06

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively handle the correlation between discrete and continuous variables in cloud-edge-device collaborative networks, resulting in high data synchronization deviations and processing latency. They are unable to adapt to dynamic changes and heterogeneous characteristics, thus affecting the timeliness of digital twin virtual entities.

Method used

A Markov decision model is constructed by integrating multi-agent and heterogeneous graph reinforcement learning. A data synchronization optimization algorithm based on multi-agent hybrid decision-making and heterogeneous graph reinforcement learning is designed to optimize data synchronization and processing decisions between devices and servers. Joint optimization is performed by combining discrete and continuous variables.

Benefits of technology

It significantly reduces the average data synchronization deviation and packet processing latency between the device and the server, improving the timeliness of data processing and the decision-making capabilities of the collaborative network.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121486985A_ABST
    Figure CN121486985A_ABST
Patent Text Reader

Abstract

The invention provides a digital twin data synchronization optimization method fusing multi-agent and heterogeneous graph reinforcement learning, and the method comprises the following steps: firstly, constructing a system model related to data synchronization and synchronous data processing, and formulating an optimization target; and secondly, designing a data synchronization optimization algorithm based on multi-agent mixed decision, and generating decision action auxiliary equipment of a mixed action space to perform reliable data synchronization. And then, the successfully synchronized data packet is distributed to a target server for data processing by a heterogeneous graph reinforcement learning data processing decision algorithm, and the processed data is forwarded to a server for deploying a digital twin model to update a virtual entity. Finally, through experiments, the average data synchronization deviation between the equipment end and the server can be remarkably reduced in the data synchronization stage, and the average data packet processing time delay can be remarkably reduced in the data processing stage.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of digital twin data synchronization optimization, and particularly relates to a digital twin data synchronization optimization method fusing multi-agent and heterogeneous graph reinforcement learning. BACKGROUND

[0002] Cloud-edge-terminal collaborative network: In real life, there are all kinds of businesses. In order to process these businesses, it is difficult to meet the processing needs of some businesses by relying on a single level (edge, cloud server, terminal). In particular, some businesses need to be coordinated. In order to meet the needs of most business processing, the cloud-edge-terminal collaborative network, as a new architecture, integrates the resources of cloud servers, edge servers and terminal devices, promotes the close cooperation between the cloud layer, the edge layer and the terminal device layer, realizes business collaborative processing, and injects new vitality into numerous application scenarios.

[0003] Digital twin: The key to digital twin is to create a virtual twin environment consistent with the real physical environment, which can reflect the state of the physical environment, realize the monitoring and real-time optimization of the whole life cycle, and then provide the corresponding environment state information or perform some simulation optimization for other businesses. The cloud-edge-terminal collaborative network uses the related characteristics of digital twin to realize more accurate and intelligent collaborative strategy formulation.

[0004] Non-orthogonal multiple access: Non-orthogonal multiple access (NOMA) is considered to be a very promising multiple device access scheme, which can meet the needs of multiple devices or users sharing frequency domain and time domain resources in the power domain, and using different powers for data transmission. The base station uses sequential interference cancellation technology (SIC) to decode the received superimposed power domain device signal in order from high to low according to the received power, which meets the device data synchronization scenario.

[0005] Deep reinforcement learning: Reinforcement learning learns experience from environmental interaction, realizes better decision performance, and deep reinforcement learning, as an important branch of reinforcement learning, gives reinforcement learning the ability to fit functions, and expands reinforcement learning to more complex business scenarios. Therefore, in the scenario of the present application, deep reinforcement learning provides a feasible solution for data synchronization and synchronization data processing decision-making, which not only enables the strategy to have environmental adaptability, but also guarantees the timeliness of the decision.

[0006] Graph neural network: The graph neural network (GNN) takes the adjacency matrix composed of the features of the nodes and the connection relationship between the nodes as input, and transmits the feature information contained in the neighbor nodes to the target node through the message passing mechanism. The node processed by the GNN not only contains its own feature information, but also aggregates the features of the neighbor nodes, providing more complete node embedding for downstream tasks.

[0007] The data generating device needs to select a suitable channel, base station and transmission power for data synchronization to reduce data synchronization deviation. The optimization variables include both discrete and continuous parameters, and optimizing these parameters belongs to a mixed integer nonlinear programming problem, which is proved to be NP-Hard. The existing schemes face the following challenges: first, the traditional global optimization scheme faces the challenge of decision timeliness; second, deep reinforcement learning has become a common solution to such problems, but existing schemes mainly handle discrete variables or continuous variables separately, ignoring the correlation between discrete and continuous variables. Therefore, in scenarios involving both discrete and continuous variables, a scheme that jointly optimizes discrete and continuous variables is needed to achieve better results.

[0008] Efficient use of resources of each server to reduce data processing delay is an important part of ensuring the timeliness of digital twin virtual entities. However, servers will face the following challenges in making data processing decisions: first, how to adapt to dynamically changing data packet quantities and server states; second, how to use the heterogeneous characteristics of data packets and servers to improve data processing decisions. In summary, designing a data processing decision scheme that can adapt to dynamic changes in the environment and integrate heterogeneous features is key to reducing data processing delay. SUMMARY

[0009] Therefore, the purpose of the present application is to provide a digital twin data synchronization optimization method that combines multi-agent and heterogeneous graph reinforcement learning, to ensure that the digital twin of the cloud-edge-end collaborative network can accurately reflect the state information of the physical environment in a timely manner, further improving the collaborative network decision-making and business processing capabilities.

[0010] To achieve the above purpose, the present application adopts the following technical scheme: a digital twin data synchronization optimization method that combines multi-agent and heterogeneous graph reinforcement learning, comprising the following steps:

[0011] Step S1: Construct a system model for data synchronization and synchronous data processing of cloud-edge-end collaborative network and digital twin;

[0012] Step S2: Define the goal of reducing the average data synchronization deviation between the device and the server in the data synchronization stage, and the goal of reducing the average data packet processing delay in the synchronous data processing stage;

[0013] Step S3: Construct the data synchronization optimization decision process as a Markov decision model;

[0014] Step S4: Design a data synchronization optimization algorithm based on multi-agent hybrid decision;

[0015] Step S5: Construct a heterogeneous graph structure and Markov decision model related to synchronous data processing;

[0016] Step S6: design a data processing decision algorithm based on heterogeneous graph reinforcement learning.

[0017] In a preferred embodiment, the system model of data synchronization and synchronous data processing in step S1 is constructed as follows:

[0018] Step 11: system model

[0019] Assuming that the system contains environmental perception devices, denoted as These devices will perform environmental perception and service generation, and will synchronize data with the base station through one of channels during their active time slots; the system also includes servers, denoted as where 1 to denotes an edge server and deploys a local digital twin (DT) environment mapping, a cloud server deploys a global DT environment mapping; in addition, the edge server serves as a supporting facility for the base station, and the edge server is called a base station when receiving device-synchronized data; this system considers optimizing data synchronization and data processing under discrete time slots, which are divided into equal-length time slots, including data synchronization and synchronous data processing;

[0020] Step 12: device model

[0021] For device , the state of generating a new data packet at time slot is denoted by , the active state is denoted by , and the data synchronization state is denoted by If the data packet fails to synchronize successfully, the device will continue to try until it succeeds.

[0022] Step 13: NOMA transmission model

[0023] Assuming that device is preparing to synchronize data to base station through channel at time slot , the signal-to-interference-and-noise ratio (SINR) is denoted as:

[0024]

[0025] where denotes the channel gain, is the transmission power of the device; due to the shared channel, the interference experienced by the device includes not only environmental noise Also includes the intra-cluster interference generated by the undecoded device signals transmitted to the same base station using the same channel And the inter-cluster interference generated by the device signals using the same channel for data synchronization to other base stations

[0026] The base station uses SIC technology to decode the device signals in order from high to low according to the received signal strength; the data packet of the device synchronization is small; if the previous signal occurs error in the SIC decoding process, the data of the device will be affected and fail to decode; and in the case that the previous SIC decoding process does not occur error, the influence of the data packet error probability needs to be considered, wherein the device The error probability of the data transmitted by the device in the t time slot is:

[0027]

[0028] Wherein, is the error probability of data packet transmission, is the data synchronization time slot size, is the channel bandwidth, is the data packet size transmitted by the device in the t time slot, is the Gaussian Q function; therefore, the data synchronization state of the device in the t time slot is defined as:

[0029]

[0030] Step 14: Data synchronization model

[0031] AoI measures the freshness and timeliness of data, and the AoI of the device at the device end is expressed as:

[0032]

[0033] The AoI of the device at the server end is related to whether the data packet is successfully received by the base station, and the AoI of the device at the server end is expressed as:

[0034]

[0035] The data synchronization deviation between the device at the server end and the device end is defined as:

[0036]

[0037] Step 15: Synchronization data processing model​

[0038] The data packet server that successfully synchronizes with each time slot device formulates a data processing strategy, processes and forwards to other servers to update the DT model; the device The server set of data is , wherein represents the server whether the data of the device is needed;

[0039] The server computing resources available for data processing in the time slot is represented as , and the link bandwidth set for data migration and forwarding is represented as :

[0040]

[0041] wherein represents the uplink bandwidth from the server to the server ; the base station is deployed in the same area, the communication distance is short, the cloud server is deployed in the remote end, and the propagation delay is ;

[0042] In the time slot , the device set successfully received by the edge server is ; if the data of the device is directly processed in the server , the migration delay is 0; then the time delay of the data synchronized by the device in the time slot is:

[0043]

[0044] wherein is the device , represents the device set planned by the server to migrate to the server for data processing, and the data of these devices will be integrated and migrated, represents the propagation delay if the device data needs to be transmitted to the cloud server , and the N+1th server is the cloud server;

[0045] Assuming that the device successfully synchronizes with the server in the time slot with a size of The number of CPU cycles required to process each bit of the data is ; device The time delay generated by processing the data in time slot is:

[0046]

[0047] wherein, represents whether the data of the device is processed by the server , represents the proportion of computing power allocated by the server to the device ; is the total computing power of the server , represents the number of CPU cycles required to process the data of the device .

[0048] Assuming that the data of the device is processed by the server , i.e. , the size of the processed data is , and the average time delay generated by forwarding the data of the device in time slot is:

[0049]

[0050] wherein represents the proportion of transmission bandwidth from the server to the server ;

[0051] The time delay of the data of the device when processing the data in time slot is represented as:

[0052]

[0053] In a preferred embodiment, the optimization target definition in step S2 is as follows:

[0054] Step 21: Data synchronization optimization target definition: in the data synchronization stage, the target is to reduce the average data synchronization deviation:

[0055]

[0056] represents the transmission power of device m in time slot t when transmitting data to base station n through thought c,​ representing the maximum transmission power of the device;

[0057] Step 22: Synchronization data processing optimization target definition, in the data processing stage, the target is to reduce the average data packet processing delay:

[0058] .

[0059] In a preferred embodiment: the construction of data synchronization optimization decision process in step S3 is a Markov decision model, the steps are as follows:

[0060] Step 31: Define the state space of the device data synchronization MDP model, including channel gain set, data size, and AoI of data.

[0061] Step 32: Define the action space of the device data synchronization MDP model, the action includes discrete and continuous two parts.

[0062] Step 33: Define the reward function of the device data synchronization MDP model as the negative number of data synchronization deviation.

[0063] In a preferred embodiment: the step S4 in the design of a data synchronization optimization algorithm based on multi-agent hybrid decision, the steps are as follows:

[0064] Step 41: Construct the device agent model, for example, the agent model of device includes the policy network and target policy network of continuous action space , D3QN network and target D3QN network of discrete action space

[0065] ;

[0066] Step 42: After the device gets continuous and discrete actions from the agent model, it will synchronize data according to the strategy given by the agent;

[0067] Step 43: Define the training mechanism of the agent model, apply the idea of maximum entropy reinforcement learning to update the policy network, D3QN network and entropy regularization coefficient in the way of reducing loss.

[0068] Step 51: data processing heterogeneous graph structure construction, device data packets and servers are graph nodes, and the connection relationship between them is an edge. Among them, the relationship "in" is used to represent that the data packet is successfully received by a certain edge server; the relationship "need" is used to represent that the server needs some device data for digital twin model updating; and the link between servers is represented by "link".

[0069] Step 52: define the state space of the synchronous data processing MDP model, and the state of data processing in each time slot is represented by a heterogeneous graph structure.

[0070] Step 53: define the action space of the synchronous data processing MDP model, including a discrete action space (target server for processing data) and a continuous action space (allocation weight of computing power and bandwidth).

[0071] Step 54: define the reward function of the synchronous data processing MDP model.

[0072] In a preferred embodiment: the data processing decision algorithm based on heterogeneous graph reinforcement learning in step S6 is designed, and the steps are as follows:

[0073] Step 61: graph structure feature extraction, feature aggregation according to different relationships, to obtain node embedding of servers and data packets;

[0074] Step 62: data processing decision model definition, after being processed by three graph neural network feature aggregation operators, the node embedding of device data packets and servers is obtained, and the data processing decision is generated by the reinforcement learning model according to the node embedding;

[0075] Step 63: loss function definition of the data processing decision model, the graph state of each time slot, the data processing decision of the device data packet and the reward information are integrated into interaction experience, and the model parameters are updated through experience replay.

[0076] Compared with the prior art, the present application has the following beneficial effects: the present application can significantly reduce the average data synchronization deviation between the device end and the server in the data synchronization stage, and can significantly reduce the average data packet processing delay in the data processing stage. BRIEF DESCRIPTION OF DRAWINGS

[0077] Figure 1 is a flowchart of a preferred embodiment of the present application.

[0078] Figure 2 is a schematic diagram of the overall system model in the preferred embodiment of the present application.

[0079] Figure 3 is a schematic diagram of the heterogeneous graph structure in the preferred embodiment of the present application.

[0080] Figure 4 is a schematic diagram of a data synchronization optimization algorithm based on multi-agent hybrid decision in a preferred embodiment of the application.

[0081] Figure 5 is a schematic diagram of a data processing decision algorithm based on heterogeneous graph reinforcement learning in a preferred embodiment of the application. DETAILED DESCRIPTION

[0082] The application will be further described below with reference to the accompanying drawings and embodiments.

[0083] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present application. Unless otherwise indicated, all technical and scientific terms used herein have the same meaning as would be commonly understood by one of ordinary skill in the art to which this application belongs.

[0084] It should be noted that the terms used herein are only for the purpose of describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application; as used herein, the singular form is intended to include the plural form unless the context clearly indicates otherwise, and furthermore, it should be understood that when the terms "comprise" and / or "include" are used in the specification, there is a presence of a feature, step, operation, device, component and / or combinations thereof.

[0085] The present embodiment provides a digital twin data synchronization optimization method that fuses multi-agent and heterogeneous graph reinforcement learning, as shown in Figures 1-5 , the detailed steps are as follows:

[0086] Step 1: Construct a system model for cloud-edge-end collaborative network and digital twin data synchronization and synchronous data processing, the overall system model diagram is shown in Figure 2 The specific steps are as follows:

[0087] Step 11: Construct the system model, which contains environmental perception devices, denoted as These devices will perform environmental perception and business generation, and in their active time slots, they will synchronize data to the base station through one of the channels. servers, denoted as , where 1 to denotes an edge server, and a local digital twin (DT) environment map is deployed, a global DT environment map is deployed on the cloud server. In addition, the edge server is a supporting facility for the base station, and the edge server is called a base station when receiving device synchronized data. This system considers optimizing data synchronization and data processing under discrete time slots, and the time is divided into one equal-length time slot, containing both data synchronization and synchronization data processing;

[0088] Step 12: Constructing the device model, the device in the time slot The state of generating a new data packet is denoted by , the active state is denoted by , and the data synchronization state is denoted by If the data packet fails to be successfully synchronized, the device will continue to try until success.

[0089] Step 13: Constructing the NOMA transmission model, assuming that the device in the time slot is ready to synchronize data to the base station through the channel , which experiences a signal-to-interference-and-noise ratio (SINR) denoted by

[0090]

[0091] where denotes the channel gain, and is the transmission power of the device. Due to the shared channel, the interference signal experienced by the device not only contains environmental noise but also intra-cluster interference generated by undecoded device signals using the same channel to transmit to the same base station:

[0092]

[0093] and inter-cluster interference generated by device signals using the same channel to synchronize data to other base stations:

[0094]

[0095] The base station uses SIC technology to decode in order from high to low according to the received signal strength. If the previous signal is wrong in the SIC decoding process, the data of the device will be affected and fail to be decoded; in the case where the previous SIC decoding process is correct, the influence of the data packet error probability needs to be considered, in which the error probability of the device transmitting data in time slots is:

[0096]

[0097] where is the data synchronization time slot size, is the channel bandwidth, is the Gaussian Q function, denotes the device in Packet size of time slot transmission; therefore, the device In time slot Data synchronization state Is defined as:

[0098]

[0099] Step 14: Build a data synchronization model, AoI measures the freshness and timeliness of data. The device The AoI at the device end is expressed as:

[0100]

[0101] The device The AoI at the server end is expressed as:

[0102]

[0103] Time slot The device The data synchronization deviation between the server end and the device end is expressed as:

[0104]

[0105] Step 15: Build a server synchronization data processing model, after the base station receives the data synchronized by the environment perception device, the server will formulate a corresponding data processing strategy, and after processing, it will be forwarded to other servers to update the DT model. For example, the device The server set of data is , where Indicates whether the server needs the data of the device . In addition, considering that the server resources mainly serve collaborative services, the resources available for data processing are relatively small and change dynamically with time slots. For any server In time slot The computing resources available for data processing are expressed as The link bandwidth set available for data migration and forwarding is expressed as:

[0106]

[0107] Where Indicates the uplink bandwidth from server to server . The base station in the scenario is deployed in the same area, and the communication distance between edge servers is short, so the propagation delay can be ignored. The cloud server is deployed in a remote place, and the propagation delay is .

[0108] In time slot , the edge server , the set of devices that successfully received the data . These device data can not only be left in the server for processing, but also be migrated to other servers for processing. If the data of the device is processed directly in the server , the migration delay is 0. For the data synchronized by the device in the time slot , the time delay generated by migration is:

[0109]

[0110] wherein is the device , represents the set of devices whose data is planned to be migrated to the server for data processing by the server , and the data of these devices will be integrated and migrated, represents the propagation delay if the device data needs to be transmitted to the cloud server (the N+1 server is the cloud server) ;

[0111] If the device successfully synchronizes data of size to the server in the time slot , the number of CPU cycles required to process each bit of the device data is . The time delay generated by processing the data synchronized by the device in the time slot is:

[0112]

[0113] wherein, represents whether the data of the device is processed by the server , represents the proportion of computing power allocated by the server to the device , and is the total computing power (total CPU cycles) of the server , represents the number of CPU cycles required to process the data of the device .

[0114] The average time delay generated by forwarding the data of the device in the time slot is:

[0115]

[0116] wherein, is the device processed data size. represents the transmission bandwidth ratio of forwarding the processed data to the server .

[0117] In summary, the data of the device is processed in the time slot , and the data processing delay is represented as:

[0118]

[0119] Step 2: Define the target of reducing the average data synchronization deviation between the device and the server in the data synchronization stage, and the target of reducing the average data packet processing delay in the synchronization data processing stage. The specific steps are as follows:

[0120] Step 21: Construct the data synchronization optimization target. To ensure that device data can be synchronized to the server in time, the target of reducing the average data synchronization deviation is set in the data synchronization stage:

[0121]

[0122] Step 22: Construct the synchronization data processing optimization target. To ensure that the synchronized data can be efficiently processed and the digital twin model can be updated in time, the target of reducing the average data packet processing delay is set in the data processing stage:

[0123]

[0124] Step 3: Construct the data synchronization optimization decision process as a Markov decision model. The specific steps are as follows:

[0125] Step 31: Define the state space of the device data synchronization MDP model. The state of the active device in the time slot is represented as:

[0126]

[0127] wherein is the set of channel gains, is the data size, is the data synchronization deviation.

[0128] Step 32: Define the action space of the device data synchronization MDP model. The action includes discrete and continuous parts. The continuous action is the transmission power, and the discrete action includes target base stations, channels, and backoff actions. For example, the device In time slot The data synchronization action is , where is a discrete action, is a continuous action. When , it means that the device selects backoff; when , the device will transmit data with transmission power through the channel to synchronize to the base station .

[0129] Step 33: Define the reward function of the device data synchronization MDP model, which is the negative of the data synchronization deviation for the device after performing the policy action in time slot :

[0130]

[0131] Step 4: Design a data synchronization optimization algorithm based on multi-agent hybrid decision-making, the specific steps are as follows:

[0132] Step 41: Construct the device agent model, as shown in Figure 3 , the agent model of any device contains a policy network and a target policy network in the continuous action space; a D3QN network and a target D3QN network in the discrete action space.

[0133] Step 42: Obtain the data synchronization continuous and discrete actions from the agent model, input the state to the policy network to obtain the continuous action distribution, sample the transmission power set from the distribution:

[0134]

[0135] Concatenate the state and the transmission power set to input the D3QN network, and select the discrete action through the greedy strategy:

[0136]

[0137] Where is the value of the discrete action , which is obtained through the state value function and the advantage function The calculation is as follows:

[0138]

[0139] When the device agent gives the exact continuous and discrete actions, the device will synchronize data according to the strategy given by the agent.

[0140] Step 43: Define the agent model training mechanism. After the device agent interacts with the environment, it will form interaction experience. The device agent will perform experience replay in the training stage to obtain the experience For example, the loss calculation of the D3QN network is as follows:

[0141]

[0142] wherein represents the target value.

[0143] If the device does not synchronize data packets in the next time slot, the future reward is expected to be 0, consistent with Therefore, the target value is represented as:

[0144]

[0145] wherein is the reward discount, is a continuous action sampled from the distribution , and the discrete action is selected by the D3QN network :

[0146]

[0147] Avoiding the problem of overestimation of Q values of the agent due to the use of outdated target discrete actions.

[0148] The update of the policy network aims to maximize the sum of the evaluation value of the D3QN network on the policy network and the entropy of the policy, wherein the entropy of the policy measures the randomness of the distribution given by the policy network, and the greater the entropy, the more random the policy. Therefore, the loss calculation of the policy network is as follows:

[0149]

[0150] wherein is the entropy regularization coefficient, used to control the proportion of exploration and utilization; is a set of continuous actions reparameterized from the policy network.

[0151] The entropy regularization coefficient The adjustment of the target entropy is related to the exploration and exploitation ratio of the policy network. In actual operation, The adjustment of the target entropy is realized by reducing the loss, and the calculation process is:

[0152]

[0153] wherein is the target entropy, is the entropy of the policy.

[0154] The two target networks update the network parameters through soft update, and the corresponding calculation process is as follows:

[0155]

[0156]

[0157] wherein and are the soft update coefficients.

[0158] Step 5: Construct a heterogeneous graph structure and a Markov decision model related to synchronous data processing, and the specific steps are as follows:

[0159] Step 51: Graph structure data construction, as shown in Figure 4 Device data packets and servers will be graph nodes, and the connection relationship between them is an edge. Among them, the successful reception of a device data packet by a certain edge server is represented by the relationship "in"; the need of a server for some device data packets for digital twin model update is represented by the relationship "need"; and the link connection between servers is represented by "link".

[0160] At time slot , the set of devices that successfully synchronize data packets is represented as If device , device synchronize at time slot The features contained in the synchronized data packet can be represented as:

[0161]

[0162] The feature representation of the server is:

[0163]

[0164] All the connection edges in the graph structure data are:

[0165]

[0166] wherein, represents the edge set of the relationship "link". Similarly, the set with set respectively represent the edge set of the relationship "in" and the relationship "need".

[0167] Step 52: Define the state space of the data processing MDP model. The state of data processing in each time slot is represented by a graph structure. Then the state of time slot t is represented as:

[0168]

[0169] Step 53: Define the action space of the data processing MDP model. For the data packet of device i, the processing decision is , which means that the synchronized data packet of device i will be processed by server j, where . The proportion of the computing power allocated by the server j can be calculated as:

[0170]

[0171] where wj is the weight of computing power allocation. The processed data is forwarded from the server j to any server k, and the proportion of bandwidth is:

[0172]

[0173] Step 54: Define the reward function of the data processing MDP model. For the data packet of device i, the reward feedback information it obtains can be represented as:

[0174]

[0175] Step 6: Design a data processing decision algorithm based on heterogeneous graph reinforcement learning. The overall algorithm structure is shown in Figure 5 , and the specific steps are as follows:

[0176] Step 61: Graph structure feature extraction. Among the three relationships, the relationships with the server node as the target are "link" and "in". For any server node i, the source node and the target node of the relationship "link" have the same feature dimension, and the feature aggregation process is represented as:

[0177]

[0178] where, , ​​​​​​​​​​​are weight and bias matrix respectively, is an activation function. is a server to the server link feature, and . In addition, is the degree of the server node .

[0179] Since the nodes connected by the relationship "in" are heterogeneous nodes, two different dimensions of weights , the mean features of the server node itself and its adjacent device packet nodes are dimensionally transformed, and the bias is added to obtain the node embedding of the server node under the relationship "in". The feature aggregation process of the server node is:

[0180]

[0181] Similarly, the target node of the relationship "need" is the packet node, so the feature aggregation process of the device packet node is represented as:

[0182]

[0183] After obtaining the node embedding of each node, the feature aggregation will be performed according to the category of the target node. Then the server node is obtained by summing the aggregated node embedding:

[0184]

[0185] Step 62: data processing decision process, after processing by three graph neural network feature aggregation operators, the node embedding of device packets and servers will be obtained, and the node embedding will be integrated. For the device packet, the feature input to the Actor-Critic network is represented as:

[0186]

[0187] After completing the feature aggregation, the feature set will be input to the Actor network to obtain the continuous action set of all packets. After splicing with the feature set, it obtains , which is input to the Critic network to obtain the size of Q value matrix, each row of which corresponds to the value of each data packet in data processing of each server, and each data packet will be processed by the server corresponding to the maximum Q value. Then the processing device The selection process of the server of the data packet can be expressed as:

[0188]

[0189] When the continuous action and the discrete action are determined, the computing power and bandwidth weight of the corresponding server are selected from the continuous action set, for example, The first element in the set is the computing power weight of the device for the data packet, and the bandwidth weight of the forwarded data can also be obtained from the set .

[0190] Step 63: Loss function definition of the data processing decision model, the graph state, decision and reward of each time slot are integrated into interactive experience, and the model is optimized through experience replay. For the experience , the loss calculation of the Critic network and its feature extraction module through the experience is expressed as:

[0191]

[0192] The Actor network is a random policy network, and the parameter update is performed through the idea of maximum entropy reinforcement learning, and the entropy is combined with the evaluation of the Critic network. The loss calculation of the Actor network and its feature extraction module can be expressed as:

[0193]

[0194] Wherein, is the entropy regularization coefficient, is the continuous action reparameterized from the output of the Actor network.

[0195] The entropy regularization coefficient is adjusted according to the target entropy , and the corresponding loss calculation can be expressed as:

[0196]

Claims

1. A digital twin data synchronization optimization method integrating multi-agent and heterogeneous graph reinforcement learning, characterized in that, Includes the following steps: Step S1: Construct a system model for data synchronization and synchronous data processing between cloud-edge-device collaborative networks and digital twins; Step S2: Define the goal of reducing the average data synchronization deviation between the device and the server during the data synchronization phase, and the goal of reducing the average data packet processing latency during the data synchronization processing phase. Step S3: Construct a Markov decision model for the data synchronization optimization decision process; Step S4: Design a data synchronization optimization algorithm based on multi-agent hybrid decision-making; Step S5: Construct the heterogeneous graph structure and Markov decision model related to synchronous data processing; Step S6: Design a data processing decision algorithm based on heterogeneous graph reinforcement learning.

2. The digital twin data synchronization optimization method integrating multi-agent and heterogeneous graph reinforcement learning according to claim 1, characterized in that: The system model construction for data synchronization and synchronized data processing in step S1 is as follows: Step 11: System Model Assume the system contains An environmental sensing device, represented as These devices will perform environmental awareness and business generation, and will transmit during their active time slots. One of the channels synchronizes data with the base station; the system also includes... One server, represented as , of which 1 to This is represented as an edge server, and a localized digital twin (DT) environment mapping is deployed. Deploy a global DT environment mapping for the cloud server; in addition, the edge server serves as a supporting facility for the base station. When receiving data synchronized by receiving devices, it is called a base station. This system considers optimizing data synchronization and data processing under discrete time slots, with the time divided into... Each time slot is of equal length and includes two parts: data synchronization and synchronous data processing. Step 12: Equipment Model For equipment In the time slot The state used to generate new data packets Indicates the active state using Indicates the data synchronization status using This means that if data packets fail to synchronize successfully, the device will continue to try until it succeeds; Step 13: NOMA transmission model Assuming the device In the time slot Preparing to pass through the channel Synchronize data to base station The signal-to-interference-plus-noise ratio (SINR) is expressed as: in Indicates channel gain. This refers to the device's transmission power; due to the shared channel, the interference experienced by the device includes not only environmental noise. It also includes intra-cluster interference caused by undecoded device signals transmitted to the same base station using the same channel. Inter-cluster interference caused by equipment signals that use the same channel to synchronize data with other base stations. ; The base station uses SIC (Signal Indicator Capability) technology to decode device signals in descending order of received signal strength; the data packets synchronized by the device are relatively small; if an error occurs in the preceding signal during SIC decoding, the device... The data will be affected and decoding will fail; however, if no errors occur in the preceding SIC decoding process, the impact of data packet error probability must be considered, including the device... exist The error probability of data transmission in time slots is: in, The probability of an error occurring during data packet transmission. For the size of the data synchronization time slot, For channel bandwidth, For equipment The size of the data packet transmitted in time slot t. It is a Gaussian Q-function; therefore, the device In the time slot Data synchronization status Defined as: Step 14: Data Synchronization Model AoI measures the freshness and timeliness of data, and devices The AoI on the device side is represented as: The device's AoI on the server side is related to whether the data packet is successfully received by the base station. The AoI representation on the server side is as follows: equipment The data synchronization deviation between the server and the device is defined as: Step 15: Synchronizing the data processing model For data packets successfully synchronized by devices in each time slot, the server will formulate a data processing strategy, process them, and forward them to other servers to update the DT model; devices The data server set is ,in Indicates server Is equipment needed? Data; server In the time slot Computational resources available for data processing are represented as The set of link bandwidth used for data migration and forwarding Represented as: in Indicates from server To the server The uplink bandwidth; base stations are deployed in the same area with a short communication distance, while cloud servers are deployed at a remote location with a propagation delay of [missing information]. ; In the time slot Edge server The set of devices that successfully received data is If the equipment The data is directly on the server The process is performed during migration, with a migration delay of 0; therefore, the device... In the time slot Latency caused by data migration during synchronization for: in For equipment , Indicates server Plan to migrate to server A collection of devices that perform data processing; the data from these devices will be integrated and migrated. This means that if device data needs to be transmitted to the cloud server, a propagation delay must also be added. The (N+1)th server is a cloud server; Assuming the device In the time slot Successfully synchronized size to the server The number of CPU cycles required to process each bit of this data is... ;equipment In the time slot The latency caused by the processing of synchronized data is: in, Indicates equipment Is the data from the server? deal with, Indicates server For equipment The proportion of computing power allocated; For server Total computing power Indicates processing equipment The number of CPU cycles required to process the data. Assuming the device The data is from the server Processing, i.e. The size of the processed data is In the time slot repeater Average latency generated by the data for: in Indicates data from server Forwarded to server The proportion of transmission bandwidth; equipment Data in time slots Data processing latency is expressed as:

3. The digital twin data synchronization optimization method integrating multi-agent and heterogeneous graph reinforcement learning according to claim 1, characterized in that: The optimization objective in step S2 is defined as follows: Step 21: Data Synchronization Optimization Goal Definition: The goal of the data synchronization phase is to reduce the average data synchronization deviation. This represents the transmission power of device m transmitting data to base station n via data transfer c in time slot t. Indicates the device's maximum transmission power; Step 22: Define the optimization target for synchronous data processing. The goal during the data processing phase is to reduce the average packet processing latency. 。 4. The digital twin data synchronization optimization method integrating multi-agent and heterogeneous graph reinforcement learning according to claim 1, characterized in that: The data synchronization optimization decision-making process in step S3 is a Markov decision model, and the steps are as follows: Step 31: Define the state space of the device data synchronization MDP model, including the channel gain set, data size, and data AoI. Step 32: Define the action space of the device data synchronization MDP model. The actions include both discrete and continuous parts. Step 33: Define the reward function of the device data synchronization MDP model as the negative of the data synchronization deviation.

5. The digital twin data synchronization optimization method integrating multi-agent and heterogeneous graph reinforcement learning according to claim 1, characterized in that: The step S4 involves designing a data synchronization optimization algorithm based on multi-agent hybrid decision-making, and the steps are as follows: Step 41: Construct a device intelligent agent model, such as the device The agent model includes a policy network with a continuous action space. and target policy network ; D3QN network in discrete action space and the target D3QN network ; Step 42: After the device obtains continuous and discrete actions from the agent model, it will synchronize data according to the strategy given by the agent. Step 43: Define the training mechanism for the agent model, and apply the idea of ​​maximum entropy reinforcement learning to update the policy network, D3QN network and entropy regularization coefficients in a way that reduces loss.

6. The digital twin data synchronization optimization method integrating multi-agent and heterogeneous graph reinforcement learning according to claim 1, characterized in that: The steps in step S5, which involve constructing the heterogeneous graph structure and Markov decision model related to synchronous data processing, are as follows: Step 51: Data Processing Heterogeneous Graph Structure Construction. Device data packets and servers are graph nodes, and the connections between them are edges. The relationship "in" represents a data packet being successfully received by an edge server; the relationship "need" represents a server needing certain device data for digital twin model updates; and the link connection between servers is represented by "link". Step 52: Define the state space of the synchronous data processing MDP model. The state of data processing in each time slot is represented by a heterogeneous graph structure. Step 53: Define the action space of the synchronous data processing MDP model, including the discrete action space (the target server for processing data) and the continuous action space (the allocation weight of computing power and bandwidth). Step 54: Define the reward function for the synchronous data processing MDP model.

7. The digital twin data synchronization optimization method integrating multi-agent and heterogeneous graph reinforcement learning according to claim 1, characterized in that: The step S6 involves designing a data processing decision algorithm based on heterogeneous graph reinforcement learning, and the steps are as follows: Step 61: Graph structure feature extraction, feature aggregation according to different relationships, to obtain the node embeddings of the server and data packets; Step 62: Define the data processing decision model. After processing by three graph neural network feature aggregation operators, the result is... Individual device data packets and The node embedding of each server is used to generate data processing decisions based on the node embedding of the reinforcement learning model. Step 63: Define the loss function of the data processing decision model. The graph status of each time slot, the data processing decision of the device data packet, and the reward information will be integrated into interactive experience. The model parameters will be updated through experience playback.