A cloud-edge-end architecture and a sense-through algorithm resource scheduling method for vehicle networking digital twin synchronization
By designing a cloud-edge-device architecture and a resource scheduling method based on a hybrid multi-agent model in the Internet of Vehicles (IoV), the collaborative decision-making between vehicle terminals and edge servers is optimized. This solves the problems of limited computing resources and insufficient synchronization timeliness in digital twin synchronization in IoV, achieving low latency, low energy consumption, and high precision synchronization.
Patent Information
- Application Number
- CN202511349483.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-22
- Publication Date
- 2025-12-16
- Estimated Expiration
- 2045-09-22
AI Technical Summary
In the Internet of Vehicles (IoV), existing technologies suffer from limited computing resources, insufficient synchronization timeliness and reliability during digital twin synchronization, especially when edge network resources are limited, making it difficult to achieve low latency, low energy consumption and high precision synchronization.
Design a cloud-edge-device architecture for vehicle-to-everything (V2X) communication, including a terminal layer, an edge layer, and a cloud layer. By using a hybrid multi-agent model and a resource scheduling method based on hybrid deep reinforcement learning, optimize the collaborative decision-making between the vehicle terminal and the edge server, and achieve efficient processing of perception data and high-precision construction and updating of digital twin models.
It significantly reduces the computing cost of the terminal, improves the synchronization reliability and accuracy of the digital twin model, solves the conflict problem in resource scheduling, and achieves lower synchronization latency and energy consumption.
Smart Images

Figure CN120856762B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application relates to the technical field of communication, in particular to a cloud-edge-terminal architecture for digital twin synchronization in vehicle networking and a sensing-computing resource scheduling method. BACKGROUND
[0002] Digital twin (DT) is a virtual model of a physical entity in digital space. In the vehicle networking environment, a regional digital twin model can be constructed to simulate traffic flow changes in real time, predict congestion trends, and assist traffic management. Maintaining low latency, low energy consumption, and high precision digital twin synchronization in vehicle networking requires efficient processing of multi-modal sensing data from vehicle terminals. However, the above process usually requires a large amount of overhead. Deploying a digital twin in vehicle networking in the cloud center and allowing vehicle terminals to offload sensing data to edge servers deployed in roadside units can improve response speed and synchronization efficiency to a certain extent. However, due to the heterogeneity of vehicle devices and the dynamic changes of the on-site environment, it is still a significant challenge to ensure synchronization timeliness and reliability under the condition of limited edge network resources.
[0003] In addition, the current application of digital twin in vehicle networking is still limited to the end and edge, and the computing power of the cloud has not been fully utilized. Due to the dynamic nature of the cloud-edge-terminal system in vehicle networking, an architecture scheme that integrates digital twin and cloud-edge-terminal system needs to be designed. At the same time, a large amount of sensing data will be generated during the running process of vehicle terminals in vehicle networking, but the computing and storage capacity of the terminals themselves is limited. Therefore, relying on an efficient cloud-edge-terminal architecture to complete large-scale computing-intensive and delay-sensitive sensing tasks in vehicle networking has become an urgent problem to be solved. SUMMARY
[0004] The purpose of the application is to provide a cloud-edge-terminal architecture for digital twin synchronization in vehicle networking and a sensing-computing resource scheduling method. The cloud-edge-terminal architecture can reduce the computing cost of the terminal and maintain high-precision digital twin model construction and update maintenance in the cloud. A hybrid multi-agent model based on AC architecture is used to improve the accuracy of the digital twin model, reduce synchronization latency and energy consumption.
[0005] Technical solution: A cloud-edge-end architecture for vehicle networking digital twin synchronization, including a terminal layer, an edge layer and a cloud layer, the terminal layer includes a vehicle terminal cluster based on geographical area division, each vehicle terminal collects road environment information in real time through a vehicle-mounted sensor, including traffic flow, pedestrian state and real-time road conditions; and transmits the compressed environment data to the edge server through a wireless link; the edge layer is composed of distributed edge servers deployed in roadside units, after receiving the perception data sent from the terminal layer, the edge server makes autonomous cooperation decision according to the current computing load; the cloud layer includes a central cloud server, which performs real-time synchronization update of the digital twin model after receiving the edge processing data for digital twin construction from the edge layer in the next time slot; wherein the digital twin model is constructed and maintained based on the edge processing data for digital twin construction uploaded by the edge layer.
[0006] Further, when the single-node computing load exceeds the threshold, redundant data is sent to a low-load edge server for cooperative processing, and the low-load edge server uploads the edge processing data for digital twin construction to the cloud after processing the environment data.
[0007] A sensing and computing resource scheduling method for vehicle networking digital twin synchronization, based on any of the above cloud-edge-end architectures, the sensing and computing resource scheduling is performed according to real-time business requirements, including the following steps:
[0008] S1, constructing a cloud-edge-end architecture for vehicle networking digital twin synchronization;
[0009] In the cloud-edge-end architecture, the terminal layer has V0 vehicle terminals, which are distributed in M0 regions, and the set of vehicle terminals is represented as ; each region is configured with an edge server deployed in a roadside unit, and the set of edge servers is represented as ; the total time length of the cloud-edge-end architecture for digital twin synchronization is divided into T0 time slots with a length of , and the set is ;
[0010] S2, the vehicle terminal sensing device perceives and collects the environment data around the vehicle terminal, and constructs a cooperative perception model;
[0011] S3, according to the communication process of the vehicle terminal and the edge server, and the edge server and the edge server, a communication model is established;
[0012] S4, according to the process of the vehicle terminal and the edge server processing perception data, a computing model is established;
[0013] S5, construct an optimization objective function according to the demand of digital twin synchronization, and convert the optimization problem into a partially observable Markov decision problem through time slot division, and construct a Markov model;
[0014] S6, based on the Markov model, a hybrid multi-agent model based on hybrid deep reinforcement learning is established to solve the optimal resource scheduling strategy.
[0015] Further, the vehicle terminal surrounding environment data includes traffic flow, pedestrian and real-time road condition information; the specific steps of constructing the cooperative perception model are as follows:
[0016] S21, determine the perception data amount of each vehicle terminal;
[0017] According to the perception frequency of the vehicle terminal, the perception data amount of each vehicle terminal is determined, and the perception data is integrated in the cloud layer after being processed by the vehicle terminal and the edge server, and the construction of the digital twin model is completed; after the vehicle terminal collects the perception data, the perception data is compressed and encoded to obtain effective data, assuming that the vehicle terminal The perception frequency of the vehicle terminal at time t is , the amount of environmental data perceived each time is d0, and the compression encoding rate of the data is , then the perception data amount is:
[0018] ,
[0019] S22, calculate the perception energy consumption of the vehicle terminal v in time slot t , the expression is:
[0020] ,
[0021] wherein, is the energy consumption of the vehicle terminal v for one perception.
[0022] Further, the specific steps of establishing the communication model are as follows:
[0023] S31, calculate the transmission rate of the vehicle terminal v, the wireless transmission delay of the vehicle terminal unloading the perception data to the edge server m, and the wireless transmission energy consumption of the vehicle terminal v at time t;
[0024] If the vehicle terminal v communicates with the edge server m by using OFDMA, the transmission rate is:
[0025] ,
[0026] wherein, is the uplink channel bandwidth of the vehicle terminal, , is the channel power gain of the vehicle terminal v, denotes the channel power gain of the vehicle terminal , denotes the noise power, is the transmission power of the vehicle terminal v, is the transmission power of the vehicle terminal ;
[0027] The vehicle terminal autonomously decides to offload part of the data to the edge server associated with the area where the vehicle terminal is located according to the computing capability, and the offloading ratio is The wireless transmission delay of the vehicle terminal when offloading the perception data to the edge server m is
[0028] ,
[0029] The wireless transmission energy consumption of the vehicle terminal v at time slot t is
[0030] ,
[0031] wherein, denotes the amount of perception data of the vehicle terminal v;
[0032] S32, after the edge server m receives the perception data, it autonomously makes a cooperation decision according to the current computing load;
[0033] The cooperation decision is as follows: when and , it means that the edge server m transmits the perception data from the vehicle terminal v to other edge servers i for processing through a wired link; when and i=m, it means that the perception data is directly processed by the edge server m; wherein, ;
[0034] If the edge server m and other edge servers i are connected by optical fibers, the optical fiber transmission delay is:
[0035] ,
[0036] wherein, hops (m, i) denotes the shortest path hop count between edge servers, v0 denotes the transmission rate of a unit of data under one hop; when the perception data is directly processed by the edge server m associated with the vehicle terminal v, there is no information transmission between edge servers, which is denoted as hops (m, m) = 0.
[0037] Further, the specific steps of establishing the total perception quality calculation model of each vehicle terminal are as follows:
[0038] S41, the vehicle terminal calculates the time delay of partial perception data;
[0039] The vehicle terminal processes the time delay of partial perception data For:
[0040] ,
[0041] Wherein, is the number of CPU cycles required to process unit perception data, is the calculation frequency of the vehicle terminal v;
[0042] Local calculation energy consumption For:
[0043] ,
[0044] Wherein, is the energy consumption coefficient of the calculation chip;
[0045] S42, calculate the waiting time delay of the perception data of the vehicle terminal v;
[0046] Suppose there are some perception data of vehicle terminals entering the cache queue before the vehicle terminal v, and the set of these vehicle terminals is In the queue, the perception data that arrives first is processed first, and the perception data that arrives later needs to wait until the perception data that arrives first is processed before being processed. The waiting time delay of the perception data of the vehicle terminal v For:
[0047] ,
[0048] Wherein, represents the queue state of the edge server m at the edge of the time slot t; represents the calculation frequency of the edge server m; represents the amount of perception data of the vehicle terminal ; the offloading ratio of the partial data to the edge server associated with the area where the vehicle terminal is located is ;
[0049] The perception data in the queue of the edge server m that has not been processed in the next time slot is represented as:
[0050] ,
[0051] The perception data of the vehicle terminal v is processed by the edge server after queuing and waiting, and the calculation time delay of the edge server is expressed as:
[0052] ,
[0053] wherein, is the total data amount calculated at the edge server m in the time slot t;
[0054] S43, the total delay of the vehicle terminal and the edge terminal in cooperation with the calculation of the current time slot sensing data is counted;
[0055] The delay of the vehicle terminal v in the offloading stage includes wireless transmission delay , wired transmission delay , waiting delay and calculation delay , expressed as:
[0056] ,
[0057] According to the vehicle terminal simultaneously realizing the calculation and offloading of sensing data, the synchronization delay of digital twinning is expressed as:
[0058] ,
[0059] wherein, represents the delay of the vehicle terminal v in the offloading stage;
[0060] S44, the total energy consumption of the vehicle terminal v in the digital twinning synchronization process is calculated;
[0061] In a complete digital twinning synchronization process, the total energy consumption of the vehicle terminal v includes sensing energy consumption , transmission energy consumption and local calculation energy consumption , expressed as:
[0062] ,
[0063] wherein, represents the total energy consumption of the vehicle terminal v;
[0064] S45, the sensing quality of the vehicle terminal v in the time slot t is calculated;
[0065] The sensing quality of the vehicle terminal v in the time slot t is:
[0066] ,
[0067] wherein, is a sensing performance parameter, and e represents a natural constant;
[0068] S46, calculate the total perception quality of each vehicle terminal;
[0069] At time slot t, the overall digital twin model accuracy is denoted as:
[0070] ,
[0071] wherein, is the contribution weight of vehicle terminal v in the perception environment data.
[0072] Further, the specific steps of constructing the Markov model are as follows:
[0073] S51, construct an optimization objective function;
[0074] By jointly considering the vehicle terminal perception frequency , the perception data offloading ratio , the wireless transmission power and the edge server cooperation strategy , the delay and energy consumption in the digital twin synchronization process of the Internet of Vehicles are minimized while the digital twin model accuracy is maximized under the condition of limited edge network resources, and the multi-objective optimization problem is denoted as:
[0075] ,
[0076] wherein, is the maximum perception frequency of the vehicle terminal, denotes the maximum wireless transmission power of the vehicle terminal, denotes the maximum delay tolerance time of the digital twin synchronization, denotes the battery capacity of the vehicle terminal, denotes the minimum requirement of the digital twin model accuracy;
[0077] Constraint C1 indicates that the perception frequency of the vehicle terminal cannot exceed the maximum value ; constraint C2 indicates that the transmission power of the vehicle terminal cannot exceed the maximum value ; constraint C3 indicates that the offloaded perception data is only processed by one edge server; constraint C4 indicates that the digital twin synchronization delay must be within the tolerance range; constraint C5 indicates that the energy consumption of the vehicle terminal within the service time cannot exceed the battery capacity ; constraint C6 indicates that the accuracy of the overall digital twin model cannot be lower than the threshold ;
[0078] S52, by time slot division, the optimization problem P1 is converted into an optimization problem P2 conforming to the Markov decision process:
[0079] ,
[0080] in, , and These represent the weights for latency, energy consumption, and model accuracy, respectively.
[0081] Furthermore, a hybrid multi-agent model based on hybrid deep reinforcement learning is established as follows:
[0082] Define the MDP state space corresponding to the vehicle terminal intelligent agent. With action space The MDP state space corresponding to the edge server intelligent agent With action space Joint reward function ;
[0083] Vehicle terminal intelligent agent state space The state space of each vehicle terminal agent is determined by the remaining battery power of the vehicle terminal agent. Channel gain and edge server agent queue status Composition, represented as ;
[0084] Vehicle terminal intelligent agent motion space In which each vehicle terminal intelligent agent is in The action space at any given moment includes the vehicle terminal sensing frequency. Perception data unloading ratio and wireless transmission power , represented as ;
[0085] Edge server agent state space The state space of each edge server agent includes the computational load. Vehicle terminal intelligent agent status Actions of vehicle terminal intelligent agents If the edge server agent m has a state space at time t, then the state space of the edge server agent m is represented as follows: ;
[0086] Edge server intelligent agent action space In this context, the action space of each edge server agent at time t is the edge server collaboration strategy. , represented as ;
[0087] reward function ,in It is punishment, expressed as , , , is a penalty value in different cases; when the digital twin model precision , the penalty identifies , otherwise 0; when the synchronization time delay , the penalty identifies , otherwise 0; when the vehicle terminal agent energy consumption , the penalty identifies , otherwise 0.
[0088] Compared with the prior art, the present application has the following remarkable effects:
[0089] 1. The present application proposes a cloud-edge-end architecture for vehicle networking digital twin synchronization, which can efficiently schedule sensing and computing resources according to real-time business needs, fully utilize the computing power of edge nodes, significantly reduce the computing cost of terminals, and maintain high-precision digital twin model construction and update maintenance in the cloud; it can improve synchronization reliability and solve the problem of insufficient global model precision and excessive terminal computing load caused by the deployment of digital twins limited to the end and edge in the prior art.
[0090] 2. In the sensing and computing resource scheduling method of the present application, for sensing data, the edge server can make autonomous collaboration decisions according to the current computing load, avoiding excessive computing load on some edge servers while others are idle, and achieving efficient processing of sensing data.
[0091] 3. In the sensing and computing resource scheduling method of the present application, the resource scheduling problem in vehicle networking digital twin synchronization is modeled as a hybrid Markov decision process, and a hybrid multi-agent model based on hybrid deep reinforcement learning is designed. Through collaborative decision-making of vehicle terminal agents and edge server agents, the present application can more accurately handle continuous variables (sensing frequency, transmission power and offloading ratio) and discrete variables (edge server collaboration decision) in resource scheduling, achieving lower synchronization time delay, lower vehicle terminal energy consumption and higher digital twin model precision in a dynamically changing vehicle networking environment, effectively solving the resource scheduling conflict problem caused by vehicle device heterogeneity and environmental dynamics. BRIEF DESCRIPTION OF DRAWINGS
[0092] Figure 1 is a cloud-edge-end architecture for digital twin synchronization in a vehicle networking scenario constructed by the present application;
[0093] Figure 2 is a sensing and computing resource scheduling method flowchart for vehicle networking digital twin synchronization constructed by the present application;
[0094] Figure 3 is a hybrid multi-agent model framework based on hybrid deep reinforcement learning constructed by the present application;
[0095] Figure 4 This is a simulation experiment diagram showing the change in digital twin accuracy under four different algorithm scenarios with different edge server bandwidths according to an embodiment of the present invention;
[0096] Figure 5 Simulation diagrams showing the average synchronization latency variation under four different algorithm scenarios with varying numbers of edge servers in this invention embodiment;
[0097] Figure 6 This is the average energy consumption of the vehicle terminal under four different algorithm scenarios with different edge server computing frequencies in this embodiment of the invention. Detailed Implementation
[0098] The present invention will now be described in further detail with reference to the accompanying drawings and specific embodiments.
[0099] like Figure 1 As shown, this invention presents a cloud-edge-device architecture for digital twin synchronization in the Internet of Vehicles. This architecture includes a terminal layer, an edge layer, and a cloud layer. The terminal layer includes a cluster of vehicle terminals based on geographical region division. Each vehicle terminal collects road environment information in real time through onboard sensors, including traffic flow, pedestrian status, and real-time road conditions. After data compression and partial computational processing, the system autonomously executes computational task offloading decisions based on a preset strategy, transmitting the compressed environmental data to the edge server via a wireless link. The edge layer consists of distributed edge servers deployed on roadside units. After receiving the sensing data sent from the terminal layer, the edge servers can autonomously make collaborative decisions based on the current computing load to ensure efficient processing of the sensing data: when the computing load of a single node exceeds a threshold, redundant data is sent to a low-load edge server for collaborative processing. After processing the environmental data, the low-load edge server uploads the edge-processed data for building the digital twin to the cloud. The cloud layer includes a central cloud server, which builds and maintains an overall digital twin model containing all sub-regions based on the edge-processed data for building the digital twin uploaded from the edge layer. When the cloud center receives the edge-processed data for building the digital twin from the edge layer in the next time slot, it performs real-time synchronous updates to the digital twin model.
[0100] The application proposes a sensing and communication algorithm resource scheduling method for vehicle networking digital twin synchronization, models the resource scheduling problem in vehicle networking digital twin synchronization as a hybrid Markov decision process, designs a hybrid multi-agent model based on hybrid deep reinforcement learning, and through the collaborative decision of the vehicle terminal agent and the edge server agent, simultaneously combines double Q networks to evaluate the action value and select actions respectively, avoiding the problem of overestimation of Q value. Compared with the single-agent reinforcement learning algorithm used in the prior art or the algorithm that does not fully integrate continuous and discrete action decision, the application can more accurately process continuous variables (sensing frequency, transmission power and offloading ratio) and discrete variables (edge server cooperation strategy) in resource scheduling, realize lower synchronization delay, lower vehicle terminal energy consumption and higher digital twin model accuracy in a dynamically changing vehicle networking environment, and effectively solve the resource scheduling conflict problem caused by the heterogeneity of vehicle equipment and the dynamicity of the environment. As shown in FIG. 1, it is a flowchart of the sensing and communication algorithm resource scheduling method for vehicle networking digital twin synchronization proposed by the application, and the application aims to realize high-precision digital twin synchronization with low latency and low energy consumption, and the specific steps are as follows: Figure 2
[0101] Step 1, build a cloud-edge-terminal collaborative architecture for vehicle networking digital twin synchronization;
[0102] The terminal layer has V0 vehicle terminals, which are distributed in M0 areas, and the set of vehicle terminals is represented as ; Each area is configured with an edge server deployed on a roadside unit to provide data processing services and reduce data transmission delay, and the edge servers transmit data to each other through optical fibers, and the set of edge servers is represented as , which are located in the edge layer; The cloud layer is configured with a central cloud server for building and maintaining the digital twin body of the entire scene using the digital twin uploaded by the edge layer; The total time length of the cloud-edge-terminal architecture for digital twin synchronization is divided into T0 time slots with a length of , and the set is .
[0103] Step 2, the vehicle terminal sensing device senses and collects the environmental data around the vehicle terminal, including traffic flow, pedestrians and real-time road condition information, and establishes a collaborative sensing model;
[0104] The specific steps are as follows:
[0105] Step 21, determine the sensing data volume of each vehicle terminal;
[0106] Because the traffic flow, pedestrian, and real-time road condition information in multiple sub-regions within the constructed overall digital twin area are highly dynamic, multiple vehicle terminals need to collaboratively perceive the data to build a refined overall digital twin model. Specifically, this involves determining the amount of sensing data for each vehicle terminal based on its sensing frequency. After processing by the vehicle terminals and edge servers, the sensing data is integrated at the cloud layer to complete the construction of a refined overall digital twin model. Let the vehicle terminals... In the time slot The perceived frequency at that time is Each time, the amount of environmental data sensed is d0. After collecting the sensed data, the vehicle terminal compresses and encodes it to obtain valid data. The compression coding rate of the sensed data is... Then the amount of perceived data for:
[0107] ,
[0108] Step 22, calculate the sensing energy consumption of vehicle terminal v within time slot t. The expression for sensing energy consumption is:
[0109] ,
[0110] in, This represents the sensing energy consumption of vehicle terminal v within time slot t; It is the energy consumption of a single sensor operation at the vehicle terminal (v).
[0111] Step 3: Based on the communication process between the vehicle terminal and the edge server, and between edge servers, establish the communication model. The specific steps are as follows:
[0112] Step 31: Calculate the transmission rate of vehicle terminal v, the wireless transmission delay of vehicle terminal offloading sensing data to edge server m, and the wireless transmission energy consumption of vehicle terminal v in time slot t.
[0113] Due to limited computing power, the vehicle terminal cannot process all the perceived data. Therefore, it autonomously decides to offload a portion of the data to an edge server associated with the area where the vehicle terminal (v) is located. The offloading ratio of perceived data is [missing information]. In the constructed system, the vehicle terminal v communicates with the edge server m using the widely used Orthogonal Frequency Division Multiple Access (OFDMA) technology. The transmission rate is... for:
[0114] ,
[0115] in, It is the uplink channel bandwidth of the vehicle terminal. , It is the channel power gain of the vehicle terminal v. denotes the channel power gain of the vehicle terminal , denotes the noise power, is the transmission power of the vehicle terminal v, is the transmission power of the vehicle terminal , and is autonomously decided by the respective vehicle terminal.
[0116] Thus, the wireless transmission latency of the vehicle terminal offloading the perception data to the edge server m is:
[0117] ,
[0118] Correspondingly, the wireless transmission energy consumption of the vehicle terminal v at time slot t is:
[0119] .
[0120] Step 32, after the edge server m receives the perception data, it autonomously makes a cooperative decision according to the current computing load situation;
[0121] There are certain differences in the internal environment of each area, resulting in some edge servers with heavy computing load, while the other part is idle. In order to guarantee the efficient processing of perception data, after the edge server m receives the perception data, it can autonomously make a cooperative decision according to the current computing load situation, set denotes the edge server cooperation strategy, then:
[0122] When and , it means that the edge server m transmits the perception data from the vehicle terminal v to other edge servers i for processing through a wired link;
[0123] When and , it means that the perception data is directly processed by the edge server m;
[0124] If the edge server m and other edge servers i are connected by optical fiber, then the optical fiber transmission latency is:
[0125] ,
[0126] where hops (m, i) represents the shortest path hop count between edge servers, and v0 represents the rate of transmitting unit data in one hop. When the perception data is directly processed by the associated edge server m of the vehicle terminal v, there is no information transmission between edge servers, which is represented as hops (m, m) = 0.
[0127] Step 4, according to the flow of processing perception data by vehicle terminals and edge servers, a calculation model is established;
[0128] The specific steps are as follows:
[0129] Step 41, the time delay of the vehicle terminal processing part of the perception data is calculated;
[0130] In addition to performing perception tasks, the vehicle terminal also has certain computing power, so the time delay of processing part of the perception data (i.e. calculating part of the perception data) is :
[0131] ,
[0132] wherein, is the number of CPU cycles required to process unit perception data, is the computing frequency of the vehicle terminal v.
[0133] Correspondingly, the local computing energy consumption is:
[0134] ,
[0135] wherein, is the energy consumption coefficient of the computing chip.
[0136] Step 42, the waiting time delay of the perception data of the vehicle terminal v is calculated;
[0137] After receiving the perception data unloaded by the vehicle terminal, the edge server stores it in the task buffer queue. represents the queue state of the edge server m at time slot t, i.e. the perception data that has not been processed in the queue of the edge server m at time slot t. In addition, there are some perception data of vehicle terminals entering the buffer queue before the vehicle terminal v, and the set of these vehicle terminals is In the queue, the perception data that arrives first is processed first, and the perception data that arrives later needs to wait until the perception data that arrives first is processed before being processed. Correspondingly, the waiting time delay of the perception data of the vehicle terminal v is:
[0138] ,
[0139] wherein, represents the computing frequency of the edge server m; representing the amount of perception data of the vehicle terminal ; part of the data is offloaded to the edge server of the area where the vehicle terminal is located, and the offloading ratio is .
[0140] The perception data in the queue of the edge server m that has not been processed in the next time slot can be represented as:
[0141] ,
[0142] The perception data of the vehicle terminal v is processed by the edge server after queuing, and the computing delay of the edge server is represented as:
[0143] ,
[0144] wherein, is the total amount of data computed at the edge server m in the time slot t.
[0145] Step 43, the total delay of the vehicle terminal and the edge server in cooperatively computing the perception data of the current time slot is counted.
[0146] Since the computing capability of the cloud center is strong enough, the delay of the digital twin model of the cloud center in updating accounts for a very small proportion in the total digital twin synchronization delay, and thus can be ignored; since the transmission between the edge server and the cloud center is wired, the data transmission delay between the edge server and the cloud center accounts for a very small proportion in the total digital twin synchronization delay, and thus can be ignored. Therefore, when counting the total digital twin synchronization delay, the present application does not consider the delay of the above two parts, but only counts the total delay of the vehicle terminal and the edge server in cooperatively computing the perception data of the current time slot.
[0147] Since the edge server has continuous power supply, its computing energy consumption is not considered. The delay of the vehicle terminal v in the offloading stage is composed of the wireless transmission delay , the wired transmission delay , the waiting delay and the computing delay , and is represented as:
[0148] ,
[0149] wherein, is the task offloading decision of the edge server m.
[0150] The vehicle terminal can simultaneously realize the computing and offloading of the perception data, and thus the synchronization delay of the digital twin depends on the maximum value of the two, and is represented as:
[0151] ,
[0152] wherein, denotes the latency of the vehicle terminal v in the unloading phase.
[0153] Step 44, calculate the total energy consumption of the vehicle terminal v in the digital twin synchronization process;
[0154] In a complete digital twin synchronization process, the total energy consumption of the vehicle terminal v is composed of the perception energy consumption , the transmission energy consumption and the local calculation energy consumption , which is expressed as:
[0155] ,
[0156] wherein, denotes the total energy consumption of the vehicle terminal v.
[0157] Step 45, calculate the perception quality of the vehicle terminal v in the time slot t;
[0158] Due to the heterogeneity of sensors carried by the vehicle terminal, there are differences in perception ability, and the perception quality of the vehicle terminal v in the time slot t is:
[0159] ,
[0160] wherein, is the perception performance parameter, and e denotes the natural constant.
[0161] Step 46, calculate the total perception quality of each vehicle terminal;
[0162] At time t, the overall digital twin model accuracy is calculated in real time by the cloud center according to the total perception quality of each vehicle terminal in the constructed digital twin model , which is specifically expressed as:
[0163] ,
[0164] wherein, is the contribution weight of the vehicle terminal v in the perception of environmental data.
[0165] Step 5, construct an optimization objective function according to the requirements of low latency, low energy consumption and high accuracy of digital twin synchronization, and convert the problem into a partially observable Markov decision problem through time slot division, and construct a Markov model;
[0166] The specific steps are as follows:
[0167] Step 51, construct an optimization objective function;
[0168] By jointly considering the vehicle terminal sensing frequency Perception data unloading ratio Wireless transmission power Collaboration strategy with edge servers Minimize latency during the synchronization process of vehicle-to-everything (V2X) digital twins, given limited edge network resources. and energy consumption At the same time, maximize the accuracy of the digital twin model Then the multi-objective optimization problem P1 is expressed as:
[0169] ,
[0170] in, It is the maximum sensing frequency of the vehicle terminal. This indicates the maximum wireless transmission power of the vehicle terminal. This represents the maximum latency tolerance time for digital twin synchronization. This indicates the battery capacity of the vehicle's terminal. This represents the minimum accuracy requirement for a digital twin model.
[0171] Constraint C1 means that the sensing frequency of the vehicle terminal cannot exceed the maximum value. C2 indicates that the transmission power of the vehicle terminal cannot exceed the maximum value. Constraint C3 indicates that the offloaded sensing data is processed by only one edge server; Constraint C4 indicates that the digital twin synchronization latency must be within the tolerance range, otherwise it will affect subsequent services; Constraint C5 indicates that the energy consumption of the vehicle terminal during service time cannot exceed the battery capacity. Constraint C6 indicates that the accuracy of the overall digital twin model cannot be lower than a threshold. .
[0172] Step 52: Transform the multi-objective optimization problem P1 into a partially observable Markov decision problem and construct a Markov model;
[0173] The total time required for the cloud-edge-device architecture to complete digital twin synchronization is divided into T0 time intervals of length T. The time slots, the time slot set is By dividing the time slots, the multi-objective optimization problem P1 is transformed into an optimization problem P2 that conforms to MDP (Markov Decision Process):
[0174] ,
[0175] in, , and These represent the weights for latency, energy consumption, and model accuracy, respectively.
[0176] Step 6, based on the Markov model, a hybrid multi-agent model based on hybrid deep reinforcement learning is established to solve the optimal resource scheduling strategy;
[0177] As shown in Figure 3 , a vehicle terminal agent and an edge server agent are respectively established based on a hybrid deep reinforcement learning algorithm, and for continuous action decision 、 and , each vehicle terminal agent is interacted with the environment as an Actor to learn the optimal strategy; for discrete action decision , the edge server agent is used as a Critic to evaluate the continuous action and output the best discrete action. Based on the Markov model, the hybrid multi-agent is trained and the current strategy is updated, and a hybrid multi-agent model based on hybrid deep reinforcement learning is established, which includes the following steps:
[0178] Step 61, a hybrid multi-agent model based on hybrid deep reinforcement learning is established;
[0179] The MDP (Markov Decision Process) state space and action space of the vehicle terminal agent are defined, the MDP state space and action space of the edge server agent are defined, and the joint reward function is defined;
[0180] The state space of the vehicle terminal agent is , wherein the state space of each vehicle terminal agent is composed of the residual power of the vehicle terminal agent, the channel gain and the queue state of the edge server agent , and is represented as ;
[0181] The action space of the vehicle terminal agent is , wherein the action space of each vehicle terminal agent at time t includes the vehicle terminal sensing frequency , the sensing data offloading ratio and the wireless transmission power , and is represented as ; the target action space derived from the target network in the Actor is , which includes the target vehicle terminal sensing frequency , the target sensing data offloading ratio and the target wireless transmission power ;
[0182] Edge server agent state space , where the state space of each edge server agent is composed of the computing load , vehicle terminal agent state In addition to making discrete action decisions, edge server agents also need to evaluate the quality of continuous action decisions, so the actions of the vehicle terminal agent are also input to the edge server agent, and the state space of the edge server agent m at time t is represented as ;
[0183] Edge server agent action space , where the action space of each edge server agent at time t is the edge server cooperation strategy , represented as ;
[0184] Reward function , where is the penalty, specifically represented as , where , , are the penalty values in different cases, all of which are constants. When the digital twin model accuracy , the penalty identifies , otherwise 0; when the synchronization delay , the penalty identifies , otherwise 0; when the vehicle terminal agent energy consumption , the penalty identifies , otherwise 0.
[0185] Step 62, training the hybrid multi-agent model;
[0186] The specific steps of agent training are as follows:
[0187] Step 621, set hyperparameters and initialize the agent model;
[0188] In the hybrid multi-agent model based on hybrid deep reinforcement learning, the main parameters are set as follows: the agent Actor learning rate is set to 1e-4, the Critic learning rate is set to 1e-3, the discount factor is set to 3e-4, the experience pool capacity is set to 100000, the sampling size is 128, the soft update rate is set to 0.005, and the training rounds are set to 4000, each round containing 100 steps.
[0189] Step 622, initialize the environment parameters at the beginning of each round;
[0190] It includes: each edge server agent state initialization, vehicle terminal agent state initialization, edge server agent queue state initialization, task cumulative index reset to 0, time slot set to 0, Markov chain state initialization.
[0191] Step 623, the environment state of the vehicle terminal agent v in the current time slot is obtained , and the agent makes a continuous action decision . .
[0192] Step 624, the environment state of the edge server agent m in the current time slot is obtained , the edge server agent m will evaluate the continuous action decision of the vehicle terminal agent v, and select the optimal discrete action decision .
[0193] Step 625, after the two types of agents make decisions respectively, the environment gives a joint reward value of the current time slot , and enters the next state;
[0194] Step 626, the state, action and reward of the current time slot are stored in the experience pool, and when the sample amount in the experience pool reaches the preset number, the saved samples are used to train the two types of agents respectively, and the current policy is updated.
[0195] In order to verify the superiority of the hybrid multi-agent model based on the hybrid deep reinforcement learning in the application, the average value of the cumulative digital twin synchronization delay of all sensing devices, the average value of the energy consumption and the highest digital twin accuracy value maintained by the cloud edge architecture are selected as evaluation indexes, and are compared with the following three benchmark solutions:
[0196] F1, DDPG algorithm combined with adversarial double deep Q network (DDPG-D3QN, Deep Deterministic Policy Gradients combined with dueling double Deep Q-Network): deploy DDPG algorithm on the vehicle terminal as the Actor module to output continuous decision, and deploy D3QN algorithm on the edge server as the Critic module to output discrete decision. The D3QN algorithm adopts an adversarial structure, and the Q value is synthesized by the value function and the advantage function. However, unlike the algorithm adopted in the application, there is only one main network and one target network in the Critic module, which may cause the problem of Q value overestimation.
[0197] F2, twin delayed deep deterministic policy gradient (TD3): the TD3 algorithm module is deployed on the vehicle terminal to output continuous decisions, and the edge server outputs discrete decisions using a random policy. The TD3 algorithm also uses an AC architecture, and uses two main networks and a target network to evaluate continuous decisions to solve the Q value overestimation problem. However, unlike the design of the Critic module of the hybrid multi-agent model of the present application, the network in TD3 directly outputs the Q value when evaluating actions.
[0198] F3, DDPG: the DDPG algorithm is deployed on the vehicle terminal to output continuous decisions, and the edge server outputs discrete decisions using a random policy as a Critic module. DDPG is a classic AC structure algorithm that uses one main network and one target network to evaluate continuous decisions and directly outputs the Q value.
[0199] As shown in Figure 4 To verify the performance of the hybrid deep reinforcement learning algorithm in the present application under limited channel bandwidth, the digital twin accuracy of the hybrid multi-agent model of the present application and other three DRL algorithms under different edge server communication bandwidth conditions is compared. The number of vehicle terminals is set to 30, the number of edge servers is set to 6, and the communication bandwidth conditions are set to 7.5 MHz, 10.0 MHz, 12.5 MHz, 15.0 MHz, and 17.5 MHz. The simulation results show that as the communication bandwidth increases, the digital twin accuracy shows a clear upward trend. In addition, the hybrid deep reinforcement learning of the present application always maintains the highest performance improvement trend, and even in the case of the smallest bandwidth, it can ensure the highest digital twin accuracy by formulating the optimal sensing algorithm resource scheduling.
[0200] As shown in Figure 5 To verify the adaptability of the present application in terms of resource scheduling in a cloud-edge-end collaborative scenario, the average synchronization delay of the four algorithms under different edge server numbers is compared. The number of vehicle terminals is set to 30, the communication bandwidth is set to 10.0 MHz, and the number of edge servers is set to 4, 5, 6, 7, and 8. The simulation results show that as the number of edge servers increases, the synchronization delay of all algorithms decreases, and the hybrid deep reinforcement learning algorithm of the present application always maintains the lowest synchronization delay throughout the process, and shows a more significant downward trend as the number of servers increases.
[0201] As shown in Figure 6As shown, in order to verify the adaptability of the application in the resource scheduling of the cloud edge end collaborative scene, the average energy consumption of the vehicle terminal under different edge server computing frequencies of four DRL algorithms is compared. Among them, the number of vehicle terminals is set to 30, the communication bandwidth is set to 10.0MHz, the number of edge servers is set to 6, and the edge server computing frequency is set to 8GHz, 12GHz, 16GHz, 20GHz and 24GHz respectively. The simulation results show that with the increase of the computing frequency, the average energy consumption of each algorithm shows a downward trend, and the hybrid deep reinforcement learning algorithm of the application always maintains the lowest energy consumption under all frequency conditions. Compared with DDPG-D3QN, TD3 and DDPG algorithms, the average energy consumption of the hybrid deep reinforcement learning algorithm of the application is reduced by about 5.2%, 9.0% and 8.0% respectively. This result shows that the hybrid deep reinforcement learning of the application has higher energy efficiency adaptability and optimization ability in the edge end collaborative scene.
Claims
1. A method for sensing and computing resource scheduling for vehicle networking digital twin synchronization, which schedules sensing and computing resources according to real-time business needs, characterized in that, The steps include the following: S1, a cloud edge system for vehicle networking digital twin synchronization is constructed; In the cloud-edge-terminal system, the terminal layer has vehicle terminals, which are distributed in areas, and a set of vehicle terminals is represented as An edge server is configured in each area and deployed in a roadside unit, and a set of edge servers is represented as The total time length of the cloud-edge-terminal system for completing digital twin synchronization is divided into time slots with a length of , and a set of time slots is represented as S2, vehicle terminal sensing devices perceive and collect vehicle terminal surrounding environment data, and a collaborative perception model is constructed; S3, a communication model is established according to the communication process of the vehicle terminal and the edge server, and the edge server and the edge server; S4, a calculation model is established according to the process of the vehicle terminal and the edge server processing sensing data; S5, an optimization objective function is constructed according to the requirements of digital twin synchronization, and the optimization problem is converted into a partially observable Markov decision problem through time slot division, and a Markov model is constructed; the specific steps of constructing the Markov model are as follows: S51, an optimization objective function is constructed; By jointly considering the vehicle terminal sensing frequency , the sensing data offloading ratio , the wireless transmission power and the edge server cooperation strategy , in the case of limited edge network resources, the delay and energy consumption in the digital twin synchronization process of the Internet of Vehicles are minimized, while the accuracy of the digital twin model is maximized, and the multi-objective optimization problem is represented as: wherein, is the maximum sensing frequency of the vehicle terminal, denotes the maximum wireless transmission power of the vehicle terminal, denotes the maximum delay tolerant time for digital twin synchronization, denotes the battery capacity of the vehicle terminal, denotes the minimum requirement for the digital twin model accuracy; constraint representing that the sensing frequency of the vehicle terminal cannot exceed a maximum value ; representing that the transmission power of the vehicle terminal cannot exceed a maximum value ; constraint representing that the offloaded sensing data is only processed by one edge server; constraint representing that the digital twin synchronization latency must be within a tolerated range; constraint representing that the energy consumption of the vehicle terminal cannot exceed the battery capacity within the service time ; constraint representing that the accuracy of the overall digital twin model cannot be below a threshold value ; S52, by time slot division, the optimization problem is converted into an optimization problem conforming to a Markov decision process : wherein, , and denote the weights of latency, energy consumption and model accuracy, respectively. S6, a hybrid multi-agent model based on hybrid deep reinforcement learning is established based on the Markov model, and an optimal resource scheduling strategy is solved; wherein the hybrid multi-agent model based on hybrid deep reinforcement learning is as follows: Defining the MDP state space for the vehicle terminal agent and action space Defining the MDP state space for the edge server agent and action space Joint reward function ; Vehicle terminal agent state space wherein the state space of each vehicle terminal agent is composed of the vehicle terminal agent remaining battery level , channel gain and edge server agent queue state and is represented as ; Vehicle terminal intelligent agent motion space In which each vehicle terminal intelligent agent is in The action space at any given moment includes the vehicle terminal sensing frequency. Perception data unloading ratio and wireless transmission power , represented as ; Edge server agent state space The state space of each edge server agent includes the computational load. Vehicle terminal intelligent agent status Actions of vehicle terminal intelligent agents Then the edge server intelligent agent exist The state space representation at time t is as follows ; Edge server agent action space wherein the action space of each edge server agent at a time instant is an edge server coordination policy denoted as ; reward function wherein is a penalty, expressed as , , , is a penalty value in different cases; when the digital twin model accuracy , the penalty identifies , otherwise 0; when the synchronization latency , the penalty identifies , otherwise 0; when the vehicle terminal agent energy consumption , the penalty identifies , otherwise 0.
2. The method of claim 1, wherein, The vehicle terminal surrounding environment data includes traffic flow, pedestrian and real-time road condition information; the specific steps of constructing the collaborative perception model are as follows: S21, the sensing data amount of each vehicle terminal is determined; The sensing data amount of each vehicle terminal is determined according to the sensing frequency of the vehicle terminal, and after the sensing data is processed by the vehicle terminal and the edge server, it is integrated in the cloud layer to complete the construction of the digital twin model; The vehicle terminal compresses and encodes the perception data to obtain effective data after collecting the perception data, and the vehicle terminal performs perception in a time slot with a perception frequency of , the amount of environmental data perceived each time is , the compression and encoding rate of the data is , and the amount of perception data is: , S22, calculating a time slot In-vehicle terminal of the sensing energy consumption , the expression is: , wherein is a vehicle terminal sensing the energy consumption once.
3. The method of claim 1, wherein, The specific steps of establishing the communication model are as follows: S31, calculate the transmission rate of the vehicle terminal , the vehicle terminal unloads the perception data to the edge server , the wireless transmission delay of the vehicle terminal , the wireless transmission energy consumption of the vehicle terminal at the time slot ; Vehicle terminal Orthogonal frequency division multiple access and edge server Communication, the transmission rate Is: , wherein is an uplink channel bandwidth of the vehicle terminal, , is a channel power gain of the vehicle terminal , denotes a channel power gain of the vehicle terminal , denotes a noise power, is a transmission power of the vehicle terminal , is a transmission power of the vehicle terminal ; The vehicle terminal autonomously decides to unload part of the data to an edge server associated with the area where the vehicle terminal is located according to the computing capability, and sets an unloading ratio as Then, the vehicle terminal unloads the perception data to the edge server The wireless transmission delay is: , Vehicle terminal In time slots Wireless transmission energy consumption Is: wherein, represents the amount of perception data of the vehicle terminal ; S32, edge server After receiving the perception data, autonomous cooperation decision is made according to current computing load condition; the cooperation decision is as follows: when and , it means that the edge server transmits the perception data from the vehicle terminal to other edge servers for processing through a wired link; when and , it means that the perception data is directly processed by the edge server ; wherein, ; Edge server and other edge servers connected by optical fiber, the optical fiber transmission delay is: , wherein, denotes the shortest path hop count between edge servers, denotes the rate of transmitting unit data in one hop; when the perception data is directly processed at the vehicle terminal associated edge server , there is no information transmission between edge servers, denoted as .
4. The method of claim 3, wherein, The specific steps of establishing the total sensing quality calculation model of each vehicle terminal are as follows: S41, the vehicle terminal calculates the time delay of part of the sensing data; Vehicle terminal processing section perceives data time delay To: , wherein, is the number of CPU cycles required to process the unit perception data, is the computing frequency of the vehicle terminal . Local computing energy consumption Is: , wherein, is the energy consumption coefficient of the computing chip; S42, calculate a waiting latency of the perception data of the vehicle terminal S42, calculate a waiting latency of the perception data of the vehicle terminal Suppose there are some vehicle terminals' perception data prior to the vehicle terminal entering the cache queue, the set of these vehicle terminals is ; in the queue, the perception data that arrives first is processed first, and the data that arrives later has to wait until the data that arrives first is processed before it can be processed, then the waiting time delay of the vehicle terminal 's perception data is : , in, Represents the edge server in time slot t The queue state; Represents edge server The calculation frequency; Indicates vehicle terminal The amount of perceived data; offloading some data to the vehicle terminal. The offloading ratio of the edge servers associated with the region is ; then the next time slot edge server unprocessed perception data in the queue of the edge server is represented as: , Vehicle terminal The perception data is processed by the edge server after queuing, and the computing latency of the edge server is represented as: , wherein, is a time slot Intrinsic edge server total amount of data computed at the edge server; S43, the total time delay of the vehicle terminal and the edge terminal for cooperating to calculate the current time slot sensing data is counted; Vehicle terminal The latency in the offloading phase includes a wireless transmission latency , a wired transmission latency , a waiting latency , and a computing latency , expressed as: , According to the vehicle terminal, the calculation and unloading of the perception data are simultaneously implemented, and the synchronization delay of digital twinning is represented as: , wherein, representing a vehicle terminal latency in the offload phase; S44, calculating total energy consumption of the vehicle terminal in the digital twin synchronization process ; In a complete digital twin synchronization process, the total energy consumption of the vehicle terminal includes perception energy consumption , transmission energy consumption and local calculation energy consumption , expressed as: , wherein represents the total energy consumption of the vehicle terminal ; S45, calculate the perceived quality of the vehicle terminal in the time slot S45, calculate the perceived quality of the vehicle terminal in the time slot In a time slot The perceived quality of a vehicle terminal is: , wherein, is a perceptual performance parameter, e denotes the natural constant; S46, the total sensing quality of each vehicle terminal is calculated; At the time slot The overall digital twin model accuracy is represented as: , wherein is a vehicle terminal a contribution weight in the perception environment data.
5. A cloud-edge-terminal system for vehicle networking digital twin synchronization, comprising a terminal layer, an edge layer and a cloud layer, for executing the method of sensing and computing resource scheduling according to any one of claims 1-4, characterized in that, The terminal layer includes a vehicle terminal cluster based on geographical area division, each vehicle terminal collects road environment information in real time through a vehicle-mounted sensor, including traffic flow, pedestrian state and real-time road condition; and the compressed environment data is transmitted to the edge server through a wireless link; the edge layer is composed of distributed edge servers deployed in roadside units, and the edge server makes autonomous cooperation decisions according to the current computing load after receiving the sensing data sent from the terminal layer; the cloud layer includes a central cloud server, which performs real-time synchronization update of the digital twin model after receiving the edge processing data for digital twin construction from the edge layer in the next time slot; wherein the digital twin model is constructed and maintained based on the edge processing data for digital twin construction uploaded by the edge layer.
6. The cloud-edge-end system for digital twin synchronization of Internet of Vehicles according to claim 5, wherein, When the single node computing load exceeds the threshold, redundant data is sent to the low-load edge server for cooperative processing, and the low-load edge server uploads the edge processing data for digital twin construction to the cloud after processing the environment data.
Citation Information
Patent Citations
Vehicle digital twinborn body edge deployment method used in Internet of Vehicles scene
CN116980424A
Digital twinning-based car networking task unloading method
CN118075265A