A method for edge deployment of vehicle digital twins in connected vehicle scenarios
By optimizing the edge deployment of vehicle digital twins through multi-agent deep reinforcement learning, the problems of high-speed vehicle mobility and limited edge server resources are solved, achieving low-latency and efficient transportation system deployment.
Patent Information
- Application Number
- CN202311113123.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-30
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2043-08-30
AI Technical Summary
The dynamic deployment of vehicle digital twins faces challenges such as high-speed mobility, low latency requirements, and limited edge server resources, making it difficult to achieve efficient edge deployment.
By employing a multi-agent deep reinforcement learning approach combined with the Actor-Critic algorithm, the deployment location of vehicle digital twins on edge servers is optimized. By constructing a digital twin-driven intelligent transportation vehicle network, the age of cloud information and the migration cost of the digital twins are minimized.
In situations where edge server resources are limited, the deployment efficiency of vehicle digital twins is improved, cloud information age and migration costs are reduced, and the real-time performance and efficiency of the transportation system are enhanced.
Smart Images

Figure CN116980424B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of mobile communication technology and relates to a method for edge deployment of vehicle digital twins in vehicle networking scenarios. Background Technology
[0002] Current intelligent transportation systems lag behind in their ability to perceive real traffic environments and analyze and process big data, which is insufficient to support the development of emerging transportation applications in the B5G / 6G era that have higher requirements for the accuracy, comprehensiveness, and real-time nature of traffic information.
[0003] Digital twin technology refers to the consistent replication of objects in physical space across time and space into a twin space, and the study and control of these objects are achieved through observation, analysis, deduction, and manipulation of the digital twin. Utilizing digital twin technology to construct a digital model of the entire road traffic scenario provides more advanced technical support for the intelligent and digital transformation of transportation, improving road traffic efficiency and safety.
[0004] Mobile edge computing technology reduces system latency and network transmission pressure by performing computations at the network edge. Deploying digital twins at the network edge can ensure low-latency interaction between the digital twin and its corresponding physical entity. However, the high-speed mobility of vehicles, the low-latency requirements for digital twin synchronization, and the limited heterogeneous resources of edge servers all pose significant challenges to the dynamic deployment of vehicle digital twins. Summary of the Invention
[0005] In view of this, the purpose of this invention is to provide a method for edge deployment of vehicle digital twins in the context of vehicle networking, so as to achieve optimal edge deployment of vehicle digital twins and minimize the average information age in the cloud and the average migration cost of the twin.
[0006] To achieve the above objectives, the present invention provides the following technical solution:
[0007] A method for edge deployment of vehicle digital twins in connected vehicle scenarios includes the following steps:
[0008] S1: Construct a digital twin-driven intelligent transportation vehicle network, including a physical terminal layer, an edge twin layer, and a cloud application layer;
[0009] S2: The base station is responsible for determining the deployment location of each vehicle digital twin within its service range on the edge side; the vehicle digital twin is initially deployed on the edge server closest to the currently accessing base station;
[0010] S3: The base station forwards the real-time status data uploaded by the vehicle to the edge server where its twin is located for digital twin synchronization. The data is then processed and provided to the cloud application layer.
[0011] S4: Calculate the average information age in the cloud and the twin migration cost under the current deployment method;
[0012] S5: Establish an objective function that minimizes the average age of information in the cloud and migration costs;
[0013] S6: The objective function is transformed into a partially observable Markov decision process by multiple agents, and the optimal vehicle digital twin deployment scheme is solved by using an Actor-Critic-based multi-agent deep reinforcement learning method.
[0014] Furthermore, in step S1, the physical terminal layer consists of mobile vehicle nodes with limited computing and storage resources, and the set of vehicle nodes is denoted as V = {1,2,…,i,…,I}; the vehicle nodes perceive surrounding environmental information and vehicle status data in real time through on-board sensors, and upload the perceived data to the edge twin layer in real time using wireless communication technology, providing data support for the construction of the digital twin;
[0015] The edge twin layer consists of base stations providing access services and edge servers providing computing and storage services. Each edge server is associated with any base station. The set of base stations is denoted as BS = {1,2,…,j,…,J}, and the set of edge servers is denoted as ES = {1,2,…,k,…,K}.
[0016] The cloud application layer includes a digital twin-based intelligent traffic management platform that collects and stores vehicle status data, road environment data, and other data from real traffic scenarios throughout their entire lifecycle, constructing a digital model of the entire road traffic scenario to support the operation of cloud-based traffic applications. It also includes a large number of digital twin-based traffic application servers, which perform efficient verification and low-cost trial and error of innovative technologies on the digital twin platform.
[0017] Furthermore, in step S3, the digital twin synchronization process follows a timeout retransmission synchronization mechanism. Specifically, the vehicle status data upload interval is dynamically determined based on the upload delay of the previous data packet and the processing delay of the edge server. During a synchronization process, after the edge server processes the status data uploaded by the vehicle node, it returns an ACK confirmation message. The vehicle uploads the next status data packet only after receiving the confirmation message. At the same time, a digital twin synchronization delay threshold is set. If the vehicle has not received the ACK confirmation message when the synchronization interval exceeds the threshold, a timeout retransmission is performed to ensure the consistency and real-time performance of the twin synchronization.
[0018] Furthermore, in step S4, the average information age in the cloud refers to the average age of all vehicle data information in the current environment in the cloud, specifically calculated as follows: Define Z = [z i,k ]I×K Deployment matrix for DT, where DT represents the vehicle twin. When the digital twin corresponding to vehicle i is deployed on edge server k, z i,k =1, otherwise z i,k =0; Assuming that at time t0, the cloud happens to receive the DT information about vehicle i uploaded by edge server k, then from time t0 until the next data update, the cloud's DT information... i Information age Δ i (t) can be written as:
[0019]
[0020] in, T represents the synchronization delay between vehicle i and its twin, i.e., the transmission delay of vehicle i's state data uploading to edge server k. k,Cloud This represents the transmission latency of edge server k forwarding the status data of vehicle i to the cloud.
[0021] Define the average information age during vehicle i-synchronization. for:
[0022]
[0023] The average age of all vehicles is obtained from the cloud. It can be represented as:
[0024]
[0025] Furthermore, in step S4, the twin migration cost includes the acquisition cost and the transmission cost, specifically calculated as follows: when the digital twin DT corresponding to vehicle i... i Migration cost when the deployment location is migrated from edge server k1 to k2 Represented as:
[0026]
[0027] Among them, c ou c represents the development cost. mig DT represents the unit migration cost, which is the cost of transmitting a unit of data over a unit distance; i Disk This represents the storage resources required to maintain the digital twin of vehicle i, and dis(k1,k2) represents the distance between edge server k1 and edge server k2.
[0028] Define Z′=[z′ i ] I×1 Let DT be the migration matrix; for vehicle i, use the binary variable z′. iThis indicates whether the digital twin of vehicle i has migrated at a given moment; if the deployment location in the current time slot is the same as the previous time slot, the migration process will not be triggered, z′ i =0, otherwise z′ i =1; then the average twin migration cost of the entire system is Cost. mig It can be represented as:
[0029]
[0030] Furthermore, in step S5, the objective function for minimizing the average information age in the cloud and the migration cost is expressed as:
[0031]
[0032] stC1:
[0033] C2:
[0034] C3:
[0035] C4:
[0036] C5:
[0037] C6:z i,k ∈{0,1}
[0038] C7:β1,β2∈(0,1),β1+β2=1
[0039] Where Z = [z i,k ] I×K This represents the vehicle twin deployment matrix. Cost represents the average age of all vehicles obtained from the cloud. mig DT represents the average twin migration cost. i CPU and DT i Disk These represent maintaining DT. i The required CPU computing resources and disk storage resources, and These represent the total CPU computing resources and disk storage resources equipped on edge server k, respectively. Among the constraints, C1 ensures that each vehicle digital twin can only be deployed on a single edge server at any given time; C2 to C3 ensure that deploying vehicle digital twins will not exhaust the computing and storage resources of any edge server; C4 to C5 ensure that the age requirements of the edge information are met simultaneously during DT synchronization. Information age requirements in the cloud in, Δ represents the time it takes for twins to synchronize once. m This indicates the maximum age of information in the cloud during synchronization; C6 explains the deployment of the variable z for the vehicle twin. i,k For binary variables, z is the value when the digital twin of vehicle i is deployed on edge server k. i,k =1, otherwise z i,k =0; β1 and β2 in C7 represent weighting factors, namely the weighting coefficients of the average information age in the cloud and the twin migration cost.
[0040] Furthermore, in step S6, the objective function is transformed into a partially observable Markov decision process involving multiple agents, specifically including:
[0041] (1) Global state space Real-time status information of all vehicles and their twin deployment locations, sub-channel occupancy status of all base stations and vehicle-related information, resource usage of all edge servers, and cloud-based information on the age of all vehicles.
[0042] (2) Local state space of agent j Real-time status information of vehicles currently associated with this base station, their twin deployment locations, information age indicators, and the location and remaining resource information of all edge servers;
[0043] (3) Action space Deployment strategy Z of digital twins for all vehicles within the coverage area of each base station;
[0044] (4) Rewards The negative of the weighted sum of the average cloud information age and average twin migration cost of all vehicles associated with the base station;
[0045] The reward r for agent j in time slot t t j Defined as:
[0046]
[0047] The reward consists of two parts, the first part being... That is, the average AoI index of all vehicles associated with base station j; the second part is That is, the migration cost of all vehicles associated with base station j; the total reward of agent j is obtained by weighting and summing the two parts using coefficients β1 and β2 respectively.
[0048] Furthermore, in step S6, an Actor-Critic-based multi-agent deep reinforcement learning method is used to solve for the optimal vehicle digital twin deployment scheme. Specifically, this includes employing an Actor-Critic-based centralized training-distributed execution deep reinforcement learning framework, where the policy of each agent is represented by a deep neural network, i.e., an Actor network, denoted as: Where, π j Let θ be the policy of agent j. j For the network parameters of a deep neural network, a j Let s be the action space of agent j. j For the observable state of agent j, set a virtual central node to deploy a DNN with parameters φ, i.e., a Critic network;
[0049] The loss function is calculated and Adam is used to update the network parameters. The training continues until all Actor networks converge or the maximum number of training rounds is reached, ultimately yielding the optimal vehicle digital twin deployment strategy.
[0050] Furthermore, in step S6, the algorithm training process specifically includes the following steps:
[0051] S61: Initialize the parameters of each Actor network and Critic network, and initialize the experience pool;
[0052] S62: The data sampling process is executed cyclically, and agent j acquires the observed state. Then, execute actions according to the strategy. Receive reward r t i Obtain the state of the next time slot. The sampled data is placed into the experience pool.
[0053] S63: Calculate the estimate of the advantage function Calculate the target value V j (s t );
[0054] S64: For each agent j, take a sample from the experience pool. Calculate the loss function of the Actor network and update its parameters. Calculate the loss function of the Critic network and update its parameters. Repeat this process until the reward function converges or the maximum number of training epochs is reached.
[0055] Furthermore, in step S6, the loss function is calculated and combined with Adam to update the network parameters, which includes: the Actor network updates the network parameters through the policy loss function, and the Actor network loss function L corresponding to agent j. Actor (θj The calculation expression for ) is:
[0056]
[0057] in, and Let a represent the old policy and the current new policy of agent j, respectively. t and s t These represent the executed action and observed state at time t, respectively. and These represent the action and observation state of agent j at time t, respectively; clip() is a truncation function that ensures the difference between the updated and old parameters is not too large; ε is a hyperparameter used to set the limit range of the clip() function. It is the dominant function The estimated value is obtained through generalized dominance estimation, specifically defined as:
[0058]
[0059] Where, δ t =r t +γV φ (s t+1 )-V φ (s t ), δ t Let r be the time series difference error at time t. t Let V be the reward at time t, γ be a discount factor that measures the importance of future rewards, and V be the reward at time t. φ (s t Let ) be the state value function at time t, where T is a sufficiently large time, and λ∈[0,1] is a hyperparameter used to balance bias and variance;
[0060] The Actor network parameters θ corresponding to agent j j Iterative updates via stochastic gradient ascent:
[0061]
[0062] Where, α Actor It is the learning rate of the Actor network. The Actor loss function with respect to network parameters θ j The partial derivatives;
[0063] The Critic network updates its parameters using a global state value loss function; the Critic network loss function L... Critic The expression for calculating (φ) is:
[0064]
[0065] Among them, V φ (s t V represents the output of the state-value function estimated by the Critic network. j (s t () represents the target value;
[0066] The Critic network parameters φ are updated iteratively using stochastic gradient descent:
[0067]
[0068] Where, α Critic It is the learning rate of the Critic network. Let φ be the partial derivative of the Critic loss function with respect to the network parameter φ.
[0069] Furthermore, in step S6, the advantage function is defined as follows: Used to evaluate the merits of taking an action relative to the average level in a given state; the global action value function is defined as follows: Among them U t The cumulative discounted return; the global state value function is
[0070] The beneficial effects of this invention are as follows: This invention enables the reasonable deployment of vehicle digital twins under the limitation of computing and storage resources of edge servers, and fully considers the impact of the deployment location of vehicle digital twins on the latency of the digital twin synchronization process. This reduces the migration cost of the twins while minimizing the average information age in the cloud, and improves the deployment efficiency of vehicle digital twins on the edge side.
[0071] Other advantages, objectives, and features of the invention will be set forth in part in the description which follows, and in part will be apparent to those skilled in the art from the following examination, or may be learned from practice of the invention. The objectives and other advantages of the invention can be realized and obtained through the following description. Attached Figure Description
[0072] To make the objectives, technical solutions, and advantages of the present invention clearer, the preferred embodiments of the present invention will be described in detail below with reference to the accompanying drawings, wherein:
[0073] Figure 1 This is a diagram of the intelligent transportation vehicle network architecture driven by digital twins constructed in this invention;
[0074] Figure 2 This is a training flowchart of the method for edge deployment of vehicle digital twins in the Internet of Vehicles scenario according to the present invention;
[0075] Figure 3This is a diagram illustrating the vehicle digital twin synchronization process and cloud information age description of the present invention;
[0076] Figure 4 This is a diagram of the multi-agent deep reinforcement learning framework based on Actor-Critic in this invention. Detailed Implementation
[0077] The following specific examples illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of the present invention. Unless otherwise specified, the following embodiments and features can be combined with each other.
[0078] Please see Figures 1-4 This invention constructs a digital twin-driven intelligent transportation vehicle network, the architecture of which is shown in the diagram below. Figure 1 As shown, this invention provides a method for edge deployment of vehicle digital twins in vehicle networking scenarios, the process of which is as follows: Figure 2 As shown, implementing this method specifically includes the following steps:
[0079] S1: Construct a digital twin-driven intelligent transportation vehicle network architecture, which includes three layers: physical terminal layer, edge twin layer, and cloud application layer.
[0080] The physical terminal layer comprises numerous mobile vehicle nodes, each with limited computing and storage resources, denoted as V = {1, 2, ..., I}. The edge twin layer consists of multiple base stations and edge servers, denoted as BS = {1, 2, ..., J} and ES = {1, 2, ..., K}, respectively. Edge servers provide computing and storage services to maintain the digital twins of the vehicles. Each edge server is associated with any base station. Vehicle nodes utilize wireless communication technology to upload status data to the edge twin layer in real time, providing data support for the construction of the digital twins. The cloud application layer contains a core cloud-based digital twin platform and maintains numerous intelligent transportation applications, such as route planning, autonomous driving, traffic flow prediction, and high-precision positioning. The cloud-based digital twin platform provides real-world data support for these intelligent transportation applications, offering efficient verification and low-cost trial-and-error for innovative technologies on the twin platform.
[0081] S2: When formulating a vehicle digital twin deployment strategy, we target the base stations in the edge twin layer. Each base station is responsible for determining the deployment location of each vehicle digital twin within its coverage area on the edge side. Each vehicle digital twin is initially deployed on the edge server closest to the currently accessing base station.
[0082] In step S2, each base station provides access service to the vehicle, with a coverage range of 500-600 meters. Each base station is responsible for determining the deployment location of the vehicle's digital twin at the edge within its coverage area. Considering that after each vehicle finishes its service, the relevant data of the digital twin maintained in the edge server will be automatically uploaded to the cloud for storage, and then the computing and storage resources of the edge server will be released. When the vehicle reconnects to the network, the vehicle's digital twin will be initialized and deployed on the edge server closest to the currently accessing base station. This edge server directly retrieves some relevant data from the cloud to initialize the digital twin.
[0083] S3: The vehicle digital twin in the edge server needs to synchronize its status with the real vehicle in the physical terminal layer. The base station forwards the real-time status data uploaded by the vehicle to the edge server where its twin is located. After the edge server processes the twin data, it forwards it to the cloud application layer to support a large number of cloud-based traffic applications based on digital twins.
[0084] In step S3, digital twin synchronization specifically includes the following steps:
[0085] S31: The communication method between the vehicle and the base station is wireless communication. Assume the vehicle's transmission power is p. i Given a channel bandwidth of B, a wireless channel gain of h, a channel white Gaussian noise power spectral density of N0, a distance of dis(i,j) between vehicle i and base station j, and a path loss factor of γ, the achievable uplink transmission rate between vehicle i and base station j is:
[0086]
[0087] At a certain moment, the amount of data transmitted by vehicle i is D. i The state data is transmitted to base station j, with a transmission delay of:
[0088]
[0089] In this embodiment, the vehicle's transmission power ranges from [500, 600] mW, the wireless channel bandwidth is 50 MHz, the wireless channel gain is -30 dB / m, and the noise power spectral density is -127 dBm / Hz.
[0090] S32: If the digital twin of vehicle i is maintained in edge server k, then the wired transmission delay for base station j to upload data to edge server k is:
[0091]
[0092] Where ψ is the wired transmission unit delay, i.e., the transmission delay required to transmit a unit size of data per unit distance, and dis(j,k) is the distance between base station j and edge server k. In this embodiment, the wired transmission unit delay is taken as 1×102. -12 .
[0093] S33: After receiving the status data uploaded by vehicle node i, edge server k needs to perform calculations such as compression, aggregation, optimization, and caching on the original data. The processing time is:
[0094]
[0095] in, f(D) represents the processing time for edge server k to handle data uploaded by vehicle i. i ) indicates that the amount of data to be processed is D. i The computational cost of DT i CPU The computing resources required to maintain the digital twin of vehicle i. In this embodiment, the computing resources DT required for the digital twin are... i CPU and storage resources DT i DISK The value ranges are [1,2]GHz and [2,3]GB, respectively.
[0096] S34: After processing the data, edge server k will return an ACK confirmation message to vehicle node i. Since the ACK confirmation message data is small, the transmission delay is negligible. Therefore, when the DT corresponding to vehicle i... i When deployed on edge server k, the time for a single synchronization is denoted as:
[0097]
[0098] S35: While processing the data, edge server k forwards the status data of vehicle i to the cloud. The transmission latency is denoted as:
[0099] T k,Cloud =ψ·D i ·dis(k,Cloud)
[0100] S4: During each digital twin synchronization process, the base station records the synchronization latency of vehicles within its coverage area under the current deployment method, the age of vehicle status data in the cloud, and the twin migration cost.
[0101] In step S4, the specific calculation steps for the average age of all vehicles in the cloud are as follows:
[0102] S41: Assume that at time t0, the cloud receives the DT information about vehicle i uploaded by edge server k. Then, from time t0 until the next data update, the cloud's DT information... i Information about age can be written as:
[0103]
[0104] S42: The average information age during vehicle synchronization is:
[0105]
[0106] S43: The average age of all vehicles obtained from the cloud can be represented as:
[0107]
[0108] In step S4, the twin migration cost calculation steps are as follows:
[0109] S44: When the digital twin DT corresponding to vehicle i i If the migration strategy changes at some point, say from edge server k1 to edge server k2, then the target edge server k2 first needs to allocate a new memory space to maintain the twin and perform initialization operations. The cost incurred in this process is called the allocation cost, denoted as c. ou In this embodiment, the development cost is set to 1×10. -5 .
[0110] S45: The original edge server k1 will transfer all the data of the twin to be migrated to the new edge server k2. This process will incur a transmission cost, denoted as c. mig ·DT i Disk ·dis(k1,k2). Where, c mig This represents the unit migration cost, that is, the cost of transmitting a unit of data over a unit distance. DT i Disk Indicates twin DT i The total amount of data included. In this embodiment, the unit migration cost is taken as 1×10. -10 .
[0111] S46: When the digital twin DT corresponding to vehicle i i The total migration cost when the deployment location is migrated from edge server k1 to k2 is expressed as:
[0112]
[0113] S47: Define Z′=[z′ i ] I×1 This is the DT migration matrix. For vehicle i, use the binary variable z′. i This indicates whether the digital twin of vehicle i has migrated at a given time. If the deployment location in the current time slot is the same as the previous time slot, the migration process will not be triggered. i =0, otherwise z′ i =1.
[0114] S48: The migration cost of the entire system is:
[0115]
[0116] S5: Establish a vehicle digital twin deployment strategy optimization model to minimize the average information age and migration cost in the cloud.
[0117] In addressing the deployment problem of vehicle digital twins, this invention considers the real-time data requirement of the cloud for vehicle node status information in the real traffic environment, using an information age index to measure the freshness of the real vehicle status data received by the cloud. Simultaneously, considering the high mobility of vehicle nodes, vehicles may continuously adjust the deployment location of their digital twins during movement, triggering a migration process between edge servers and incurring corresponding migration costs. Therefore, this invention balances migration costs while optimizing the cloud information age target, avoiding triggering a large number of digital twin migration processes. The optimization model in step S5 is specifically expressed as follows:
[0118]
[0119] stC1:
[0120] C2:
[0121] C3:
[0122] C4:
[0123] C5:
[0124] C6:z i,k ∈{0,1}
[0125] C7:β1,β2∈(0,1),β1+β2=1
[0126] In the formula, Z = [z i,k ] I×K This represents the vehicle twin deployment matrix. Cost represents the average age of information in the cloud. mig This represents the average twin migration cost. Among the constraints, C1 ensures that each vehicle twin can only be deployed on a single edge server at any given time. C2-C3 ensure that deploying the vehicle digital twin will not exhaust the computing and storage resources of any edge server. In this embodiment, the total computing and storage resources of each edge server range from [80,90] GHz and [2,3] TB. C4-C5 ensure that the information age requirements of both the edge and cloud are met simultaneously during DT synchronization. C6 describes the deployment of variable z for the vehicle twin. i,k For binary variables, z is the value when the digital twin of vehicle i is deployed on edge server k. i,k =1, otherwise z i,k =0. β1 and β2 in C7 represent weighting factors, namely the weighting coefficients of the average information age in the cloud and the twin migration cost. In this embodiment, the values of β1 and β2 are 0.6 and 0.4, respectively.
[0127] S6: Transform the optimization model into a multi-agent partially observable Markov decision process, and solve it using an Actor-Critic-based multi-agent deep reinforcement learning method to obtain the optimal deployment scheme for the vehicle digital twin.
[0128] In step S6, the optimization model is first transformed into a partially observable Markov decision process involving multiple agents. This partially observable Markov decision process mainly consists of four parts, defined in detail as follows:
[0129] (1) Global state space Real-time status information of all vehicles and their twin deployment locations, sub-channel occupancy status of all base stations and vehicle-related information, resource usage of all edge servers, and cloud-based information on the age of all vehicles.
[0130] (2) Local state space of agent j Real-time status information of vehicles currently associated with this base station, their twin deployment locations, information age indicators, and the location and remaining resource information of all edge servers.
[0131] (3) Action space Deployment strategy Z of digital twins for all vehicles within the coverage area of each base station.
[0132] (4) Rewards The negative of the weighted sum of the average cloud information age and average twin migration cost of all vehicles associated with the base station. The reward r for agent j in time slot t. t j Defined as
[0133]
[0134] After transforming it into a partially observable Markov decision process involving multiple agents, a centralized training-distributed execution deep reinforcement learning framework based on Actor-Critic is employed, such as... Figure 4 As shown.
[0135] In step S6, the policy π of each agent j is... j This can be represented by a deep neural network, namely an Actor network, denoted as: Set up a virtual central node to deploy a DNN with parameter φ, namely the Critic network.
[0136] The algorithm training process in step S6 specifically includes the following steps:
[0137] S61: Initialize the parameters of each Actor network and Critic network, and initialize the experience pool;
[0138] S62: The data sampling process is executed cyclically, and agent j acquires the observed state. Then, execute actions according to the strategy. Receive reward r t i Obtain the state of the next time slot. The sampled data is placed into the experience pool.
[0139] S63: Obtain the estimate of the dominance function through generalized dominance estimation. The specific calculation method is as follows:
[0140]
[0141] In the formula, δ t =r t +γV φ (s t+1 )-V φ (s t In this embodiment, the discount factor γ is 0.9, and the hyperparameter λ is 0.95. The target value V of the Critic network is then calculated. j (s t ).
[0142] S64: For each agent j, take a sample from the experience pool. Calculate the Actor network loss function and update the Actor network parameters. The Actor network loss function for agent j is calculated as follows:
[0143]
[0144] In the formula, and These represent the old policy and the current new policy of agent j, respectively. `clip()` is a truncation function that ensures the difference between the updated and old parameters is not too large. `ε` is a hyperparameter used to set the limit range of the `clip()` function. In this embodiment, `ε` is set to 0.2. The Actor network parameters θ corresponding to agent j... j Iterative updates via stochastic gradient ascent:
[0145]
[0146] Where, α Actor It is the learning rate of the Actor network. The Actor loss function with respect to network parameters θ j The partial derivatives of .
[0147] S65: Calculate the Critic network loss function and update the Critic network parameters. The Critic network updates its parameters using the global state value loss function. The Critic network loss function is calculated as follows:
[0148]
[0149] Among them, V φ (s t V represents the output of the state-value function estimated by the Critic network. j (s t ) represents the target value.
[0150] The Critic network parameters φ are updated iteratively using stochastic gradient descent:
[0151]
[0152] Where, α Critic It is the learning rate of the Critic network. Let be the partial derivative of the Critic loss function with respect to the network parameter φ.
[0153] The training process is repeated until the reward function converges or the maximum number of training rounds is reached.
[0154] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A method for edge deployment of vehicle digital twins in vehicle-to-everything (V2X) scenarios, characterized in that, The method specifically includes the following steps: S1: Construct a digital twin-driven intelligent transportation vehicle network, including a physical terminal layer, an edge twin layer, and a cloud application layer; S2: The base station is responsible for determining the deployment location of each vehicle digital twin within its service range on the edge side; the vehicle digital twin is initially deployed on the edge server closest to the currently accessing base station; S3: The base station forwards the real-time status data uploaded by the vehicle to the edge server where its twin is located for digital twin synchronization. The data is then processed and provided to the cloud application layer. S4: Calculate the average information age in the cloud and the twin migration cost under the current deployment method; S5: Establish an objective function that minimizes the average age of information in the cloud and migration costs; S6: The objective function is transformed into a multi-agent partially observable Markov decision process, and the optimal vehicle digital twin deployment scheme is solved by using an Actor-Critic-based multi-agent deep reinforcement learning method. In step S1, the physical terminal layer consists of mobile vehicle nodes with limited computing and storage resources, and the set of vehicle nodes is denoted as . Vehicle nodes perceive surrounding environmental information and vehicle status data in real time through onboard sensors, and use wireless communication technology to upload the perceived data to the edge twin layer in real time, providing data support for the construction of digital twins. The edge twin layer consists of base stations providing access services and edge servers providing computing and storage services. Each edge server is associated with any one base station; the set of base stations is denoted as . The set of edge servers is denoted as ; The cloud application layer includes a digital twin-based intelligent traffic management platform and multiple digital twin-based traffic application servers. In step S4, the average information age in the cloud refers to the average age of all vehicle data information in the current environment in the cloud. The specific calculation method is as follows: [Definition...] Deploy a matrix for DT. DT Represents a vehicle twin, when the vehicle The corresponding digital twin is deployed on an edge server. hour, ,otherwise ; Assuming in At that moment, the cloud happened to receive a message from the edge server. Uploaded information about vehicles The DT information, then from From the start of this moment until the next data update, the cloud-based information about... Information Age writing: in, Indicates vehicle Synchronization delay with the vehicle twin, i.e., vehicle Status data is uploaded to the edge server Transmission delay, Represents edge server Vehicle Status data is forwarded to the cloud. Transmission delay; Define vehicle Average information age during synchronization for: The average age of all vehicles is obtained from the cloud. Represented as: The twin migration cost includes acquisition cost and transmission cost, and is calculated as follows: when the vehicle Corresponding digital twin Deployment location from edge server Migrate to At that time, migration costs Represented as: in, Indicates development costs, This represents the unit migration cost, which is the cost of transmitting a unit of data over a unit distance. Indicates vehicle maintenance The storage resources required for a digital twin Represents edge server With edge servers The distance between them; definition For vehicles, the DT migration matrix is used. Using binary variables To indicate vehicles Whether the digital twin migrates at a certain moment; if the deployment location of the current time slot is the same as that of the previous time slot, the migration process will not be triggered. ,otherwise Then the average twin migration cost of the entire system Represented as: In step S5, the objective function for minimizing the average information age in the cloud and the migration cost is expressed as: in, This represents the vehicle twin deployment matrix. This indicates the average age of all vehicles obtained from the cloud. This represents the average twin migration cost. and They represent maintenance. The required CPU computing resources and disk storage resources, and These represent edge servers. The system is equipped with a total of CPU computing resources and disk storage resources. Among the constraints, C1 ensures that each vehicle digital twin can only be deployed on a single edge server at any given time; C2-C3 ensure that deploying vehicle digital twins will not exhaust the computing and storage resources of any edge server; and C4-C5 ensure that the age requirements of the edge information are met simultaneously during DT synchronization. Information age requirements in the cloud ,in, This indicates the time it takes for twins to synchronize once. This indicates the maximum age of information in the cloud during synchronization; C6 explains the variables deployed for the vehicle twin. It is a binary variable, when the vehicle Digital twins deployed on edge servers When, ,otherwise ;C7 and This represents the weighting factor, which is the weighting coefficient between the average age of information in the cloud and the migration cost of the twin. In step S6, the objective function is transformed into a partially observable Markov decision process involving multiple agents, specifically including: (1) Global state space Real-time status information of all vehicles and their twin deployment locations, sub-channel occupancy status of all base stations and vehicle-related information, resource usage of all edge servers, and cloud-based information on the age of all vehicles; (2) Intelligent agent local state space Real-time status information of vehicles currently associated with this base station, their twin deployment locations, information age indicators, and the location and remaining resource information of all edge servers; (3) Action space Digital twin deployment strategy for all vehicles within the coverage area of each base station ; (4) Rewards : The negative of the weighted sum of the average cloud information age and average twin migration cost of all vehicles associated with the base station; intelligent agents exist Time slot rewards Defined as: The reward consists of two parts, the first part being... That is, with base station The average AoI metric for all associated vehicles; the second part is That is, with base station Migration costs for all associated vehicles; intelligent agents The total reward is calculated using a coefficient. and We obtain the result by weighted summation of the two parts.
2. The vehicle digital twin edge deployment method according to claim 1, characterized in that, In step S3, the digital twin synchronization process follows a timeout retransmission synchronization mechanism. Specifically, the vehicle status data upload interval is dynamically determined based on the upload delay of the previous data packet and the processing delay of the edge server. During a synchronization process, after the edge server processes the status data uploaded by the vehicle node, it returns an ACK confirmation message. The vehicle uploads the next status data packet only after receiving the confirmation message. At the same time, a digital twin synchronization delay threshold is set. If the vehicle has not received the ACK confirmation message when the synchronization interval exceeds the threshold, a timeout retransmission is performed to ensure the consistency and real-time performance of the twin synchronization.
3. The vehicle digital twin edge deployment method according to claim 1, characterized in that, In step S6, an Actor-Critic-based multi-agent deep reinforcement learning method is used to solve for the optimal vehicle digital twin deployment scheme. Specifically, this includes employing an Actor-Critic-based centralized training-distributed execution deep reinforcement learning framework, where the policy of each agent is represented by a deep neural network, i.e., an Actor network, denoted as: ,in, For intelligent agents strategy, For the network parameters of a deep neural network, For intelligent agents The space of motion For intelligent agents The observable state; set a virtual central node to deploy a parameter of The DNN, namely the Critic network; The loss function is calculated and Adam is used to update the network parameters. The training continues until all Actor networks converge or the maximum number of training rounds is reached, ultimately yielding the optimal vehicle digital twin deployment strategy.
4. The vehicle digital twin edge deployment method according to claim 3, characterized in that, In step S6, the loss function is calculated and combined with Adam to update the network parameters, which includes: the Actor network updating network parameters through the policy loss function, and the agent... Corresponding Actor network loss function The calculation expression is: in, and Representing intelligent agents respectively The old strategy and the current new strategy and They are respectively t The actions performed and the observed state at any given moment. and respectively intelligent agents exist t The actions performed and the observed status at any given moment; It is a truncation function that ensures that the difference between the updated parameters and the old parameters is not too large; It is a hyperparameter used to set Function range restrictions; It is the dominant function The estimated value is obtained through generalized dominance estimation, specifically defined as: in, , for The timing difference error at time 10:
00. for t Momentary rewards A discount factor to measure the importance of future rewards. for The state value function at time t, T For a sufficiently large moment, It is a hyperparameter used to balance bias and variance; intelligent agent Corresponding Actor network parameters Iterative updates via stochastic gradient ascent: in, It is the learning rate of the Actor network. The Actor loss function with respect to network parameters The partial derivatives; The Critic network updates its parameters using a global state value loss function; the Critic network loss function... The calculation expression is: in, This represents the output of the state-value function estimated by the Critic network. Indicates the target value; Critic network parameters Iterative updates via stochastic gradient descent: in, It is the learning rate of the Critic network. The Critic loss function with respect to network parameters The partial derivatives of .
5. The vehicle digital twin edge deployment method according to claim 4, characterized in that, In step S6, the advantage function is defined as follows: It is used to evaluate how good or bad an action is relative to the average level in a given state; Define the global action value function as follows: ,in The cumulative discounted return; the global state value function is .
Citation Information
Patent Citations
User state sensing method in mobile edge computing network and related equipment
CN113472842A
Dynamic deployment method for digital twin servers in edge Internet of Vehicles
CN115209426A