A migration method of a full-quantity and light-weight digital twin model
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GUANGDONG UNIV OF TECH
- Filing Date
- 2026-04-27
- Publication Date
- 2026-08-07
AI Technical Summary
[0005]基于此,本发明针对现有技术中缺乏合理的数字孪生迁移决策导致数字孪生迁移过程出现不稳定的问题,提供一种全量和轻量数字孪生模型的迁移方法,具体的技术方案如下:
本发明通过在迁移决策过程中引入双时间尺度协同优化机制,在不同时间尺度上分别对两类数字孪生体的迁移行为进行决策与控制。短时间尺度侧重于快速响应车辆位置变化和局部服务需求,保障数字孪生的实时性;长时间尺度则从系统整体运行状态出发,对迁移行为进行全局优化,同时将边缘服务器的计算资源约束和容量限制显式纳入决策空间,对可能引发资源超载的迁移动作进行限制,提升了数字孪生迁移过程的稳定性和可靠性,进而提升数字孪生服务系统的运行效率。
Smart Images

Figure CN122526795A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of vehicle-to-everything (V2X) digital twins, and in particular to a method for migrating full and lightweight digital twin models. Background Technology
[0002] As a crucial foundational support system for smart cities, the Internet of Vehicles (IoV) integrates onboard sensing devices, roadside infrastructure, wireless communication networks, and edge computing resources to achieve real-time information interaction and collaborative control among multiple elements: people, vehicles, roads, and the cloud. In a typical urban IoV architecture, numerous vehicles continuously generate high spatiotemporal resolution operational data through onboard terminals. This data is then accessed, processed, and fed back using Roadside Units (RSUs) and Edge Servers (ES), thereby supporting applications such as intelligent driving assistance, traffic condition awareness, signal control, and emergency dispatch.
[0003] To enhance the perception and decision-making capabilities of vehicle-to-everything (V2X) systems in complex traffic environments, digital twin technology has been introduced into the fields of intelligent transportation and V2X. This technology constructs a high-fidelity dynamic mapping of physical vehicles and their operating environment in virtual space. By integrating vehicle sensor data, roadside perception information, and historical traffic patterns, the digital twin can reflect real-time vehicle operating status and changes in the traffic environment, providing data support for traffic situation prediction, collaborative decision-making, and system optimization. Furthermore, addressing the differentiated needs of vehicle services in terms of real-time response and global optimization, existing research has gradually proposed and explored a two-layer digital twin architecture. The lightweight digital twin is typically deployed on the RSU node close to the vehicle to support high-frequency perception updates, instant interaction, and local decision-making. Its characteristics include small model size, low computational complexity, and fast response speed. The full digital twin, on the other hand, is deployed on the ES node with stronger computing and storage capabilities to handle complex model computations, cross-regional collaboration, and city-level traffic optimization tasks.
[0004] However, in a two-layer digital twin architecture, the continuous movement of vehicles necessitates dynamic adjustments to the deployment locations of both lightweight and full digital twins across different edge nodes to maintain spatial proximity between the twin and the physical vehicle, as well as synchronization consistency of the twin's state. This leads to a digital twin migration process that typically includes model state packaging, network transmission, and reconstruction and recovery at the target node. This process inevitably incurs additional communication latency, computational overhead, and energy costs. Especially in densely populated areas or high-speed driving scenarios, lightweight digital twins may require frequent migrations between RSUs, while full digital twin migrations are more costly and less frequent, exhibiting significant differences in migration cycles, resource requirements, and service impact. Without a reasonable migration decision-making mechanism, frequent or unnecessary migrations can easily trigger resource competition between edge nodes, exacerbate system load fluctuations, and even lead to twin service interruptions or performance degradation. Therefore, coordinating the migration behaviors of lightweight and full digital twins under resource-constrained and highly dynamic environmental conditions, while ensuring service continuity and low latency and improving overall system efficiency, has become a key technical challenge in the application of vehicle-to-everything (V2X) digital twins. Summary of the Invention
[0005] Based on this, the present invention addresses the problem of instability in the digital twin migration process caused by the lack of reasonable digital twin migration decisions in the prior art, and provides a migration method for full and lightweight digital twin models. The specific technical solution is as follows: This invention proposes a transfer method for full and lightweight digital twin models, comprising the following steps: S1: Construct a digital twin migration model, which includes a lightweight digital twin migration model and a full digital twin migration model; S2: Construct the first optimization target model based on the digital twin transfer model; S3: The transfer decision framework based on dual-timescale joint optimization transforms the first optimization objective model into the second optimization objective model; S4: Use an alternating iterative strategy and a preset algorithm to solve the second optimization objective model and obtain the optimal migration decision.
[0006] Furthermore, the specific content of the lightweight digital twin transfer model described in step S1 is as follows: The migration process of a lightweight digital twin consists of two steps. The first step is for the RSU to package the vehicle's digital twin. The packaging latency at this stage is defined as follows:
[0007] in, It is a binary correlation variable representing vehicles. v Lightweight digital twint- Is slot 1 deployed on RSU? superior, Indicates deployment on RSU superior, Indicates vehicle v The number of CPU cycles required for lightweight digital twin packaged computing. RSU Assigned to vehicles v The computing power required for lightweight digital twin packaging; the relationship between the computing resource requirements of lightweight digital twin packaging and the data volume of its digital twin is as follows: The second process involves lightweight digital twins in link transmission. The transmission latency and power consumption of this process are defined as follows:
[0008]
[0009] in, Indicates the transmission latency of lightweight digital twins. Indicates the power consumption of lightweight digital twin transmission. Indicates vehicle v The data volume of a lightweight digital twin This represents the time coefficient for transmitting a unit of data per unit distance. Indicates roadside unit and roadside units The Euclidean distance between them Average wired transmission power; vehicle v The latency and transfer utility function required for lightweight digital twin migration are defined as follows:
[0010]
[0011] in, This represents the maximum threshold for lightweight digital twin migration latency. Then, this represents the utility function for lightweight digital twin transfers; The specific content of the full digital twin transfer model described in step S1 is as follows: The migration process of the full digital twin consists of two steps. The first step is that ES packages the vehicle's digital twin, and the packaging latency at this stage is defined as:
[0012] in, Indicates vehicle v The full digital twin is deployed in time slot t-1 on ES superior, Indicates vehiclev The number of CPU cycles required for full digital twin packaged computation. ES Assigned to vehicles v The computing power required for packaging the entire digital twin model, and the relationship between the computing resource requirements for packaging the entire digital twin model and the data volume of its digital twin, are as follows: , The first process is the packetization period coefficient; the second process is the full digital twin transmission over the link, where the transmission delay and energy consumption are defined as follows:
[0013]
[0014] in, Indicates the full digital twin transmission latency. This indicates the energy consumption of the full digital twin transmission. Indicates vehicle v The relationship between the total data volume of the full digital twin, the lightweight digital twin data volume of the vehicle, and the total data volume of the full digital twin is as follows: , This represents the time coefficient for transmitting a unit of data per unit distance. Indicates ES and ES Euclidean distance between vehicles v The latency and transfer utility function required for a full digital twin transfer are defined as follows:
[0015]
[0016] in, This represents the maximum threshold for the full digital twin migration latency. This represents the full digital twin transfer utility function.
[0017] Furthermore, before constructing the first optimization target model, it is necessary to construct migration triggering conditions, which include triggering conditions for full digital twin migration and lightweight digital twin migration. The triggering condition for lightweight digital twin migration is: t- Vehicles in time slot 1 v With the original RSU Communication latency Exceeding the threshold And in t Vehicles in time slots v With target RSU Communication latency At the threshold Within the range, the expression is:
[0018] The trigger condition for a full digital twin migration is: vehicle v The full digital twin in time slots t-1 Deployed on the original ES Lightweight digital twin deployment on RSU Up, ES and RSU distance Exceeded the threshold And the target ES with RSU distance At the threshold Within the range, the expression is: .
[0019] Further, the first optimization objective model in step S2 is:
[0020]
[0021]
[0022]
[0023]
[0024]
[0025]
[0026]
[0027]
[0028]
[0029] ;
[0030] The optimization objective considers the migration utility function of the full digital twin and the lightweight digital twin. Regarding constraints, constraints C1 and C2 state that the full digital twin of each vehicle can only be deployed on one ES (Executable System) in the same time slot, and the lightweight digital twin of each vehicle can only be deployed on one RSU (Remote Unit) in the same time slot. Constraints C3 and C4 ensure that the computing power allocated to the ES and RSU cannot exceed their respective maximum computing power resources. Constraints C5 and C6 state that the migration latency of the full digital twin and the lightweight digital twin cannot exceed their respective thresholds. Constraints C7 and C8 state that the digital twins deployed on the ES or RSU cannot exceed the node's space capacity. Constraints C9 and C10 ensure the range of values for the binary correlation variables of the full digital twin migration decision and the lightweight digital twin migration decision. Constraint C11 ensures that the computing power allocation is a non-negative number.
[0031] Furthermore, the specific content of the dual-timescale joint optimization migration decision framework is as follows: optimize the migration and RSU resource allocation of lightweight digital twins in short time slots, and optimize the migration and ES resource allocation of full digital twins in long time slots. The second optimization objective model is:
[0032]
[0033]
[0034]
[0035]
[0036]
[0037]
[0038]
[0039]
[0040]
[0041] ;
[0042] The optimization objective is to maximize the lightweight digital twin transfer utility function within a short time slot t. In long time slots Maximize the transfer utility function of the entire digital twin .
[0043] Furthermore, the specific content of the second optimization objective model to be solved is as follows: construct a first subproblem to optimize the migration of the vehicle's lightweight digital twin between RSUs in short time slots; construct a second subproblem to optimize the migration of the vehicle's full digital twin between ESs in long time slots; solve the first and second subproblems using a preset algorithm and iterate alternately based on the time scale until the global objective function converges.
[0044] The specific content of the first subproblem is as follows: By determining the deployment strategy of the vehicle's lightweight digital twin in short time slot t and the allocation of migration computing resources, the relationship between latency and energy consumption during the migration of the lightweight digital twin between RSUs is balanced. The optimization objective and related constraints are defined as follows:
[0045]
[0046]
[0047]
[0048]
[0049] ;
[0050] The optimization objective is to maximize the long-term average migration utility of lightweight digital twins. C2 represents the lightweight digital twin deployment constraint per vehicle per RSU in short time slot t, C4 represents the upper limit of RSU computing power, C6 represents the migration latency threshold constraint, and C8 represents the upper limit constraint of RSU storage capacity.
[0051] The specific steps for solving the first subproblem are as follows: The first subproblem is modeled using MDP (Markov Decision Process) to obtain a global state space, action space, and reward function. The global state space is the set of the vehicle's observation space. The action space includes two actions: discrete migration decision and continuous computing power allocation decision. The reward function includes rewards for two performance indicators: migration delay and transmission energy. The MAPPO algorithm based on RSU action awareness is used to solve the MDP model to obtain the optimal lightweight digital twin transfer strategy. The steps of the MAPPO algorithm based on RSU action awareness include: Initialize the parameters of the Actor network responsible for outputting the transition action, and the parameters of the Critic network responsible for evaluating the long-term value of the state; In each training round, the short time slots are traversed one by one. In each time slot, all vehicles obtain their own local observations and calculate the RSU action feasibility mask. The Actor network combines the mask to output the action of each vehicle. After all vehicle actions are output, they are executed in a unified manner to obtain the global state of the next time slot. The reward value of each vehicle is calculated, and the current global state, action, global state of the next time slot, and reward value are recorded and stored in the sample set. Update the Actor network parameters and Critic network parameters based on the collected samples; Repeat the above training rounds until the network converges to obtain the optimal lightweight digital twin transfer strategy.
[0052] The specific content of the second subproblem is as follows: By determining long time slots The deployment strategy for the full digital twin of vehicles and the allocation of migration computing resources are used to balance the relationship between latency and energy consumption during the migration of the full digital twin between ES. The optimization objectives and constraints are defined as follows:
[0053]
[0054]
[0055]
[0056]
[0057]
[0058]
[0059] The optimization objective of the second subproblem is to maximize the long-term average migration utility of the entire digital twin, where C1 represents the migration utility of each ES in a long time slot. Constraints for the full digital twin deployment of each vehicle: C3 represents the upper limit constraint of RSU computing power, C5 represents the migration latency threshold constraint, and C8 represents the upper limit constraint of ES storage capacity.
[0060] The specific steps for solving the second subproblem are as follows: The second subproblem is modeled using MDP (Markov Decision Process) to obtain a global state space, action space, and reward function. The global state space is the set of the vehicle's observation space. The action space includes two actions: discrete migration decision and continuous computing power allocation decision. The reward function includes rewards for two performance indicators: migration delay and transmission energy. The MAPPO algorithm based on lightweight twin guidance is used to solve the MDP model to obtain the optimal full digital twin migration strategy. The steps of the MAPPO algorithm based on lightweight twin guidance include: Initialize the parameters of the Actor network responsible for outputting the transition action, and the parameters and long time slots of the Critic network responsible for evaluating the long-term value of the state. ; In each training round, short time slots are traversed one by one. When the cumulative number of short time slots traversed reaches a preset threshold, all vehicles acquire their own local observations. The Actor network, combined with a mask, outputs the migration target selection (ES) action and the actions of packetized computing power allocation. After all vehicle actions are output, they are executed uniformly to obtain the global state of the next long time slot. The reward value of each vehicle is calculated, and the current global state, action, global state of the next long time slot, and reward value are recorded and stored in the sample set. Add 1; After all time slots of a training round have been traversed, the standard dominance function is calculated based on the collected samples using the GAE method. Then, the spatial consistency index between the full twin and the lightweight twin is calculated. The standard dominance function is weighted based on the consistency index to obtain the weighted dominance function guided by the lightweight twin. The parameters of the Actor network and the Critic network are then updated. Repeat the above training rounds until the network converges to obtain the optimal full-scale digital twin transfer strategy.
[0061] Beneficial effects: This invention introduces a dual-timescale collaborative optimization mechanism into the migration decision-making process, enabling decision-making and control of the migration behavior of two types of digital twins at different time scales. The short-term time scale focuses on rapid response to changes in vehicle location and local service needs, ensuring the real-time performance of the digital twin. The long-term time scale, however, considers the overall system operation status, performing global optimization of the migration behavior while explicitly incorporating the computing resource constraints and capacity limitations of edge servers into the decision space. This restricts migration actions that may cause resource overload, improving the stability and reliability of the digital twin migration process and ultimately enhancing the operational efficiency of the digital twin service system. Attached Figure Description
[0062] Figure 1 This is a flowchart of a migration method for full and lightweight digital twin models in this embodiment; Figure 2 This is a schematic diagram of the dual time scale relationship in this embodiment; Figure 3 This is a diagram illustrating the digital twin migration scenario in this embodiment; Figure 4 This is a schematic diagram illustrating the relationship between the digital twin migration scenario and the dual-timescale framework in this embodiment; Figure 5This is a flowchart illustrating the specific implementation method of a migration method for full and lightweight digital twin models in this embodiment. Detailed Implementation
[0063] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0064] Example 1 This embodiment provides a method for transferring full and lightweight digital twin models, and the flowchart of the method is as follows. Figure 1 As shown, it includes the following steps: S1: Construct a digital twin migration model, which includes a lightweight digital twin migration model and a full digital twin migration model; It should be noted that the specific content of the lightweight digital twin transfer model described in step S1 is as follows: Since the vehicle's lightweight digital twin is deployed on the RSU, the migration of the vehicle's lightweight digital twin occurs between RSUs. The migration process consists of two steps. The first step is for the RSU to package the vehicle's digital twin. The packaging latency at this stage is defined as follows:
[0065] in, It is a binary correlation variable representing vehicles. v Lightweight digital twin t- Is slot 1 deployed on RSU? superior, This indicates a vehicle v Lightweight digital twin t- 1 slot deployed on RSU Up, otherwise it indicates the vehicle v The lightweight digital twin t-1 time slot was not deployed on the RSU. Similarly, It is also a binary correlation variable. Indicates vehicle v Lightweight digital twin t Time slots deployed on RSU Above, therefore if and only if and Time (i.e.) ), indicating the vehicle v The lightweight digital twin between time slot t-1 and time slot t is from RSU Migration to RSU superior, Indicates vehiclev The number of CPU cycles required for lightweight digital twin-packed computing. RSU Assigned to vehicles v The computing power (cycles / s) of lightweight digital twin packaging; the relationship between the computing resource requirements of lightweight digital twin packaging and the data volume of its digital twin is as follows: The second process involves lightweight digital twins in link transmission. The transmission latency and power consumption of this process are defined as follows:
[0066]
[0067] in, Indicates the transmission latency of lightweight digital twins. Indicates the power consumption of lightweight digital twin transmission. Indicates vehicle v The data size (in bytes) of a lightweight digital twin. This represents the time coefficient for transmitting a unit of data per unit distance. Indicates roadside unit and roadside units The Euclidean distance between them Average wired transmission power; vehicle v The latency and transfer utility function required for lightweight digital twin migration are defined as follows:
[0068]
[0069] in, This represents the maximum threshold for lightweight digital twin migration latency. Then, this represents the utility function for lightweight digital twin transfers; The specific content of the full digital twin transfer model described in step S1 is as follows: Since the full digital twin of the vehicle is deployed on Elasticsearch (ES), the migration of the full digital twin occurs between ES instances. The migration process consists of two steps. The first step involves ES packaging the vehicle's digital twin; the packaging latency is defined as follows:
[0070] in, and It is a binary correlation variable. Indicates vehicle v The full digital twin is deployed in time slot t-1 on ES superior, Indicates vehicle v The full digital twin is deployed in time slot t on ES Above, therefore if and only if and Time (i.e.) ), indicating the vehicle v The full digital twin between time slot t-1 and time slot t is from ES Migrate to ES superior, Indicates vehicle v The number of CPU cycles required for full digital twin packaged computation. Indicates ES Assigned to vehicles v The computing power (cycles / s) of packaging the full digital twin model, and the relationship between the computing resource requirements of packaging the full digital twin model and the data volume of its digital twin, are as follows: , The first process is the packetization period coefficient; the second process is the full digital twin transmission over the link, where the transmission delay and energy consumption are defined as follows:
[0071]
[0072] in, Indicates the full digital twin transmission latency. This indicates the energy consumption of the full digital twin transmission. Indicates vehicle v The relationship between the data volume (bytes) of the full digital twin and the data volume of the vehicle's lightweight digital twin is as follows: , This represents the time coefficient for transmitting a unit of data per unit distance. Indicates ES and ES Euclidean distance between vehicles v The latency and transfer utility function required for a full digital twin transfer are defined as follows:
[0073]
[0074] in, This represents the maximum threshold for the full digital twin migration latency. This represents the full digital twin transfer utility function.
[0075] S2: Construct the first optimization target model based on the digital twin transfer model; It should be noted that before constructing the first optimization target model, it is necessary to construct migration triggering conditions, which include triggering conditions for full digital twin migration and lightweight digital twin migration. The triggering condition for lightweight digital twin migration is: t- Vehicles in time slot 1 v With the original RSU Communication latency Exceeding the threshold And in t Vehicles in time slots v With target RSU Communication latency At the threshold Within the range, the expression is:
[0076] The trigger condition for a full digital twin migration is: vehicle v The full digital twin in time slots t-1 Deployed on the original ES Lightweight digital twin deployment on RSU Up, ES and RSU distance Exceeded the threshold And the target ES with RSU distance At the threshold Within the range, the expression is: .
[0077] The two formulas above give the triggering conditions for full digital twin migration and lightweight digital twin migration. Since full digital twins and lightweight digital twins have real-time synchronization requirements, and the data synchronization between the two is achieved through wired transmission between ES and RSU, it is reasonable to determine whether the full digital twin of vehicle v needs to be migrated by measuring the Euclidean distance between the ES deployed in the full digital twin of vehicle v and the RSU deployed in its lightweight digital twin. Finally, it should be clarified that the triggering conditions are mainly used to limit the action space, and the final migration decision is optimized through reinforcement learning.
[0078] Step S2: The first optimization objective model is:
[0079]
[0080]
[0081]
[0082]
[0083]
[0084]
[0085]
[0086]
[0087]
[0088] ;
[0089] The optimization objective considers the migration utility function of the full digital twin and the lightweight digital twin. Regarding constraints, constraints C1 and C2 state that the full digital twin of each vehicle can only be deployed on one ES (Executable System) in the same time slot, and the lightweight digital twin of each vehicle can only be deployed on one RSU (Remote Unit) in the same time slot. Constraints C3 and C4 ensure that the computing power allocated to the ES and RSU cannot exceed their respective maximum computing power resources. Constraints C5 and C6 state that the migration latency of the full digital twin and the lightweight digital twin cannot exceed their respective thresholds. Constraints C7 and C8 state that the digital twins deployed on the ES or RSU cannot exceed the node's space capacity. Constraints C9 and C10 ensure the range of values for the binary correlation variables of the full digital twin migration decision and the lightweight digital twin migration decision. Constraint C11 ensures that the computing power allocation is a non-negative number.
[0090] S3: The transfer decision framework based on dual-timescale joint optimization transforms the first optimization objective model into the second optimization objective model; It should be noted that the specific content of the dual-timescale joint optimization migration decision framework is as follows: optimize the migration and RSU resource allocation of lightweight digital twins in short time slots, and optimize the migration and ES resource allocation of full digital twins in long time slots. This framework achieves global optimization of resource utilization, migration latency, and migration utility through decisions at different time scales. When a vehicle moves on the road, its lightweight digital twin is deployed on the RSU node closest to the vehicle for real-time perception and immediate decision-making; simultaneously, the full digital twin is deployed on ES nodes with base stations to support global-level traffic optimization, collaborative decision-making, and city-level control. Vehicle movement triggers rapid migration of the lightweight digital twin between RSUs and may trigger migration of the full digital twin between ES nodes. The migration of the lightweight digital twin follows small-scale vehicle movement events, thus involving short-time domain decisions; the migration of the full digital twin involves changes in ES-RSU distance and global consistency maintenance, thus involving long-time domain decisions, such as... Figure 2 As shown, the two have a mutually influential relationship: the deployment location of the full digital twin affects the migration triggering and task load scale of the lightweight digital twin, while the cumulative migration utility of the lightweight digital twin affects the global utility and triggering of the full digital twin.
[0091] The second optimization objective model is:
[0092]
[0093]
[0094]
[0095]
[0096]
[0097]
[0098]
[0099]
[0100]
[0101] ;
[0102] The optimization objective is to maximize the lightweight digital twin transfer utility function within a short time slot t. In long time slots Maximize the transfer utility function of the entire digital twin Among them, long time slots and short time slots should meet the following requirements. The relationship is complex. However, the above optimization objective is a mixed-integer nonlinear problem, belonging to the NP-hard category, and also has dual time-scale coupling. Therefore, it is necessary to decompose the optimization objective to reduce computational complexity and ensure the maximization of the system's long-term benefits through alternative solutions.
[0103] S4: Use an alternating iterative strategy and a preset algorithm to solve the second optimization objective model and obtain the optimal migration decision.
[0104] It should be noted that the specific content of the second optimization objective model to be solved is as follows: construct the first subproblem to optimize the migration of the vehicle's lightweight digital twin between RSUs in short time slots; construct the second subproblem to optimize the migration of the vehicle's full digital twin between ESs in long time slots; solve the first and second subproblems using a preset algorithm and iterate alternately based on the time scale until the global objective function converges.
[0105] The specific content of the first subproblem is as follows: By determining the deployment strategy of the vehicle's lightweight digital twin in short time slot t and the allocation of migration computing resources, the relationship between latency and energy consumption during the migration of the lightweight digital twin between RSUs is balanced, so as to achieve the optimal lightweight digital twin migration strategy under short time slots. The optimization objective and related constraints are defined as follows:
[0106]
[0107]
[0108]
[0109]
[0110] ;
[0111] The optimization objective is to maximize the long-term average migration utility of lightweight digital twins. C2 represents the lightweight digital twin deployment constraint per vehicle per RSU in short time slot t, C4 represents the upper limit of RSU computing power, C6 represents the migration latency threshold constraint, and C8 represents the upper limit constraint of RSU storage capacity.
[0112] The specific steps for solving the first subproblem are as follows: The first subproblem is modeled using MDP to obtain a global state space, an action space, and a reward function. The global state space is the set of the vehicle's observation space. The action space includes two actions: discrete migration decision and continuous computing power allocation decision. The reward function includes rewards for two performance indicators: migration delay and transmission energy. In a specific embodiment, the specific modeling process of MDP is as follows: First, the observation space is set. The observation state of each agent is a 5-dimensional vector, mainly observing factors such as the communication latency between the vehicle and the RSU, the deployment location of the vehicle's lightweight digital twin, the computing resources and spatial capacity of the RSU, etc., to ensure that the observation space of each agent can fully reflect the state of the agent in the current time slot. Therefore, the vehicle The observation space is defined as:
[0113] in, Indicates vehicle Compared with the currently deployed RSU ( The normalized value of the communication delay can reflect the connection quality between the vehicle and the current RSU; in addition, Indicates the index of the currently deployed RSU, for example when The time indicates the vehicle in the current time slot. Lightweight digital twin deployment on RSU Above, that is The third observation value is This indicates the RSUs (denoted as RSUs) deployed in the current vehicle lightweight digital twin. The normalized result of the used computing power as a percentage of its maximum computing power, with values strictly limited to [0,1], is used to quantify the current computing power resource stress of the RSU. This indicates that under the current time slot t, all lightweight digital twins will be deployed in vehicles (collection) The total computing power used for packaging (unit: cycles / s). Represents a single vehicle exist The computing power obtained from the packaging This represents the maximum computing power limit for a single RSU; similarly, the fourth observation is... The current vehicle lightweight digital twin deployment of RSU (denoted as The normalized result of the used memory space of the RSU is the proportion of its maximum memory, with the value strictly limited to [0,1]. It is used to quantify the memory space tightness of the current RSU. This indicates that under the current time slot t, all lightweight digital twins will be deployed in vehicles (collection) The amount of memory occupied, Indicates vehicle exist Lightweight digital twin data volume, Indicates the maximum memory limit for a single RSU; finally, Indicates vehicle v The number of CPU cycles required to compute a lightweight digital twin model. Based on the definition of the observation space, the global state space is the set of all agents' observation spaces, defined as:
[0114] Next is the design of the action space, which includes two actions: discrete migration decision and continuous computing power allocation decision. Therefore, the action space is defined as follows:
[0115] in, Indicates vehicle v exist t The migration decision action for a time slot is a discrete value. For example, when When, it indicates the vehicle v Lightweight digital twins in t Time slot determines migration to RSU In other words, it refers to vehicles in the T+1 time slot. v Lightweight digital twin deployment on RSU Above, that is Equivalent to In particular, when This indicates that regardless of whether it is in time slot t or time slot t+1, the vehicle v Lightweight digital twin deployment on RSU That means the vehicles in these two time slots... v The lightweight digital twins are all deployed on the same RSU, i.e., the vehicle. v The lightweight digital twin did not migrate. This represents the decision action for allocating computing power in a package, and is a continuous value. For example, when a vehicle... v Lightweight digital twin decision from t time slot RSU RSUs migrated to slot t+1 At that time, that is At this time, RSU The allocated computing power is In particular, when When, it indicates the vehicle v The lightweight digital twin did not migrate; it was deployed on the RSU in both time slots t and t+1. Above, the computing power allocation at this time is .
[0116] Finally, regarding the design of the reward function, which includes rewards for two performance metrics: migration latency and transmission power consumption, the reward function is defined as follows:
[0117] The first item is a lightweight digital twin migration latency reward. , among them The first term represents the maximum latency threshold for lightweight digital twin migration; the lower the latency, the closer the reward value is to 1. Similarly, the second term represents the energy consumption reward for lightweight digital twin transmission. , among them The maximum energy consumption threshold for lightweight digital twin transmission, where This represents the largest data volume among all lightweight digital twins. The first term represents the maximum Euclidean distance between RSUs; the lower the energy consumption, the closer the reward value is to 1. The third term characterizes the effectiveness of lightweight digital twin migration.
[0118] in, Indicates short time slots t In the middle, vehicles v The lightweight digital twin deployed RSU at the current moment With vehicles v Communication delay between them This indicates that the RSU and the vehicle v The maximum communication latency between the lightweight digital twin and the vehicle. This reward is designed based on the communication latency between the lightweight digital twin and the vehicle after the migration decision, and to some extent represents the communication quality between the lightweight digital twin and the vehicle, while also further indicating the effectiveness of the lightweight digital twin migration. Additionally, weights... , and satisfy This normalizes the reward value.
[0119] After completing the MDP modeling, the MAPPO algorithm based on RSU action perception is used to solve the MDP model to obtain the optimal lightweight digital twin transfer strategy. The steps of the MAPPO algorithm based on RSU action awareness include: Initialize the parameters of the Actor network responsible for outputting the transition action, and the parameters of the Critic network responsible for evaluating the long-term value of the state; In each training round, the short time slots are traversed one by one. In each time slot, all vehicles obtain their own local observations and calculate the RSU action feasibility mask. The Actor network combines the mask to output the action of each vehicle. After all vehicle actions are output, they are executed in a unified manner to obtain the global state of the next time slot. The reward value of each vehicle is calculated, and the current global state, action, global state of the next time slot, and reward value are recorded and stored in the sample set. Update the Actor network parameters and Critic network parameters based on the collected samples; Repeat the above training rounds until the network converges to obtain the optimal lightweight digital twin transfer strategy.
[0120] In a specific embodiment, the design details of the MAPPO algorithm based on RSU motion awareness are as follows: Subproblem P1 addresses frequent migration and computational allocation decisions for lightweight digital twins within short time slots, exhibiting the following typical characteristics: it's a medium-to-large-scale multi-agent scenario; the action space is a hybrid of discrete RSU migration and continuous computational allocation; significant RSU computation and storage constraints render many migration actions infeasible in most time slots; and finally, the high decision frequency places stringent demands on algorithm convergence speed and stability. Therefore, MAPPO is chosen as the basic framework, and to address the resource-constrained nature of P1, an RSU-aware action masking mechanism is introduced into the discrete action decision-making of the Actor network to reduce invalid exploration and improve sample utilization efficiency. The definition of the Actor network policy function is as follows:
[0121]
[0122] in, This refers to discrete lightweight digital twin transfer actions. To facilitate the visualization of the mathematical formulas below, variables are used. (in To represent discrete lightweight digital twin migration actions, a feasibility mask is defined, taking into account the deterministic nature of RSU resource constraints. As shown below:
[0123] If and only if RSU r In the time slot t The value is 1 when the constraints of computing power, storage, and latency are met. Based on this mask, the discrete migration strategy is defined as:
[0124] in, It follows the Softmax strategy. For continuous computing power allocation actions, Adopting a Gaussian strategy, following a Gaussian distribution The definition of the centralized value function of the Critic network is as follows:
[0125] in, This is the discount factor for short time slots. The Actor's training method is the same as the PPO algorithm, updating network parameters by calculating the advantage function. Advantage function This indicates the degree of superiority or inferiority of the action. This indicates that the action is superior, therefore the output probability of that action is increased; the larger the advantage value, the greater the increase. Therefore, it is necessary to reduce its output probability. Based on the TD error, the advantage function is calculated using GAE (Generalized Advantage Estimation). The function is defined as follows:
[0126] in, The smoothing coefficient for the advantage estimate. The single-step time error within a short time slot can be expressed as:
[0127] The strategy is continuously iterated and updated during training. The data used for training in this process is sampled from the old strategy. If the differences between the old and new strategies are significant, reusing this old data will lead to large policy updates and oscillations. To reuse the sampled data and improve data utilization, the PPO algorithm employs a pruning function to limit the range of change between the old and new strategies. The loss function of the Actor network is defined as:
[0128] in, This represents the difference in strategy before and after the update. (Pruning function) Indicates will Limited to Between, where parameters The pruning factor is used to limit the update magnitude of the policy, thus achieving small-step updates. The core function of the Critic network's loss function is to minimize the deviation between the network's estimate of the current state's value and the Bellman objective value. Through gradient descent optimization using mean squared error, the Critic network continuously learns and fits a more accurate true value of the state (i.e., the mathematical expectation of the long-term cumulative reward in the state), providing a reliable value benchmark for the Actor network to calculate the advantage function and update the policy, ultimately achieving policy optimization and convergence. Therefore, the Critic's loss function is defined as:
[0129] The pseudocode for the algorithm is as follows: 1: Initialize Actor network parameters and Critic network parameters
[0130] 2: for episode=0, 1, 2, ..., MAX_EPISODE do: 3: for do: 4: for do: 5: Obtain local observations Calculate the RSU action feasibility mask
[0131] 6: Output Action
[0132] 7: endfor 8: Perform the action To obtain global observations at the next moment.
[0133] 9: Calculate Rewards and record , , ,
[0134] 10:endfor 11: According to the definition renew
[0135] 12: According to the definition renew
[0136] 13: endfor It should be noted that the specific content of the second sub-problem is as follows: By determining long time slots ( The deployment strategy and migration computing resource allocation for the full digital twin of vehicles in the system are used to balance the latency and energy consumption of migrating the full digital twin between ES. The optimization objectives and constraints are defined as follows:
[0137]
[0138]
[0139]
[0140]
[0141]
[0142]
[0143] The optimization objective of the second subproblem is to maximize the long-term average migration utility of the entire digital twin, where C1 represents the migration utility of each ES in a long time slot. Constraints for the full digital twin deployment of each vehicle: C3 represents the upper limit constraint of RSU computing power, C5 represents the migration latency threshold constraint, and C8 represents the upper limit constraint of ES storage capacity.
[0144] The specific steps for solving the second subproblem are as follows: The second subproblem is modeled using MDP to obtain a global state space, an action space, and a reward function. The global state space is the set of the vehicle's observation space. The action space includes two actions: discrete migration decision and continuous computing power allocation decision. The reward function includes rewards for two performance indicators: migration delay and transmission energy. In a specific embodiment, the MDP modeling of the second sub-problem is as follows: First, the observation space is set. The observation state of each agent is a 5-dimensional vector. The main observations include the distance between the RSU deployed in the lightweight digital twin of the vehicle and the ES deployed in the full digital twin, the deployment location of the full digital twin of the vehicle, the computing resources and spatial capacity of the ES, and other related factors. This ensures that the observation space of each agent can fully reflect the state of the agent in the current time slot. Therefore, the vehicle... The observation space is defined as:
[0145] in, Indicates vehicle Lightweight digital twins currently deploy RSUs ( ) and the ES deployed by the full digital twin ( The Euclidean distance normalized value can reflect the real-time synchronization status between the full digital twin and the lightweight digital twin. This is the maximum Euclidean distance between ES and RSU; additionally, Indicates vehicle v The full digital twin of the currently deployed Elasticsearch index, for example when The time indicates the vehicle in the current time slot. Full digital twin deployment on ES Above, that is The third observation value is This indicates the ES (denoted as) deployed in the current full-scale digital twin of the vehicle. The normalized result of the used computing power as a percentage of its maximum computing power, with values strictly limited to [0,1], is used to quantify the current computing power resource stress of the Elasticsearch (ES). Indicates the current time slot Below, all will be fully deployed digital twins on vehicles (collection) The total computing power used for packaging (unit: cycles / s). Represents a single vehicle exist The computing power obtained from the packaging This represents the maximum computing power limit for a single Elasticsearch Engine (ES); similarly, the fourth observation is... The current full-scale digital twin deployment of vehicles uses ES (denoted as The normalized result of the used memory space of ) is the maximum memory space occupied by its maximum memory, with the value strictly limited to [0,1], and is used to quantify the current memory space tension of ES. Indicates the current time slot Below, all will be fully deployed digital twins on vehicles (collection) The amount of memory occupied, Indicates vehicle exist The total amount of digital twin data on the platform Indicates the maximum memory limit for a single Elasticsearch instance; finally, Indicates vehicle v The number of CPU cycles required to compute the full digital twin model. Based on the definition of the observation space, the global state space is the set of all agents' observation spaces, defined as:
[0146] Next is the design of the action space, which includes two actions: discrete migration decision and continuous computing power allocation decision. Therefore, the action space is defined as follows:
[0147] in, Indicates vehicle v exist The full digital twin migration decision action for a time slot is a discrete value. For example, when When, it indicates the vehicle v The full digital twin in Time slot determines migration to ES Above, in other words, is at Time-slot vehicles v Full digital twin deployment on ES Above, that is Equivalent to In particular, when At that time, it indicates that regardless of The time slot is still there Time slot, vehicle v Full digital twin deployment on ES That means the vehicles in these two time slots... v All digital twins are deployed on the same ES, i.e., the vehicle v The full digital twin did not migrate. This represents the decision action for allocating computing power in a package, and is a continuous value. For example, when a vehicle... v The full digital twin decision from ES of time slots Migrate to ES of time slots At that time, that is At this time, ES The allocated computing power is In particular, when When, it indicates the vehicle v The lightweight digital twin did not migrate, in Time slots and The time slots are all deployed on ES Above, the computing power allocation at this time is .
[0148] Finally, regarding the design of the reward function, which includes rewards for two performance metrics: migration latency and transmission power consumption, the reward function is defined as follows:
[0149] The first item is a reward for latency during the full digital twin migration. , among them The maximum threshold for migration latency is defined; the smaller the latency, the closer the reward value is to 1. Similarly, the second term is the energy consumption reward for full digital twin transmission. , among them The maximum threshold for transmission energy consumption, where This represents the largest data volume among all full-scale digital twins. The first term represents the maximum Euclidean distance between ESs; the lower the energy consumption, the closer the reward value is to 1. The third term characterizes the effectiveness of the full digital twin migration.
[0150] in, Indicates long time slot In the middle, vehicles v The full digital twin deployed at the current moment in ES Its lightweight digital twin deployed RSU The Euclidean distance between them This represents the maximum Euclidean distance between ES and RSU. This reward is designed based on the distance between the full digital twin and the lightweight digital twin after the migration decision, representing to some extent the synchronization quality between the full and lightweight digital twins, and further indicating the migration effectiveness of the full digital twin. Additionally, the weight... , and satisfy This normalizes the reward value.
[0151] After completing the MDP modeling, the MAPPO algorithm based on lightweight twin guidance is used to solve the MDP model to obtain the optimal full digital twin migration strategy. The steps of the MAPPO algorithm based on lightweight twin guidance include: Initialize the parameters of the Actor network responsible for outputting the transition action, and the parameters and long time slots of the Critic network responsible for evaluating the long-term value of the state. ; In each training round, short time slots are traversed one by one. When the cumulative number of short time slots traversed reaches a preset threshold, all vehicles acquire their own local observations. The Actor network, combined with a mask, outputs the migration target selection (ES) action and the actions of packetized computing power allocation. After all vehicle actions are output, they are executed uniformly to obtain the global state of the next long time slot. The reward value of each vehicle is calculated, and the current global state, action, global state of the next long time slot, and reward value are recorded and stored in the sample set. Add 1; After all time slots of a training round have been traversed, the standard dominance function is calculated based on the collected samples using the GAE method. Then, the spatial consistency index between the full twin and the lightweight twin is calculated. The standard dominance function is weighted based on the consistency index to obtain the weighted dominance function guided by the lightweight twin. The parameters of the Actor network and the Critic network are then updated. Repeat the above training rounds until the network converges to obtain the optimal full-scale digital twin transfer strategy.
[0152] In a specific embodiment, the MAPPO algorithm based on lightweight twin bootstrapping is designed as follows: Subproblem P2 focuses on the migration decisions of the full digital twin over long time slots. Its goal is to maximize the long-term operational benefits of the system and maintain the stability of the full digital twin deployment, considering migration overhead and resource constraints. Unlike lightweight twins, which change rapidly within short time slots, the migration cost of a full digital twin is significantly higher. Its optimal decision depends not only on the system state in the current time slot but also on the spatial distribution and load evolution of lightweight twins on the RSU side across multiple short time slots. However, the standard MAPPO algorithm, in its long-time slot decision modeling, typically estimates long-term returns based on instantaneous or aggregated states, making it difficult to explicitly characterize the cumulative impact of frequent migrations of the lightweight twin layer within short time slots on the long-term benefits and deployment stability of the full digital twin. If this cross-time scale and cross-layer coupling is not effectively modeled, the full digital twin migration strategy can become overly sensitive to short-term fluctuations, leading to unnecessary migration oscillations and reduced overall system performance. Based on the above considerations, this paper introduces a lightweight twin-guided advantage-weighted MAPPO algorithm in subproblem P2. By integrating the state consistency and evolutionary characteristics of the lightweight twin layer into the construction process of the advantage function, the algorithm guides policy updates to focus more on decision directions that are conducive to the long-term stable deployment of the full twin, thereby improving policy convergence stability and decision robustness while ensuring long-term gains. The definition of the Actor network policy function is as follows:
[0153] in, It is a discrete, fully quantized digital twin transfer action. It also follows the Softmax strategy. Similarly, for the allocation of continuous computing power, Adopting a Gaussian strategy, following a Gaussian distribution The definition of the centralized value function of the Critic network is as follows:
[0154] in, This is the discount factor for long time slots. The Actor training method is the same as the PPO algorithm, updating network parameters by calculating the advantage function. Based on the TD error, the advantage function is also calculated using GAE. The function is defined as follows:
[0155] in, The smoothing coefficient for the advantage estimate. The single-step time error in a long time slot can be expressed as:
[0156] However, the dominance function in this form In policy updates, all long-slot decision samples are treated equally, without distinguishing their consistency differences with the evolution results of the lightweight digital twin layer. To explicitly characterize the impact of the deployment state formed by the lightweight digital twin (P1) within short slots on the long-term decisions of the full digital twin (P2), this paper introduces a consistency index between the lightweight and full digital twins:
[0157] in, Indicates vehicle v In long time slots The associated full-scale digital twin deployment of ES location, Indicate that subproblem P2 is in a long time slot Observable vehicles during decision-making v The immediate deployment status of the lightweight twin, which is the location of the latest RSU deployed by the lightweight twin. This represents the Euclidean distance between ES and RSU. This is the largest Euclidean distance between ES and RSU, an metric used to quantify the spatial and network consistency between full twin deployment decisions and lightweight twin service relationships. In the standard dominance function... Based on this, a lightweight twin-guided weighted advantage function is constructed. :
[0158] in, The standard GAE advantage is that it can reflect the degree of improvement in long-term returns relative to the baseline value function. The consistency weights between lightweight and full twins are defined. Based on the weighted advantage, the PPO policy optimization objective (loss function of the Actor network) for subproblem P2 is modified as follows:
[0159] in, This represents the difference in strategy before and after the update. (Pruning function) Indicates will Limited to Between, where parameters This is the pruning factor, used to limit the update magnitude of the strategy, thus achieving small-step updates. Through... The deployment statistics generated by the lightweight twin in P1 within a short time slot are explicitly injected into the policy update process of P2, compensating for the shortcomings of standard MAPPO in cross-timescale modeling. Furthermore, the full twin transfer decision, inconsistent with the lightweight twin deployment, is given less weight in gradient updates, thereby reducing unnecessary transfers caused by short-term fluctuations. Finally, policy updates focus more on samples conducive to long-term stable deployment, effectively reducing the variance of policy gradients over long time scales and improving training stability. Similar to the loss function definition for short-time-slot Critic networks, the loss function for long-time-slot Critic networks is defined as follows:
[0160] The pseudocode for the algorithm is as follows: 1: Initialize Actor network parameters and Critic network parameters ,
[0161] 2: for episode=0, 1, 2, ..., MAX_EPISODE do: 3: for do: 4:if do: 5: for do: 6: Obtain local observations Output action
[0162] 7:endfor 8: Execution of actions To obtain global observations at the next moment.
[0163] 9: Calculate Rewards and record , , ,
[0164] 10:
[0165] 11:endif 12:endfor 13: Calculate the dominance function based on GAE Calculate the consistency index
[0166] 14: According to Calculate the weighted advantage function of lightweight twin guidance
[0167] 15: According to the definition and Update separately and
[0168] 16:endfor This invention models the transfer decision-making of vehicle edge digital twins as a multi-agent reinforcement learning problem, allowing each vehicle twin to participate in the decision-making process as an independent agent. Through collaborative learning among multiple agents, the transfer strategy can continuously self-adjust in complex environments such as high-speed vehicle movement, dynamic changes in network state, and concurrent competition for edge resources among multiple vehicles, avoiding the negative impact of local optimal decisions on the overall system performance. Simultaneously, by employing a centralized training and distributed execution approach, the strategy fully utilizes global information to improve policy convergence quality during the training phase, while relying solely on local observations for decision-making during the execution phase, enhancing the deployability and robustness of the transfer strategy in real-world urban vehicle-to-everything (V2X) environments.
[0169] To address the significant differences between lightweight digital twins and full digital twins in terms of migration frequency, resource consumption, and service objectives, this invention introduces a dual-timescale collaborative optimization mechanism in the migration decision-making process. This mechanism makes decisions and controls the migration behavior of the two types of digital twins at different time scales. The short-term time scale focuses on quickly responding to changes in vehicle location and local service needs, ensuring the real-time performance of the digital twin. The long-term time scale, however, considers the overall system operation and performs global optimization of the migration behavior, avoiding resource waste caused by frequent migrations. Compared to traditional migration strategies that do not differentiate between time scales, this method effectively reduces the oscillation of migration decisions, meeting real-time service requirements while also considering long-term system performance improvements.
[0170] In real-world urban vehicle-to-everything (V2X) scenarios, the computing resources and capacity of edge servers are typically limited. Existing digital twin migration methods often fail to adequately consider server overload during the decision-making process, potentially leading to decreased service quality or even twin failure after migration. This invention explicitly incorporates the computing resource constraints and capacity limitations of edge servers into the decision space during multi-agent reinforcement learning decision-making, restricting migration actions that might trigger resource overload. This ensures that every migration decision meets system resource feasibility requirements. This approach reduces the probability of infeasible migrations at the decision-making mechanism level, improving the stability and reliability of the digital twin migration process in multi-vehicle concurrent scenarios, enabling the system to operate continuously and stably under resource-constrained conditions.
[0171] Example 2 This embodiment provides a description of a two-layer digital twin migration scenario for vehicles and the relationship between the MAPPO framework for vehicle digital twin migration based on dual time scales and the migration scenario, including the following: like Figure 3 As shown, the full digital twin of vehicle V1 is deployed in the ES (Edge Array) of the base station, while the lightweight digital twin of the vehicle is deployed in the RSU (Roadside Unit). As vehicle V1 moves, in order to ensure continuous high-quality service from the digital twin, the digital twin needs to migrate to a new edge node to ensure real-time interaction with the vehicle. From time slot t to time slot t+1, vehicle V1 moves away from RSU i and continuously approaches RSU j. At this time, the lightweight digital twin of vehicle V1 needs to migrate from RSU i to RSU j, while the full digital twin of V1 does not migrate. From time slot t+1 to time slot t+2, vehicle V1 gradually moves away from RSU j and approaches RSU k. At this time, the lightweight digital twin of V1 needs to migrate from RSU j to RSU k. Simultaneously, to ensure the digital twin's ability to dynamically optimize and intelligently decide on vehicles, traffic flow, and urban road networks, the full digital twin of V1 needs to migrate from ES i to ES j. In the above process, two processes need to be considered: the decision-making process for the migration of the vehicle lightweight digital twin and the decision-making process for the migration of the full digital twin.
[0172] Figure 4 This paper presents the relationship between the MAPPO algorithm architecture and the migration environment for two-layer digital twin migration of vehicles under dual time scales. It includes collaborative optimization between two MAPPO modules based on the dual time scales, corresponding to the migration decision-making processes for full digital twins and lightweight digital twins, respectively. The upper layer is a full-twin MAPPO module guided by lightweight twins, oriented towards long time slots. The entire twin state sequence is stored through an experience buffer. After sampling, the data is input into a centralized Critic network and a multi-agent Actor network. A weighted advantage function is calculated by combining the consistency index of lightweight twins and full twins, and the hybrid action of the full twin transfer decision is output. This drives its on-demand migration between edge servers; the lower layer is the lightweight twin MAPPO module (MAPPO Based on RSU Action Perception), which is designed for short time slots t and stores lightweight twin state sequences through an empirical cache. After sampling, the input is combined with the Critic network and the Actor network which incorporates an RSU action feasibility mask, and the output is a hybrid action of lightweight twin transfer decision. This enables high-frequency migration between roadside units. Furthermore, the long-slot decision integrates the lightweight twin deployment states of n short-slots as input, and the deployment results of the full twin influence the migration triggering conditions of the lightweight twin, forming a closed loop of real-time response in short time slots and global optimization in long time slots.
[0173] Example 3 This embodiment provides a specific implementation method for the migration of full and lightweight digital twin models, and the implementation flowchart is as follows. Figure 5 As shown, it includes the following steps: Step 1: Construct the migration utility function for lightweight digital twin migration of vehicles and the migration utility function for full digital twin migration of vehicles. Both functions comprehensively consider migration latency and migration energy consumption during the digital twin migration process, so as to improve the migration quality of all digital twins in the long-term operation process as much as possible.
[0174] Step 2: Lightweight digital twins typically migrate frequently as vehicles move at small scales, characterized by high decision-making frequency and strict response time requirements. Full-scale digital twin migration is costly and serves more global collaboration and long-term optimization, with a relatively low decision-making frequency. A dual-timescale collaborative optimization framework is introduced to decouple the original optimization problem into a lightweight digital twin migration decision subproblem P1 and a full-scale digital twin migration decision subproblem P2.
[0175] Step 3: For the high-frequency migration characteristics between RSUs in a short time slot, subproblem P1 is modeled using MDP, and a MAPPO-based lightweight digital twin migration decision algorithm based on RSU action perception is designed. This algorithm uses the vehicle as the agent, with each agent employing an independent policy network. The algorithm iteratively updates the value of each global state using a centralized value network (Critic) and the migration decision is output by the policy network (Actor), until convergence, ultimately obtaining the optimal migration strategy for the vehicle's lightweight digital twin.
[0176] Step 4: In long time slots, considering the relatively low migration frequency of the full digital twin of the vehicle between ES and the characteristics of global coordination, MDP modeling is performed on subproblem P2, and a MAPPO digital twin migration decision algorithm based on lightweight twin guidance is designed. Similarly, through continuous iteration and updates until convergence, the optimal migration strategy of the vehicle's lightweight digital twin is finally obtained.
[0177] Step 5: The edge nodes will promptly construct a digital twin of the device from the migrated digital twin, and combine it with the data uploaded by the corresponding vehicle to update the digital twin for the matching vehicle, thereby completing the digital twin migration.
[0178] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, and are not intended to limit the implementation of the present invention. Those skilled in the art can make other variations or modifications based on the above description. It is neither necessary nor possible to exhaustively describe all embodiments here. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the claims of the present invention.
Claims
1. A transfer method for full and lightweight digital twin models, characterized in that, Includes the following steps: S1: Construct a digital twin migration model, which includes a lightweight digital twin migration model and a full digital twin migration model; S2: Construct the first optimization target model based on the digital twin transfer model; S3: The transfer decision framework based on dual-timescale joint optimization transforms the first optimization objective model into the second optimization objective model; S4: Use an alternating iterative strategy and a preset algorithm to solve the second optimization objective model and obtain the optimal migration decision.
2. The transfer method for full and lightweight digital twin models according to claim 1, characterized in that, The specific content of the lightweight digital twin transfer model described in step S1 is as follows: The migration process of lightweight digital twins consists of two steps. The roadside unit is denoted as RSU. The first step is for the RSU to package the vehicle's digital twin. The packaging latency at this stage is defined as: in, It is a binary correlation variable representing vehicles. v Lightweight digital twin t- Is slot 1 deployed on RSU? superior, Indicates deployment on RSU superior, Indicates vehicle v The number of CPU cycles required for lightweight digital twin packaged computing. RSU Assigned to vehicles v The computing power required for lightweight digital twin packaging; the relationship between the computing resource requirements of lightweight digital twin packaging and the data volume of its digital twin is as follows: ; The second process involves lightweight digital twins in link transmission. The transmission latency and power consumption of this process are defined as follows: in, Indicates the transmission latency of lightweight digital twins. Indicates the power consumption of lightweight digital twin transmission. Indicates vehicle v The data volume of a lightweight digital twin This represents the time coefficient for transmitting a unit of data per unit distance. RSU and RSU The Euclidean distance between them Average wired transmission power; vehicle v The latency and transfer utility function required for lightweight digital twin migration are defined as follows: in, This represents the maximum threshold for lightweight digital twin migration latency. Then, this represents the utility function for lightweight digital twin transfers; The specific content of the full digital twin transfer model described in step S1 is as follows: The migration process of the full digital twin is divided into two steps. The edge server is denoted as ES. The first step is that ES packages the vehicle's digital twin. The packaging latency at this stage is defined as: in, Indicates vehicle v The full digital twin is deployed in time slot t-1 on ES superior, Indicates vehicle v The number of CPU cycles required for full digital twin packaged computation. Indicates ES Assigned to vehicles v The computing power required for packaging the entire digital twin model, and the relationship between the computing resource requirements for packaging the entire digital twin model and the data volume of its digital twin, are as follows: , The first process is the packetization period coefficient; the second process is the full digital twin transmission over the link, where the transmission delay and energy consumption are defined as follows: in, Indicates the full digital twin transmission latency. This indicates the energy consumption of the full digital twin transmission. Indicates vehicle v The relationship between the total data volume of the full digital twin, the lightweight digital twin data volume of the vehicle, and the total data volume of the full digital twin is as follows: , This represents the time coefficient for transmitting a unit of data per unit distance. Indicates ES and ES Euclidean distance between vehicles v The latency and transfer utility function required for a full digital twin transfer are defined as follows: in, This represents the maximum threshold for the full digital twin migration latency. This represents the full digital twin transfer utility function.
3. The method for transferring full and lightweight digital twin models according to claim 2, characterized in that, Before constructing the first optimization target model, it is necessary to construct migration triggering conditions, which include triggering conditions for full digital twin migration and triggering conditions for lightweight digital twin migration. The triggering condition for lightweight digital twin migration is: t- Vehicles in time slot 1 v With the original RSU Communication latency Exceeding the threshold And in t Vehicles in time slots v With target RSU Communication latency At the threshold Within the range, the expression is: The trigger condition for a full digital twin migration is: vehicle v The full digital twin in time slots t-1 Deployed on the original ES Lightweight digital twin deployment on RSU Up, ES and RSU distance Exceeded the threshold And the target ES with RSU distance At the threshold Within the range, the expression is: 。 4. The method for transferring full and lightweight digital twin models according to claim 3, characterized in that, Step S2: The first optimization objective model is: ; The optimization objective considers the migration utility function of the full digital twin and the lightweight digital twin. Regarding constraints, constraints C1 and C2 state that each vehicle's full digital twin can only be deployed on one ES (Executable System) in the same time slot, and each vehicle's lightweight digital twin can only be deployed on one RSU (Remote Unit) in the same time slot. Constraints C3 and C4 ensure that the computing power allocated to ES and RSU cannot exceed their respective maximum computing power resources. Constraints C5 and C6 state that the migration latency of the full digital twin and the lightweight digital twin cannot exceed their respective thresholds. Constraints C7 and C8 state that the digital twins deployed on ES or RSU cannot exceed the corresponding node space capacity. Constraints C9 and C10 ensure the range of values for the binary correlation variables of the full digital twin migration decision and the lightweight digital twin migration decision. Constraint C11 ensures that the computing power allocation is a non-negative number.
5. The method for transferring full and lightweight digital twin models according to claim 4, characterized in that, The specific content of the migration decision framework jointly optimized by the two time scales is as follows: optimize the migration and RSU resource allocation of lightweight digital twins in short time slots, and optimize the migration and ES resource allocation of full digital twins in long time slots. The second optimization objective model is: ; The optimization objective is to maximize the lightweight digital twin transfer utility function within a short time slot t. In long time slots Maximize the transfer utility function of the entire digital twin .
6. A method for transferring full and lightweight digital twin models according to claim 1 or 5, characterized in that, The specific content of the second optimization objective model to be solved is as follows: construct the first subproblem to optimize the migration of the vehicle's lightweight digital twin between RSUs in short time slots; construct the second subproblem to optimize the migration of the vehicle's full digital twin between ESs in long time slots; solve the first and second subproblems using a preset algorithm and iterate alternately based on the time scale until the global objective function converges.
7. The transfer method for full and lightweight digital twin models according to claim 6, characterized in that, The specific content of the first subproblem is as follows: By determining the deployment strategy of the vehicle's lightweight digital twin in short time slot t and the allocation of migration computing resources, the relationship between latency and energy consumption during the migration of the lightweight digital twin between RSUs is balanced. The optimization objective and related constraints are defined as follows: ; The optimization objective is to maximize the long-term average migration utility of lightweight digital twins. C2 represents the lightweight digital twin deployment constraint per vehicle per RSU in short time slot t, C4 represents the upper limit of RSU computing power, C6 represents the migration latency threshold constraint, and C8 represents the upper limit constraint of RSU storage capacity.
8. The method for transferring full and lightweight digital twin models according to claim 7, characterized in that, The specific steps for solving the first subproblem are as follows: A Markov decision process model is established for the first subproblem to obtain the global state space, action space and reward function. The global state space is the set of the vehicle's observation space. The action space includes two actions: discrete migration decision and continuous computing power allocation decision. The reward function includes rewards for two performance indicators: migration delay and transmission energy. The Markov decision process model was solved using the MAPPO algorithm based on RSU action awareness to obtain the optimal lightweight digital twin transfer strategy. The steps of the MAPPO algorithm based on RSU action awareness include: Initialize the parameters of the Actor network responsible for outputting the transition action, and the parameters of the Critic network responsible for evaluating the long-term value of the state; In each training round, the short time slots are traversed one by one. In each time slot, all vehicles obtain their own local observations and calculate the RSU action feasibility mask. The Actor network combines the mask to output the action of each vehicle. After all vehicle actions are output, they are executed in a unified manner to obtain the global state of the next time slot. The reward value of each vehicle is calculated, and the current global state, action, global state of the next time slot, and reward value are recorded and stored in the sample set. Update the Actor network parameters and Critic network parameters based on the collected samples; Repeat the above training rounds until the network converges to obtain the optimal lightweight digital twin transfer strategy.
9. The method for transferring full and lightweight digital twin models according to claim 6, characterized in that, The specific content of the second subproblem is as follows: By determining long time slots The deployment strategy for the full digital twin of vehicles and the allocation of migration computing resources are used to balance the relationship between latency and energy consumption during the migration of the full digital twin between ES. The optimization objectives and constraints are defined as follows: The optimization objective of the second subproblem is to maximize the long-term average migration utility of the entire digital twin, where C1 represents the migration utility of each ES in a long time slot. Constraints for the full digital twin deployment of each vehicle: C3 represents the upper limit constraint of RSU computing power, C5 represents the migration latency threshold constraint, and C8 represents the upper limit constraint of ES storage capacity.
10. The transfer method for full and lightweight digital twin models according to claim 9, characterized in that, The specific steps for solving the second subproblem are as follows: A Markov decision process model is established for the second subproblem to obtain the global state space, action space and reward function. The global state space is the set of the vehicle's observation space. The action space includes two actions: discrete migration decision and continuous computing power allocation decision. The reward function includes rewards for two performance indicators: migration delay and transmission energy. The Markov decision process model was solved using the MAPPO algorithm based on lightweight twin guidance to obtain the optimal full digital twin transfer strategy. The steps of the MAPPO algorithm based on lightweight twin guidance include: Initialize the parameters of the Actor network responsible for outputting the transition action, and the parameters and long time slots of the Critic network responsible for evaluating the long-term value of the state. ; In each training round, short time slots are traversed one by one. When the cumulative number of short time slots traversed reaches a preset threshold, all vehicles acquire their own local observations. The Actor network, combined with a mask, outputs the migration target selection (ES) action and the actions of packetized computing power allocation. After all vehicle actions are output, they are executed uniformly to obtain the global state of the next long time slot. The reward value of each vehicle is calculated, and the current global state, action, global state of the next long time slot, and reward value are recorded and stored in the sample set. Add 1; After all time slots of a training round have been traversed, the standard dominance function is calculated based on the collected samples using the GAE method. Then, the spatial consistency index between the full twin and the lightweight twin is calculated. The standard dominance function is weighted based on the consistency index to obtain the weighted dominance function guided by the lightweight twin. The parameters of the Actor network and the Critic network are then updated. Repeat the above training rounds until the network converges to obtain the optimal full-scale digital twin transfer strategy.