Optimization Method for Vehicle Task Offloading and Computing Resource Allocation Assisted by Digital Twin

By building a digital twin network and A3C algorithm to optimize vehicle task offloading and resource allocation, the problems of limited and dynamic changes in vehicle computing resources are solved, low latency and low cost computing resource allocation are achieved, and the efficiency and accuracy of vehicle edge computing are improved.

CN120111575BActive Publication Date: 2025-07-25SHANDONG NORMAL UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510573289.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-06
Publication Date
2025-07-25
Estimated Expiration
2045-05-06

AI Technical Summary

Technical Problem

In the field of intelligent transportation and vehicle networking, vehicle computing resources are limited and dynamically changing. The existing task offloading and computing resource allocation strategies are difficult to meet real-time requirements, and there is a lack of an effective incentive mechanism, resulting in a long delay in task processing.

Method used

Build a digital twin network, obtain information through interaction between vehicles and RSU twins, establish a task calculation and offload model, use the A3C algorithm to optimize resource allocation, and combine vehicle twin assisted computing to achieve optimal offload decisions and resource allocation.

Benefits of technology

Significantly reduce task processing delay and computing costs, adapt to dynamic environments, optimize computing resource allocation, improve computing efficiency and accuracy, and ensure efficient operation of vehicle edge computing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120111575B_ABST
    Figure CN120111575B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of vehicle networking wireless communication, and specifically relates to an optimization method for vehicle task offloading and computing resource allocation assisted by digital twins, including: obtaining information of vehicles and RSUs, constructing vehicle twins and RSU twins, and building a digital twin network; determining the total delay and total cost of the system based on the digital twin network, where the total delay includes local computing delay, vehicle twin-assisted computing delay, GAP error, uplink delay, and RSU computing delay, and the total cost includes RSU computing cost and vehicle twin-assisted computing cost; establishing an optimization problem model according to minimizing the total delay and total cost of the system; and the digital twin network solving the optimization problem model. The method of the present invention can effectively reduce the average system delay, reduce the resource usage cost, and enhance the collaborative advantage of digital twin technology and the reinforcement learning framework in edge computing resource optimization in typical urban traffic scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of vehicle networking wireless communication, and particularly to an optimization method for vehicle task offloading and computing resource allocation assisted by digital twins. Background Art

[0002] In the context of the booming development of intelligent transportation and vehicle networking, the data generated by vehicles has grown exponentially. According to statistics, in the context of autonomous driving, the amount of data generated by a single vehicle per hour can reach several GB. The huge data processing requirements make it difficult for the traditional cloud computing model to meet the strict requirements of vehicles for computing real-time performance, and vehicle edge computing technology has emerged as the times require. This technology executes computing tasks at the network edge close to the vehicle, effectively reducing data transmission latency and improving system response speed, providing key support for applications such as autonomous driving decision-making and real-time road condition monitoring, and becoming an important force driving the development of intelligent transportation.

[0003] Vehicle edge computing faces many challenges in practical applications. On the one hand, the computing resources of vehicles and RSU (Road Side Unit servers) are both limited and in dynamic change. When a vehicle is driving, the continuous changes in its position and speed will affect the communication quality with the RSU and the computing resources that can be obtained; the differences in the hardware configurations of different vehicles and the load changes during operation also make the computing capabilities of vehicles uneven. On the other hand, the existing task offloading and computing resource allocation strategies have deficiencies and are difficult to make full use of computing resources, resulting in longer task processing delays. Under the traditional offloading strategy, the average processing delay of complex tasks is high, which cannot meet the real-time requirements, and there is a lack of effective incentive mechanisms to mobilize the enthusiasm of vehicles to participate in auxiliary computing. Summary of the Invention

[0004] In view of this, the present invention proposes an optimization method for vehicle task offloading and computing resource allocation assisted by digital twins, constructs a digital twin network architecture, uses the interaction between vehicle and RSU twins and the physical layer to obtain multi-dimensional information, establishes a task computing, partial offloading and cost model, uses the digital twin network to obtain the optimal offloading strategy, uses the characteristics of the vehicle twin having vehicle state information and task information to assist the vehicle in task computing, establishes an optimization model, uses the A3C algorithm to solve the optimization problem model, and then obtains the optimal offloading decision and resource allocation plan, realizing the optimization of vehicle edge offloading and providing strong guarantee for the efficient operation of the intelligent transportation system.

[0005] In a first aspect, an embodiment of the present invention provides an optimization method for vehicle task offloading and computing resource allocation assisted by digital twins, including:

[0006] Obtain the information of vehicles and RSU, construct vehicle twins and RSU twins, and build a digital twin network;

[0007] Determine the total delay and total cost of the system based on the digital twin network. The total delay includes local computing delay, vehicle twin-assisted computing delay, GAP error, uplink delay, and RSU computing delay. The total cost includes RSU computing cost and vehicle twin-assisted computing cost;

[0008] Establish an optimization problem model according to minimizing the total delay and total cost of the system;

[0009] The digital twin network solves the optimization problem model.

[0010] In a possible implementation, the total delay The calculation formula is:

[0011] ,

[0012] ,

[0013] where, N represents the total number of vehicles, T represents the total number of time slots, represents the task computing delay, represents the uplink delay, represents the RSU computing delay, represents the total delay of vehicle users.

[0014] In a possible implementation, the uplink delay is:

[0015] ,

[0016] where, represents the task transmission rate, represents the vehicle n in the time slot t the offloading ratio within, represents the vehicle n in the time slot t the task size generated within;

[0017] The RSU computing delay is:

[0018] ,

[0019] where, represents the CPU cycles required to complete the task generated by vehicle n in the time slot t within, represents the computing resources allocated by the RSU to task ;

[0020] In a possible implementation, the total vehicle user delay is:

[0021] ,

[0022] ,

[0023] ,

[0024] ,

[0025] wherein, represents the local computing delay, represents the vehicle twin assisted computing delay, Gap represents the GAP error, represents the vehicle n in the time slot t the offloading ratio within, represents the vehicle n in the time slot t the ratio of vehicle twin assisted computing within, represents the CPU cycles required to complete the tasks generated by the vehicle n in the time slot t within, represents the computing power of the vehicle n , represents the computing power of the vehicle mapped by the vehicle twin through virtual representation, represents the computing power error generated during the vehicle twin mapping process.

[0026] In a possible implementation, the total cost is:

[0027] ,

[0028] ,

[0029] ,

[0030] wherein, represents the RSU computing cost, represents the vehicle twin assisted computing cost, N represents the total number of vehicles, T represents the total number of time slots, represents the vehicle n in the time slot t the offloading ratio within, represents the CPU cycles required to complete the tasks generated by the vehicle n in the time slot t within represents the unit price calculated by the RSU, represents the vehicle n in the time slot t the ratio of vehicle twin-assisted computing, represents the unit price of vehicle twin-assisted computing.

[0031] In a possible implementation, the optimization problem model is:

[0032] ,

[0033] where, represents the delay weight factor, represents the price weight factor, x represents the offloading decision, represents the computing resources allocated by the RSU to task ; represents the vehicle n in the time slot t the offloading ratio, represents the vehicle n in the time slot t the ratio of vehicle twin-assisted computing, represents the total delay, represents the total cost;

[0034] The model satisfies the following constraints:

[0035] ,

[0036] ,

[0037] ,

[0038] ,

[0039] ,

[0040] where, represents the task computing delay, represents the maximum tolerable delay, represents the vehicle n the m th task offloading strategy, represents the maximum computing resources of the RSU.

[0041] In a possible implementation, the digital twin network solves the optimization problem model, specifically:

[0042] Step S41, construct the state space, action space, and reward function;

[0043] Step S42: Initialize the global Actor network parameters and the global Critic network parameters, and initialize the local Actor network parameters and the local Critic network parameters within each parallel asynchronous thread to the global Actor network parameters and the global Critic network parameters;

[0044] Step S43: Execute all threads in parallel. During the running of each thread, obtain the current state of the thread , and the local Actor network generates an action probability distribution based on the current state, and randomly samples an action from the action probability distribution ; Execute the sampled action in the environment to obtain an immediate reward and the next state ; Store the experience ;

[0045] Step S44: The local Critic network calculates the temporal difference error based on the experience;

[0046] Step S45: Calculate the Actor network loss and the Critic network loss respectively based on the temporal difference error, and update the local Actor network parameters and the local Critic network parameters respectively;

[0047] Step S46: Loop through steps S43 - S45 until the loop termination condition is met, and output the optimal policy represented by the global Actor network.

[0048] In a possible implementation, the state space is: , where represents the task generated by vehicle n within time slot t , represents the twin of vehicle n , represents the n th vehicle and the k th RSU distance;

[0049] The action space is: ;

[0050] The reward function is: .

[0051] In a possible implementation, the temporal difference error is:

[0052] ,

[0053] where represents the agent att The immediate reward obtained after performing an action at a moment, denotes the discount factor, denotes the value estimate of the local Critic network in the next state, denotes the value estimate of the local Critic network in the current state.

[0054] In a possible implementation, the calculation formula for updating the local Actor network parameters is:

[0055] ,

[0056] ,

[0057] where, denotes the local Actor network parameters, denotes the learning rate of the Actor network, denotes the gradient of the Actor network loss with respect to the network parameters, denotes the Actor network loss, denotes the output action probability distribution of the local Actor network;

[0058] The calculation formula for updating the local Critic network parameters is:

[0059] ,

[0060] ,

[0061] where, denotes the local Critic network parameters, denotes the learning rate of the Critic network, denotes the gradient of the Critic network loss with respect to the network parameters, denotes the Critic network loss.

[0062] The present invention applies digital twin technology to a mobile edge computing network. By constructing a digital twin network architecture, a task offloading model, and a task computing model, and combining the information interaction between the twin layer and the physical layer, the computing task parameters are comprehensively and accurately obtained. The partial offloading mode is adopted, an offloading and cost model is designed, and price incentives are set for vehicle twins to assist in computing. The vehicle twins and the RSU twins jointly form a digital twin network, and jointly execute the A3C algorithm to collaboratively achieve the optimal decision-making for task offloading, realize the optimal offloading decision and computing resource allocation. After obtaining the optimal decision, the vehicle twins are used as computing nodes to assist the vehicle in computing the tasks that are not offloaded to the RSU. The method of this application effectively improves the computing efficiency and accuracy, significantly reduces the task processing delay and computing cost, can adapt to the dynamically changing environment, optimizes the computing resource allocation, and provides a strong guarantee for the efficient operation of vehicle edge computing. Compared with the prior art, the present invention has achieved the following remarkable technical effects.

[0063] (1) The present invention sets the vehicle twins and the RSU twins in the cloud. The vehicle twins and the RSU twins jointly form a digital twin network, and information sharing is carried out between the vehicle twins and the RSU twins. By constructing a virtual representation method in the digital twin network, the characteristic states of physical vehicles and RSUs and the prediction of the system are reflected in real time, so as to achieve better computing resource scheduling, and achieve the efficient utilization of computing resources and the maximization of the interests of both parties.

[0064] (2) The present invention innovatively proposes to use the twins constructed by digital twin technology to assist the physical layer ontology in task computing, and uses the "total cost" as the optimization index. By integrating the computing cost of the RSU and the computing cost assisted by the vehicle twins, the vehicle twins are incentivized to participate in vehicle task computing. Different from the limitations of only considering energy consumption or resource occupancy in the prior art, the present invention comprehensively covers the cost elements of the computing process from the dimension of comprehensive cost, and while ensuring the task completion delay, fully utilizes each computing node in the system to complete task computing.

[0065] (3) The present invention combines digital twin technology with the A3C algorithm to achieve the optimal offloading decision and computing resource allocation. Digital twin technology constructs a virtual model highly similar to the physical entity and maps the state of the physical system in real time. As an asynchronous parallel deep reinforcement learning algorithm, the A3C algorithm can utilize the digital twin environment, and according to a large amount of real-time data provided by different twins, simultaneously perform policy evaluation and update in multiple threads. The parallel learning characteristic of the A3C algorithm can significantly accelerate the convergence speed of the model and shorten the time for policy optimization. The high consistency between the twins constructed by digital twin technology and the physical entity, and the experience and strategies learned by the A3C algorithm in the digital twin environment, the combination of the two can more accurately balance the delay and cost factors, effectively reduce the overall system loss, and can be effectively migrated to the actual physical system. Description of the Drawings

[0066] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.

[0067] Figure 1 It is a schematic flowchart of an optimization method for vehicle task offloading and computing resource allocation assisted by digital twin provided by an embodiment of the present invention;

[0068] Figure 2 It is a scenario model diagram of an optimization method for vehicle task offloading and computing resource allocation assisted by digital twin provided by an embodiment of the present invention;

[0069] Figure 3 It is a simulation diagram of the optimal offloading scheme solved by the A3C algorithm under the digital twin framework in the optimization method for vehicle task offloading and computing resource allocation assisted by digital twin provided by an embodiment of the present invention and obtaining the minimum system cost;

[0070] Figure 4 It is a comparison diagram of the system simulation costs of different calculation methods based on obtaining the optimal offloading scheme provided by an embodiment of the present application;

[0071] Figure 5 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. Specific Embodiments

[0072] To better understand the technical solutions of the present invention, the following will describe the embodiments of the present invention in detail with reference to the accompanying drawings.

[0073] It should be clear that the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0074] The terms used in the embodiments of the present invention are only for the purpose of describing specific embodiments, and are not intended to limit the present invention. The singular forms of "a", "the" and "said" used in the embodiments of the present invention and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise.

[0075] It should be understood that the term "and / or" used herein is merely a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, both A and B exist simultaneously, and B exists alone. Additionally, the character " / " in this text generally represents an "or" relationship between the preceding and following associated objects.

[0076] See Figure 1 , which is a schematic flowchart of the digital twin-assisted vehicle task offloading and computing resource allocation optimization method provided by the embodiment of the present invention. As Figure 1 shown, it mainly includes the following steps.

[0077] Step S01, obtain the information of the vehicle and the RSU, construct the vehicle digital twin and the RSU digital twin, and build the digital twin network. Specifically:

[0078] Apply the digital twin technology to the mobile edge computing network, map the vehicle and the RSU (roadside unit server) to construct the vehicle digital twin and the RSU digital twin, and through the virtual representation method, set the data update frequency so that the digital twin can reflect the characteristics of its corresponding physical device in real time and predict the state of the physical device, and establish the digital twin network and the digital twin-assisted vehicle edge computing model.

[0079] The vehicle digital twin continuously interacts with the physical-layer vehicle to obtain information such as the vehicle's position, speed, and computing power; the RSU digital twin obtains information such as the remaining computing resources, channel resources, and road conditions of the physical-layer RSU, uses the digital twin network to assist the physical layer in making offloading decisions, the vehicle digital twin stores all the information of the tasks generated by the vehicle and has computing resources, and after the digital twin network assists the vehicle in offloading tasks to the RSU, uses the vehicle digital twin to assist the vehicle in task computing and establish a task computing model.

[0080] It should be particularly noted that regardless of whether the vehicle offloads tasks to the RSU digital twin, the vehicle digital twin always assists the vehicle in task computing, and the assistance ratio is jointly restricted by the reward mechanism and the latency constraint.

[0081] See Figure 2 , which is a scenario model diagram of the digital twin-assisted vehicle task offloading and computing resource allocation optimization method provided by the embodiment of the present invention. As Figure 2 shown, set the vehicle digital twin and the RSU digital twin in the cloud, and the vehicle digital twin and the RSU digital twin together constitute the digital twin network, and information sharing occurs between the digital twins. By constructing a virtual representation method in the digital twin network, the characteristic state of the physical layer and the prediction of the system are reflected in real time. The vehicle set is The vehicle digital twin set is The RSU set is , the RSU twin set is .

[0082] In the time slot set , each vehicle generates a delay-sensitive task in each time slot , n represents the vehicle number, t represents the time slot number, , where represents the task size generated by vehicle n in time slot t (which can be simply referred to as "task size"), represents the CPU cycles required to complete the task generated by vehicle n in time slot t (which can be simply referred to as "CPU cycles required to complete the task"), represents the maximum tolerable delay of the task generated by vehicle n in time slot t (which can be simply referred to as "maximum tolerable delay").

[0083] The vehicle twin continuously interacts with the physical-layer vehicle through the digital twin network to obtain the information of the physical-layer vehicle , where and respectively represent the position and speed of vehicle n , represents the computing power of vehicle n , represents the computing power of the vehicle mapped by the vehicle twin through virtual representation, represents the computing power error generated during the mapping process of the vehicle twin. While the vehicle twin provides vehicle information to the digital twin network, it uses the computing power of the mapped physical-layer vehicle to assist the vehicle in task calculation.

[0084] Step S02, determine the total delay and total cost of the system based on the digital twin network. The total delay includes local computing delay, vehicle twin assisted computing delay, GAP error, uplink delay, and RSU computing delay. The total cost includes RSU computing cost and vehicle twin assisted computing cost. Specifically:

[0085] Establish a task calculation model. In the task offloading scenario, calculate the task transmission rate of the uplink delay based on the task transmission rate; adopt a partial offloading mode, respectively construct the delay formulas for tasks in RSU computing, local computing, and vehicle twin assisted computing, and at the same time consider the errors of the computing capabilities of the twin and the ontology to obtain the task calculation delay formula; comprehensively consider the local, vehicle twin, and RSU computing resources to determine the total delay formula for the computing task.

[0086] Design the offloading and cost model. Partial offloading is adopted, and the task delay is jointly affected by local computing, the proportion of vehicle twin-assisted computing, and the proportion of RSU partial offloading. Given the unit price of RSU computing and the unit price of vehicle twin-assisted computing, construct the cost formula for RSU to complete the task and the cost formula for vehicle twin-assisted computing to complete the task respectively. Further, to incentivize vehicle twin-assisted computing, make the unit price of vehicle twin-assisted computing higher than the unit price of RSU computing.

[0087] For the offloading decision, partial offloading is considered. For the offloading action, the present invention considers that 0 means not offloading to the RSU, and the task delay of the system is affected by local computing and the proportion of vehicle twin-assisted computing; 1 means offloading to the RSU, and the task delay is affected by three parts: local computing, the proportion of vehicle twin-assisted computing, and the RSU offloading proportion. Let the n offloading proportion within the time slot t be (which can be simply referred to as the "offloading proportion"), and it satisfies . Then represents the proportion of vehicle n user computing within the time slot t . Let the proportion of vehicle n performing vehicle twin-assisted computing within the time slot t be (which can be simply referred to as the "vehicle twin-assisted computing proportion"), and it satisfies . Then represents the proportion of local computing of vehicle n within the time slot t .

[0088] The local computing delay of the partial task of local computing is:

[0089] ,

[0090] where represents the local computing delay, represents the offloading proportion of vehicle n within the time slot t , represents the proportion of vehicle n performing vehicle twin-assisted computing within the time slot t , represents the CPU cycles required to complete the tasks generated by vehicle n within the time slot t (which can be simply referred to as the "CPU cycles required to complete the tasks"), represents the computing power of vehicle n .

[0091] When the vehicle digital twin performs auxiliary calculations, the latency of the vehicle digital twin's auxiliary calculations is:

[0092] ,

[0093] where, represents the computing power of the vehicle mapped by the vehicle digital twin through virtual representation method.

[0094] It should be noted that although digital twin technology can map physical layer parameters in real time, in actual situations, there is a GAP error in the computing power between different digital twins and the physical entities Gap , and the GAP error Gap has the following calculation formula:

[0095] ,

[0096] where, represents the computing power error generated during the mapping process of the vehicle digital twin.

[0097] In summary, the total latency of the vehicle user considering the vehicle digital twin's auxiliary calculations is:

[0098] ,

[0099] When performing partial task offloading to the RSU, the uplink latency is considered first , and the calculation formula is:

[0100] ,

[0101] where, represents the uplink latency, represents the task transmission rate, represents the vehicle n in the time slot t the offloading ratio within, represents the vehicle n in the time slot t the task size generated within.

[0102] Furthermore, the calculation formula for the task transmission rate is:

[0103] ,

[0104] where, represents the bandwidth, represents the signal-to-noise ratio.

[0105] The RSU computing latency of the partial tasks offloaded to the RSU is:

[0106] ,

[0107] Among them, represents the RSU computing delay, represents the vehicle n in the time slot t unloading ratio, represents the CPU cycles required to complete the tasks generated by the vehicle n in the time slot t and represents the computing resources allocated by the RSU to the task (which can be simply referred to as "computing resources allocable by the RSU").

[0108] In summary, when the computing task is offloaded to the RSU for computing and assisted by the vehicle digital twin, the task computing delay is:

[0109] ,

[0110] From this, it can be obtained that the total delay of the system is:

[0111] ,

[0112] Among them, N represents the total number of vehicles, T represents the total number of time slots.

[0113] Based on the above embodiments, this embodiment further provides a method for calculating the total cost of completing tasks.

[0114] The vehicle digital twin and the RSU digital twin are placed in the cloud to form a global virtual layer, where decisions and unloading ratio optimization are performed, and at the same time, according to the optimization results, the vehicle digital twin is used to assist in completing vehicle task computing.

[0115] Assume that the unit price of RSU computing is , then the RSU computing cost required for the RSU to complete the task is:

[0116] ,

[0117] Utilizing the ability of digital twin technology to map physical layer information, when a vehicle generates a task, the vehicle digital twin can obtain all the information of the task. The present invention considers using the vehicle digital twin to assist in computing to reduce the delay of task completion. Assume that the unit price of vehicle digital twin assisted computing is , then the vehicle digital twin assisted computing cost is:

[0118] ,

[0119] wherein, represents the offloading ratio of the vehicle n within the time slot t , and represents the ratio of the vehicle n performing vehicle twin-assisted calculation within the time slot t .

[0120] Subsequently, the total cost required for all vehicle tasks in the system to be calculated and completed is obtained as:

[0121] ,

[0122] It should be noted that since the cloud digital twin network enables vehicle twin-assisted calculation while obtaining the optimal offloading decision, therefore, the present invention considers that digital twin technology requires huge computing resources and construction costs, such as cloud service fees, physical device maintenance costs, etc., so it is considered that the price of vehicle twin-assisted calculation is higher than the price of RSU offloading calculation to encourage cloud vehicle twin-assisted calculation tasks.

[0123] Step S03, establish an optimization problem model according to minimizing the total delay and total cost of the system. Specifically:

[0124] In order to minimize the total delay and total cost of the system, an optimization problem model is established by jointly optimizing the offloading decision of vehicle users, the computing resource allocation of RSU, the twin-assisted calculation ratio, and the RSU offloading ratio. The optimization problem model is:

[0125] ,

[0126] wherein, represents the delay weight factor, represents the price weight factor, x represents the offloading decision, represents the computing resource allocated by RSU to task , represents the vehicle n within the time slot t , represents the vehicle n within the time slot t performing vehicle twin-assisted calculation, represents the total delay, represents the total cost.

[0127] The model satisfies the following constraint equations:

[0128] ,

[0129] ,

[0130] ,

[0131] ,

[0132] ,

[0133] wherein, represents the task calculation delay; represents the maximum tolerable calculation delay; represents the offloading strategy of n the m tasks of vehicle . If n the m tasks of vehicle are not offloaded. If n the m tasks of vehicle are all offloaded;

[0134] It should be noted that the first constraint means that the task calculation delay cannot exceed the maximum tolerable calculation delay; the second constraint means whether the task is offloaded; the third constraint means partial offloading of the task, and the offloading ratio is [0, 1]; the fourth constraint means the ratio of vehicle twin-assisted task calculation, and the offloading ratio is [0, 1]; the fifth constraint means that the total calculation resources allocated by the RSU to the offloaded tasks shall not exceed the maximum calculation resources of the RSU.

[0135] Step S04, the digital twin network solves the optimization problem model. Specifically:

[0136] In this embodiment, the A3C algorithm is used to solve the optimization problem model.

[0137] The A3C (Asynchronous Advantage Actor-Critic) algorithm is an asynchronous parallel algorithm based on the Actor-Critic framework, which avoids the problem that the DQN algorithm has errors due to overestimating the value function, thereby affecting the search for the optimal strategy and algorithm convergence. The A3C algorithm runs multiple threads simultaneously, and each thread has a copy of the agent.

[0138] Step S04 specifically includes: constructing a state space, an action space, and a reward function, and constructing an A3C algorithm model; initializing the parameters of the global Actor network and the global Critic network of the A3C algorithm model, as well as the parameters of the local Actor network and the local Critic network in multiple parallel asynchronous threads; during the operation of each thread, the agent obtains the environmental state, generates an action probability distribution through the local Actor network and samples to obtain an action, interacts with the environment after executing the action to obtain a reward and a new state, calculates the temporal difference error using the local Critic network, updates the parameters of the local Actor network and the Critic network, and loops until the training termination condition is met, finally obtaining the optimal offloading decision and resource allocation plan to achieve the optimization of vehicle edge computing.

[0139] In this embodiment, the process of using the A3C algorithm to solve the optimization problem model is as follows:

[0140] Step S41, construct a state space, an action space, and a reward function, and construct an A3C algorithm model.

[0141] Construct the state space. The state set is , where represents the task generated by vehicle n in time slot t . represents the twin of vehicle n . represents the n th vehicle and the k th distance between the vehicle and the RSU. The state includes the state information of vehicle users in the entire network, the computing power of the vehicle, reflecting the local computing resource level; the computing power and computing power error of the vehicle twin, determining the effectiveness of the vehicle twin's auxiliary computing. The task-related information of vehicle users, such as the task size, the CPU cycles required to complete the task, and the maximum tolerable latency of the task, these parameters are directly related to the task processing difficulty and time requirements. The environmental information in the scenario, such as the position coordinates of the nearest RSU, can be used to calculate the distance between the vehicle and the RSU; the remaining computing resources, channel resources, and road conditions of the RSU, these environmental factors are crucial for task offloading and computing resource allocation decisions.

[0142] Construct the action space. The action space covers the key decision actions that the agent can execute. The offloading decision action is . The offloading decision is partial offloading, , where 0 represents not offloading the task to the RSU, and the vehicle relies only on local computing resources and twin-assisted computing; 1 represents offloading the task to the RSU, and the offloading ratio needs to be further determined.

[0143] Construct a reward function. As an agent, the vehicle makes corresponding decisions by maximizing the reward of the algorithm through interacting with the environment. To minimize the total delay and total cost of the system, the defined reward function r is .

[0144] Step S42: Initialize the global Actor network parameters and the global Critic network parameters, and initialize the local Actor network parameters and the local Critic network parameters within each parallel asynchronous thread to the global Actor network parameters and the global Critic network parameters.

[0145] Specifically, perform initialization. Initialize the global Actor network parameters and the global Critic network parameters . At the same time, initialize the network parameters of multiple parallel asynchronous threads. Each thread has independent local Actor network parameters and local Critic network parameters , and set the local Actor network parameters and the local Critic network parameters to be consistent with the global Actor network parameters and the global Critic network parameters , that is . ; Further, it is also necessary to initialize the environment and the experience buffer.

[0146] Step S43: Execute all threads in parallel. During the running process of each thread, obtain the current state of the thread . The local Actor network generates an action probability distribution according to the current state, and randomly samples a sampling action from the action probability distribution ; Execute the sampling action in the environment to obtain an immediate reward and the next state ; Store the experience .

[0147] Specifically, enter the asynchronous thread execution stage. Each thread operates independently. For each thread loop, obtain the current state of the thread . The local Actor network outputs an action probability distribution according to the current state , and randomly samples and selects a sampling action from the action probability distribution ; The agent executes the action in the environment, obtains an immediate reward and a new state , and stores the experience . into the experience buffer.

[0148] Step S44: The local Critic network calculates the temporal difference error based on the experience.

[0149] Specifically, the experience is input into the local Critic network, and the local Critic network calculates the temporal difference error according to the temporal difference error algorithm. , where represents the temporal difference error, which measures the deviation between the current state value estimate and the actual return; represents the immediate reward obtained by the agent after executing the action at t time; represents the discount factor, which is used to balance the importance of immediate rewards and future rewards; represents the value estimate of the local Critic network in the next state, represents the value estimate of the local Critic network in the current state.

[0150] Step S45: Based on the temporal difference error, calculate the Actor network loss and the Critic network loss respectively, and update the local Actor network parameters and the local Critic network parameters respectively. Specifically,

[0151] Update the Actor network: Calculate the loss function of the Actor network , where represents the output action probability distribution of the local Actor network, and use the gradient ascent method to update the parameters of the local Actor network , calculate the gradient of the local Actor loss function with respect to the network parameters , update the parameters , where is the learning rate of the Actor network;

[0152] Update the Critic network: Calculate the loss function of the Critic network , and use the gradient descent method to update the parameters of the local Critic network , calculate the gradient of the local Critic loss function with respect to the network parameters , update the parameters , where is the learning rate of the Critic network.

[0153] Step S46: Loop steps S43 - S45 until the loop termination condition is met, and output the optimal policy represented by the global Actor network.

[0154] Specifically, steps S43 - S45 are looped until the training termination condition is met (such as reaching the set number of training steps or the agent's performance reaching the expectation), and finally the optimal policy represented by the global Actor network is output.

[0155] Under the digital twin network architecture, the vehicle twin collects and synchronizes the running speed of the user's vehicle in real time, monitors the computing resources, task information, network status, and cost - benefit of the user's vehicle; the RSU twin collects and shares the status data of roadside devices in real time and collaborates with the vehicle twin to complete data interaction. The vehicle twin and the RSU twin together constitute the digital twin network. The vehicle twin and the RSU twin, as intelligent agents, jointly execute the A3C algorithm to collaboratively achieve the optimal decision - making for task offloading. After obtaining the optimal decision, considering that the vehicle twin has abundant computing resources and all the information of the user's vehicle tasks, the vehicle twin is further used as a computing node to assist the user's vehicle in calculating the tasks that are not offloaded to the RSU.

[0156] The present invention combines digital twin technology with the A3C algorithm to achieve optimal offloading decision - making and computing resource allocation. The digital twin technology constructs a virtual model highly similar to the physical entity and maps the state of the physical system in real time. The A3C algorithm, as an asynchronous parallel deep reinforcement learning algorithm, can utilize the digital twin environment and, based on a large amount of real - time data provided by different twins, simultaneously perform policy evaluation and update in multiple threads. The parallel learning feature of the A3C algorithm can significantly accelerate the convergence speed of the model and shorten the time for policy optimization. See Figure 3 , which is a simulation diagram of the optimal offloading scheme solved by the A3C algorithm under the digital twin framework and obtaining the minimum system cost in the digital twin - assisted vehicle task offloading and computing resource allocation optimization method provided by the embodiment of the present invention. See Figure 4 , which is a comparison diagram of the system simulation costs of different computing methods based on obtaining the optimal offloading scheme. Among them, the DT + A3C (Proposed) curve represents the system simulation cost curve of the computing method combining digital twin technology and the A3C algorithm provided by the embodiment of the present application, the DT + SAC curve represents the system simulation cost curve of the computing method combining digital twin technology and the SAC algorithm, the DT + DDPG curve represents the system simulation cost curve of the computing method combining digital twin technology and the DDPG algorithm, No - DT + A3C represents the system simulation cost curve of the single A3C algorithm, the Local Only curve represents the system simulation cost curve of the local - only computing method, and the RSU Only represents the system simulation cost curve of the computing method of offloading only to the RSU. As Figure 4As shown in the figure, the digital twin technology can update the status information of users and various physical devices in the scenario in real time, and the digital twins in the digital twin network act as agents to make decisions, with lower latency compared to traditional solutions. In the scenarios of different vehicles, the objective function value of the proposed solution is always at a lower level compared to the cases of only offloading to the RSU, only local computing, and no digital twin. This means that this technology can more accurately balance the delay and cost factors and effectively reduce the overall system loss. Similarly, although the DDPG and SAC algorithms are also significantly better than some traditional algorithms, highlighting the universal advantage of the digital twin technology in optimizing the objective function under a multi-algorithm architecture and being able to provide a more economical and efficient solution for the system operation, their objective function is still higher than that of the method of this application. Facing the increasing complexity of traffic scenarios brought about by the increase in the number of vehicles, the digital twin technology endows the algorithm with excellent scalability and adaptability. As Figure 4 shown, as the number of vehicles increases from 5 to 30, for each algorithm based on the digital twin, the growth trend of the objective function value is gentle, while for the solution without the digital twin, the increase in the objective function value is relatively large. This indicates that the digital twin technology enables the algorithm to maintain stable performance when dealing with a large number of vehicles and complex traffic flows, ensuring the efficient operation of the system in a dynamically changing environment and effectively expanding the application boundary and scope of application of the algorithm.

[0157] Corresponding to the above embodiments, the embodiment of the present invention also provides an electronic device.

[0158] See Figure 5 , which is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. As Figure 5 shown, the electronic device 500 may include: a processor 501, a memory 502, and a communication unit 503. These components communicate through one or more buses. Those skilled in the art can understand that the structure of the electronic device shown in the figure does not constitute a limitation on the embodiment of the present invention. It can be a bus structure, a star structure, and may also include more or fewer components than shown in the figure, or combine some components, or have different component arrangements.

[0159] Among them, the communication unit 503 is used to establish a communication channel so that the electronic device can communicate with other devices.

[0160] The processor 501 is the control center of the electronic device, connecting various parts of the entire electronic device through various interfaces and circuits. By running or executing software programs and / or modules stored in the memory 502, and by invoking the data stored in the memory, it executes various functions of the electronic device and / or processes data. The processor may be composed of an integrated circuit (IC). For example, it may be composed of a single packaged IC, or may be composed of multiple packaged ICs with the same or different functions connected together. For example, the processor 501 may only include a central processing unit (CPU). In the embodiment of the present invention, the CPU may be a single arithmetic core or may include multiple arithmetic cores.

[0161] The memory 502 is used to store the execution instructions of the processor 501. The memory 502 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disc.

[0162] When the execution instructions in the memory 502 are executed by the processor 501, the electronic device 500 is enabled to execute some or all of the steps in the above method embodiments.

[0163] Corresponding to the above embodiments, the embodiment of the present invention further provides a computer-readable storage medium. The computer-readable storage medium can store a program. When the program runs, it can control the device where the computer-readable storage medium is located to execute some or all of the steps in the above method embodiments. Specifically, the computer-readable storage medium can be a magnetic disk, an optical disc, a read-only memory (ROM) or a random access memory (RAM), etc.

[0164] Corresponding to the above embodiments, the embodiment of the present invention further provides a computer program product. The computer program product contains executable instructions. When the executable instructions are executed on a computer, the computer is enabled to execute some or all of the steps in the above method embodiments.

[0165] In the embodiments of the present invention, "at least one" means one or more, and "a plurality of" means two or more. "And / or" describes the association relationship of associated objects and indicates that there can be three relationships. For example, A and / or B can represent the cases of A existing alone, A and B existing simultaneously, and B existing alone. Here, A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after. "At least one of the following" and its similar expressions refer to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, and c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c can be single or multiple.

[0166] Those of ordinary skill in the art can realize that the various units and algorithm steps described in the embodiments disclosed herein can be implemented by a combination of electronic hardware, computer software, and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present invention.

[0167] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0168] In several embodiments provided by the present invention, if any function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM for short), random access memory (RAM for short), magnetic disks, or optical discs that can store program codes.

[0169] The above are only specific embodiments of the present invention. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present invention and should be covered by the protection scope of the present invention. The protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. An optimization method for vehicle task offloading and computing resource allocation assisted by digital twins, characterized in that, Including: Obtain the information of the vehicle and the RSU, construct the vehicle twin and the RSU twin, and build a digital twin network; Determine the total delay and total cost of the system based on the digital twin network. The total delay includes local computing delay, vehicle twin assisted computing delay, GAP error, uplink delay, and RSU computing delay. The total cost includes RSU computing cost and vehicle twin assisted computing cost; Establish an optimization problem model according to minimizing the total delay and total cost of the system; The digital twin network solves the optimization problem model; The optimization problem model is: Q tot = Q k + Z dt where, α is the delay weight factor, β is the price weight factor, x is the offloading decision, is the computing resource allocated by the RSU to task m n,t and m n,t is the task generated by vehicle n within time slot t, and ρ n,t is the offloading ratio of vehicle n within time slot t, is the ratio of vehicle n performing vehicle twin-assisted computing within time slot t, T tot is the total delay, Q tot is the total cost; N is the total number of vehicles, T is the total number of time slots, is the task computing delay, is the uplink delay, is the RSU computing delay, T loct is the total delay of vehicle users, r n is the task transmission rate, S n,t is the task size generated by vehicle n within time slot t, C n,t is the number of CPU cycles required to complete the task generated by vehicle n within time slot t, is the local computing delay, is the vehicle twin-assisted computing delay, Gap is the GAP error, f n is the computing power of vehicle n, f′ n is the computing power of the vehicle twin mapping the vehicle through virtual representation method, Δf′ n is the computing power error generated during the vehicle twin mapping process, Q k is the RSU computing cost, Z dt is the vehicle twin-assisted computing cost, q k is the unit price of RSU computing, z dt is the unit price of vehicle twin-assisted computing; The model satisfies the following constraint equations: x n,m ∈{0,1} ρ n,t ∈ [0, 1] Among them, T max is the maximum tolerable delay, and x n,m is the offloading strategy for the m-th task of vehicle n, and is the maximum computing resource of the RSU.

2. The digital-twin-assisted vehicle task offloading and computing resource allocation optimization method according to claim 1, wherein The digital twin network solves the optimization problem model, specifically: Step S41, construct the state space, action space, and reward function; Step S42, initialize the global Actor network parameters and global Critic network parameters, and initialize the local Actor network parameters and local Critic network parameters within each parallel asynchronous thread to the global Actor network parameters and global Critic network parameters; Step S43, execute all threads in parallel. During the running of each thread, obtain the current state S of the thread t , the local Actor network generates an action probability distribution based on the current state, and randomly samples an action a from the action probability distribution t ; execute the sampled action in the environment to obtain an immediate reward r t and the next state S t+1 ; store the experience (S t , a t , r t , S t+1 ); Step S44, the local Critic network calculates the temporal difference error based on experience; Step S45, calculate the Actor network loss and Critic network loss respectively based on the temporal difference error, and update the local Actor network parameters and local Critic network parameters respectively; Step S46, loop steps S43 - S45 until the loop termination condition is met, and output the optimal policy represented by the global Actor network.

3. The digital twin-assisted vehicle task offloading and computing resource allocation optimization method according to claim 2, wherein The state space is: where m n,t is the task generated by vehicle n in time slot t, v’ n is the twin of vehicle n, and d k,n is the distance between the nth vehicle and the kth RSU; The action space is: The reward function is as follows:

4. The digital twin-assisted vehicle task offloading and computing resource allocation optimization method according to claim 2, wherein Temporal difference error E t is as follows: where r t is the immediate reward obtained by the agent after executing an action at time t, γ is the discount factor, is the value estimate of the local Critic network in the next state, is the value estimate of the local Critic network in the current state.

5. The digital twin-assisted vehicle task offloading and computing resource allocation optimization method according to claim 4, characterized in that The calculation formula for updating the local Actor network parameters is: Among them, is the local Actor network parameter, α π is the learning rate of the Actor network, is the gradient of the Actor network loss with respect to the network parameter, L π is the Actor network loss, is the output action probability distribution of the local Actor network; The calculation formula for updating the local Critic network parameters is: Among them, are the local Critic network parameters, and α v is the learning rate of the Critic network, is the gradient of the Critic network loss with respect to the network parameters, and L v is the Critic network loss.

Citation Information

Patent Citations

  • Vehicle digital twinborn body edge deployment method used in Internet of Vehicles scene

    CN116980424A

  • Vehicle-mounted task unloading and resource allocation method and device, equipment and storage medium

    CN118972804A

  • Construction and resource allocation method of digital twinborn in Internet of Vehicles

    CN119012390A