Resource allocation method for digital twinborn maintenance and computing task processing

By establishing vehicle mobility, communication and computing models and optimizing resource allocation using the MADDPG algorithm, the resource management problem between the vehicle and the VEC server is solved, and more efficient resource utilization and task processing is achieved.

CN120256101APending Publication Date: 2025-07-04XINKONG (SUQIAN) TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510318138.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

With the increasing demand for vehicle computing tasks, it is difficult for the prior art to effectively manage resource allocation between the vehicle and the VEC server, resulting in delays in computing tasks and waste of resources.

Method used

By establishing vehicle mobility, communication and computing models, combined with the multi-agent reinforcement learning algorithm MADDPG, resource allocation strategies are optimized to meet the simultaneous processing requirements of digital twin maintenance and computing tasks.

Benefits of technology

Improve resource utilization efficiency, reduce task delays, and achieve higher resource efficiency in multi-vehicle environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120256101A_ABST
    Figure CN120256101A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of Internet of Vehicles, and particularly relates to a resource allocation method for digital twinborn maintenance and computing task processing, which comprises the following steps of: firstly, modeling a scene, modeling a communication model, a mobile model and a computing model, and then setting a state, an action space and a reward function; and finally, solving an optimal resource allocation scheme through multi-agent reinforcement learning so as to maximize the resource efficiency. According to the method, the resource allocation problem which is neglected in the existing research and exists when twin maintenance and calculation task processing are carried out at the same time is considered, the research on the aspect is supplemented for the resource allocation problem, and the twin maintenance under the problem is preliminarily explored, namely, a general model is established. Simulation results show that compared with other algorithms, the scheme provided by the invention can obtain higher resource efficiency, thereby obtaining higher long-term benefits.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of vehicle networking, and particularly relates to a resource allocation method for digital twin maintenance and computing task processing based on multi-agent reinforcement learning. Background Art

[0002] With the rapid development of the fifth-generation (5G) technology, there are more and more in-vehicle applications and multimedia services in the fields of autonomous driving, navigation, high-definition video, etc. These advancements aim to enhance the overall driving experience but also lead to an increase in the number of vehicle computing tasks that need to be addressed. However, the storage capacity and computing power of vehicles are often insufficient to meet such demands, which poses a challenge to the real-time processing of computationally intensive tasks. To address this challenge, vehicle edge computing has emerged as a promising solution. By deploying VEC (Vehicle Edge Computing) servers at the roadside, computing tasks can be transferred from vehicles to these servers for processing and then the results can be effectively returned.

[0003] Although VEC can provide computing services for vehicles, it also faces problems such as vehicle mobility and environmental dynamics during implementation. Digital twin technology (DT), as an emerging innovative technology, helps to create a virtual model that accurately represents a physical object. With the development of related technologies such as 5G, edge computing, and artificial intelligence, the capabilities of DT are continuously enhanced. It has now gone beyond one-way mirror simulation and achieved two-way information interaction. This two-way mapping provides rich information for the VEC network, making it possible to extract features and make predictions about the physical vehicle and its surrounding environment. Specifically, a vehicle can use V2X (Vehicle-to-Everything) technology to transmit its data to the server, enabling it to establish a digital copy on the server based on historical data, thereby creating a DT model.

[0004] The combination of DT and VEC can not only collect the real-time operation data of vehicles in the VEC network but also perform real-time control and change on the vehicle state. The real-time information of the vehicle can be accessed and compared with historical data to proactively identify and solve potential problems or risks during vehicle operation. However, this integration also brings challenges, especially in terms of resource management. The establishment and maintenance of the DT model require computing resources for information synchronization, and processing in-vehicle computing tasks on the server further consumes resources. Therefore, it is crucial to explore how to formulate a reasonable resource allocation strategy to prevent task execution delays. Summary of the Invention

[0005] In view of the problems in the background art, the present invention considers the DT scenario where multiple vehicles are associated with a single VEC server in a mobile edge network, and proposes a resource allocation method for digital twin maintenance and computing task processing, which maximizes the resource utility of computing resource allocation while ensuring the time limit requirements of twin maintenance and vehicle computing tasks.

[0006] To achieve the above object, the present invention provides the following technical solutions:

[0007] A resource allocation method for digital twin maintenance and computing task processing, comprising the following steps:

[0008] (1) Establish a system model

[0009] (1.1) Vehicle movement model modeling:

[0010] Taking the base station as the origin, a spatial orthogonal coordinate system is established. The positive direction of the x-axis is the direction in which the vehicle travels along the lane, the positive direction of the y-axis is the direction along which the number of lanes increases, and the positive direction of the z-axis is the direction along the base station antenna. The position of vehicle i on lane j at time slot t is expressed as:

[0011] PO i (t)=(X i,j (t),Y i,j (t),0), i ∈ {1,2,3,…,N} (1)

[0012] In the formula, X i,j (t) is the abscissa of vehicle i traveling on lane j at time slot t, and Y i,j (t) is the ordinate of vehicle i;

[0013] The abscissa X i,j (t) of vehicle i traveling on lane j at time slot t is further expressed as:

[0014] X i,j (t)=X i,j (t - 1)+τv i (2)

[0015] In the formula, τ is the duration of each time slot, and v i is the traveling speed of vehicle i;

[0016] Let ω0 represent the width of each lane, and L0 represent the distance between the first lane and the base station. Therefore, the ordinate Y i,j (t) of vehicle i is written as:

[0017] Y i,j (t)=L0 + jω0, j ∈ (0,1,2,3,…,J - 1) (3)

[0018] (1.2) Communication model modeling:

[0019] The communication rate between vehicle i and the base station is expressed as:

[0020]

[0021] where b refers to the base station, W is the channel bandwidth, p i is the transmission power of vehicle i, and σ 2 is the channel noise;

[0022] The channel gain between vehicle i and the base station is calculated as:

[0023]

[0024] where is the small-scale fading, is the large-scale path fading and shadow; for the small-scale fading, a general fading model is adopted and calculated as:

[0025]

[0026] where obeys the circularly symmetric complex Gaussian distribution with unit variance, κ is the correlation coefficient, PO i (t) is the position of vehicle i at time slot t, and P B (t) is the position of the base station, and β is the path loss exponent;

[0027] (1.3) Computation model modeling:

[0028] Let represent the information that vehicle i needs to synchronously update at time slot t, where is the size of the information that vehicle i needs to synchronously update at the current time t, is the CPU frequency required to update the information of vehicle i per unit size, and T i dt is the maximum delay limit for updating the information of vehicle i; the updated information needs to be first transmitted from vehicle i to the server equipped at the base station, and then the server allocates computing resources to maintain the digital twin of vehicle i;

[0029] Denote the total computing resources on the server as F, and the computing resources required to maintain the twin of vehicle i as f i dt Therefore, the time required to complete the maintenance of the digital twin of vehicle i is calculated as:

[0030]

[0031] Denote the computing task generated by vehicle \(i\) at time slot \(t\) as where is the size of the computing task generated by vehicle \(i\) at time slot \(t\), is the CPU frequency required for a computing unit size task, and \(T\) i tk is the maximum latency limit for processing the computing task; Vehicle \(i\) first offloads the computing task to the server, and then the server allocates computing resources to process the computing task. Therefore, the time required to complete the computing task is calculated as:

[0032]

[0033] where \(f\) i tk represents the computing resources requested by vehicle \(i\) to process the computing task transmitted to the server;

[0034] Define the satisfaction function \(Q\) to represent the satisfaction of the digital twin maintenance latency and the computing task processing latency of vehicle \(i\) under a certain resource allocation strategy. The satisfaction function of the digital twin maintenance latency of vehicle \(i\) and the satisfaction function of the computing task processing latency of vehicle \(i\) are respectively expressed as:

[0035]

[0036] Through equations (9) and (10), evaluate the resource utility \(U\) i obtained by vehicle \(i\) under the resource allocation strategy \(\omega\) i (\(\omega\) i ):

[0037]

[0038] where \(\rho\) is the weight factor between twin maintenance and computing task processing, and its value is \(0 < \rho < 1\), which is used to measure the importance of the two tasks;

[0039] (2) Establish the state, action space, and reward function;

[0040] (3) Use the MADDPG algorithm (Multi-Agent Deep Deterministic Policy Gradient, a multi-agent reinforcement learning algorithm) to find the optimal resource allocation method to maximize the long-term reward.

[0041] Furthermore, in step (1.3), maintaining the digital twin and processing the computing task are carried out simultaneously. Therefore, the resources \(f\) allocated to the vehicle for maintaining the digital twin idt and the sum of the resources f allocated for processing the computing tasks generated by the vehicle i tk shall not exceed the total computing resources F of the server, that is

[0042]

[0043] Furthermore, the step (2) specifically includes the following steps:

[0044] (2.1) State space:

[0045] Select the update information Γ n (t), the computing task ζ n (t), the vehicle position PO n (t), the speed v n and the channel gain which is expressed as:

[0046]

[0047] (2.2) Action space:

[0048] Each vehicle is regarded as an agent, and according to the observed state s n (t), it makes an action a n (t), and the action includes the computing resources for twin maintenance and the computing resources for processing the computing tasks which is expressed as:

[0049]

[0050] (2.3) Reward function:

[0051]

[0052] On this basis, the long-term discounted reward is:

[0053]

[0054] where γ n is the discount factor, which is between 0 and 1, and t0 represents the past moment.

[0055] Further, the MADDPG algorithm structure in step (3) includes an actor network and a critic network, where the actor network consists of an estimated actor network and a target actor network, and a critic network consists of an estimated critic network and a target critic network; the algorithm is divided into centralized training and distributed execution; in the centralized training phase, the server first obtains the states and actions of all agents, so the server can train the estimated critic network of each agent to maximize the Q value.

[0056] Furthermore, step (3) specifically includes the following steps:

[0057] For a single agent n, the method for obtaining and updating the Q value of the critic network is:

[0058]

[0059] In the formula are respectively the symbolic references for the environmental state, the overall actions of all agents, and calculating the expectation of the policy network; and are the parameters of the estimated critic network and the target critic network respectively;

[0060] The loss function is expressed as:

[0061]

[0062] In the formula, E represents calculating the expectation; δ is the time difference error, and the calculation formula is

[0063] Update the critic network parameters through stochastic gradient descent

[0064]

[0065] In the formula is the loss function of the critic network parameters, represents the gradient descent when updating the parameters;

[0066] Update the actor network parameters through stochastic gradient descent:

[0067]

[0068] represents the action policy taken by agent n based on the current state;

[0069] Adopt the soft update method for the parameters of the target actor network and the target critic network:

[0070]

[0071] where ε is the parameter update rate, which is between 0 and 1; and are the parameters of the estimated target actor network and the estimated actor network, respectively.

[0072] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0073] The present invention first models the scenario, models the communication model, the movement model, and the computing model, and proposes an optimization problem by analyzing the scenario. To solve the optimization problem, states, action spaces, and reward functions are set to transform the problem. Finally, the MADDPG algorithm is used to obtain the optimal resource allocation scheme to maximize the resource utility.

[0074] The present invention takes into account the resource allocation problem that exists when twin maintenance and computing task processing are carried out simultaneously, which is ignored in existing research. This aspect of research is supplemented, and a preliminary exploration of twin maintenance under this problem is carried out, that is, a general model is established. The simulation results show that, compared with other algorithms, the scheme proposed by the present invention maximizes the resource utility of each vehicle. BRIEF DESCRIPTION OF THE DRAWINGS

[0075] The drawings described herein are used to provide a further understanding of the embodiments of the present invention, form a part of this application, and do not limit the embodiments of the present invention. In the drawings:

[0076] Figure 1 is the system scenario diagram of the present invention.

[0077] Figure 2 is the movement model diagram of the present invention.

[0078] Figure 3 is the reward curve diagram of the present invention.

[0079] Figure 4 is the resource utilization rate diagram of the present invention.

[0080] Figure 5 is the resource utility comparison diagram of the present invention.

[0081] Figure 6 is the twin maintenance delay and computing task processing delay diagram of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0082] To further elaborate on the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the following detailed description is provided in conjunction with the accompanying drawings and preferred embodiments regarding the specific implementation manners, structures, features, and effects of the present invention.

[0083] A resource allocation method for digital twin maintenance and computing task processing specifically includes the following steps:

[0084] Step (1): Establish a system model. As Figure 1 shown, there are N vehicles and a base station equipped with a server. For clarity, let represent the number of vehicles traveling on the lane, and their initial positions follow a Poisson clustering process. Specifically, these vehicles will travel in one direction along the lane and generate computing tasks during the travel; due to limited computing resources, the vehicles will connect to the base station (BS) through the V2I link using the cellular interface (Uu), offload the computing tasks to the server for processing, and the calculation results will be returned to the vehicles. In addition, a DT model of these vehicles is established on the server to obtain the real-time operating status of the vehicles. At the same time, in order to maintain the digital twin models of these vehicles, the server needs to synchronize information with the vehicles to ensure that the models can accurately reflect the driving conditions of the vehicles.

[0085] 1) Mobility model:

[0086] Establish a spatial orthogonal coordinate system with the base station as the origin, as Figure 2 shown. The positive direction of the x-axis is the direction in which the vehicles travel along the lane, i.e., eastward, the positive direction of the y-axis is southward, and the positive direction of the z-axis is along the direction of the base station antenna. Therefore, the position of vehicle i on lane j at time slot t can be expressed as:

[0087] PO i (t) = (X i,j (t), Y i,j (t), 0), i ∈ {1, 2, 3, …, N} (1)

[0088] where X i,j (t) is the abscissa of vehicle i traveling on lane j at time slot t, and Y i,j (t) is the ordinate of vehicle i.

[0089] When the value τ of each time slot is small enough, it can be approximately considered that the position of the vehicle is constant in each time slot. At the same time, the position of the vehicle in the current time slot t is related to the position in the previous time slot t - 1. Therefore, the abscissa X i,j (t) of vehicle i traveling on lane j at time slot t can be further expressed as:

[0090] X i,j (t) = X i,j (t - 1) + τv i (2)

[0091] where τ is the duration of each time slot, and v i is the traveling speed of vehicle i.

[0092] Since vehicle i is traveling on lane j, in the present invention, ω0 represents the width of each lane, and L0 represents the distance between the first lane and the BS. Therefore, the ordinate Y of vehicle i i,j (t) can be written as:

[0093] Y i,j (t) = L0 + jω0, j ∈ (0, 1, 2, 3, …, J - 1) (3)

[0094] 2) Communication model:

[0095] In this model, the present invention considers applying orthogonal frequency division multiple access (OFDMA) to N V2I links, that is, N V2I links in the present invention are pre-allocated to orthogonal sub-bands, where N V2I links occupy N sub-bands. Therefore, the communication rate between vehicle i and the base station can be expressed as:

[0096]

[0097] where b refers to the base station, W is the channel bandwidth, p i is the transmission power of vehicle i, σ 2 is the channel noise.

[0098] The channel gain between vehicle i and the base station can be calculated as:

[0099]

[0100] where, is small-scale fading, is large-scale path fading and shadowing. For small-scale fading, a general fading model is adopted and calculated as:

[0101]

[0102] where obeys a circularly symmetric complex Gaussian distribution with unit variance, κ is the correlation coefficient, PO i (t) is the position of vehicle i at time slot t, P B (t) is the position of the base station, and β is the path loss exponent.

[0103] 3) Calculation model:

[0104] After establishing the digital twin model of vehicle i, in order to enable the digital twin model to accurately reflect the current situation of the vehicle, the digital twin of vehicle i in the VEC server needs to be synchronized with the information of vehicle i in the physical space. The present invention makes represent the information that vehicle i needs to synchronize and update at time slot t, where is the size of the information that vehicle i needs to synchronously update at the current moment t. is the CPU frequency required to update the information of vehicle i with a unit size, T i dt is the maximum delay limit for updating the information of vehicle i. The updated information needs to be first transmitted from vehicle i to the server equipped at the base station, and then the server allocates computing resources to maintain the digital twin of vehicle i. Due to hosting limitations, the maintenance of the digital twin on vehicle i is also related to the total number of digital twins that need to be maintained on its server. Additionally, in the present invention, the total amount of computing resources on the server is denoted as F, and the computing resources required for maintaining the digital twin of vehicle i are denoted as f i dt . Therefore, the time required to complete the maintenance of the digital twin of vehicle i can be calculated as:

[0105]

[0106] During the driving process, a vehicle may generate computing tasks. Due to insufficient computing power, the vehicle may choose to offload the computing tasks to an edge server for processing. In the present invention, the computing tasks generated by vehicle i at time slot t are represented as where is the size of the computing tasks generated by vehicle i at time slot t, is the CPU frequency required for a computing unit size task, T i tk is the maximum delay limit for processing the computing tasks. Similar to the previous digital twin maintenance, vehicle i first offloads the computing tasks to the server, and then the server allocates computing resources to process the computing tasks. Therefore, the time required to complete the computing tasks can be calculated as:

[0107]

[0108] where f i tk represents the computing resources requested by vehicle i for processing the computing tasks transmitted to the server.

[0109] It should be noted that the maintenance of the digital twin and the processing of the computing tasks are carried out simultaneously. Therefore, the sum of the resources f i dt allocated to the vehicle for maintaining the digital twin and the resources f i tk allocated for processing the computing tasks generated by the vehicle cannot exceed the total computing resources F of the server. That is to say, In the formula, 1 means that the vehicle number i starts from 1 for calculation and summation.

[0110] Since vehicle i performs digital twin maintenance and offloading calculation tasks simultaneously in time slot t, the VEC server will allocate computing resources for the two respectively, i.e., f i dt and f i tk . By adopting different computing resource allocation strategies, different information synchronization delays and task calculation delays can be obtained, thus affecting the maintenance delay of the digital twin and the processing delay of the calculation task The present invention defines a satisfaction function Q, which respectively represents the satisfaction of the digital twin maintenance delay and the calculation task processing delay of vehicle i under a certain resource allocation strategy. Therefore, the satisfaction function of the digital twin repair delay of vehicle i and the satisfaction function of the calculation task processing delay of vehicle i can be respectively expressed as:

[0111]

[0112] Through Equation (9) and Equation (10), the resource utility U i obtained by vehicle i under the resource allocation strategy ω i (ω i ) can be evaluated:

[0113]

[0114] In the formula, ρ is the weight factor between twin maintenance and calculation task processing, and its value is 0 < ρ < 1, which is used to measure the importance of the two tasks.

[0115] Step (2): Establish the state, action space and reward function.

[0116] 1) State space:

[0117] The updated information, calculation tasks, vehicle position, speed and channel gain can be expressed as:

[0118]

[0119] 2) Action space:

[0120] Each vehicle is regarded as an agent, and according to the observed state s n (t), it makes an action a n (t). The actions include the computing resources for twin maintenance and the computing resources for processing calculation tasks can be expressed as:

[0121]

[0122] 3) Reward function:

[0123]

[0124] Based on this, the long-term discount reward can be obtained as follows:

[0125]

[0126] where γ n is the discount factor, which is between 0 and 1, and t0 represents a past moment.

[0127] The MADDPG algorithm structure includes an actor network and a critic network, where the actor network consists of an estimated actor network and a target actor network. The actor network is the actor network, and the critic network is the critic network. These two networks are common networks in multi-agent algorithms. Similarly, a critic network consists of an estimated critic network and a target critic network. The algorithm is divided into centralized training and distributed execution. Specifically, since the server manages the key networks of each agent, the states and actions of all agents can be obtained. In the centralized training phase, the server first obtains the states and actions of all agents. Therefore, the server can train the estimated critical network of each agent to maximize the Q value.

[0128] Distributed execution means that agent n observes and obtains its local state and executes an action according to the policy π n Note that in the early stage, due to lack of experience, agent n will take actions randomly to explore more possible actions. When there is enough experience, agent n will take actions to maximize the reward. For a single agent n, the method of obtaining and updating the Q value of the critic network is as follows:

[0129]

[0130] where, are the environmental state, the overall actions of all agents, and the symbolic representation for calculating the expectation of the policy network, respectively. and are the parameters of the estimated critic network and the target critic network, respectively.

[0131] The critic network is controlled by parameters. To obtain the optimal parameters, the loss function needs to be known. Here, the loss function can be expressed as:

[0132]

[0133] where E represents calculating the expectation; δ is the time difference error, and the calculation formula is To minimize the loss function, the present invention updates the critic network parameters through stochastic gradient descent

[0134]

[0135] where is the loss function of the critic network parameters, represents the gradient descent when updating the parameters.

[0136] Due to the distributed execution of the actor network (deployed on each vehicle), the agent n takes actions by observing the local state. For the parameter update method of the actor network, the present invention selects the gradient descent method:

[0137]

[0138] At represents the action policy taken by the agent n based on the current state.

[0139] During the training process, to ensure the stability of the algorithm, the present invention adopts the method of soft update for the parameters of the target actor network and the target critic network:

[0140]

[0141] where ε is the parameter update rate, which is between 0 and 1. and are the parameters of the estimated target actor network and the estimated actor network respectively.

[0142] Embodiment:

[0143] In Figure 3 the average rewards of a resource allocation method for digital twin maintenance and computing task processing proposed by the present invention are compared with those of the SAC algorithm, the PPO algorithm, and the random algorithm. As the number of iterations increases, the average rewards of the other three algorithms except the random algorithm will eventually converge to a larger value. Specifically, when the number of iterations is small, due to lack of experience, the total amount of resources requested by the agent from the server may exceed the computing resources of the server itself. In addition, the agent may also be penalized for violating the constraints due to improper early resource allocation strategies. As the number of iterations gradually increases, the agent can explore appropriate resource allocation strategies within the total computing resources of the VEC server, resulting in the gradual convergence of the reward value. The fluctuations after convergence are due to the dynamic randomness of the environment, which will affect learning. Among the three algorithms, the average convergence reward value of the present invention is greater than those of the SAC algorithm and the PPO algorithm.

[0144] Figure 4 The resource utilization rates of different algorithms under different numbers of vehicles are compared. From Figure 4As can be seen from (a), compared with SAC and PPO, the present invention always maintains a high resource utilization rate. On the contrary, when the number of vehicles gradually increases, the resource utilization rates of SAC and PPO will gradually increase. For the present invention, the reason why the resource utilization rate is higher when the number of vehicles is 6 than other numbers is that the algorithm achieves the optimal resource allocation balance in this case, that is, there are neither idle resources nor overloaded resources. Specifically, when the number of vehicles is less than 6, the number of twin models to be maintained in the server is small, and the resource demand is small, so the resource utilization rate is slightly low. When the number of vehicles is more than 6, the resource demand increases, but the total amount of resources in the server is limited, resulting in a decrease in the overall resource utilization rate. This numerical result also shows that when the resources of the twin model are limited, there is an optimal number of service vehicles, which needs to be theoretically determined in the future. In fact, due to Figure 4 the centralized training mentioned in Figure 4 (b) and Figure 4 (c).

[0145] Figure 5 shows the highest efficiency that each vehicle can achieve under different numbers of vehicles. As the number of vehicles gradually increases, the efficiency that each vehicle can achieve also decreases. Obviously, the increase in the number of vehicles means an increase in the demand for twin maintenance and task processing. Due to the resource competition among vehicles, vehicles always request more resources to avoid the failure of these two tasks. However, the total resources of the server remain unchanged, so under different algorithm schedules, the resources that each vehicle can request will also decrease, thereby reducing the performance. In addition, since the present invention can be based on global information, it can better handle the situation of multi-vehicle competition.

[0146] Figure 6 compares the twin maintenance delay and task processing delay of different algorithms under different numbers of vehicles. It can be seen that both delays gradually increase with the increase in the number of vehicles. Specifically, in Figure 6In (a), except for the random algorithm, the other three algorithms can all complete the twin maintenance task before the maximum delay limit. Here, when the number of vehicles increases from 5 to 6, the reason for the slow growth of the delay is that in this case, the algorithm can allocate resources more effectively and reduce resource competition. When the number of vehicles increases from 4 to 5, the demand for resources increases, and the resource allocation has not reached the optimal state, resulting in a higher delay growth rate, but the total delay is still less than the delay when the number of vehicles is 5 and 6. Similarly, when the number of vehicles increases from 6 to 8, the resource demand further increases, and the resource competition intensifies, resulting in a significant increase in the delay time again. And in Figure 6 (b), all four types of algorithms meet the maximum time tolerance of the computing task. In addition, compared with the other three types of algorithms, the present invention can always achieve lower delay. This is because, on the one hand, as shown in Figure 4 , the resource utilization rate of the present invention is higher, so there are more schedulable resources. On the other hand, as described in Figure 5 , the present invention can better handle resource competition under multiple vehicles.

[0147]

[0148] Here, a supplementary explanation of the algorithm settings is given. The number of hidden layer neurons in the actor network and the critic network is 300, 100, 2 and 300, 100, 1 respectively. The exploration noise samples Gaussian noise.

[0149] The content not described in detail in the specification of the present invention belongs to the prior art well-known to those skilled in the art.

[0150] The above are only the preferred embodiments of the present invention, and do not impose any form of limitation on the present invention. Although the present invention has been disclosed above with the preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some changes or modifications to the equivalent embodiments by using the above-disclosed technical content within the scope of the technical solution of the present invention. However, as long as it does not depart from the content of the technical solution of the present invention, any brief modification, equivalent change and modification made to the above embodiments according to the technical essence of the present invention still fall within the scope of the technical solution of the present invention.

Claims

1. A resource allocation method for digital twin maintenance and computing task processing, characterized in that: It includes the following steps: (1) Establish a system model (1.1) Vehicle movement model modeling: Establish a spatial orthogonal coordinate system with the base station as the origin. The positive direction of the x-axis is the direction in which the vehicle travels along the lane, the positive direction of the y-axis is the direction along the increasing number of lanes, and the positive direction of the z-axis is the direction along the base station antenna. The position of vehicle i on lane j at time slot t is expressed as: PO i (t) = (X i,j (t), Y i,j (t), 0), i ∈ {1, 2, 3, …, N} (1) where X i,j (t) is the abscissa of vehicle i traveling on lane j at time slot t, and Y i,j (t) is the ordinate of vehicle i; The abscissa X of vehicle i traveling on lane j at time slot t i,j (t) is further expressed as: X i,j X(t) = i,j X(t - 1)+τv i (2) where τ is the duration of each time slot, and v i is the driving speed of vehicle i; Let ω0 denote the width of each lane and L0 denote the distance between the first lane and the base station. Thus, the ordinate Y i,j (t) is written as: Y i,j (t) = L0 + jω0, j ∈ (0, 1, 2, 3, …, J - 1) (3) (1.2) Communication model modeling: Communication rate between vehicle i and the base station Expressed as: where b refers to the base station, W is the channel bandwidth, p i is the transmission power of vehicle i, and σ 2 is the channel noise; Channel gain between vehicle i and the base station Calculated as: where is small-scale fading, is large-scale path fading and shadowing; for small-scale fading, a general fading model is adopted and calculated as follows: where obeys a circularly symmetric complex Gaussian distribution with unit variance, κ is the correlation coefficient, and PO i (t) is the position of vehicle i at time slot t, and P B (t) is the position of the base station, and β is the path loss exponent; (1.3) Computation model modeling: Let represent the information that vehicle i needs to synchronously update in time slot t, where is the size of the information that vehicle i needs to synchronously update at the current time t, is the CPU frequency required to update the information of vehicle i per unit size, is the maximum delay limit for updating the information of vehicle i; the updated information needs to be first transmitted from vehicle i to the server equipped at the base station, and then the server allocates computing resources to maintain the digital twin of vehicle i; Let the total amount of computing resources on the server be denoted as F, and the computing resources required for the twin maintenance of vehicle i be denoted as f i dt , so the time required to complete the digital twin maintenance of vehicle i is calculated as: Denote the computing task generated by vehicle \(i\) at time slot \(t\) as where is the size of the computing task generated by vehicle \(i\) at time slot \(t\), is the CPU frequency required for the computing unit size task, is the maximum latency limit for processing the computing task; vehicle \(i\) first offloads the computing task to the server, and then the server allocates computing resources to process the computing task. Therefore, the time required to complete the computing task is calculated as: where f i tk represents the computing resources requested by vehicle i for processing the computing tasks transmitted to the server; Define the satisfaction function \(Q\) to represent the satisfaction of the digital twin maintenance delay and the computing task processing delay of vehicle \(i\) under a certain resource allocation strategy, and the satisfaction function of the digital twin maintenance delay of vehicle \(i\). And the satisfaction function of the computing task processing delay of vehicle \(i\). They are respectively expressed as: Through equations (9) and (10), evaluate the resource utility U i obtained by vehicle i under the resource allocation strategy ω i (ω i ): In the formula, ρ is the weight factor between twin maintenance and computation task processing, and its value is 0 < ρ < 1, which is used to measure the importance of the two tasks; (2) Establish the state, action space, and reward function; (3) Obtain the optimal resource allocation method through the MADDPG algorithm to maximize the long-term reward.

2. The resource allocation method for digital twin maintenance and computing task processing according to claim 1, wherein: In the step (1.3), the maintenance of the digital twin and the processing of the computing tasks are carried out simultaneously. Therefore, the resource f allocated to the vehicle for maintaining the digital twin i dt and the resource f allocated for processing the computing tasks generated by the vehicle i tk The sum of cannot exceed the total computing resource F of the server, that is 3. The resource allocation method for digital twin maintenance and computing task processing according to claim 1, characterized in that: The specific steps of step (2) include the following steps: (2.1) State space: Select update information Γ n (t), computing task ζ n (t), vehicle position PO n (t), speed v n and channel gain Expressed as: (2.2) Action space: Each vehicle is regarded as an agent, and based on the observed state s n (t), it makes an action a n (t). The actions include computing resources for twin maintenance and computing resources for processing computing tasks which is expressed as: (2.3) Reward function: On this basis, the long-term discounted reward is obtained as: where γ n is the discount factor, which ranges from 0 to 1, and t0 represents a past moment.

4. The resource allocation method for digital twin maintenance and computing task processing according to claim 3, wherein: The MADDPG algorithm structure in step (3) includes an actor network and a critic network. The actor network consists of an estimated actor network and a target actor network, and a critic network consists of an estimated critic network and a target critic network; the algorithm is divided into centralized training and distributed execution; In the centralized training stage, the server first obtains the states and actions of all agents. Therefore, the server can train the estimated critic network of each agent to achieve the purpose of maximizing the Q value.

5. The resource allocation method for digital twin maintenance and computing task processing according to claim 4, wherein: The specific steps of step (3) include the following steps: For a single agent n, the method for obtaining and updating the Q value of the critic network is: where are the symbolic references for the environmental state, the overall actions of all agents, and the expected value of the computational policy network, respectively; and are the parameters of the estimated critic network and the target critic network, respectively; The loss function is expressed as: where E represents the calculated expectation; δ is the time difference error, and the calculation formula is Update the critic network parameters by stochastic gradient descent In the formula is the loss function of the critic network parameters, represents the gradient descent when updating the parameters; Update the actor network parameters through stochastic gradient descent: Denotes the action strategy adopted by agent n based on the current state; Adopt a soft update method for the parameters of the target actor network and the target critic network: where ε is the parameter update rate, which is between 0 and 1; and are the parameters of the estimated target actor network and the estimated actor network, respectively.