Resource allocation method for balancing quality and cost of Internet of Vehicles scene based on digital twinning
By building a digital twin-driven vehicle edge computing system and multi-agent deep reinforcement learning algorithm, the problem of unstable resource allocation in the Internet of Vehicles is solved, quality and cost optimization is achieved, and resource utilization efficiency is improved.
Patent Information
- Application Number
- CN202510595877.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-07-01
AI Technical Summary
In the scenario of Internet of Vehicles, the existing technology is difficult to achieve joint optimization of multi-dimensional resources, and cannot control system costs while ensuring the quality of digital twins, and the time-varying environment is insufficient, resulting in unstable resource allocation and waste of resources.
Build a digital twin-driven vehicle edge computing architecture, design new evaluation indicators of quality and cost, establish multi-objective optimization problems, and use multi-agent multi-objective deep reinforcement learning algorithm for resource allocation. Through distributed actor network storage and playback experience, evaluate agent actions to solve the optimal resource allocation strategy.
It has achieved significant reduction in resource overhead while ensuring the quality of digital twins, improved resource utilization efficiency, solved the problem of unstable resource allocation, and met the needs of multi-objective optimization.
Smart Images

Figure CN120238920A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of mobile communications and relates to a resource allocation method for quality and cost trade-off in a vehicle-to-everything (V2X) scenario based on digital twin. Background Art
[0002] In the V2X scenario, digital twin technology provides core support for traffic flow simulation, network resource management, and vehicle collaborative decision-making by mapping physical entities to the virtual space in full space-time. With the help of mobile communication, edge computing, and multi-source sensing technologies, vehicles can upload sensor data to the edge server in real time to construct a dynamic digital twin to reduce system latency and bandwidth pressure.
[0003] However, this process faces the following key challenges:
[0004] 1. Multi-dimensional resource coupling constraints: The modeling accuracy of the digital twin (such as update frequency, data granularity) directly depends on the coordinated allocation of computing, communication, and storage resources. For example, high-precision modeling requires high-frequency data upload and intensive computing, but limited by the limited computing resources of the edge server and the dynamic nature of the wireless channel, excessive resource occupancy will lead to a sharp increase in costs (such as energy consumption, hardware load, transmission overhead). Existing methods often optimize a single resource dimension separately and lack a quantification mechanism for quality-cost joint trade-off.
[0005] 2. Insufficient adaptability to time-varying environments: The high-speed mobility of vehicles in the V2X network leads to rapid changes in network topology and channel state, and static resource allocation strategies are difficult to respond in a timely manner. Traditional threshold-triggered adjustments are prone to cause resource allocation oscillations, resulting in fluctuations in modeling quality or resource waste, and it is difficult to ensure the consistency of the digital twin.
[0006] 3. Heterogeneous demand conflicts: Different traffic applications (such as collision warning, path planning) have significantly different quality requirements for the digital twin. Existing edge resource allocation mechanisms lack fine-grained quality of service (QoS) classification and are difficult to achieve a balance between cross-service priorities and global resource efficiency.
[0007] Current research mostly focuses on a single optimization goal (such as minimizing latency or maximizing accuracy), does not establish a quality-cost dynamic trade-off model, and rarely considers the impact of vehicle mobility on the stability of resource allocation. Therefore, there is an urgent need for a dynamic allocation method that supports multi-resource joint optimization and environmental adaptability to control system costs while ensuring the quality of the digital twin, thereby promoting the large-scale deployment of V2X digital twin technology. Summary of the Invention
[0008] In view of this, the purpose of the present invention is to provide a resource allocation method for quality and cost trade-off in a vehicle networking scenario based on digital twin, which achieves an optimal trade-off through reasonable system resource allocation. First, a digital twin-driven vehicle edge computing architecture is proposed, and data perception and uploading are modeled to reflect the real-time state of the vehicle's physical environment. Second, to implement the digital twin-driven vehicle edge computing architecture, new metrics for quantitatively evaluating system quality and cost are designed. On this basis, a multi-objective optimization problem is constructed to maximize system quality and minimize cost. Finally, a multi-agent multi-objective deep reinforcement learning algorithm is proposed, and a distributed actor for storing and replaying experiences and a learner with a dueling evaluation network for evaluating agent actions are designed.
[0009] To achieve the above object, the present invention provides the following technical solutions:
[0010] A resource allocation method for quality and cost trade-off in a vehicle networking scenario based on digital twin, the method comprising the following steps:
[0011] S1. Construct a digital twin-driven vehicle edge computing architecture, including at least a set of discrete time slots, a set of vehicles, and a set of edge nodes;
[0012] S2. Dynamically obtain the sensing decision, sensing frequency, and uploading priority of the vehicle in each time slot, and model the data sensing and uploading of the real-time state of the vehicle's physical environment;
[0013] S3. Establish a comprehensive quality evaluation function and a total cost function of the digital twin. Among them, the quality metrics include timeliness metrics and consistency metrics, and the cost metrics include redundancy, sensing cost, and transmission cost;
[0014] S4. Establish a multi-objective optimization problem, where the multi-objective optimization problem aims to maximize system quality and minimize cost;
[0015] S5. Solve the problem based on the multi-agent multi-objective deep reinforcement learning algorithm. Among them, the distributed actor network is used to store and replay experiences, and a learner with a dueling evaluation network is used to evaluate agent actions to solve the optimal resource allocation strategy.
[0016] Further, in step S1, in the process of constructing the set of discrete time slots, the set of discrete time slots is represented by T = {1, 2,..., t,..., T}, where T is the number of time slots; let D represent the set of heterogeneous information, and each piece of information d ∈ D is characterized by a triple d = (typed, ud, |d|), where typed, ud, and |d| are the type, update interval, and data size respectively;
[0017] In the established vehicle set S, each vehicle s is characterized by a triple where D s and π s are the position, the set of sensing information, and the transmission power respectively; for each piece of information d ∈ D s , the sensing cost in vehicle s is represented by ;
[0018] In the established edge node set E, each edge node e ∈ E is characterized by the triple e = (l e , r e , b e ), where l e , r e , and b e represent the position, the communication range, and the bandwidth respectively. The distance between vehicle s and edge node e is represented by , where distance(·) represents calculating the Euclidean distance;
[0019] At time t, the set of vehicles within the radio coverage of edge node e is represented by ; the sensing decision represents whether the information d is sensed by vehicle s at time t; the set of information sensed by vehicle s at time t is represented by , where the information types are different; the sensing frequency of vehicle s for information d at time t is represented by , which should meet the requirements of vehicle s' sensing ability, that is where represent the minimum and maximum sensing frequencies of vehicle s respectively;
[0020] At time t, the upload priority of information d in vehicle s is represented by , where the upload priorities of different types of information are different; the transmission power of vehicle s at time t is represented by , and it cannot exceed the maximum power of vehicle s; the V2I bandwidth allocated by edge node e to vehicle s at time t is represented by , and the total V2I bandwidth allocated by edge node e cannot exceed its capacity b e , that is
[0021] Let v′ represent physical entities, including vehicles, pedestrians, and roadside infrastructure, and represent the set of physical entities as V′; the set of information related to entity v′ is represented by D v′ , the size of D v′ is |D v′; For each entity, model the corresponding digital twin v in the edge node, use V to represent the set of digital twins, and use to represent the set of digital twins modeled in the edge node e at time t; Therefore, the set of information received by the edge node e and required by the digital twin v is represented as indicating the number of information.
[0022] Furthermore, in step S2, first model the collaborative perception based on the multi-class M / G / 1 priority queue to complete vehicle perception. Among them, assume that the information upload times of multiple data types follow a general distribution with a mean of and a variance of . Then the upload task data of vehicle s is represented as:
[0023]
[0024] In the formula, represents the perception frequency of vehicle s for information d at time t; According to the multi-class M / G / 1 priority queuing principle, the constraint
[0025] The arrival time of information d before time t is represented by and is calculated by the following formula:
[0026]
[0027] The update time of information d before time t is represented by and is calculated as:
[0028]
[0029] In the formula, u d is the update interval of information d;
[0030] At time t, the set of information in vehicle s with a higher upload priority than information d is represented by , where d* represents the information with high upload priority; Then the upload task data before information d is calculated by the following formula:
[0031]
[0032] In the formula, and are the perception frequency and average transmission time of information d* in vehicle s at time t, respectively;
[0033] According to the Pollaczek-Khintchine formula, the queuing time of information d in vehicle s is calculated as:
[0034]
[0035] In the formula, represents the arrival time of information d before time t for vehicle s, represents the distribution, represents the perception frequency of information d by vehicle s at time t, represents the task data that vehicle s needs to upload before task d at time t.
[0036] Furthermore, in step S2, then model the V2I upload based on the channel fading distribution and the SNR threshold. Among them, at time t, the SNR of the V2I communication between vehicle s and edge node e is calculated as:
[0037]
[0038] In the formula, N0 is the additive white Gaussian noise; h s,e is the channel fading gain; τ is a constant depending on the antenna design, is the path loss exponent;
[0039] Let |h s,e | 2 follow a distribution with a mean of μ s,e and a variance of σ s,e , which is expressed as:
[0040]
[0041] In the formula, represents the distribution, represents the expectation;
[0042] Measure the transmission reliability by the possibility that the successful transmission probability exceeds the reliability threshold:
[0043]
[0044] In the formula, and δ are the target SNR threshold and the reliability threshold respectively, it represents taking the minimum value for all possible probability distributions, represents the probability under the probability distribution P;
[0045] According to Shannon's theorem, at time t, the transmission rate of the V2I communication between vehicle s and edge node e is represented by and is calculated by the following formula:
[0046]
[0047] In the formula, represents the transmission bandwidth of vehicle s at time t;
[0048] Therefore, the duration for uploading information d from vehicle s to edge node e is calculated by the following formula:
[0049]
[0050] In the formula, is the moment when vehicle s starts to transmit information d.
[0051] Furthermore, in step S3, the digital twin quality assessment includes timeliness and consistency. Among them, timeliness is defined as the total time delay from information update to reception, and consistency is measured by comparing the update time differences of different information within the same twin; the system quality is defined as the average value of QDT of each digital twin modeled on the edge node within the scheduling period T, and the calculation is as follows:
[0052]
[0053] In the formula, QDT v is the quality of the digital twin corresponding to physical entity v;
[0054] The digital twin cost quantification includes the following parts: redundancy, sensing consumption, and transmission effect. Redundancy refers to the additional quantity of each information type repeatedly sensed by multiple vehicles; the system cost is defined as the average value of CDT of each digital twin model on the edge node within the scheduling period T, and the calculation is as follows:
[0055]
[0056] In the formula, CDT v is the total digital twin cost.
[0057] Furthermore, in step S4, a multi-objective optimization problem is established with the aim of simultaneously maximizing the system quality and minimizing the system cost, which is expressed as:
[0058]
[0059] In the formula, C1 to C5 are the corresponding parameter constraints; C6 ensures the queue steady state; C7 ensures the transmission reliability; C8 restricts that the sum of the V2I bandwidths allocated to edge node e cannot exceed its capacity b e .
[0060] Given a solution x = (C, Λ, P, Π, B), where C represents the determined sensed information, Λ represents the determined sensing frequency, P represents the determined upload priority, Π represents the determined transmission power, and Q represents the determined V2I bandwidth allocation:
[0061]
[0062] In the formula, and are the perception decision, perception frequency, and upload priority of information d in vehicle s at time t, respectively; and are the transmission power and V2I bandwidth of vehicle s at time t, respectively;
[0063] According to the definition of CDT, the profit PDT of the digital twin ν ∈(0, 1) is defined as the compensation of the CDT of digital twin v:
[0064] PDT ν = 1 - CDT ν
[0065] Define the system profit as the average value of PDT of each digital twin modeled in the edge node within the scheduling period T:
[0066]
[0067] Therefore, problem P1 is rewritten as follows:
[0068]
[0069] s.t. C1~C8
[0070] The rewritten optimization problem is transformed into solving to maximize the system quality and maximize the profit.
[0071] Furthermore, in step S5, the multi-agent multi-objective deep reinforcement learning algorithm includes the following steps:
[0072] Deploy a local policy network and an evaluation network, which are used to generate action policies and evaluate state values, respectively;
[0073] Adopt distributed actor execution experience replay, and store the interaction data in the replay buffer;
[0074] Design an action advantage network and a state value network to evaluate the advantage function and value function of the agent's behavior.
[0075] Furthermore, based on the above algorithm, the vehicle and the edge node determine their respective behaviors in a distributed manner through the local policy network. The local observation value of the system state of vehicle s at time t is expressed as:
[0076]
[0077] In the formula, t is the time slot index; s is the vehicle index; is the location where vehicle s is located; D sDenote the information set that vehicle s can perceive; Φ s Denote D s The perception cost of the information in D; Denote the cache information set of edge node e at time t; Denote the information set required for the digital twin modeled in edge node e at time t; w t Denote the weight vector of each target, which is randomly generated in each iteration. In particular, w t = [w (1),t w (2),t , where w (1),t ∈(0,1) and w (2),t ∈(0,1) are the weights of system quality and system cost respectively;
[0078] The local observation of the system state of edge node e at time t is expressed as:
[0079]
[0080] In the formula, e is the edge node index, Denote the distance set between the vehicle and edge node e;
[0081] Then the system state at time t is expressed as
[0082] The action of vehicle s is expressed as:
[0083]
[0084] In the formula, is the perception decision; and are the perception frequency and upload priority of the information respectively, is the transmission power of vehicle s at time t;
[0085] The action of the vehicle is generated by the local vehicle policy network according to its local observation of the system state:
[0086]
[0087] In the formula, is the exploration noise to increase the diversity of vehicle actions; ∈ s is the exploration constant of vehicle s;
[0088] The set of vehicle actions is expressed as
[0089] The action of edge node e is expressed as:
[0090]
[0091] In the formula, is the V2I bandwidth allocated by the edge node e to the vehicle s at time t;
[0092] The local edge policy network obtains the action of the edge node e based on the system state and the action of the vehicle:
[0093]
[0094] In the formula, and ∈ e are the detection noise and detection constant of the edge node e, respectively;
[0095] The joint action of the vehicle and the edge node is expressed as
[0096] The environment obtains the system reward vector by executing the joint action, which is expressed as:
[0097]
[0098] In the formula, and are the rewards of two objectives, namely the system quality and the system cost, and are calculated by the following formula:
[0099]
[0100] Then the reward of the vehicle s in the jth objective is obtained by reward allocation based on the difference reward, and this difference reward is the difference between the system reward and the reward obtained without taking action, which is expressed as:
[0101]
[0102] In the formula, is the system reward obtained without the contribution of the vehicle s, which is obtained by setting the zero action set of the vehicle s;
[0103] The reward vector of the vehicle s at time t is represented by and is
[0104] The set of difference rewards of the vehicle is represented as
[0105] The system reward is further transformed into the normalized reward of the edge node through min-max normalization. The formula for calculating the reward of the edge node e at time t in the jth objective is:
[0106]
[0107] In the formula, and are the same system state o respectivelyt Vehicle actions below The minimum and maximum values of the system rewards obtained when the vehicle actions remain unchanged; the reward vector of the edge node e at time t is represented by denoted as The interaction experience includes the system state o t and vehicle actions Edge actions Vehicle rewards Edge rewards Weight w t and the next system state o t+1 These interaction experiences are stored in the replay buffer and the interaction continues until the training process of the learner is completed.
[0108] The beneficial effects of the present invention are as follows:
[0109] First, by constructing a digital twin-driven vehicle edge computing architecture, the present invention realizes the accurate perception and efficient upload and modeling of the real-time state of the vehicle's physical environment. This architecture effectively integrates the computing resources of in-vehicle terminals, edge nodes, and cloud platforms, significantly improving the real-time performance and accuracy of data processing and analysis, and providing strong technical support for the construction and update of digital twins.
[0110] Secondly, the present invention designs a new quality and cost evaluation index, providing a scientific basis for the quantitative evaluation of system performance. By defining the timeliness and consistency indexes of digital twins, and quantifying cost elements such as redundancy, perception cost, and transmission cost, the present invention constructs a comprehensive quality evaluation function and a total cost function, providing a clear goal orientation for system optimization. By jointly optimizing the quality of digital twins and system costs, the present invention solves the contradiction between the two, significantly reducing resource overhead while ensuring the accuracy of the twin model, and its performance is superior to traditional benchmark schemes.
[0111] The present invention constructs a multi-objective optimization problem aiming to simultaneously maximize system quality and minimize cost. By introducing a multi-agent multi-objective deep reinforcement learning algorithm, the scheme successfully solves the optimal resource allocation strategy, effectively solving the problem that it is difficult to balance quality and cost in traditional methods. Through a distributed Actor-Critic architecture, this algorithm realizes the collaborative optimization of vehicles and edge nodes, significantly improving the system resource utilization efficiency.
[0112] Other advantages, objectives, and features of the present invention will be described to some extent in the subsequent specification, and to some extent, will be obvious to those skilled in the art based on the study of the following text, or can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be realized and obtained through the following specification. Brief Description of the Drawings
[0113] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be described in detail and preferably below in conjunction with the accompanying drawings, where:
[0114] Figure 1 is a digital twin network architecture diagram for the vehicle networking under the embodiments of the present invention;
[0115] Figure 2 is a training flow chart of a multi-agent multi-objective deep reinforcement learning algorithm under the embodiments of the present invention. Detailed Embodiments
[0116] The following uses specific specific examples to illustrate the embodiments of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the drawings provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0117] Among them, the drawings are only for illustrative purposes, showing only schematic diagrams, rather than physical diagrams, and cannot be construed as limitations on the present invention; in order to better illustrate the embodiments of the present invention, some components in the drawings will be omitted, enlarged, or reduced, which do not represent the dimensions of actual products; for those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.
[0118] In the drawings of the embodiments of the present invention, the same or similar reference numerals correspond to the same or similar components; in the description of the present invention, it should be understood that if there are terms such as "upper", "lower", "left", "right", "front", "rear", etc. indicating the orientation or positional relationship, they are based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, the terms describing the positional relationship in the drawings are only for illustrative purposes and cannot be construed as limitations on the present invention. For those of ordinary skill in the art, the specific meanings of the above terms can be understood according to specific circumstances.
[0119] Please refer to Figures 1 to 2 , which is a resource allocation method for quality and cost trade-off in vehicle networking scenarios based on digital twin.
[0120] Embodiment
[0121] This embodiment details the specific process of a resource allocation method for quality and cost trade-off in a vehicle networking scenario based on digital twins, which mainly includes the following steps:
[0122] S1. Construct a digital twin-driven vehicle edge computing architecture, including at least a discrete time slot set, a vehicle set, and an edge node set;
[0123] S2. Dynamically obtain the sensing decisions, sensing frequencies, and upload priorities of vehicles in each time slot, and perform data sensing and upload modeling on the real-time state of the vehicle physical environment;
[0124] S3. Establish a comprehensive quality evaluation function and a total cost function for the digital twin. Among them, the quality indicators include timeliness indicators and consistency indicators, and the cost indicators include redundancy, sensing cost, and transmission cost;
[0125] S4. Establish a multi-objective optimization problem, where the multi-objective optimization problem aims to maximize system quality and minimize cost;
[0126] S5. Solve the problem based on the multi-agent multi-objective deep reinforcement learning algorithm. Among them, replay experiences are stored through a distributed actor network, and a learner with a dueling evaluation network is used to evaluate agent actions to solve the optimal resource allocation strategy.
[0127] In step S1 of this embodiment, as Figure 1 shown, it includes three parts: constructing a discrete time slot set, establishing a vehicle set S, and deploying an edge node set E. Among them, in the process of constructing the discrete time slot set, the set of discrete time slots is represented by T = {1, 2,..., t,..., T}, where T is the number of time slots. Let D represent the set of heterogeneous information, and each piece of information d ∈ D is characterized by a triple d = (typed, ud, |d|), where typed, ud, and |d| are the type, update interval, and data size, respectively. In the established vehicle set S, each vehicle s is characterized by a triple characterized, where D s and π s are the location, the set of sensing information, and the transmission power, respectively. For each piece of information d ∈ D s , the sensing cost (i.e., energy consumption) in vehicle s is represented by . In the established edge node set E, each edge node e ∈ E is characterized by a triple e = (l e , r e , b e ), where l e , r e and b erespectively represent the location, communication range, and bandwidth. The distance between vehicle s and edge node e is represented by where distance(·) represents calculating the Euclidean distance. At time t, the set of vehicles within the radio coverage of edge node e is represented by Perception decision represents whether information d is perceived by vehicle s at time t; the set of information perceived by vehicle s at time t is represented by where the information types are different. The perception frequency of information d by vehicle s at time t is represented by which should meet the requirements of vehicle s' perception ability, that is where respectively represent the minimum and maximum perception frequencies of vehicle s. At time t, the upload priority of information d in vehicle s is represented by where the upload priorities of different types of information are different. The transmission power of vehicle s at time t is represented by and it cannot exceed the maximum power of vehicle s. The V2I bandwidth allocated by edge node e to vehicle s at time t is represented by and the total V2I bandwidth allocated by edge node e cannot exceed its capacity b e , that is Let v′ represent physical entities such as vehicles, pedestrians, and roadside infrastructure, and represent the set of physical entities as V′. The set of information related to entity v′ is represented by D v′ and the size of D v′ is |D v′ |. For each entity, a corresponding digital twin v may be modeled in the edge node. Let V represent the set of digital twins, and the set of digital twins modeled in edge node e at time t is represented by . Therefore, the set of information received by edge node e and required by digital twin v can be represented as represents the number of information.
[0128] In step S2 of this embodiment, multi-class M / G / 1 priority queues are used to model collaborative perception to complete vehicle perception. Among them, let the upload time of type typed information follow a general distribution with a mean of and a variance of , then the upload task data of vehicle s is expressed as:
[0129]
[0130] In the formula, represents the perception frequency of information d by vehicle s at time t;
[0131] According to the multi-class M / G / 1 priority queuing principle, for the queue to have a steady state The arrival time of information d before time t is given by and is calculated by the following formula:
[0132]
[0133] The update time of information d before time t is given by and is calculated as:
[0134]
[0135] where u d is the update interval of information d.
[0136] At time t, the set of information in vehicle s with a higher upload priority than information d is given by where d* represents high-upload-priority information. Therefore, the upload task data before information d (i.e., the amount of elements that vehicle s needs to upload before d at time t) is calculated by the following formula:
[0137]
[0138] where and are the sensing frequency and average transmission time of information d* in vehicle s at time t, respectively. According to the Pollaczek-Khintchine formula, the queuing time of information d in vehicle s is calculated as:
[0139]
[0140] Then, V2I upload is modeled based on the channel fading distribution and SNR threshold. At time t, the SNR of V2I communication between vehicle s and edge node e is calculated as:
[0141]
[0142] where N0 is the additive white Gaussian noise; h s,e is the channel fading gain; τ is a constant depending on the antenna design, is the path loss exponent. Assume that |h s,e | 2 follows a distribution with a mean μ s,e and a variance σ s,e and is expressed as:
[0143]
[0144] where represents the probability distribution, Denotes the expectation under the distribution P.
[0145] The transmission reliability is measured by the likelihood that the successful transmission probability exceeds the reliability threshold:
[0146]
[0147] where and δ are the target SNR threshold and the reliability threshold respectively, denotes taking the minimum over all possible probability distributions, denotes the probability under the probability distribution P.
[0148] According to Shannon's theorem, the transmission rate of V2I communication between vehicle s and edge node e at time t is given by denoted and calculated by the following formula:
[0149]
[0150] where denotes the transmission bandwidth of vehicle s at time t.
[0151] Therefore, the duration for uploading information d from vehicle s to edge node e is calculated by the following formula:
[0152]
[0153] where is the time when vehicle s starts to transmit information d.
[0154] In step S3 of this embodiment, the digital twin quality assessment includes timeliness and consistency. Among them, timeliness is defined as the total time delay from information update to reception. For example, if a piece of information is updated at multiple time points, its timeliness is the sum of the differences between all update times and the reception time. Consistency is measured by comparing the update time differences of different information within the same twin. Specifically, the maximum value of the update time differences for all information pairs is taken to reflect the data synchronization degree.
[0155] The system quality is defined as the average value of QDT of each digital twin modeled on the edge node within the scheduling period T, and is calculated as follows:
[0156]
[0157] The quantification of digital twin cost includes the following parts: redundancy, sensing consumption, and transmission effect. Redundancy refers to the additional quantity of each information type being repeatedly sensed by multiple vehicles. From this, the total cost is obtained, comprehensively considering redundancy, sensing energy consumption, and transmission energy consumption.
[0158] The system cost is defined as the average value of each digital twin model CDT on the edge node within the scheduling period T, and is calculated as follows:
[0159]
[0160] In step S4 of this embodiment, given a solution x = (C, Λ, P, Π, B), where C represents the determined sensing information, Λ represents the determined sensing frequency, P represents the determined upload priority, Π represents the determined transmission power, and Q represents the determined V2I bandwidth allocation:
[0161]
[0162] In the formula, and are the sensing decision, sensing frequency, and upload priority of information d in vehicle s at time t, respectively; and are the transmission power and V2I bandwidth of vehicle s at time t, respectively.
[0163] A multi-objective optimization problem is established with the aim of simultaneously maximizing the system quality and minimizing the system cost, which is expressed as:
[0164]
[0165] In the formula, C1 to C5 are the corresponding parameter constraints; C6 ensures the queue stability; C7 ensures the transmission reliability; C8 limits that the sum of the V2I bandwidths allocated to the edge node e cannot exceed its capacity b e .
[0166] According to the definition of CDT, the profit of the digital twin PDT ν ∈(0, 1) is defined as the compensation for the CDT of the digital twin v:
[0167] PDT ν = 1 - CDT ν
[0168] Furthermore, the system profit is defined as the average value of PDT of each digital twin modeled in the edge node within the scheduling period T:
[0169]
[0170] Therefore, problem P1 can be rewritten as follows:
[0171]
[0172] s.t. C1 to C8
[0173] In step S5 of this embodiment, as Figure 2 shown, the multi-agent multi-objective deep reinforcement learning algorithm includes the following steps:
[0174] Deploy local policy network and evaluation network, which are used to generate action policies and evaluate state values respectively;
[0175] Adopt distributed actor to perform experience replay, and store interaction data into the replay buffer;
[0176] Design action advantage network and state value network to evaluate the advantage function and value function of the agent's behavior.
[0177] Specifically, in the algorithm, the vehicle and the edge node determine their respective behaviors in a distributed manner through the local policy network. The local observation value of the system state of vehicle s at time t is expressed as:
[0178]
[0179] In the formula, t is the time slot index; s is the vehicle index; is the location where vehicle s is located; D s represents the information set that vehicle s can perceive; Φ s represents D s the perception cost of the information in; represents the cache information set of edge node e at time t; represents the information set required for the digital twin modeled in edge node e at time t; w t represents the weight vector of each objective, which is randomly generated in each iteration. In particular, w t =[w (1),t w (2),t , where w (1),t ∈(0,1) and w (2),t ∈(0,1) are the weights of system quality and system cost respectively. On the other hand, the local observation value of the system state of edge node e at time t is expressed as:
[0180]
[0181] In the formula, e is the edge node index, represents the distance set between the vehicle and edge node e. Therefore, the system state at time t can be expressed as
[0182] The action of vehicle s can be expressed as:
[0183]
[0184] In the formula, is the perception decision; and are the perceived frequency and upload priority of the information respectively, and is the transmission power of vehicle s at time t. The actions of the vehicle are generated by the local vehicle policy network based on its local observation of the system state:
[0185]
[0186] where is the exploration noise for increasing the diversity of vehicle actions; ∈ s is the exploration constant of vehicle s. The set of vehicle actions is denoted as The action of edge node e is denoted as:
[0187]
[0188] where is the V2I bandwidth allocated by edge node e to vehicle s at time t. Similarly, the local edge policy network can obtain the action of edge node e based on the system state and the action of the vehicle:
[0189]
[0190] where and ∈ e are the exploration noise and exploration constant of edge node e respectively. Further, the joint action of the vehicle and the edge node is denoted as
[0191] The environment obtains the system reward vector by executing the joint action, denoted as:
[0192]
[0193] where and are the rewards for two objectives (i.e., system quality and system cost) respectively, which can be calculated by the following formula:
[0194]
[0195] Therefore, the reward of vehicle s in the jth objective is obtained by reward allocation based on the differential reward, which is the difference between the system reward and the reward obtained without taking action, and can be expressed as:
[0196]
[0197] where is the system reward obtained without the contribution of vehicle s, which can be obtained by setting the zero action set of vehicle s. The reward vector of vehicle s at time t is denoted by and is The difference reward set of the vehicle is represented as
[0198] On the other hand, the system reward is further transformed into the normalized reward of the edge node through min-max normalization. The formula for calculating the reward of the edge node e at time t in the jth target is as follows:
[0199]
[0200] In the formula, and are respectively the minimum and maximum values of the system rewards obtained when the vehicle actions t remain unchanged under the same system state o. The reward vector of the edge node e at time t is represented by The interaction experience includes the system state o The vehicle actions The edge actions t The vehicle rewards The edge rewards The weight w The edge rewards The weight w t and the next system state o t+1 , and these interaction experiences are stored in the replay buffer . This interaction will continue until the training process of the learner is completed.
[0201] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the present technical solution, and they should all be covered by the scope of the claims of the present invention.
Claims
1. A resource allocation method based on digital twins to balance the quality and cost of vehicle networking scenarios, characterized by: The method comprises the following steps: S1. Build a digital twin-driven vehicle edge computing architecture, including at least a discrete time slot set, a vehicle set, and an edge node set; S2, dynamically obtain the vehicle's perception decision, perception frequency and upload priority in each time slot, and perform data perception and upload modeling on the real-time status of the vehicle's physical environment; S3. Establish a comprehensive quality evaluation function of the digital twin and a total cost function of the digital twin, where the quality indicators include timeliness indicators and consistency indicators, and the cost indicators include redundancy, perception cost, and transmission cost; S4. Establish a multi-objective optimization problem, wherein the multi-objective optimization problem takes maximizing system quality and minimizing cost as the goal; S5. Solve problems based on a multi-agent multi-objective deep reinforcement learning algorithm, in which the experience is replayed through a distributed actor network, and a learner with a duel evaluation network is used to evaluate the agent's actions to solve the optimal resource allocation strategy.
2. According to claim 1, a resource allocation method for balancing the quality and cost of vehicle networking scenarios based on digital twins is characterized by: In step S1, in the process of constructing a discrete time slot set, the set of discrete time slots is represented by T = {1, 2, ..., t, ..., T}, where T is the number of time slots; let D represent a set of heterogeneous information, and each information d∈D is characterized by a triple d = (typed, ud, |d|), where typed, ud, |d| are type, update interval and data size respectively; In the established vehicle set S, each vehicle s consists of a triple portrayal, among which D s and π s They are location, sensing information set and transmission power respectively; for each information d∈D s , the perception cost in vehicle s is given by express; In the established edge node set E, each edge node e∈E is composed of a triplet e=(l e ,r e ,b e ) characterizes, where l e 、r e and b e denote the location, communication range and bandwidth respectively; the distance between vehicle s and edge node e is given by Indicates that, where distance(·) means to find the Euclidean distance; At time t, the set of vehicles within the radio coverage of edge node e is given by Representation; Perception Decision Indicates whether information d is perceived by vehicle s at time t; the information set perceived by vehicle s at time t is represented by , where the information types are different; the perception frequency of vehicle s to information d at time t is expressed as It means that it should meet the requirements of vehicle s perception ability, that is, in They represent the minimum and maximum values of the perception frequency of vehicle s respectively; At time t, the upload priority of information d in vehicle s is given by Indicates that different types of information upload have different priorities; the transmission power of vehicle s at time t is expressed as , and it cannot exceed the maximum power of vehicle s; the V2I bandwidth allocated by edge node e to vehicle s at time t is given by Indicates that the total V2I bandwidth allocated to edge node e cannot exceed its capacity b e ,Right now Let v′ represent a physical entity, including vehicles, pedestrians, and roadside infrastructure, and denote the set of physical entities as V′; the information set related to entity v′ is represented by D v′ Indicates that D v′ The size is |D v′ |; For each entity, the corresponding digital twin v is modeled in the edge node, V is used to represent the set of digital twins, and the set of digital twins modeled in the edge node e at time t is represented by Therefore, the set of information received by the edge node e and required by the digital twin v is expressed as Indicates the amount of information.
3. The resource allocation method for balancing the quality and cost of vehicle networking scenarios based on digital twins according to claim 2 is characterized by: In step S2, collaborative perception is first modeled based on multi-class M / G / 1 priority queues to complete vehicle perception, where the information upload time for each data type is The mean is The variance is The general distribution of , the uploaded task data of vehicle s is expressed as: In the formula, represents the frequency of vehicle s’s perception of information d at time t; according to the multi-class M / G / 1 priority queuing principle, the constraint The arrival time of information d before time t is given by It is expressed as follows: The update time of information d before time t is given by It is expressed as: In the formula, u d is the update interval of information d; At time t, the set of information in vehicle s that has a higher upload priority than information d is given by Indicates, where d* represents high upload priority information; then the upload task data before information d is calculated by the following formula: In the formula, and are the perception frequency and average transmission time of the information d* in vehicle s at time t, respectively; According to the Polachek-Khinchin formula, the queuing time of information d in vehicle s is calculated as: In the formula, represents the arrival time of information d of vehicle s before time t, represents the variance, represents the perception frequency of vehicle s to information d at time t, Represents the task data that vehicle s needs to upload before task d at time t.
4. The resource allocation method for balancing the quality and cost of vehicle networking scenarios based on digital twins according to claim 3 is characterized by: In step S2, the V2I upload is then modeled based on the channel fading distribution and the SNR threshold, where the SNR of the V2I communication between vehicle s and edge node e at time t is calculated as: Where N0 is additive white Gaussian noise; h s,e is the channel fading gain; τ is a constant that depends on the antenna design, is the path loss exponent; Assume h s,e | 2 is a class with mean μ s,e and variance σ s,e The distribution of is expressed as: In the formula, represents the distribution, express expectations; Transmission reliability is measured by the probability that the successful transmission probability exceeds a reliability threshold: In the formula, and δ are the target SNR threshold and reliability threshold respectively, It means taking the minimum value of all possible probability distributions. represents the probability under the probability distribution P; According to Shannon’s theorem, at time t, the transmission rate of V2I communication between vehicle s and edge node e is given by It is expressed as follows: In the formula, represents the transmission bandwidth of vehicle s at time t; Therefore, the duration of uploading information d from vehicle s to edge node e is Calculated by the following formula: In the formula, is the moment when vehicle s starts to transmit information d.
5. The resource allocation method for balancing the quality and cost of vehicle networking scenarios based on digital twins according to claim 4 is characterized by: In step S3, the quality assessment of digital twins includes timeliness and consistency, where timeliness is defined as the total time delay from information update to reception, and consistency is measured by comparing the update time differences of different information in the same twin; system quality is defined as the average value of the quality QDT of each digital twin modeled on the edge node within the scheduling period T, which is calculated as follows: Where, QDT v is the mass of the digital twin corresponding to the physical entity v; The quantification of digital twin cost includes the following parts: redundancy, perception consumption and transmission effect. Redundancy refers to the additional number of each information type repeatedly perceived by multiple vehicles. The system cost is defined as the average CDT of each digital twin model on the edge node within the scheduling period T, which is calculated as follows: In the formula, CDT v is the total cost of the digital twin.
6. The resource allocation method for balancing the quality and cost of vehicle networking scenarios based on digital twins according to claim 5 is characterized by: In step S4, a multi-objective optimization problem is established, the goal is to maximize the system quality and minimize the system cost at the same time, which is expressed as: Where C1~C5 are corresponding parameter constraints; C6 ensures the steady state of the queue; C7 ensures transmission reliability; C8 limits the sum of the V2I bandwidth allocated to edge node e to not exceed its capacity b e ; Given a solution x = (C, Λ, P, Π, B), where C represents the determined sensing information, Λ represents the determined sensing frequency, P represents the determined upload priority, Π represents the determined transmission power, and Q represents the determined V2I bandwidth allocation: In the formula, and They are the perception decision, perception frequency and upload priority of information d in vehicle s at time t respectively; and are the transmission power and V2I bandwidth of vehicle s at time t, respectively; According to the definition of CDT, the profit of digital twin is PDT ν ∈(0, 1) is defined as the compensation of the CDT of the digital twin v: PDT ν =1-CDT ν Defining System Profit is the average value of PDT of each digital twin modeled in the edge node within the scheduling period T: Therefore, problem P1 can be rewritten as follows: stC1~C8 The rewritten optimization problem is transformed into solving the problem of maximizing system quality and maximizing profit.
7. The resource allocation method for balancing the quality and cost of vehicle networking scenarios based on digital twins according to claim 6 is characterized by: In step S5, the multi-agent multi-objective deep reinforcement learning algorithm includes the following steps: Deploy a local policy network and evaluation network to generate action strategies and evaluate state values respectively; Use distributed actors to perform experience replay and store interaction data in the replay buffer; Design action advantage network and state value network to evaluate the advantage function and value function of the agent's behavior.
8. The resource allocation method for balancing the quality and cost of vehicle networking scenarios based on digital twins according to claim 7 is characterized by: Based on the above algorithm, vehicles and edge nodes determine their respective behaviors in a distributed manner through the local policy network. The local observation value of the system state of vehicle s at time t is expressed as: Where t is the time slot index; s is the vehicle index; is the location of vehicle s; D s represents the information set that vehicle s can perceive; Φ s Indicates D s the perceived cost of information; represents the cache information set of edge node e at time t; represents the information set required for modeling the digital twin in the edge node e at time t; w t Represents the weight vector of each target, which is randomly generated in each iteration; w t =[w (1),t w (2),t ], where w (1),t ∈(0,1) and w (2),t ∈(0,1) are the weights of system quality and system cost respectively; The local observation value of the system state of edge node e at time t is expressed as: Where e is the edge node index, represents the distance set between the vehicle and the edge node e; Then the system state at time t is expressed as The action of vehicle s is expressed as: In the formula, for perceived decision making; and Information Perceived frequency and upload priority, is the transmission power of vehicle s at time t; The vehicle's actions are generated by the local vehicle policy network based on its local observations of the system state: In the formula, Exploration noise to increase the diversity of vehicle actions; ∈ s is the exploration constant of vehicle s; The set of vehicle actions is represented as The action of edge node e is expressed as: In the formula, is the V2I bandwidth allocated by edge node e to vehicle s at time t; The local edge strategy network obtains the action of edge node e according to the system state and the action of the vehicle: In the formula, and ∈ e are the detection noise and detection constant of edge node e respectively; The joint action of the vehicle and the edge node is expressed as The environment obtains the system reward vector by performing joint actions, expressed as: In the formula, and are the rewards for the two objectives, namely system quality and system cost, and are calculated as follows: Then the reward of vehicle s in the jth goal is obtained by the reward distribution based on the difference reward, which is the difference between the system reward and the reward obtained when no action is taken, expressed as: In the formula, is the system reward obtained without the contribution of vehicle s, obtained by setting the zero action set of vehicle s; The reward vector of vehicle s at time t is Indicates that The difference reward set of the vehicle is expressed as The system reward is further converted into the normalized reward of the edge node through minimum and maximum normalization. The reward calculation formula of the edge node e at time t in the jth target is: In the formula, and The same system state is o t Vehicle movement The minimum and maximum values of the system reward obtained when it remains unchanged; the reward vector of the edge node e at time t is expressed as express, Interaction experience includes system state o t , Vehicle Action Edge Action Vehicle Rewards Edge Rewards Weight w t and the next system state o t+1 , these interaction experiences are stored in the replay buffer The interaction continues until the training process of the learner is completed.
Citation Information
Cited By
A traffic event digital twin reproduction method based on vehicle-road cloud perception data
CN122616352A