A method for allocating sensing and computing resources and deploying vehicle digital twins in a vehicle networking

By constructing a digital twin-assisted vehicle-to-everything (V2X) network integrating sensing, communication, and computing, designing a dynamic frame structure, and employing multi-agent deep reinforcement learning, the system optimizes the utility of perception data and the deployment of digital twins. This solves the problems of uneven allocation of perception and communication resources and unbalanced edge computing load in V2X, thereby improving system performance.

CN119110315BActive Publication Date: 2025-10-17CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411091402.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-09
Publication Date
2025-10-17
Estimated Expiration
2044-08-09

AI Technical Summary

Technical Problem

In the Internet of Vehicles, how to balance the perception and communication time resource allocation of telepresence integrated devices to ensure perception accuracy and communication efficiency, while reasonably allocating edge computing server resources to avoid load imbalance.

Method used

We construct a vehicle-to-everything (V2X) network with digital twin assistance, design a dynamic frame structure based on integrated sensing technology, optimize the utility of perception data and the deployment of digital twins through multi-agent deep reinforcement learning, and establish an optimization model using information age and energy consumption evaluation functions to minimize the information age of perception data and the deployment and migration costs of digital twins.

Benefits of technology

It has improved the utility of perception data and balanced the load of edge computing resources, reduced the migration cost of digital twins, and improved the quality of system services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119110315B_ABST
    Figure CN119110315B_ABST
Patent Text Reader

Abstract

The present application relates to a kind of vehicle networking in the method for allocating and vehicle digital twin deployment of sensing algorithm resources, belong to mobile communication technical field.The method includes: S1: construct digital twin assisted vehicle networking sensing algorithm integration network;S2: design vehicle dynamic frame based on sensing integration technology;S3: construct the utility evaluation function of vehicle perception data, solve the utility of perception information;S4: build vehicle digital twin deployment strategy, calculate vehicle digital twin deployment cost and migration utility;S5: establish the optimization model of minimizing perception data information age and vehicle digital twin deployment and migration cost, adopt multi-agent deep reinforcement learning to solve optimal vehicle sensing algorithm resource allocation and vehicle digital twin deployment scheme.The present application reduces the average information age of perception data and digital twin migration rate by dynamic frame design and vehicle digital twin deployment strategy, effectively improves the utility of perception data and edge computing resource load balancing.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of mobile communication, and relates to a method for allocating sensing and computing resources and deploying a digital twin of a vehicle in a vehicle networking. BACKGROUND

[0002] Integrated sensing, communication and computing (ISCC) is one of the important technologies for realizing intelligent transportation, and a main implementation scheme is to combine the integrated sensing and communication technology with mobile edge computing. The integrated sensing and communication technology realizes sensing and communication functions through sharing wireless spectrum resources. Mobile edge computing (MEC) can provide computing and storage resources for mobile vehicles by sinking computing capability to the edge, thereby reducing computing delay.

[0003] Digital twin (DT) technology is a technology for mapping a physical entity to a digital space in real time, and can capture dynamic state information of the physical entity in real time. Meanwhile, the digital twin technology can quickly and accurately provide computing and communication resources in a physical space according to requirements of a user. Deploying the digital twin on an edge server can realize low-latency interaction with the corresponding entity.

[0004] When the integrated sensing and communication device is used, although a long sensing time can ensure the accuracy and effectiveness of sensing, a large amount of sensing data can cause longer transmission delay and reduce data utility, and a too short sensing time can cause inaccurate sensing and cannot guarantee safety. How to balance the allocation of sensing and communication time resources in the integrated sensing and communication needs to be solved. Similarly, an edge computing server usually provides services for multiple terminal users, and multiple vehicles can compete for the same server resource, while other servers are in an idle state. How to reasonably allocate server resources to terminal users to ensure effective load balancing needs to be solved. SUMMARY

[0005] Therefore, the application aims to provide a method for allocating sensing and computing resources and deploying a digital twin of a vehicle in a vehicle networking, so as to realize optimal allocation of communication and sensing resources of the vehicle and optimal deployment of the digital twin, thereby maximizing sensing data utility and edge computing resource load balancing and improving system service quality.

[0006] To achieve the above-mentioned purpose, the application provides the following technical scheme.

[0007] A method for allocating sensing and computing resources and deploying a digital twin of a vehicle in a vehicle networking scenario, comprising:

[0008] S1: constructing an integrated sensing and computing network of a vehicle networking assisted by a digital twin;

[0009] S2: Design the vehicle dynamic frame structure based on the integrated sensing technology;

[0010] S3: Construct the utility evaluation function of vehicle perception data, and solve the utility of perception information;

[0011] S4: Build a vehicle digital twin deployment strategy, and calculate the vehicle digital twin deployment cost and migration utility;

[0012] S5: Establish an optimization model that minimizes the perception data information age and the vehicle digital twin deployment and migration cost;

[0013] S6: Use multi-agent deep reinforcement learning to solve the optimal sensing algorithm resource allocation and vehicle digital twin deployment scheme.

[0014] Further, in S1, the proposed digital twin assisted vehicle networking sensing algorithm integration network includes two layers: the terminal layer and the edge layer. The terminal layer in S1 includes a set of vehicles with limited computing and storage capabilities, denoted as N = {1, 2, 3,..., n}. The edge layer includes a set of base stations that provide access services and a set of edge servers that provide computing and storage services. Each base station is equipped with an edge server. The set of base stations v is denoted as V = {1, 2,..., v}, and the set of edge servers m is denoted as M = {1, 2, 3,..., m}. Consider a set of time slots K with a long life cycle, and the twin information of each vehicle is updated within a single time slot. The set of time slots k is set as K = {1, 2,.., k}. In each time slot, the vehicle and the base station update the state.

[0015] Further, in S2, following the current mainstream use scheme of integrated sensing devices, a time division dynamic frame structure based on 5GNR (5th Generation New Radio, 5GNR) is used. Compared with 4GLTE (4th Generation Long-Term Evolution, 4GLTE), the frame time slot can be adjusted, so that the vehicle can dynamically adjust the proportion of communication and perception time according to the demand.

[0016] Further, in S2, the vehicle uses Orthogonal Frequency Division Multiplexing (OFDM) signals for sensing. The conditional mutual information between the impulse response and the echo signal is used to evaluate the sensing accuracy. The throughput is used to measure the communication performance between the vehicle and the base station.

[0017] Further, in S3, the data utility of perception data in different locations is calculated according to the Age of Information (AoI). The vehicle perception information transmission process includes three parts: vehicles perform perception, upload perception information to base stations, and base stations forward information to their corresponding edge servers where digital twins are deployed. A nonlinear function is used to simulate the utility of perception information for the AoI. The specific utility function is defined as u(t) = 1 / t 2 . The average AoI utility of perception data is obtained by integrating the AoI of each location respectively, and the AoI-based utility Q u of vehicle n is obtained by accumulation.

[0018] Further, in S4, the association relationship between vehicle n and edge server m is represented by matrix X = [x nm ]. If x nm = 1, the digital twin (DT) of vehicle n is deployed at edge server m. The association matrix is represented as: x represents the association relationship between vehicle n and edge server m, n belongs to one of the vehicles in the set N, and m belongs to one of the edge servers in the set M.

[0019]

[0020] Further, in S4, the DT deployment strategy produces different delays and energy consumptions. Specifically, it includes the wired transmission delay of perception data between the base station and the edge server, the wired transmission delay of the data used by the user to build the DT uploaded to the edge server through the base station, and the delay produced by the ES processing the data uploaded by the user. Let F be the delay required to transmit one unit of data in each unit distance. Then the wired transmission delay from the nearby base station v of the perception data vehicle n to the edge server m where the digital twin is deployed is:

[0021]

[0022] where, is the size of the perception data generated during the perception duration.

[0023] Let be the DT data that vehicle n needs to upload at time slot k, then the synchronization delay of the DT transmitted by vehicle through base station v to edge server m under the above deployment strategy is

[0024]

[0025] p v,m represents the transmission power of base station v to edge server m. Without considering the loss, the transmission energy consumption is represented as:

[0026]

[0027] The number of users DT served by each edge server cannot exceed its storage and computing load, for edge server m, its storage size is C m , and its computing resource is F m , the occupied storage size of vehicle n deploying DT is set as The computing resource required for processing vehicle tasks is set as The delay required for processing uploaded data is:

[0028]

[0029] Further, in S4, the deployment cost of the vehicle digital twin is:

[0030]

[0031] wherein T s represents the time length of a time slot.

[0032] Further, in S4, the migration strategy is represented as x' n , and specifically as x' n = 1, vehicle n triggers DT migration, otherwise no migration is performed. The entire migration strategy matrix is represented as: x' = (x'1, x'2,..., x' n )

[0033] The vehicle migration utility function U(k) is defined, which quantifies the impact of DT migration of the vehicle at time slot k based on latency and energy consumption; the server node energy consumption of DT of vehicle n migrating between edge server m and edge server j is represented as:

[0034]

[0035] wherein wherein g j (k) is a binary variable representing whether edge server j is in an active state, if in an active state, represented as g j (k) = 1. Otherwise, represented as g j (k) = 0; represents the computing resource utilization rate of edge server m; F m is the computing resource of edge server m; represents the state transition energy consumption of edge server j at time slot k.

[0036] The vehicle utility function is represented as:

[0037]

[0038] wherein, denotes the impact of the satisfaction of the vehicle latency on the utility, μ n denotes the latency requirement of vehicle n, T n (k) denotes the actual interaction latency of vehicle n, μ n -T n (k) increases, also increases, denoting the increase of the utility of vehicle n. γ ∈ (0, 1) is the weight factor, denotes the digital twin migration utility of vehicle n.

[0039] Further, in S5, the proposed optimization model that minimizes the age of the perception data information and the deployment and migration cost of the vehicle digital twin is denoted as:

[0040]

[0041]

[0042]

[0043]

[0044]

[0045]

[0046]

[0047]

[0048]

[0049]

[0050]

[0051] C11: ψ1+ ψ2+ ψ3= 1, ψ1, ψ2, ψ3∈ (0, 1)

[0052] Q u , Q pl , Q trans denote the perception data utility, the digital twin deployment cost, and the digital twin migration utility, respectively; τ n and α n denote the perception and communication duration within a single time slot of the vehicle, respectively; denotes the vehicle perception mutual information, denotes the communication uplink rate between the vehicle and the base station; denotes the vehicle perception frequency; τ n , α n , βn They represent vehicle perception time, communication transmission time, and wired transmission delay between base station and edge server respectively; T s is the duration of a time slot; Indicates the computing resource utilization of the edge server; They represent the computing resources and storage resources required for the vehicle digital twin; F m ,C m Represent the total computing and storage resources of the edge server, respectively; ψ1, ψ2, and ψ3 represent weighting factors. Base station constraint C1 ensures vehicle perception quality. Constraint C2 states that the uplink communication link rate when the vehicle communicates with the base station cannot exceed the maximum value. Constraint C3 states that the vehicle's perception frequency cannot exceed the maximum perception frequency. Constraint C4 states that the total system latency cannot exceed the length of one frame. Constraint C5 states that the edge server's computing resource utilization is less than 1. Constraints C6 and C7 ensure that the vehicle's DT can only be deployed on one edge server. Constraints C8 and C9 state that the vehicle's DT deployment and computing requirements cannot exceed the server's maximum cache and computing space, respectively. Constraint C10 states that the vehicle's storage and computing resource requirements for the edge server are non-negative. Constraint C11 represents the relationship between the weighting factors.

[0053] Furthermore, in S6, the specific steps for solving the optimal vehicle perception communication resource allocation and digital twin deployment method are as follows: first, the objective function is modeled as a multi-agent partially observable Markov decision process, and the base station in the scene is regarded as an agent. The specific elements are defined as follows:

[0054] (1) State: The global state space S includes all vehicle state information, all vehicle DT deployment information, and the resource utilization of all edge servers. Vehicle state information includes the data to be transmitted by the vehicle, the available communication transmission rate between the vehicle and the base station, and the AoI during the vehicle transmission process;

[0055] (2) Action space: Each agent selects its actions, including the communication and perception time allocation and vehicle DT deployment strategy within the agent's coverage area;

[0056] (3) State transition probability: After the agent performs an action on the environment in the current state, the environment state will transition to the next state with a certain probability;

[0057] (4) Reward function: The reward function should be designed to enable each agent to make optimal decisions on DT deployment and synaesthesia resource allocation, and is therefore designed based on the agent's contribution to the system.

[0058] Further, in S6, the multi-agent "centralized training, distributed execution" method is used to complete the cooperative learning to achieve the optimal VDT deployment and communication-aware resource allocation strategy. The algorithm training process specifically includes the following steps:

[0059] S61: initialize Q network parameters and deterministic policy network parameters and experience replay pool D;

[0060] S62: collect initial state

[0061] S63: calculate continuous action parameters τ v (k) and α v (k) according to the gradient descent method;

[0062] The formula of the gradient descent method is:

[0063]

[0064] where λ d is the learning rate of the Q network, λ s and λ ve represent the weights of the local parameters and the fusion parameters of the agent, respectively;

[0065] is the network parameter used to update the target Q network, is the network parameter used to update the deterministic policy network;

[0066] S64: sample an action a v,k from the exploration probability ε and the probability distribution ψ select the discrete action that maximizes the discrete action value function

[0067] τ v (k), α v (k). Where s v,k represents the state, represents the deployment strategy of the vehicle DT in the edge server,

[0068] τ v (k), α v (k) represent the vehicle perception and communication duration, respectively.

[0069] S65: execute action a v,k , obtain instantaneous reward r v,k and next state s v,k+1

[0070] S66: tuple (s v,k,α v,k ,r v,k ,s v,k+1 ) Deposit into experience replay pool D v

[0071] S67: Calculate the loss function of the Q network and the deterministic policy network based on the loss function and

[0072] The loss functions are:

[0073]

[0074]

[0075] S68: Update Q network parameters according to gradient descent method and

[0076]

[0077]

[0078] where λ s and λ ve are the weights of the agent’s local parameters and fusion parameters respectively;

[0079] S69: Update converged network parameters: and

[0080]

[0081] Among them, ← is the update symbol, which means updating the calculated fusion network parameters. is the updated parameter of the Q network at the next moment, Update the parameters of the deterministic policy network for the next moment. represents the update gradient of the deterministic network parameters, is the update gradient of the Q network parameters, λ v ' e is the updated fusion network parameter weight. represents the global loss function of the fusion network, represents the expectation about the strategy, y ve represents the target value of the fusion network, Q me represents the global Q function.

[0082] S610: Network parameters are updated and sent to each agent.

[0083] The beneficial effects of the present invention are:

[0084] (1) The dynamic frame structure based on the perception and communication integration technology designed by the application can dynamically adjust the proportion of the time length occupied by vehicle perception and communication. The perception data utility function based on information age can improve the utility of perception data while ensuring the quality of communication.

[0085] (2) The vehicle digital twin deployment strategy based on data transmission energy consumption and delay evaluation proposed by the application can reasonably deploy vehicle digital twin in the edge server, reasonably allocate edge computing resources, and avoid edge load congestion.

[0086] (3) The digital twin migration strategy based on the user digital twin migration utility function evaluation constructed based on migration delay and energy consumption proposed by the application effectively reduces the digital twin migration cost.

[0087] Other advantages, objects and features of the application will be set forth in part in the following specification, and in part will become apparent to those skilled in the art from a reading of the following specification, or can be learned from practice of the application. The objects and other advantages of the application can be realized and attained by the methods and instrumentalities particularly pointed out in the following description. BRIEF DESCRIPTION OF DRAWINGS

[0088] In order to make the purpose, technical scheme and advantages of the application clearer, the preferred detailed description of the application will be combined with the drawings as follows, wherein:

[0089] Figure 1 Digital twin assisted car networking perception and communication integrated network architecture diagram;

[0090] Figure 2 Dynamic frame structure based on perception and communication integration design;

[0091] Figure 3 Training flowchart for communication, perception and computing resource and vehicle digital twin deployment method for car networking scene;

[0092] Figure 4 Multi-agent deep reinforcement learning block diagram based on centralized training-distributed execution. DETAILED DESCRIPTION

[0093] The following describes the embodiments of the present invention by means of specific examples, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present invention, and the following embodiments and features in the embodiments can be combined with each other without conflict.

[0094] Among them, the accompanying drawings are only for illustrative purposes and represent only schematic diagrams rather than actual pictures, and should not be understood as limiting the present invention. In order to better illustrate the embodiments of the present invention, some parts of the accompanying drawings may be omitted, enlarged or reduced, and do not represent the dimensions of actual products. For those skilled in the art, it is understandable that some well-known structures and their descriptions may be omitted in the accompanying drawings.

[0095] The same or similar numbers in the drawings of the embodiments of the present invention correspond to the same or similar parts; in the description of the present invention, it should be understood that if there are terms such as "upper", "lower", "left", "right", "front", "back", etc. indicating directions or positional relationships, they are based on the directions or positional relationships shown in the drawings. They are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific direction, be constructed and operate in a specific direction. Therefore, the terms describing the positional relationship in the drawings are only used for illustrative purposes and cannot be understood as limiting the present invention. For ordinary technicians in this field, the specific meanings of the above terms can be understood according to specific circumstances.

[0096] The network architecture diagram of the vehicle network integrated with telematics and computing assisted by digital twins in this embodiment is as follows: Figure 1 As shown, the process of the communication perception computing resource allocation and vehicle digital twin deployment method for the Internet of Vehicles scenario is as follows Figure 2 As shown, implementing this method specifically includes the following steps:

[0097] S1: Build a digital twin-assisted connected vehicle network with integrated sensing and computing capabilities;

[0098] S2: Design the vehicle dynamic frame structure based on synaesthesia integration technology;

[0099] S3: Construct the information age utility function of vehicle perception data and solve the perception information utility;

[0100] S4: Construct a vehicle digital twin deployment strategy and calculate the vehicle digital twin deployment cost and twin migration utility;

[0101] S5: Establish an optimization model that minimizes the perception data information age and the deployment and migration cost of the vehicle digital twin;

[0102] S6: Use multi-agent deep reinforcement learning to solve the optimal sensing algorithm resource allocation and vehicle digital twin deployment scheme. In S1 above, the proposed digital twin assisted vehicle networking sensing algorithm integrated network includes two layers: the terminal layer and the edge layer. The terminal layer in S1 includes a set of vehicles with limited computing and storage capabilities, denoted as N = {1, 2, 3,..., n}. Each vehicle is equipped with a sensing and communication integrated device, which can sense the surrounding environment and communicate with the base station to upload data. The edge layer in S1 includes a set of base stations that provide access services and a set of edge servers that provide computing and storage services. Each base station is equipped with an edge server. The set of base stations v is denoted as V = {1, 2,..., v}, and the set of edge servers m is denoted as M = {1, 2, 3,..., m}. Consider a set of time slots K with a long life cycle, and the twin information of each vehicle is updated within a single time slot. The set of time slots k is set as K = {1, 2,..., k}. In each time slot, the vehicle and the base station update their states.

[0103] In S2 above, a time division dynamic frame structure based on 5G NR is used, and the dynamic frame structure design is shown in Figure 3 Compared with 4G LTE, the frame time slot can be adjusted to dynamically adjust the proportion of communication and sensing time according to the business needs of the vehicle.

[0104] S21: The vehicle uses orthogonal frequency division multiplexing (OFDM) signals for sensing, denoted as

[0105]

[0106] where f c is the center frequency, T s is equal to the time length of the complete OFDM signal; represents the waveform amplitude of device n at time slot k; represents the phase code of the modulation symbol of device n at time slot k; rect[x] is a matrix function, which is 1 when 0 ≤ x ≤ 1, and 0 otherwise.

[0107] S22: Use the conditional mutual information between the impulse response and the echo signal to evaluate the sensing accuracy. MI is written as:

[0108]

[0109] τ sen represents the sensing process, represents the interference (SINR) generated by vehicle n when sensing at time slot k. Specifically, it is represented as:

[0110]

[0111] S23: The communication performance between vehicles and base stations is measured using throughput, the throughput between vehicle n and base station v at time slot k can be expressed as:

[0112]

[0113] SINR between vehicle n and base station v at time k. Specifically, it is expressed as:

[0114]

[0115] represents the channel gain between vehicle n and base station v.

[0116] In S3 above, the data utility of perception data at different locations is calculated according to the information age. The process of vehicle perception information transmission includes three parts: vehicles perform perception, upload perception information to base stations, and base stations forward information to their corresponding edge servers deployed with digital twins.

[0117] S31: Define s(t) to represent the generation time of the latest received information until the instantaneous time t, which is specifically expressed as

[0118]

[0119] S32: The AoI at the instantaneous time t can be given as

[0120] Δ(t) = t - s(t)

[0121] S33: A nonlinear function is used to simulate the utility of perception information for information age. The specific utility function is defined as u(t) = 1 / t 2 According to the evolution of information age at each location, the average information age utility of perception data is obtained by integrating respectively:

[0122] S34: For vehicles:

[0123]

[0124] S35: For base station servers:

[0125]

[0126] S36: For edge servers:

[0127]

[0128] S37: Obtain the information age-based utility function of vehicle n according to the utility function:

[0129]

[0130] In S4 above, the association relationship between vehicle n and edge server m is represented by matrix X = [x nm ; if x nm = 1, the digital twin (DT) of vehicle n is deployed at edge server m. The association matrix is represented as: x nm represents the association relationship between vehicle n and edge server m, n belongs to one of the vehicles in the vehicle set N, and m belongs to one of the edge servers in the edge server set M. Different deployment strategies result in different delays and energy consumptions. Specifically, it includes the wired transmission delay of perception data between the base station and the edge server, the wired transmission delay of the data used by the user to build the DT uploaded to the edge server through the base station, and the delay generated by the ES processing the data uploaded by the user.

[0131] S42: Let F be the delay required to transmit one unit of data in each unit distance. Then the wired transmission delay from the nearby base station v of the perception data vehicle n to the edge server m where its digital twin is deployed is:

[0132]

[0133] wherein, is the size of the perception data generated during the perception duration.

[0134] S43: Let be the DT data that vehicle n needs to upload at time slot k, then the synchronization delay of the DT transmitted by vehicle to edge server m through base station v under the above deployment strategy is

[0135]

[0136] p v,m represents the transmission power of base station v to edge server m, and the transmission energy consumption is represented as:

[0137]

[0138] S44: The number of user DTs served by each edge server cannot exceed its storage and computing load. For edge server m, its storage size is C m , the computing resource is F m , and the storage size occupied by the DT of vehicle n is set as The computing resource required for processing vehicle tasks is set as The delay required to process the uploaded data is:

[0139]

[0140] S45: The deployment cost of the vehicle digital twin is:

[0141]

[0142] Among them, T s Indicates the length of a time slot.

[0143] S46: Migration strategy is represented by x' n , specifically expressed as x' n =1, vehicle n triggers DT migration, otherwise no migration occurs. The entire migration strategy matrix is ​​expressed as:

[0144] x'=(x'1,x'2,…,x' n )

[0145] S47: Define the user migration utility function U(k) to quantify the DT migration impact of the vehicle at time slot k based on latency and energy consumption. The computing resource utilization of edge server m at time slot k can be expressed as

[0146]

[0147] Due to the migration of DTs, the running states of some edge servers may change. For example, suppose edge server j did not run any DT before, but since DTs are migrated to this node, the node will transition from dormant state to active state, resulting in state transition energy consumption. represents the state transition energy consumption of edge server j at time slot k:

[0148]

[0149] G j (k) indicates whether there is a state transition in edge server j, as follows:

[0150]

[0151] Therefore, the server node energy consumption when the DT of vehicle n migrates between edge servers m, j is expressed as:

[0152]

[0153] S48: The utility function can be expressed as:

[0154]

[0155] in, denotes the impact of the vehicle latency satisfaction on the utility, μ n denotes the latency requirement of vehicle n, T n (k) denotes the actual interaction latency of vehicle n, μ n -T n (k) increases, increases as well, denoting the increase of the utility of vehicle n. γ ∈ (0, 1) is the weight factor, denotes the digital twin migration utility of vehicle n.

[0156] Further, S5, the proposed objective function that maximizes the perception data utility and minimizes the vehicle digital twin deployment and migration cost is expressed as:

[0157]

[0158]

[0159]

[0160]

[0161]

[0162]

[0163]

[0164]

[0165]

[0166]

[0167]

[0168] C11: ψ1+ ψ2+ ψ3= 1, ψ1, ψ2, ψ3∈ (0, 1)

[0169] Q u , Q pl , Q trans denote the perception data utility, digital twin deployment cost, digital twin migration utility, respectively; τ n and α n denote the perception and communication duration within a single time slot of the vehicle, respectively; denotes the vehicle perception mutual information, denotes the communication uplink rate between the vehicle and the base station; denotes the vehicle perception frequency; τ n , α n,β n respectively represent the vehicle perception time, the communication transmission time, the base station and edge server wired transmission delay; T s is the length of a time slot; represents the computing resource utilization of the edge server; respectively represent the computing resource and storage resource required by the vehicle digital twin; F m ,C m respectively represent the total amount of computing and storage resources of the edge server; ψ1, ψ2, ψ3 respectively represent the weight factors. The base station constraint C1 guarantees the vehicle perception quality, the constraint C2 represents that the uplink communication link rate cannot exceed the maximum value when the vehicle communicates with the base station, the constraint C3 represents that the perception frequency of the vehicle cannot exceed the maximum perception frequency, the constraint C4 represents that the total delay of the system cannot exceed the length of a frame, the constraint C5 represents that the computing resource utilization of the edge server is less than 1, the constraints C6 and C7 guarantee that the DT of the vehicle can only be deployed on one edge server, and the constraints C8 and C9 respectively represent that the DT deployment and computing demand of the vehicle cannot exceed the maximum cache and computing space of the server. The constraint C10 represents that the storage and computing resource demand of the vehicle to the edge server is not negative, and the constraint C11 represents the relationship between the weight factors.

[0170] Further, S6, the solving method of the optimal vehicle perception communication resource allocation and digital twin deployment is a specific step: first, model the objective function as a multi-agent partially observable Markov decision process, and regard the base station in the scene as an agent. The specific element definitions are as follows:

[0171] (1) State: global state space including all vehicle state information ζ n , all vehicle DT deployment information X, and all resource utilization of the edge server e m . The vehicle state information includes the data to be transmitted by the vehicle and the available communication transmission rate between the vehicle and the base station, and the AoI in the vehicle transmission information process. The definitions are as follows

[0172]

[0173] (2) Action space: each agent selects its action including the communication and perception time length allocation of the vehicle in the coverage range of the agent and the vehicle DT deployment strategy. The action space of multiple agents can be written as

[0174]

[0175] (3) State transition probability: the agent executes the action a v,k in the current state s v,kAfterwards, the environmental state will transition to the next state with a certain probability, denoted as

[0176]

[0177] (4) Reward function: The design of the reward function should enable each agent to make the optimal decision on the DT deployment and the sensing resource allocation, thus it is designed according to the contribution degree of the agent to the system. The reward of agent v in time slot k is defined as

[0178]

[0179] where λ n,v is a binary variable representing the association between vehicle n and agent j, and Q n represents the contribution size of vehicle n to the system. Therefore, the global reward of the system at time slot k is

[0180]

[0181] Further, in S6, the multi-agent “centralized training, decentralized execution” method is adopted to complete the cooperative learning, so as to realize the optimal VDT deployment and communication sensing resource allocation strategy. The algorithm training process specifically includes the following steps:

[0182] S61: initialize the Q network parameters and the deterministic policy network parameters and the experience replay pool D;

[0183] S62: collect the initial state

[0184] S63: calculate the continuous action parameters τ v (k) and α v (k) according to the gradient descent method;

[0185] The formula of the gradient descent method is:

[0186]

[0187] where λ d is the learning rate of the Q network, λ s and λ ve represent the weights of the local parameters and the fusion parameters of the agent, respectively;

[0188] is the network parameter used to update the target Q network, is the network parameter used to update the deterministic policy network;

[0189] S64: sample an action a v,kselecting a discrete action value function the largest discrete action simultaneously determining a corresponding continuous action

[0190] τ v (k),α v (k) respectively represent the vehicle perception and communication duration. v,k denotes the state, denotes the deployment strategy of the vehicle DT in the edge server,

[0191] τ v (k),α v (k) respectively represent the vehicle perception and communication duration.

[0192] S65: performing action a v,k , obtaining instantaneous reward r v,k and next state s v,k+1

[0193] S66: storing tuple (s v,k ,α v,k ,r v,k ,s v,k+1 ) into experience replay pool

[0194] S67: calculating the loss functions of Q network and deterministic policy network according to loss function respectively and

[0195] The loss functions are respectively:

[0196]

[0197]

[0198] S68: updating Q network parameters according to gradient descent method and

[0199]

[0200]

[0201] where λ s and λ ve are the weights of local parameters and fusion parameters of the agent respectively;

[0202] S69: updating fusion network parameters: and

[0203]

[0204] wherein, is an update symbol, indicating updating the fusion network parameters calculated, is the updated parameter of the Q network at the next moment, is the updated parameter of the deterministic policy network at the next moment. represents the update gradient of the deterministic network parameter, is the update gradient of the Q network parameter, λ v e is the updated fusion network parameter weight. represents the global loss function of the fusion network, E represents the expectation about the policy, y ve represents the target value of the fusion network, Q me represents the global Q function.

[0205] S610: The network parameters are updated and distributed to each agent.

[0206] Finally, it should be pointed out that the above embodiments are only used to illustrate the technical solutions of the present application and are not limiting. Although the present application has been described in detail with reference to the preferred embodiments, it should be understood by those skilled in the art that the technical solutions of the present application can be modified or replaced equivalently without departing from the purpose and scope of the technical solutions, which should be covered in the scope of the claims of the present application.​

Claims

1. A method for allocating telematics resources and deploying vehicle digital twins in an Internet of Vehicles, characterized by: The method includes the following steps: S1: Build a digital twin-assisted connected vehicle network with integrated sensing and computing capabilities; S2: Design the vehicle dynamic frame structure based on synaesthesia integration technology; S3: Construct a utility evaluation function for vehicle perception data and solve the utility of perception information; S4: Construct a vehicle digital twin deployment strategy and calculate the vehicle digital twin deployment cost and digital twin migration utility; S5: Establish an optimization model to minimize the age of perception data information and the deployment and migration cost of vehicle digital twins; S6: Using multi-agent deep reinforcement learning to solve the problem of synaesthesia resource allocation and vehicle digital twin deployment; In S1, the proposed digital twin-assisted IoV integrated network includes two layers: the terminal layer and the edge layer; The terminal layer includes vehicles with limited computing and storage capabilities. The set of vehicles n is denoted as N = {1, 2, 3, ..., n}; The edge layer includes base stations that provide access services and edge servers that provide computing and storage services. Each base station is equipped with an edge server. The digital twin of the vehicle is deployed in the edge server. The set of base stations v is denoted as V = {1, 2, ..., υ}, and the set of edge servers m is denoted as M = {1, 2, 3, ..., m}. Considering a long-lifetime time slot set K, the twin information of each vehicle is updated within its single time slot k. The set of time slots is set to K = {1, 2, ..., k}. In S2, a time-division dynamic frame structure based on 5GNR is used to enable vehicles to dynamically adjust the ratio of communication and perception durations according to business needs. Vehicles use orthogonal frequency division multiplexing (OFDM) signals for perception. Conditional mutual information and throughput are used to evaluate perception accuracy and communication performance between the vehicle and the base station, respectively. In S3, the data utility of perception data at different locations is evaluated based on the age of information (AoI). The vehicle perception information transmission process includes three parts: the vehicle perceives, uploads the perception information to the base station, and the base station forwards the information to the edge server where the digital twin is deployed. A nonlinear function is used to simulate the utility of perceived information for information age. The average information age utility of the perceived data is obtained by integrating and summing the information age evolution of each position. The specific utility function is defined as u(t) = 1 / t 2 , where variable t is time; the average information age utility of the perception data is obtained by integrating the information age evolution of each location, and the information age-based utility Q of vehicle n is obtained by accumulating u ; In the above S4, the matrix X=[x nm ] represents the association relationship between vehicle n and edge server m; if x nm =1, then the digital twin DT of vehicle n is deployed at edge server m; the association matrix is ​​expressed as: x nm Represents the association between vehicle n and edge server m, where n belongs to a vehicle in the vehicle set N and m belongs to an edge server in the edge server set M. The initial digital twin deployment strategy is to deploy the vehicle digital twin at the edge server closest to it. DT deployment strategies result in different delays and energy consumption, including the wired transmission delay of sensor data between the base station and the edge server, the wired transmission delay of the vehicle's DT data uploaded to the edge server through the base station, and the delay caused by the ES processing the vehicle's uploaded data. Let Φ be the delay required to transmit one unit of data per unit distance, and d(v,m) represent the wired transmission distance between base station v and edge server m. Then, the wired transmission delay of sensor data from the base station v near vehicle n to the edge server m where its digital twin is deployed is: in, The amount of perception data generated during the perception process; set up For the constructed DT data that vehicle n needs to upload in time slot k, the synchronization delay of the DT transmitted by the vehicle to the edge server m through the base station v under the above deployment strategy is: p υ,m represents the transmission power from base station v to edge server m. Without considering the loss, the transmission energy consumption is expressed as: The number of vehicles DT served by each edge server cannot exceed its storage and computing load. For edge server m, its storage size is C m , the computing resources are F m , the storage size occupied by vehicle n deployment DT is set to The computing resources required to process the vehicle task are set as The delay required to process the uploaded data is: The deployment cost of a vehicle digital twin is: Among them, T s Indicates the length of a time slot; The migration strategy is represented by x' n , specifically expressed as x' n =1, vehicle n triggers DT migration, otherwise no migration occurs; the entire migration strategy matrix is ​​expressed as: x'=(x'1,x'2,...,x' n ) The vehicle migration utility function U(k) is defined to quantify the impact of vehicle DT migration at time slot k based on delay and energy consumption. The server node energy consumption when vehicle n's DT migrates from edge server m to edge server j is expressed as: Among them, g j (k) is a binary variable indicating whether edge server j is active. If it is active, it is represented by g j (k)=1; otherwise, it is expressed as g j (k) = 0; represents the computing resource utilization of edge server m; F m is the computing resource of edge server m; represents the state transition energy consumption of edge server j at time slot k; The vehicle utility function is expressed as: in, represents the impact of vehicle delay satisfaction on utility, μ n represents the delay requirement of vehicle n, T n (k) represents the actual interaction delay of vehicle n, μ n -T n (k) increases, Also increases, indicating that the utility of vehicle n increases; γ∈(0,1) is the weight factor, represents the digital twin migration utility of vehicle n; In S5, the optimization model for minimizing the age of perception data information and the deployment and migration cost of vehicle digital twins is expressed as: C11:ψ1+ψ2+ψ3=1,ψ1,ψ2,ψ3∈(0,1) They represent the perception data utility of vehicle n, the digital twin deployment cost, and the digital twin migration utility respectively; τ n represents the duration of vehicle perception in a single time slot, α n Indicates the communication duration within a single time slot of the vehicle; represents the mutual information of vehicle perception, Indicates the communication uplink rate between the vehicle and the base station; represents the vehicle perception frequency; τ n ,α n ,β n They represent the vehicle perception time, communication transmission time, and wired transmission delay between base station and edge server respectively; T s is the duration of a time slot; Indicates the computing resource utilization of the edge server; represents the computing resources required for the vehicle digital twin, F represents the storage resources required by the vehicle digital twin; m represents the total computing resources of the edge server, C m represents the total storage resources of the edge server; ψ1, ψ2, and ψ3 represent weight factors respectively; base station constraint C1 ensures the vehicle perception quality, constraint C2 indicates that the uplink communication link rate cannot exceed the maximum value when the vehicle communicates with the base station, constraint C3 indicates that the vehicle's perception frequency cannot exceed the maximum perception frequency, constraint C4 indicates that the total system delay cannot exceed the length of one frame, constraint C5 indicates that the computing resource utilization of the edge server is less than 1, constraints C6 and C7 ensure that the vehicle's DT can only be deployed on one edge server, constraints C8 and C9 indicate that the vehicle's DT deployment and computing requirements cannot exceed the server's maximum cache and computing space respectively; constraint C10 indicates that the vehicle's storage and computing resource requirements for the edge server are not negative, and constraint C11 represents the relationship between the weight factors; In S6, the specific steps for solving the optimal vehicle perception communication resource allocation and digital twin deployment scheme are as follows: first, the objective function is modeled as a multi-agent partially observable Markov decision process, and the base station in the scene is regarded as an agent; the specific elements are defined as follows: (1) State: Global state space Including all vehicle status information n , all vehicle DT deployment information X, and all edge server resource utilization e m Vehicle status information includes the data to be transmitted by the vehicle, the available communication transmission rate between the vehicle and the base station, and the AoI during the vehicle transmission process; it is defined as follows: (2) Action space: Each agent chooses its action This includes the communication and perception time allocation of vehicles within the coverage area of ​​the agent and the vehicle DT deployment strategy; the action space of multiple agents is written as: (3) State transition probability: The agent is in the current state s υ,k Execute action a to the environment υ,k After that, the environment state will transfer to the next state with a certain probability; it can be expressed as: (4) Reward function: The reward function should be designed to enable each agent to make optimal decisions on DT deployment and synaesthesia resource allocation. Therefore, it should be designed based on the agent's contribution to the system. The reward of agent v in time slot k is defined as: Among them, λ n,υ is a binary variable representing the relationship between vehicle n and agent j, Q n represents the contribution of vehicle n to the system; at time slot k, the global reward of the system is A multi-agent "centralized training, distributed execution" approach is used to complete collaborative learning to achieve the optimal VDT deployment and communication perception resource allocation strategy. The algorithm training process specifically includes the following steps: S61: Initialize Q network parameters and deterministic policy network parameters and experience replay pool D; S62: Collect initial status S63: Calculate the continuous action parameter τ according to the gradient descent method υ (k) and α υ (k); The gradient descent formula is: Among them, λ d is the learning rate of the Q network, λ s and λ υe Represent the weights of the local parameters and fusion parameters of the agent respectively; It is used to update the network parameters of the target Q network. It is used to update the network parameters of the deterministic policy network; S64: Sample an action a with exploration probability ε and probability distribution ψ υ,k , with probability 1-ε, select the discrete action value function Maximum discrete movement At the same time, determine the corresponding continuous action τ υ (k),α υ (k); where s υ,k Indicates status, represents the deployment strategy of vehicle DT in the edge server, τ υ (k),α υ (k) represents the vehicle perception and communication duration respectively; S65: Execute action a υ,k , get instant reward r υ,k and the next state s υ,k+1 ; S66: The tuple (s υ,k ,α υ,k ,r υ,k ,s υ,k+1 ) is stored in the experience replay pool S67: Calculate the loss function of the Q network according to the loss function and the loss function of the deterministic policy network S68: Update Q network parameters according to gradient descent method where λ s and λ υe are the weights of the agent’s local parameters and fusion parameters respectively; S69: Update converged network parameters: and Among them, ← is the update symbol, which means updating the calculated fusion network parameters. is the updated parameter of the Q network at the next moment, Update the parameters of the deterministic policy network for the next moment; represents the update gradient of the deterministic network parameters, is the update gradient of the Q network parameters, λ′ ve is the updated fusion network parameter weight; represents the global loss function of the fusion network, represents the expectation about the strategy, y ve represents the target value of the fusion network, Q ve represents the global Q function; S610: Network parameters are updated and sent to each agent.