Traffic message aware inter-vehicle communication autonomous resource selection method

By constructing a dual-agent DRL model and differentially training the DQN weight parameters of CAM and DENM, the problems of high probability of resource selection collision and low utilization in cellular vehicle network communications are solved, and the differentiated dissemination performance guarantee of traffic safety business messages under high-density traffic conditions is achieved.

CN120751493AInactive Publication Date: 2025-10-03NANTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511142350.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-15
Publication Date
2025-10-03
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In existing cellular vehicle-to-vehicle communications, the SB-SPS mechanism has problems such as high probability of resource selection collision and low resource utilization. Especially in high-density traffic conditions, it is difficult to guarantee the key performance indicators of CAM and DENM messages.

Method used

A message-aware dual-agent DRL model is constructed. By differentially training the DQN weight parameters of CAM and DENM, the ε-greedy algorithm is used to select adaptive resource blocks, establish workshop communication links, and achieve differentiated resource allocation.

Benefits of technology

It improves the packet reception success rate of CAM, reduces the packet end-to-end communication delay of DENM, improves the utilization rate of limited spectrum resources, and meets the key performance requirements of different traffic business messages.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120751493A_ABST
    Figure CN120751493A_ABST
Patent Text Reader

Abstract

The invention discloses a traffic message aware inter-vehicle communication autonomous resource selection method. In a pre-calculation stage, a vehicle-mounted user terminal (VUE) uses a differentiated experience sample pool to train a DQN weight parameter in a D-Agent based on a constructed message awareness double agent (D-Agent), and stores the trained D-Agent in a VUE storage space. When a traffic safety service generation message triggers a resource allocation mechanism, the VUE senses the attribute of the message, and simultaneously calculates the state characteristic parameters of the resource blocks in the candidate resource library in parallel. Then, enabling a corresponding single agent (Agent) according to a message attribute, and inputting a resource block state characteristic parameter matched with the message to calculate an instant reward function value of each sub-frame; and finally, the agent adopts an epsilon-greedy algorithm to select a resource block of a self-adaptive message attribute to establish an inter-vehicle communication link, so that the purposes of differentially guaranteeing key performance indexes of different messages and improving the utilization rate of limited spectrum resources are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for optimizing vehicle network communication, and in particular to an intelligent resource allocation method for traffic safety service messages. Background Art

[0002] Cellular Vehicle to Everything (C-V2X) establishes vehicle-to-vehicle (V2V) communications through vehicle user equipment (VUE) to transmit traffic safety information, including periodic broadcasts of Cooperative Awareness Messages (CAM) and event-triggered Decentralized Environment Notification Messages (DENM), to improve road efficiency and traffic safety. CAM primarily transmits traffic status information such as vehicle position, direction of movement, and speed, requiring ultra-high transmission reliability to share the relative position and movement status of vehicles. DENM transmits traffic warning information such as traffic accidents, emergency braking, and road anomalies through various methods such as broadcast (unicast / multicast), requiring ultra-low end-to-end transmission latency to respond to traffic anomalies in real time. CAM and DENM have differentiated message attributes and key performance indicators, and the required communication resources are different and complementary. Message-aware resource allocation selects communication resources that match message attributes for different traffic safety services, which helps reduce the probability of resource contention and access conflicts and differentiates and meets the differentiated performance requirements of different messages.

[0003] Currently, the 3GPP standard system uses a sensing-based semi-persistent scheduling mechanism (SB-SPS) to randomly allocate available communication resources for all V2V communication messages. This mechanism includes three processes: resource sensing, candidate resource selection, and resource selection (reselection). While the SB-SPS mechanism is simple to implement, it suffers from issues such as a high probability of resource selection collisions and low utilization of limited resources. This makes it difficult to guarantee the key performance indicators of message transmission, especially in high-density traffic conditions. Dynamic spectrum-based resource allocation methods can dynamically adjust the proportion of available resource pools based on differentiated service requirements. However, the dynamic nature of moving vehicles increases computational complexity, resulting in significant end-to-end latency for DENM messages. Summary of the Invention

[0004] Purpose of the invention: In view of the above-mentioned existing technologies, a message-aware vehicle communication autonomous resource selection method is proposed to achieve differentiated propagation performance guarantee for different traffic business messages.

[0005] Technical solution: A traffic message-aware vehicle-to-vehicle communication autonomous resource selection method, including:

[0006] In the pre-calculation stage, VUE uses a differentiated experience sample pool to train the DQN weight parameters in the constructed message-aware dual agent based on the message-aware dual agent, and stores the trained message-aware dual agent in the VUE storage space;

[0007] When a traffic service generates a message that triggers the resource allocation mechanism, the VUE perceives the message attributes and simultaneously calculates the state characteristic parameters of the resource blocks in the candidate resource library in parallel. Subsequently, the corresponding single agent is enabled based on the message attributes, and the state characteristic parameters of the resource blocks matching the message are input to calculate the immediate reward function value of each subframe. Finally, the agent uses the ε-greedy algorithm to select resource blocks with adaptive message attributes to establish a vehicle-to-vehicle communication link.

[0008] Furthermore, the construction process of the message-aware dual-agent includes:

[0009] Building a general DRL model:

[0010] At time t, let the V2V communication environment be the current candidate resource library, and the state s t is the time slot frequency characteristic parameter of the resource block in the candidate resource pool, action a t For the selection of resource blocks, immediate reward r t To perform action a t The optimization target value obtained when the environment and state s t 、Action a t 、Instant Reward t A general DRL model is established based on the five basic elements of the agent; wherein the agent contains two DQNs with identical structures but different weight parameters, namely the main DQN (Main DQN) and the target DQN (Target DQN); the main DQN is used to calculate the input state s t The approximate Q value Q(s t ,a t ; θ), θ is the weight parameter of the main DQN; the target DQN is used to calculate the target state s t+1 The Q value Q(s t+1 ,a t+1 θ - ),θ - is the weight parameter of the target DQN; the main DQN copies the θ value to θ at regular intervals - ;

[0011] Build a dual-agent DRL model:

[0012] In traffic safety service messages, the candidate resource pool state parameters that CAM and DENM are sensitive to are different, and the immediate rewards generated for the same resource block are different. Therefore, a message-aware dual-agent (D-Agent) DRL model is constructed based on the general DRL model. The D-Agent DRL model includes a CAM agent (Agent-CAM) DRL model and a DENM agent (Agent-DENM) DRL model with the same structure, and has an experience sample pool for message matching, namely the CAM experience sample pool E C and DENM experience sample pool E D .

[0013] Furthermore, the step of obtaining the state characteristic parameters of the resource blocks in the selected resource library includes the following steps:

[0014] The SB-SPS mechanism uses a 1000ms window as its resource sensing process. It measures the time-frequency characteristics of the sidelink communication signal, including the Received Signal Strength Indication (RSSI) and Reference Signal Received Power (RSRP), to assess the status of each resource block in a candidate resource pool consisting of 10 time-domain subframes and 12 frequency-domain subchannels.

[0015] The V2V communication cycle of the 3GPP standard system is 100ms. Therefore, 10 measurement values ​​are obtained in one sensing window. The RSSI average value in the i-th subframe is As shown in formula (1), it is used to evaluate the idle rate of resource blocks in the current subframe. The smaller the value, the higher the idle rate;

[0016]

[0017] Average RSRP value on the j-th subchannel As shown in formula (2), it is used to evaluate the quality of the current subchannel. The smaller the value, the worse the availability of cellular users (CU) and the higher the idle rate, but the smaller the number of available resource blocks in the candidate resource pool.

[0018]

[0019] Among them, i∈[1,10], j∈[1,12], h∈[1,10] is a measurement cycle within the perception window;

[0020] At present, the 3GPP standard system sets the transmission power per resource block P sThe noise power N0 of the unit resource block is 2.3dBm, and the noise power N0 of the unit resource block is 0.9dBm. Therefore, VUE can estimate the communication quality of the resource block by calculating the mean SINR of the unit resource block. Suppose the mean SINR of the resource block corresponding to the jth subchannel in the i-th subframe is The calculation is shown in formula (3);

[0021]

[0022] Among them, g sr is the channel gain between the transmitter s and the receiver r, g kr is the channel gain between the interference terminal k and the receiving terminal r, K is the set of VUE terminals competing for the same resource block in the measurement period, P s is the transmit power per unit resource block, and N0 is the noise power per unit resource block.

[0023] Furthermore, the training process of the dual-agent DRL model includes:

[0024] Establish a message-aware reward function:

[0025] The CAM message size is about 300 bytes, requiring ultra-high reliability. Resource blocks with high communication quality in high idle time slots are selected to establish a continuous and stable V2V communication link to obtain high communication rate. Therefore, the state space of the Agent-CAM DRL model is the mean subframe RSSI and resource block SINR of the candidate resource pool. The state space at time t is recorded as Constrained by communication synchronization, V2V communication only selects one or more consecutive sub-channel resource blocks in the same subframe to establish a communication link; if CAM occupies N resource blocks in total, the communication rate of all N resource blocks and R C As shown in formula (4):

[0026]

[0027] Then the action at time t The reward function r t C Constructed as the maximum R of N resource blocks C , as shown in formula (5):

[0028]

[0029] Among them, the communication bandwidth of the unit resource block is B RB 180kHz;

[0030] The size of DENM is about 190 bytes and usually only occupies one resource block. Ultra-low latency performance requires more access opportunities and selects higher-quality resource blocks to establish V2V communication links to obtain the minimum propagation delay. Therefore, the state space of the modeling Agent-DENM DRL model is the mean sub-channel RSRP and resource block SINR of the candidate resource pool. The Agent-DENM state space at time t is recorded as Assume that the maximum tolerable delay time of DENM is T max , the remaining transmission time is T re , then the end-to-end delay of DENM is T ED As shown in formula (6):

[0031] T ED =T max -T re (6)

[0032] Then the action at time t action The reward function r t D Constructed as a single resource block with the minimum T ED , as shown in formula (7):

[0033]

[0034] Training the dual-agent DRL model:

[0035] When training the DQN weight parameters of the CAM agent Agent-CAM DRL model and the DENM agent Agent-DENM DRL model in parallel, VUE uses the stochastic gradient descent algorithm to randomly extract N from their respective experience sample pools in a uniform sampling manner. E Experience samples are used for training; suppose the target output value y of the target DQN in the general DRL model is as shown in formula (8):

[0036] y t =r t +γmaxQ(s t+1 ,a t+1 θ - ) (8)

[0037] Among them, r t is the immediate reward, γ is the discount factor; according to the target output value y t And the approximate cumulative reward value of the main DQN, that is, the Q value Q(s t ,a t ; θ), the loss function of constructing the DQN weight parameter θ is L(θ)=[y t -Q(s t ,a t;θ)] 2 , then the loss function mean L E (θ) is shown in formula (9):

[0038]

[0039] Use the Stochastic Gradient Descent (SGD) algorithm to minimize the loss function L E (θ), that is Then the DQN weight parameter θ is updated as shown in formula (10):

[0040]

[0041] in,

[0042] During the training process, the master DQN weight parameter θ is copied to the target DQN to update the weight parameter θ every certain number of iterations M. - , improving the approximation between the target DQN and the main DQN; VUE stores the trained dual-agent DRL model and is used to online select V2V communication resource blocks that match message attributes.

[0043] Furthermore, when online selecting a V2V communication resource block that matches the message attribute, the following specific steps are included:

[0044] At time t, suppose the message grouping generated by the traffic safety service in the VUE triggers the resource allocation mechanism. The VUE unit calculates the state characteristic parameters of the resource blocks in the candidate resource library online and senses the service message type at the same time.

[0045] When the service message is CAM, VUE enables Agent-CAM and enters the state space With r t C As the reward function, N consecutive candidate resource blocks in the same subframe are regarded as a set of available resource blocks, and the main DQN calculates the Q value of each set of resource blocks;

[0046] Agent-CAM uses the ε-greedy algorithm to select a set of resource blocks with the maximum Q value as the output action a with probability (1-ε) t ; Randomly select a set of resource blocks as output action a with probability ε t , to avoid the algorithm falling into local optimality;

[0047] Agent-CAM generates experience samples C =(s t ,a t ,r t ,s t+1) into the CAM experience sample pool E C , and use the First Input First Output (FIFO) method to update and retain the last M E Experience samples;

[0048] When the business message is DENM, VUE enables Agent-DENM and enters the state space With r t D As the reward function, the main DQN calculates the Q value of each candidate resource block;

[0049] Agent-DENM uses the ε-greedy algorithm to select the resource block with the maximum Q value as the output action a with probability (1-ε) t ; Randomly select a candidate resource block as the output action a with probability ε t , to prevent the algorithm from falling into local optimum.

[0050] Agent-DENM generates experience samples e D =(s t ,a t ,r t ,s t+1 ) into the DENM experience sample pool E D , using FIFO method to update and retain the last M E An experience sample.

[0051] Beneficial effects: To address the problem of differentiated message resource selection for V2V communication under high-density traffic conditions, a dual-agent (D-Agent) model is constructed, an experience database for message perception is established, and the weight parameters of the Deep Q Network (DQN) adapted to two types of traffic business messages are trained separately. These parameters are used for autonomous resource selection for V2V communication messages, achieving the goal of differentiated protection of the propagation performance of different messages and maximizing the utilization of limited communication resources.

[0052] 1. A message-aware D-Agent DRL model was established to ensure differentiated dissemination performance for different traffic business messages.

[0053] 2. Using differentiated reward functions and experience sample pools to train D-Agents, we established DQN weight parameters that matched message attributes.

[0054] 3. The D-Agent DRL model trained in VUE storage can respond to CAM and DENM resource allocation requests online, improving the key performance indicators of different traffic business messages. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] Figure 1 It is a general DRL model;

[0056] Figure 2 It is a dual-agent DRL model;

[0057] Figure 3 It is a flow chart of the message-aware resource autonomous selection method;

[0058] Figure 4 is the CAM packet reception rate (PRR) performance curve;

[0059] Figure 5 is the DENM end-to-end delay (ED) performance curve;

[0060] Figure 6 It is the limited spectrum resource utilization (RU) curve. DETAILED DESCRIPTION

[0061] The present invention will be further explained below with reference to the accompanying drawings.

[0062] The process of this embodiment is as follows Figure 3 As shown,

[0063] 1) In high-density traffic scenarios, multiple V2V pairs should simultaneously transmit traffic safety service messages. V2V communication establishes V2V communication links by reusing unused licensed spectrum in the cellular network while preventing interference with legitimate authorized users within the cellular network. Resource blocks within subframes with low RSSI values ​​are typically selected. When traffic service messages are generated, the SB-SPS scheduling mechanism is triggered to allocate spectrum resources and detect the message type.

[0064] 2) The VUE calculation unit calculates the subframe busy state parameter of the candidate resource based on the state characteristic parameters of the resource blocks in the candidate resource pool measured during the sensing period using formula (1) Apply formula (2) to calculate the subframe quality status parameter of the candidate resource Apply formula (3) to calculate the subchannel quality state parameter of the candidate resource

[0065] 3) When the business message is perceived as CAM, VUE enables Agent-CAM intelligent agent, inputs the characteristic state parameters of the candidate resource library from the environment and The main DQN model stored after training is applied to iteratively calculate the reward function value for each set of resource blocks.

[0066] 4) Agent-CAM uses the ε-greedy algorithm with probability (1-ε) to Select a group with max{rt C} resource block set as output action a t ; Randomly select a set of resource blocks as output action a with probability ε t .

[0067] 5) When the business message is perceived as DENM, VUE enables Agent-DENM intelligent agent to input the characteristic state parameters of the candidate resource library from the environment and The main DQN model stored after training is used to iteratively calculate the reward function value of each candidate resource block.

[0068] 6) Agent-DENM uses the ε-greedy algorithm with probability (1-ε) to Select a subframe with max{r t D} resource block as output action a t ; Randomly select a candidate resource block as the output action a with probability ε t .

[0069] 7) The SB-SPS scheduling mechanism in VUE outputs action a according to the agent t Assign available resource blocks to the corresponding service messages and establish a V2V communication link.

[0070] The following is an example of implementing the present invention based on the WiLabV2Xsim simulation system on the MATLAB platform. In the simulation, the standard SB-SPS resource scheduling mechanism released by 3GPP Release 17 is used, with a communication channel bandwidth of 10 MHz, a perception window duration of 1000 ms, a resource reservation interval of 100 ms, a resource reservation probability of 0.4, a modulation and coding scheme of 11, and a transmit power of P per resource block. s =2.3dBm, the noise power per unit resource block N0 = 0.9dBm, and the channel propagation model is the WINNER+B1 model. In the traffic safety service mixed message, CAM accounts for 95%, each CAM packet size is 300 bytes, and occupies 2 resource blocks. DENM accounts for 5%, each DENM packet size is 190 bytes, and occupies 1 resource block. Assume that the learning rate of the DQN model is α = 0.01, the discount factor is γ = 0.9, and the number of experience pool samples is M. E = 10,000 samples, the number of small batch training experience samples N E =64, initial experience sample set E C and E DThe ε-greedy strategy is initialized with ε = 1 and then decreases at a rate of 0.995 until ε is less than 0.01.

[0071] Aiming at the reliability index PRR sensitive to service message CAM, the delay index ED sensitive to DENM and the performance of limited spectrum bandwidth utilization, the resource selection method of the present invention (Service Sensing resource selection method based on Deep Reinforcement Learning, SSDRL) and the multi-agent reinforcement learning based resource selection method (Multi-Agent Reinforcement Learning based resource selection method, MARL) in reference [1) are simulated and compared. 1 ) and the performance curves of the Random Selection Scheme (RSS) built into the standard SB-SPS scheduling mechanism. The reliability metric PRR is the packet reception ratio, which refers to the average ratio of the number of successfully transmitted CAM packets to the total number of transmitted CAM packets. The delay metric ED is the end-to-end delay of message transmission, which refers to the average time interval from the generation of a message packet to its successful reception. The resource utilization rate (RU) is the average ratio of the number of occupied resource blocks in the candidate resource pool to the total number of resource blocks.

[0072] For CAM messages, such as Figure 4 As shown, the SSDRL method is based on the CAM message experience database E C The training and storage agent Agent-CAM can be used in the smallest Select the maximum The resource blocks of the V2V communication link are used to obtain the resource blocks with the best quality in the current time slot, thereby reducing the probability of V2V communication interruption and improving the PRR performance.

[0073] For DENM messages, such as Figure 5 As shown, the SSDRL method is applied based on the DENM message experience database E D Training stored agent Agent-DENM, in the maximum Get the maximum number of available resource blocks in a subframe and select the maximum Obtaining the best resource block in the current subframe improves ED performance by increasing DENM packet access opportunities and improving packet transmission success rates. In particular, with fixed communication resources, increasing vehicle density increases the probability of V2V communication contention access collisions, and DENM packet retransmissions significantly degrade ED values.

[0074] Regarding the utilization of limited spectrum resources, such as Figure 6 As shown, the SSDRL method of the present invention differentially evaluates resource block characteristics and distinguishes and matches optimal subframes based on the attributes of different service class messages. This method leverages the complementarity of resource block characteristics required by different messages to reduce the probability of contention access collisions between different messages, improves the allocation ratio of available resource blocks within the effective resource block, and achieves high utilization of limited resources.

[0075] In summary, the method of the present invention establishes, trains, and applies a D-Agent DRL model that is adaptive to the attributes of traffic safety service messages. It utilizes the differences in key performance indicators and the complementarity of required resources between different service class messages to select a resource block set that matches the service message attributes for V2V communication. This meets the key performance requirements of different service class messages, improves the packet success reception rate of CAM, reduces the packet end-to-end communication delay of DENM, and improves the utilization rate of limited spectrum resources in high-density traffic scenarios.

[0076]

[0077] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as within the scope of protection of the present invention.

Claims

1. A method for autonomous resource selection of vehicle-to-vehicle communication based on traffic information perception, characterized in that: include: In the pre-calculation stage, VUE uses a differentiated experience sample pool to train the DQN weight parameters in the constructed message-aware dual agent based on the message-aware dual agent, and stores the trained message-aware dual agent in the VUE storage space; When a traffic service generates a message that triggers the resource allocation mechanism, the VUE perceives the message attributes and simultaneously calculates the state characteristic parameters of the resource blocks in the candidate resource library in parallel. Subsequently, the corresponding single agent is enabled according to the message attributes, and the state characteristic parameters of the resource blocks matching the message are input to calculate the immediate cumulative reward function value of the selected resource blocks in each subframe. Finally, the agent uses the ε-greedy algorithm to select a set of resource blocks with adaptive message attributes to establish a vehicle-to-vehicle communication link.

2. The method for autonomous resource selection of vehicle-to-vehicle communication based on traffic message awareness according to claim 1, characterized in that: The construction process of the message-aware dual-agent includes: Building a general DRL model: At time t, let the V2V communication environment be the current candidate resource library, and the state s t is the time slot frequency characteristic parameter of the resource block in the candidate resource pool, action a t For the selection of resource blocks, immediate reward r t To perform action a t The optimization target value obtained when the environment and state s t 、Action a t 、Instant Reward t A general DRL model is established with five basic elements of the agent; wherein the agent contains two DQNs with identical structures but different weight parameters, namely the main DQN and the target DQN; the main DQN is used to calculate the input state s t The approximate Q value Q(s t ,a t ; θ), θ is the weight parameter of the main DQN; the target DQN is used to calculate the target state s t+1 The Q value Q(s t+1 ,a t+1 θ - ),θ - is the weight parameter of the target DQN; the main DQN copies the θ value to θ at regular intervals - ; Build a dual-agent DRL model: A message-aware dual-agent DRL model is constructed based on the general DRL model; the dual-agent DRL model includes a CAM agent DRL model and a DENM agent DRL model with the same structure, and has an experience sample pool for message matching, namely, the CAM experience sample pool E C and DENM experience sample pool E D .

3. The method for autonomous resource selection of vehicle-to-vehicle communication based on traffic message awareness according to claim 1 or 2, characterized in that: The step of obtaining the state characteristic parameters of the resource blocks in the resource library comprises the following steps: By measuring the RSSI and RSRP of the sidelink communication signal, the state characteristics of each resource block in the candidate resource pool consisting of I time domain subframes and J frequency domain subchannels are evaluated; If m measurement values ​​are obtained in a perception window, the RSSI mean RSSI in the i-th subframe is i As shown in formula (1): Average RSRP value on the j-th subchannel As shown in formula (2): Among them, i∈[1,I], j∈[1,J], h∈[1,m] is a measurement cycle within the perception window; VUE estimates the communication quality of a resource block by calculating the SINR mean of the unit resource block; let the SINR mean of the resource block corresponding to the jth subchannel in the i-th subframe be As shown in formula (3): Among them, g sr is the channel gain between the transmitter s and the receiver r, g kr is the channel gain between the interference terminal k and the receiving terminal r, K is the set of VUE terminals competing for the same resource block in the measurement period, P s is the transmit power per unit resource block, and N0 is the noise power per unit resource block.

4. The method for autonomous resource selection of vehicle-to-vehicle communication based on traffic message awareness according to claim 3, characterized in that: The training process of the dual-agent DRL model includes: Establish a message-aware reward function: The state space of the CAM agent DRL model is the mean subframe RSSI and resource block SINR of the candidate resource pool. The state space at time t is V2V communication only selects one or more consecutive sub-channel resource blocks in the same subframe to establish a communication link; if CAM occupies N resource blocks in total, the communication rate of all N resource blocks and R C As shown in formula (4): Then the action at time t The reward function r t C Constructed as the maximum R of N resource blocks C , as shown in formula (5): Among them, B RB is the communication bandwidth per resource block; The state space of the DENM agent DRL model is the mean sub-channel RSRP and resource block SINR of the candidate resource pool. The state space at time t is Assume that the maximum tolerable delay time of DENM is T max , the remaining transmission time is T re , then the end-to-end delay of DENM is T ED As shown in formula (6): T ED =T max -T re (6) Then the action at time t The reward function r t D Constructed as a single resource block with the minimum T ED , as shown in formula (7): Training the dual-agent DRL model: When training the DQN weight parameters of the CAM agent DRL model and the DENM agent DRL model in parallel, VUE uses the stochastic gradient descent algorithm to randomly extract N from their respective experience sample pools in a uniform sampling manner. E Experience samples are used for training; suppose the target output value y of the target DQN in the general DRL model is as shown in formula (8): y t =r t +γmaxQ(s t+1 ,a t+1 ;θ - ) (8) Among them, r t is the immediate reward value, γ is the discount factor; according to the target output value y t And the approximate cumulative reward value of the main DQN, that is, the Q value Q(s t ,a t ; θ), the loss function of constructing the DQN weight parameter θ is L(θ)=[y t -Q(s t ,a t ;θ)] 2 , then the loss function mean L E (θ) is shown in formula (9): Use the stochastic gradient descent algorithm to minimize the loss function L E (θ), to update the DQN weight parameter θ; During the training process, the master DQN weight parameter θ is copied to the target DQN to update the weight parameter θ every certain number of iterations M. - ; VUE stores the trained dual-agent DRL model and is used to online select V2V communication resource blocks that match message attributes.

5. The method for autonomous resource selection of vehicle-to-vehicle communication based on traffic message awareness according to claim 3, characterized in that: When online selecting a V2V communication resource block that matches the message attributes, the following specific steps are included: At time t, suppose the message grouping generated by the traffic safety service in the VUE triggers the resource allocation mechanism. The VUE unit calculates the state characteristic parameters of the resource blocks in the candidate resource library online and senses the service message type at the same time. When the business message is CAM, VUE enables the CAM agent and enters the state space With r t C As the reward function, N consecutive candidate resource blocks in the same subframe are regarded as a set of available resource blocks, and the main DQN calculates the Q value of each set of resource blocks; The CAM agent uses the ε-greedy algorithm to select a set of resource blocks with the maximum Q value as the output action a with probability (1-ε). t ; When the business message is DENM, VUE enables the DENM agent and enters the state space With r t D As the reward function, the main DQN calculates the Q value of each candidate resource block; The DENM agent uses the ε-greedy algorithm to select the resource block with the maximum Q value as the output action a with probability (1-ε) t .

6. The method for autonomous resource selection of vehicle-to-vehicle communication based on traffic message awareness according to claim 5, characterized in that: When the business message is CAM, it also includes: CAM agent generates experience sample e C =(s t ,a t ,r t ,s t+1 ) into the CAM experience sample pool E C , and use the FIFO method to update and retain the last M E Experience samples; When the business message is DENM, it also includes: DENM agent generates experience sample e D =(s t ,a t ,r t ,s t+1 ) into the DENM experience sample pool E D , using FIFO method to update and retain the last M E An experience sample.