Enhanced Method for Information Transmission in a Senso-Communication Integrated Vehicular Network Based on Deep Reinforcement Learning

Through the method based on deep reinforcement learning, the beamforming and transmission power of roadside units are optimized, which solves the problems of high computational complexity and poor convergence performance in the transmission of vehicle network information, and achieves the improvement of information transmission rate and the guarantee of perceived performance.

CN116346860BActive Publication Date: 2025-08-01STATE GRID ANHUI ELECTRIC POWER CO LTD ELECTRIC POWER SCI RES INST +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310380282.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-11
Publication Date
2025-08-01
Estimated Expiration
2043-04-11

AI Technical Summary

Technical Problem

In the existing Internet of Vehicle Information Transmission Technology, the traditional convex optimization method has high computational complexity and poor convergence performance, making it difficult to effectively solve the beamforming and power distribution problems of ISAC signals. Especially when considering target perception constraints, the non-convex Clariano-Lower Boundary (CRLB) makes the existing methods difficult to apply.

Method used

Using a method based on deep reinforcement learning, the roadside units are modeled as agents. By designing the Krame-Royal lower bound as perceptual constraints, a sum-rate maximum model of beamforming and transmission power is established. Through iterative training of deep reinforcement learning, the beamforming and transmission power of the roadside units are optimized to maximize the information transmission rate and perceptual performance.

Benefits of technology

It has achieved the improvement of the Internet of Vehicle Information Transmission Rate under the premise of ensuring perceptual performance, with the characteristics of independent learning, dynamic adaptability, and rapid convergence, and enhanced the synesthesia fusion Internet of Vehicle Information Transmission Ability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116346860B_ABST
    Figure CN116346860B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of vehicle network information transmission, and discloses a method for enhancing the communication and sensing fusion vehicle network information transmission based on deep reinforcement learning, including: S1, designing the Cramér-Rao lower bound of the target vehicle relative to the roadside unit as the sensing constraint; S2, obtaining the sum rate of the ISAC signals between all target vehicles and the roadside unit; S3, based on the sum rate of the ISAC signals and the sensing constraint, establishing a sum rate maximization model related to the beamforming and transmit power of the roadside unit; S4, solving the sum rate maximization model based on the deep reinforcement learning method, and through jointly optimizing the beamforming and transmit power of the roadside unit, obtaining the optimal beamforming and transmit power, so as to improve the transmission rate of the vehicle network ISAC signal while ensuring the requirement of sensing performance. The present invention improves the vehicle network information transmission rate while ensuring the requirement of sensing performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of vehicle network information transmission, and particularly to a method for enhancing vehicle network information transmission through integrated communication and sensing based on deep reinforcement learning. Background Art

[0002] The integration of communication and sensing technologies (ISAC) will improve the information transmission rate and sensing accuracy performance of the sixth-generation communication network. Therefore, ISAC technology has a framework for sharing spectrum, hardware platforms, and joint signal processing, and can be used as one of the new key technologies in fields such as vehicle networks and autonomous driving.

[0003] Autonomous driving and vehicle network technologies are becoming important development trends in intelligent transportation networks. By equipping vehicles with sensors and microprocessing units, vehicles can obtain environmental information to assist driving and communicate with surrounding vehicles, infrastructure, and human users. However, problems such as high-speed vehicle movement and limited computing power can lead to a decrease in sensor accuracy, and there are many difficult problems in the processing and transmission of a large amount of sensing data. Therefore, to solve these problems, infrastructure can be installed around roads, such as small base stations and roadside units, to hand over some or all of the sensing tasks to roadside infrastructure and establish high-speed communication links to interact data in real time.

[0004] The ISAC signal deployed by this roadside infrastructure is transmitted by a radar communication transmitter-receiver. The roadside infrastructure communicates with target vehicles by transmitting ISAC signals and simultaneously performs sensing tasks. The obtained sensing data can be used to assist communication, reduce signaling overhead in communication, and improve the information transmission rate. This technology has broad application prospects in the field of intelligent transportation, can improve the efficiency and accuracy of vehicle communication and sensing, and make the intelligent transportation network safer and more convenient. Existing technologies usually adopt traditional convex optimization methods to solve beamforming and power allocation to ensure the information transmission rate, but this method has deficiencies such as high computational complexity and poor convergence performance. In addition, the constraint of target sensing must be considered in ISAC technology, and the non-convex Cramer-Rao lower bound (CRLB) as a sensing constraint will make it difficult to apply existing methods. Summary of the Invention

[0005] The present invention provides a method for enhancing vehicle network information transmission through integrated communication and sensing based on deep reinforcement learning to solve the above problems.

[0006] The present invention is achieved through the following technical solutions:

[0007] A method for enhancing vehicle network information transmission through integrated communication and sensing based on deep reinforcement learning, comprising:

[0008] S1. Design the Cramér-Rao lower bound of the target vehicle relative to the roadside unit as the sensing constraint;

[0009] S2. Obtain the sum rate of the ISAC signals between all target vehicles and the roadside unit;

[0010] S3. Based on the sum rate of the ISAC signals and the sensing constraint, establish a sum rate maximization model related to the beamforming and transmit power of the roadside unit;

[0011] S4. Solve the sum rate maximization model based on the deep reinforcement learning method. By jointly optimizing the beamforming and transmit power of the roadside unit, obtain the optimal beamforming and transmit power, so as to improve the information transmission rate of the vehicle network while ensuring the sensing performance requirements.

[0012] As an optimization, in S1, the Cramér-Rao lower bound of the target vehicle relative to the roadside unit includes the angular Cramér-Rao lower bound of the angle of the target vehicle relative to the roadside unit and the distance Cramér-Rao lower bound of the distance, where the angular Cramér-Rao lower bound is:

[0013]

[0014] The distance Cramér-Rao lower bound is:

[0015]

[0016] where θ m is the angle of the target vehicle m relative to the roadside unit, d m is the distance of the target vehicle m relative to the roadside unit, w m represents the beamforming steering vector of the ISAC signal of the roadside unit for the target vehicle m, p m represents the transmit power of the roadside unit for transmitting the ISAC signal to the target vehicle m, and represent the estimated noise variances corresponding to the angle and distance, and c represents the speed of electromagnetic wave signals propagating in vacuum.

[0017] As an optimization, in S2, the sum rate of the ISAC signals between all target vehicles and the roadside unit is specifically:

[0018]

[0019] where M is the total number of target vehicles, γ m (w m ,p m ) is the signal-to-interference-plus-noise ratio of the target vehicle m receiving the ISAC signal.

[0020] As an optimization, the signal-to-interference-noise ratio γ of the ISAC signal received by the target vehicle m m (w m ,p m )The specific formula is:

[0021]

[0022] Where G represents the gain of the antenna array of the roadside unit, a H (θ m ) is the transmission steering vector of the antenna array, is the noise power received by the vehicle, w j represents the steering vector of the ISAC signal beamforming performed by the roadside unit on vehicle j that is not the target vehicle m, p j represents the transmission power of the ISAC signal transmitted by the roadside unit to vehicle j other than target vehicle m, α m is the path loss coefficient of the communication transmission channel between the target vehicle m and the roadside unit.

[0023] As an optimization, the path loss coefficient α of the communication transmission channel between the target vehicle m and the roadside unit is m Specifically:

[0024]

[0025] where ζ and α0 are the path loss exponent and the path loss at the reference distance d0, respectively.

[0026] As an optimization, the maximum sum rate model is specifically:

[0027]

[0028]

[0029]

[0030]

[0031] Where W represents the downlink beamforming matrix between the roadside unit and the target vehicle, p is the power vector of the ISAC signal transmitted by the roadside unit, and P max is the maximum transmission power supported by the roadside unit, ∈ θ and ∈ d are the maximum tolerable Cramer-Rao lower bound thresholds to ensure perception accuracy, ∈ θ That is the maximum threshold of the angle Cramer-Rao lower bound, ∈ d That is the maximum threshold from the Cramer-Rao lower bound.

[0032] As an optimization, before S4, iterative training is performed on the neural network corresponding to the deep reinforcement learning method to optimize the strategy of the iterative process. The iterative process follows the Bellman equation. The neural network corresponding to the deep reinforcement learning method includes a target neural network and an estimation neural network. The training process is as follows:

[0033] S4.1. Regard the roadside unit as an agent, and design the action space, state space, and environmental reward of the agent;

[0034] S4.2. Store the action, state, and environmental reward information obtained by the agent interacting with the environment during the decision-making phase in the experience replay pool;

[0035] S4.3. Randomly select a batch of training data of a fixed size from the experience replay pool for training;

[0036] S4.4. Transmit the batch of training data to the target neural network and the estimation neural network respectively;

[0037] S4.5. After the batch of training data is input into the target neural network, a target Q value is output. After the small batch of training data is input into the estimation neural network, an estimated Q value is output;

[0038] S4.6. Calculate the error loss between the target Q value and the estimated Q value. The estimation neural network updates the network parameters based on the error loss in the reverse gradient direction to reduce the error;

[0039] S4.7. Regularly update the target neural network according to the parameters of the estimation neural network.

[0040] As an optimization, in S4.2, the specific process of storing the action, state, and environmental reward information obtained by the agent interacting with the environment during the decision-making phase in the experience replay pool is as follows:

[0041] S4.2.1. Initialize the parameters of the target neural network and the estimation neural network, and obtain the initial state s0 of the agent;

[0042] S4.2.2. In the current state s t of the agent, based on the estimated Q value of the estimation neural network and the greedy strategy, output the current action value of the agent;

[0043] S4.2.3. Based on the current action value of the agent, obtain a reward value r t for the agent by interacting with the environment, and generate the next state value s t+1 ; [[ID=DI=41]]

[0044] S4.2.4. Repeat S4.2.2 - S4.2.3 to obtain a number of training data, where the training data includes state s t , action a t , reward r t and the state s at the next moment t+1 .

[0045] As an optimization, the parameters of the target neural network and the estimation neural network include the learning rate η, the reward discount factor γ, the exploration probability the training step size M max , the number of iterations N max .

[0046] As an optimization, the state space of the agent is where R is the sum rate of the vehicle and the roadside unit transmitting the ISAC signal, is the estimated angle of the vehicle position obtained from the echo signal, is the estimated distance of the vehicle position obtained from the echo signal;

[0047] The action space of the agent is a t = [A θ , A p ; including the beamforming action subspace and the power allocation action subspace where d θ is the length of the beamforming action subspace, θ max is the maximum angle value of the roadside unit beamforming, P max is the maximum power of the roadside unit transmitting the ISAC signal, d p is the length of the power allocation action subspace.

[0048] The environmental reward of the agent is a combination of the transmission sum rate of the ISAC signal and the Cramér - Rao lower bound constraint, that is, r t = R - λ1CRLB(θ m , w m , p m ) - λ2CRLB(d m , w m , p m ), where λ1 and λ2 are penalty parameters.

[0049] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0050] The present invention models the roadside unit as an independent agent, constructs a reward function related to the transmission and rate and sensing constraints of the vehicle-to-everything (V2X) integrated sensing and communication (ISAC) signal based on the state space dynamically obtained by the roadside unit, and thus makes an intelligent joint decision on beamforming and transmit power according to the reward value, which not only enhances the information transmission rate performance of the V2X system, but also ensures the sensing performance of the roadside unit. In addition, the constructed intelligent V2X information transmission system has the characteristics of autonomous learning, dynamic adaptability, and fast convergence, etc., and can ensure the sensing performance of the V2X and enhance the information transmission ability of the integrated sensing and communication V2X. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] To more clearly illustrate the technical solutions of the exemplary embodiments of the present invention, the following will briefly introduce the drawings required for the embodiments. It should be understood that the following drawings only show some embodiments of the present invention, and thus should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts. In the drawings:

[0052] Figure 1 is a schematic diagram of a downlink integrated sensing and communication V2X communication system composed of roadside infrastructure and target vehicles;

[0053] Figure 2 is a framework diagram of a method for enhancing information transmission of an integrated sensing and communication V2X based on deep reinforcement learning. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0054] To make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the following will further elaborate on the present invention in combination with the embodiments and the drawings. The illustrative embodiments and descriptions of the present invention are only used to explain the present invention and do not limit the present invention.

[0055] An embodiment of a method for enhancing information transmission of an integrated sensing and communication V2X based on deep reinforcement learning mainly considers a downlink integrated sensing and communication V2X communication system composed of roadside infrastructure (such as roadside units) and target vehicles, as Figure 1 shown.

[0056] The roadside unit is respectively equipped with N T transmit and N RA root receiving uniform linear array antenna provides communication services for M single-antenna vehicles. The roadside unit adopts full-duplex technology to maintain uninterrupted downlink communication and can receive potential echo signals for perception. The roadside unit simultaneously sends separate dedicated ISAC signals to M vehicles for communication, and the sent ISAC signals are multiplexed for perception at the same time; the target vehicle reflects the echo signal to the roadside unit, and through signal processing (such as matched filtering), the position of the target vehicle can be estimated. The driving directions of all vehicles are always parallel to the roadside unit. The communication between the roadside unit and the vehicle uses a line-of-sight channel. The roadside unit transmits a downlink ISAC signal x m (t) for communication. This signal is reflected on the target vehicle, and the roadside unit receives the echo signal r m (t). The roadside unit processes the echo signal r m (t) to obtain the estimated position information of the target vehicle.

[0057] The method of the present invention mainly includes the following steps:

[0058] S1. Design the Cramér-Rao lower bound of the target vehicle relative to the roadside unit as a perception constraint.

[0059] To measure the perception performance of the roadside unit, the Cramér-Rao lower bound is introduced as a perception performance index. For the angle θ of the target vehicle m relative to the roadside unit m the angle Cramér-Rao lower bound and the distance d m of the distance Cramér-Rao lower bound are respectively where w m represents the ISAC signal beamforming steering vector of the roadside unit for the target vehicle m, p m represents the transmission power of the ISAC signal transmitted by the roadside unit to the target vehicle m, and represent the estimated noise variances corresponding to the angle and distance, c represents the speed of electromagnetic wave signals propagating in vacuum, and for the downlink communication link between the roadside unit and the target vehicle, the path loss coefficient of the communication transmission channel between the target vehicle m and the roadside unit is where ζ and α0 are the path loss exponent and the path loss at the reference distance d0, then the signal-to-interference-plus-noise ratio of the receiving end of the ISAC signal of the target vehicle m is where G represents the gain of the antenna array of the roadside unit, a H (θ m ) is the transmission steering vector of the antenna array, is the vehicle-end received noise power, w j represents the ISAC signal beamforming steering vector of the roadside unit for the vehicle j other than the target vehicle m, p jDenote the transmission power of the RSUs to vehicle j of non-target vehicle m for the ISAC signal.

[0060] S2. Obtain the sum rate of the ISAC signals between all target vehicles and the RSUs; thus, the sum rate of all target vehicles and the RSUs where M is the total number of target vehicles, and γ m (w m , p m ) is the signal-to-interference-plus-noise ratio of target vehicle m for receiving the ISAC signal.

[0061] S3. Based on the sum rate of the ISAC signals and the sensing constraint, establish a sum rate maximization model related to the beamforming and transmission power of the RSUs;

[0062] In this embodiment, the sum rate maximization model is specifically:

[0063]

[0064]

[0065]

[0066]

[0067] where W represents the downlink beamforming matrix between the RSUs and the target vehicles, p is the power vector of the RSUs for transmitting the ISAC signal, P max is the maximum transmission power supported by the RSUs, ∈ θ and ∈ d are respectively the maximum tolerable Cramér-Rao lower bound thresholds for ensuring the sensing accuracy, ∈ θ is the maximum threshold of the angular Cramér-Rao lower bound, ∈ d is the maximum threshold of the range Cramér-Rao lower bound, the maximum tolerable Cramér-Rao lower bound threshold for ensuring the sensing accuracy.

[0068] Under the requirement of ensuring the sensing performance, with the maximization of the information transmission rate as the optimization objective, but not limited to this, the optimization problem of maximizing the system sum rate by jointly optimizing the beamforming and power is as follows:

[0069]

[0070]

[0071]

[0072]

[0073] S4. Solve the sum-rate maximization model based on the deep reinforcement learning method. By jointly optimizing the beamforming and transmit power of the roadside unit, the optimal beamforming and transmit power are obtained, thereby improving the information transmission rate of the vehicle-to-everything (V2X) network while meeting the requirements of sensing performance.

[0074] Before S4, iteratively train the neural network corresponding to the deep reinforcement learning method to optimize the strategy of the iterative process. The iterative process follows the Bellman equation. The neural network corresponding to the deep reinforcement learning method includes a target neural network and an estimation neural network. The training process is as follows:

[0075] S4.1. Regard the roadside unit as an agent, and design the action space, state space, and environmental reward of the agent.

[0076] Since it is difficult to obtain the closed-form expression of the sensing constraint and the non-convex optimization objective, it is difficult to solve the problem through the standard convex optimization method. Therefore, the present invention transforms the joint beamforming and power allocation formula into a Markov decision process problem, and then uses an intelligent method based on deep reinforcement learning to seek the optimal solution to this problem.

[0077] To find the optimal beamforming and power allocation strategy, the present invention updates the actions in different states of the roadside unit by maximizing the long-term average return, as Figure 2 shown.

[0078] Specifically, in the communication and sensing integrated V2X network system, the roadside unit is modeled as an agent. The actions of the agent include beamforming and power allocation, that is, the actions are optimized by maximizing the future reward. The reward is a combination of the information transmission sum rate and the Cramér-Rao lower bound sensing constraint.

[0079] First, design the action space, state space, and reward function of the agent:

[0080] The action space of the roadside unit as an agent consists of an angle-based beamforming action subspace and a power allocation action subspace. The beamforming action subspace for the roadside unit to serve the target vehicle m is a discrete angle codebook set, for example where d θ is the length of the angle action space, and θ max is the maximum angle value of the roadside unit beamforming; the action subspace of power is a discrete power set from zero to the maximum power P max , for example, a uniform power set with a length of d p Therefore, the action space set of the agent is a =[A t , A θ , A p ;

[0081] The state space of the roadside unit as an agent consists of three parts, namely the estimated angle of the vehicle position obtained from the echo signal and the distance as well as the information transmission and rate R of the system. The state space set of the agent is

[0082] The environmental reward of the roadside unit as an agent can be designed as a combination of the information transmission and rate of the system and the Cramér-Rao lower bound constraint, that is, r t = R - λ1CRLB(θ m , w m , p m ) - λ2CRLB(d m , w m , p m ). To ensure the rate and sensing performance, the penalty parameters λ1 and λ2 can usually be set to relatively large values. For example, in the present invention, the numerical range can be defined as 10 to 10 4 , but not limited to this. If the Cramér-Rao lower bound constraint is not satisfied, the agent receives a large negative reward.

[0083] S4.2. Store the action, state, and environmental reward information obtained by the agent interacting with the environment during the decision-making phase in the experience replay pool;

[0084] The specific process is as follows:

[0085] S4.2.1. Initialization: Initialize the parameters of the target neural network and the estimation neural network, and obtain the initial state s0 of the agent, that is, set the parameters of the estimation neural network and the target neural network, namely the learning rate η, the reward discount factor γ, the exploration probability the training step size M max the number of iterations N max , define the size of the experience replay pool and the mini-batch data training, and obtain the initial state s0 of the agent;

[0086] S4.2.2. Enter the decision-making phase: The agent outputs the current action value of the agent according to the estimated Q value of the estimation neural network in the current state s t of the current environment; the agent in the current state s t of the current environment executes the beamforming and power allocation action a t according to the Q value output of the estimation neural network and the greedy policy;

[0087] S4.2.3. Obtain a reward value r t for the agent through interaction with the environment based on the current action value of the agent, and generate the next state value s t+1 ; The environment transitions to the next moment state s based on the action of the agentt+1 Meanwhile, the environment feedbacks the reward r corresponding to the action t to the intelligent agent;

[0088] S4.2.4. Repeat S4.2.2 - S4.2.3 to obtain a number of training data, where the training data includes the state s t , the action a t , the reward r t and the state s at the next moment t+1 . The parameters of the iterative process include the state s t , the action a t , the reward r t and the state s at the next moment t+1 which are stored in the experience replay pool for the training of the neural network.

[0089] S4.3. Randomly select a batch of training data with a fixed size from the experience replay pool for training. For example, the present invention uses 32 batches of data for training;

[0090] S4.4. Transmit the batch of training data to the target neural network and the estimation neural network respectively;

[0091] S4.5. After inputting the batch of training data into the target neural network, the target Q value is output. After inputting the small - batch training data into the estimation neural network, the estimated Q value is output;

[0092] S4.6. Calculate the error loss between the target Q value and the estimated Q value. The estimation neural network updates the network parameters based on the error loss in the reverse gradient direction to reduce the error;

[0093] S4.7. Regularly update the target neural network according to the parameters of the estimation neural network.

[0094] Finally, input the sum - rate maximum model into the trained estimation neural network. The roadside unit in the vehicle - to - everything network can rely on the trained estimation neural network to make decisions on the beamforming and power allocation of the roadside unit, so as to improve the information transmission rate of the vehicle - to - everything network while meeting the requirements of sensing performance.

[0095] The above - described specific implementation manners further elaborate the purpose, technical solutions, and beneficial effects of the present invention. It should be understood that the above - described are only specific implementation manners of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. An information transmission enhancement method for a cross-sensory fusion vehicle network based on deep reinforcement learning, characterized in that, Including: S1. Design the Cramér-Rao lower bound of the target vehicle relative to the roadside unit as a sensing constraint; The Cramér-Rao lower bound of the target vehicle relative to the roadside unit includes the angular Cramér-Rao lower bound of the angle of the target vehicle relative to the roadside unit and the distance Cramér-Rao lower bound of the distance, where the angular Cramér-Rao lower bound is: The distance Cramér-Rao lower bound is: Among them, θ m is the angle of the target vehicle m relative to the roadside unit, and d m is the distance of the target vehicle m relative to the roadside unit. w m represents the ISAC signal beamforming steering vector of the roadside unit for the target vehicle m, and p m represents the transmission power of the roadside unit for transmitting the ISAC signal to the target vehicle m. and represent the estimated noise variances corresponding to the angle and distance, and c represents the speed of propagation of the electromagnetic wave signal in a vacuum; S2. Obtain the sum rate of the ISAC signals between all target vehicles and the roadside unit; S3. Based on the sum rate of the ISAC signals and the sensing constraint, establish a sum rate maximization model related to the beamforming and transmit power of the roadside unit; S4. Solve the sum rate maximization model based on the deep reinforcement learning method. By jointly optimizing the beamforming and transmit power of the roadside unit, obtain the optimal beamforming and transmit power, so as to improve the information transmission rate of the vehicle network while ensuring the sensing performance requirements.

2. The information transmission enhancement method for a synaesthetic fusion vehicle networking based on deep reinforcement learning according to claim 1, wherein In S2, the sum rate of the ISAC signals between all target vehicles and the roadside unit is specifically: where M is the total number of target vehicles, γ m (w m , p m ) is the signal-to-interference-plus-noise ratio of target vehicle m receiving the ISAC signal.

3. The information transmission enhancement method for the synesthesia fusion vehicle networking based on deep reinforcement learning according to claim 2, wherein The signal-to-interference-plus-noise ratio γ of the target vehicle m receiving the ISAC signal m (w m ,p m ) The specific formula is as follows: Among them, G represents the gain of the antenna array of the roadside unit, a H (θ m ) is the transmitting steering vector of the antenna array, is the vehicle receiving noise power, w j represents the steering vector for the roadside unit to perform ISAC signal beamforming on vehicle j of non-target vehicle m, p j represents the transmission power of the ISAC signal transmitted by the roadside unit to vehicle j of non-target vehicle m, α m is the path loss coefficient of the communication transmission channel between the target vehicle m and the roadside unit.

4. An information transmission enhancement method for a synesthesia fusion vehicle networking based on deep reinforcement learning according to claim 3, characterized in that, Path loss coefficient α of the communication transmission channel between the target vehicle m and the roadside unit m Specifically: where ζ and α0 are the path loss exponent and the path loss at the reference distance d0, respectively.

5. The enhanced method for information transmission in a cross-sensory fusion vehicle network based on deep reinforcement learning according to claim 4, wherein The sum rate maximization model is specifically: Among them, \(W\) represents the downlink beamforming matrix between the roadside unit and the target vehicle, \(p\) is the power vector for the roadside unit to transmit the ISAC signal, and \(P\) max is the maximum transmit power supported by the roadside unit, and \(\sigma\) θ and \(\sigma\) d are the maximum tolerable Cramér-Rao lower bound thresholds to ensure the sensing accuracy, and \(\sigma\) θ is the maximum threshold of the angular Cramér-Rao lower bound, and \(\sigma\) d is the maximum threshold of the range Cramér-Rao lower bound.

6. The enhanced method for information transmission in a cross-sensory fusion vehicle network based on deep reinforcement learning according to claim 5, wherein Before S4, iterate and train the neural network corresponding to the deep reinforcement learning method, and optimize the strategy of the iteration process. The iteration process follows the Bellman equation. The neural network corresponding to the deep reinforcement learning method includes a target neural network and an estimation neural network. The training process is: S4.

1. Regard the roadside unit as an agent, and design the action space, state space, and environmental reward of the agent; S4.

2. Store the action, state, and environmental reward information obtained by the agent interacting with the environment in the decision-making stage in the experience replay pool; S4.

3. Randomly select a batch of training data with a fixed size from the experience replay pool for training; S4.

4. Transmit the batch of training data to the target neural network and the estimation neural network respectively; S4.

5. After inputting the batch of training data into the target neural network, output the target Q value. After inputting the training data into the estimation neural network, output the estimated Q value; S4.

6. Calculate the error loss between the target Q value and the estimated Q value. The estimation neural network updates the network parameters based on the error loss to reduce the error; S4.

7. Regularly update the target neural network according to the parameters of the estimation neural network.

7. An information transmission enhancement method for a synaesthesia fusion vehicle networking based on deep reinforcement learning according to claim 6, characterized in that, In S4.2, the specific process of storing the action, state, and environmental reward information obtained by the agent interacting with the environment in the experience replay pool is: S4.2.

1. Initialize the parameters of the target neural network and the estimation neural network, and obtain the initial state s0 of the agent; S4.2.

2. The state s of the agent in the current environment t Under this condition, based on the estimated Q-value of the estimated neural network and the greedy policy, output the current action value of the agent; S4.2.

3. Obtain a reward value r by interacting with the environment based on the current action value of the agent t Give it to the agent, and generate the next state value s t+1 ; S4.2.

4. Repeat S4.2.2 - S4.2.3 to obtain a number of training data, where the training data includes state s t , action a t , reward r t and the state s at the next moment t+1 .

8. The information transmission enhancement method for a synaesthetic fusion vehicle network based on deep reinforcement learning according to claim 7, characterized in that The parameters of the target neural network and the estimation neural network include the learning rate η, the reward discount factor γ, and the exploration probability the training step size M max , the number of iterations N max .

9. The information transmission enhancement method for a synaesthetic fusion vehicle networking based on deep reinforcement learning according to claim 6, wherein The state space of the agent is where R is the sum rate of transmitting the ISAC signal, is the estimated angle parameter of the vehicle position obtained from the echo signal, is the estimated distance parameter of the vehicle position obtained from the echo signal; The action space of the agent is a t =[A θ ,A p ; including the beamforming action subspace and the power allocation action subspace where d θ is the length of the beamforming action subspace, θ max is the maximum angle value of the roadside unit beamforming, P max is the maximum power of the roadside unit to transmit the ISAC signal, d p is the length of the power allocation action subspace; The environmental reward of the agent is a combination of the transmission and rate of the ISAC signal and the Cramér-Rao lower bound constraint, i.e., r t = R - λ1CRLB(θ m , w m , p m ) - λ2CRLB(d m , w m , p m ), where λ1 and λ2 are penalty parameters.