A Millimeter-Wave Radar Blind Spot Detection Method Based on Deep Reinforcement Learning

By improving DQN algorithm and data processing technology, the problem of insufficient decision-making of millimeter-wave radar blind spot detection system in complex traffic scenarios is solved, more efficient and reliable blind spot detection is achieved, and the safety of autonomous driving is improved.

CN120103296BActive Publication Date: 2025-07-11ANHUI FALCON WAVE TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510570828.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-06
Publication Date
2025-07-11
Estimated Expiration
2045-05-06

AI Technical Summary

Technical Problem

The existing millimeter-wave radar blind spot detection system is difficult to make real-time decisions in a timely and accurate manner in complex and changing traffic scenarios. The lack of robustness and adaptability leads to a decrease in detection accuracy and reliability, which cannot meet the safety needs of intelligent transportation and autonomous driving.

Method used

Using a deep reinforcement learning method, we optimize the reward function by improving the DQN algorithm, combining chaotic sequence modulation, Kalman filtering and DBSCAN algorithm for data processing, and introducing a priority experience playback mechanism and a gradient descent algorithm with dynamic weight allocation are introduced to optimize the decision process.

Benefits of technology

It improves real-time decision-making capabilities in complex traffic scenarios, enhances the environmental adaptability and anti-interference capabilities of signals, improves the reliability of data processing and target tracking accuracy, and ensures the accuracy and safety of blind spot detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120103296B_ABST
    Figure CN120103296B_ABST
Patent Text Reader

Abstract

The present invention relates to autonomous driving, and specifically to a millimeter-wave radar blind spot detection method based on deep reinforcement learning, which detects target information in the blind spot area of a vehicle by transmitting and receiving millimeter-wave signals; filters, clusters, and tracks the millimeter-wave radar data to obtain target data in the blind spot area; makes real-time decisions based on the target data in the blind spot area through an improved DQN algorithm, judges whether there is a dangerous situation and gives the best response strategy; provides a driving alert or completes an automatic avoidance action according to the output result of the improved DQN algorithm; the technical solution provided by the present invention can effectively overcome the defect that the prior art has difficulty in making real-time decisions in a timely and accurate manner in complex and changeable traffic scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to autonomous driving, and more particularly to a millimeter-wave radar blind spot detection method based on deep reinforcement learning. Background Art

[0002] In the fields of intelligent transportation and autonomous driving, blind spot detection has always been a key research content for ensuring traffic safety. With the rapid development of intelligent transportation systems and the gradual popularization of autonomous driving technologies, the safety issues during vehicle driving have received increasing attention. Reliable blind spot detection technologies have a crucial impact on improving vehicle driving safety and reducing the incidence of traffic accidents. Especially in complex and changing traffic scenarios, accurate blind spot detection can timely alert drivers of potential dangers, avoid collision accidents caused by visual blind spots, and provide strong guarantees for the safe operation of intelligent transportation and autonomous driving.

[0003] In traditional blind spot detection methods, most rely on simple sensor technologies or rule-based algorithms, mainly focusing on partial environmental information around the vehicle. However, these methods are significantly affected by environmental factors such as weather and lighting. Under harsh conditions such as rain, fog, and night, their detection accuracy and reliability will drop significantly. At the same time, traditional blind spot detection methods have limited capabilities in detecting and identifying targets in complex traffic scenarios, are difficult to accurately distinguish different types of targets, and cannot meet the safety requirements of intelligent transportation and autonomous driving.

[0004] Currently, although millimeter-wave radars have been somewhat applied in blind spot detection, most are based on their basic ranging and speed measurement functions, lacking in-depth mining and effective utilization of radar data. Moreover, existing millimeter-wave radar-based blind spot detection systems often adopt fixed detection strategies, making it difficult to adapt to different driving scenarios and environmental changes. In addition, in the face of complex traffic environments and dynamically changing targets, the robustness and adaptability of the system are insufficient, unable to make timely and accurate decisions, resulting in unsatisfactory blind spot detection effects, making it impossible for drivers to obtain reliable blind spot monitoring and avoidance information during driving, increasing the risk of traffic accidents, and also restricting the further development and application of autonomous driving technologies. Summary of the Invention

[0005] Aiming at the above-mentioned shortcomings of the prior art, the present invention provides a millimeter-wave radar blind spot detection method based on deep reinforcement learning, which can effectively overcome the defect of the prior art that it is difficult to make real-time decisions in a timely and accurate manner in complex and changing traffic scenarios.

[0006] To achieve the above object, the present invention is realized through the following technical solutions:

[0007] A millimeter-wave radar blind spot detection method based on deep reinforcement learning, comprising the following steps:

[0008] S1. Detect the target information in the vehicle blind spot area by transmitting and receiving millimeter wave signals;

[0009] S2. Filter, cluster, and track the millimeter wave radar data to obtain the target data in the blind spot area;

[0010] S3. Based on the target data in the blind spot area, make real-time decisions through an improved DQN algorithm, judge whether there is a dangerous situation, and give the best response strategy;

[0011] S4. Provide driving alerts or complete automatic avoidance actions according to the output results of the improved DQN algorithm;

[0012] Among them, the improved DQN algorithm optimizes the design of the reward function by considering multi-dimensional factors and dynamically adjusting the reward coefficient to adapt to complex and changeable traffic scenarios and make real-time decisions efficiently;

[0013] In the training process of the improved DQN algorithm, a prioritized experience replay mechanism is introduced to preferentially select important experience samples for learning, accelerating the model convergence speed; at the same time, a gradient descent algorithm based on dynamic weight allocation and feedback regulation is used to update the parameters of the evaluation network, and the parameters of the target network are updated regularly to reduce the Q value estimation deviation, make real-time decisions more efficiently, accurately judge dangerous situations, and give the best response strategies.

[0014] Preferably, in S1, detecting the target information in the vehicle blind spot area by transmitting and receiving millimeter wave signals includes:

[0015] S11. Use a chaotic sequence x n to perform chaotic modulation on the millimeter wave carrier signal . The millimeter wave signal after chaotic modulation s ( t ) is:

[0016] ;

[0017] Among them, A c , f c are the amplitude and frequency of the millimeter wave carrier signal respectively, t is the time, k is the modulation coefficient, n is the chaotic sequence x n 's time index, N is the chaotic sequence x n 's total length, rectis a rectangular pulse function, is the sampling period;

[0018] S12. In an actual vehicle driving scenario, to better adapt to the influence of different environmental factors on the propagation of millimeter-wave signals s ( t ), let the modulation coefficient k adaptively change with events or the environment, denoted as k ( t ). At the same time, considering the multipath fading characteristics of the channel, introduce the channel impulse response h ( t ) and convolve it with the millimeter-wave signal s ( t ). The updated millimeter-wave signal after chaotic modulation s’ ( t ) is:

[0019] .

[0020] Preferably, in S2, filter, cluster, and target-track the millimeter-wave radar data to obtain the target data in the blind spot area, including:

[0021] S21. Perform Kalman filtering on the millimeter-wave radar data to remove the noise interference in the millimeter-wave radar data;

[0022] S22. Cluster the data points with similar characteristics according to the density connectivity between data points to distinguish different targets;

[0023] S23. Associate the trajectories of the same target in the consecutive-frame millimeter-wave radar data to obtain the target data in the blind spot area.

[0024] Preferably, performing Kalman filtering on the millimeter-wave radar data in S21 to remove the noise interference in the millimeter-wave radar data includes:

[0025] The state equation and observation equation of Kalman filtering are:

[0026] ;

[0027] Among them, X k , X k-1 are the state vectors at time k and time k -1 respectively, A is the state transition matrix, U k is the control input vector at time k , B is the input matrix,W k is the process noise vector at time k ;

[0028] Z k is the observation vector at time k ; H is the observation matrix; V k is the observation noise vector at time k ;

[0029] In S22, according to the density connectivity between data points, data points with similar features are clustered to distinguish different targets, including:

[0030] According to the density connectivity between data points, the DBSCAN algorithm is used to cluster data points with similar features to distinguish different targets;

[0031] In S23, the trajectories of the same target are associated in consecutive frame millimeter-wave radar data to obtain the target data in the blind spot area, including:

[0032] The multi-object tracking MOT algorithm is used to associate the trajectories of the same target in consecutive frame millimeter-wave radar data to obtain the target data in the blind spot area.

[0033] Preferably, the improved DQN algorithm optimizes the design of the reward function by considering multi-dimensional factors and dynamically adjusting the reward coefficient, including:

[0034] S311. In the reward function, a velocity-distance reward function and an environmental interference reward function are introduced:

[0035] ;

[0036] where is the relative velocity with respect to the target, t is the time, is the relative distance with respect to the target, , are the velocity change weight coefficient and the distance change weight coefficient respectively, used to adjust the influence degree of the velocity and distance change rates on the reward. When the target approaches rapidly and the distance is too close, an additional negative reward is given;

[0037] I is the interference intensity in the environment, u is the environmental interference reward coefficient. When the interference intensity is large and the target can still be accurately detected, a positive reward compensation is given;

[0038] S312. In the reward function, introduce reward items for successfully detecting blind spots, taking correct avoidance strategies, false alarms, missed alarms, and excessive avoidance. The corresponding reward coefficients for each part are , , , , :

[0039] ;

[0040] Among them, , , , , are respectively the initial reward values in the reward coefficients for successfully detecting blind spots, taking correct avoidance strategies, false alarms, missed alarms, and excessive avoidance. And , , , , , , , , , are respectively the growth coefficients in the reward coefficients for successfully detecting blind spots, taking correct avoidance strategies, false alarms, missed alarms, and excessive avoidance. n is the number of training steps, N is the total number of training steps;

[0041] S313. The reward function is finally expressed as:

[0042] ;

[0043] Among them, s is the current state, a is the action taken in the current state s , s’ is the next state after taking the action a , S 1, S 2, S 3, S 4, S 5 are respectively the state identifiers for successfully detecting blind spots, taking correct avoidance strategies, false alarms, missed alarms, and excessive avoidance, and the values are all 0 or 1.

[0044] Preferably, the improved DQN algorithm introduces a prioritized experience replay mechanism during training, preferentially selects important experience samples for learning, and accelerates the model convergence speed, including:

[0045] S321. Let the iThe experience sample of the step e i is as follows:

[0046] ;

[0047] Among them, s i is the state of the i th step, a i is the action taken in the state s i , r i is the reward obtained after taking the action a i , s i+1 is the state of the i +1th step, d i is the flag indicating whether it ends;

[0048] S322. Calculate the TD error e i of the experience sample :

[0049] ;

[0050] Among them, is the value estimation of the evaluation network with parameter s i when taking the action a i in the state Q , is the value estimation of the target network with parameter s i+1 when taking the action a’ in the state Q , is the discount factor, and ;

[0051] S323. Prioritize all experience samples according to the magnitude of the TD error and calculate the priority e i of the experience sample p i :

[0052] ;

[0053] Among them, is a very small integer used to avoid the case where the priority is zero when

[0054] S324. Calculate the sampling probability of the empirical samples e i : P ( i ):

[0055] ;

[0056] wherein i and j both represent the number of steps, is a parameter that controls the degree of influence of the priority, and when it is equivalent to random sampling, and when it samples all empirical samples completely according to the priority, N is the size of the experience replay buffer;

[0057] S325. Sample a batch of important empirical samples according to the sampling probability of each empirical sample to form an important empirical sample set wherein n is the size of the important empirical sample set.

[0058] Preferably, the parameters of the evaluation network are updated by using a gradient descent algorithm based on dynamic weight allocation and feedback regulation, including:

[0059] S331. Calculate the target value of the important empirical sample in the important empirical sample set :

[0060] ;

[0061] wherein is the state at the b k -th step, is the action taken in the state , is the reward obtained after taking the action , is the state at the b k +1-th step, is the flag indicating whether it is over, is the target network in the state when taking the action a’ of the Q value estimate;

[0062] S332. Calculate the evaluation network At state Take action at the Q value estimation , and calculate the loss function :

[0063] ;

[0064] S333. Update the parameters of the evaluation network using the gradient descent algorithm based on dynamic weight allocation and feedback regulation of .

[0065] Preferably, in S333, update the parameters of the evaluation network using the gradient descent algorithm based on dynamic weight allocation and feedback regulation of , including:

[0066] S3331. Initialization stage: Initialize the learning rate , the dynamic weight allocation coefficient , the feedback regulation coefficient , construct the comprehensive sensitivity S for recording the loss sensitivity of important experience samples to the parameters , and set the initial value to 0;

[0067] S3332. Gradient calculation stage: Calculate the target value of the important experience sample and the evaluation network of Q value estimation The square of the difference between them with respect to the parameter gradient :

[0068] ;

[0069] Among them, is the evaluation network of Q value estimation with respect to the parameter gradient;

[0070] Calculate the average value of the above gradients for all important experience samples in the important experience sample set to obtain the loss function with respect to the parameter gradient :

[0071] ;

[0072] S3333. Dynamic Weight Allocation Phase: Calculate important experience samples For parameter Loss sensitivity , that is, for parameter Perform a small perturbation Get the change in the loss function And this small perturbation The ratio between them:

[0073] ;

[0074] Calculate the set of important experience samples The comprehensive sensitivity of all important experience samples in the set to parameter : S :

[0075] ;

[0076] According to the comprehensive sensitivity S Update the dynamic weight allocation coefficient w :

[0077] ;

[0078] S3334. Parameter Update Phase: Adjust the learning rate w According to the dynamic weight allocation coefficient , and update parameter , that is ;

[0079] S3335. Feedback Regulation Phase: After completing one round of parameter update, evaluate the performance of the evaluation network L val Through the average loss of the validation set . If the decrease in the average loss of the validation set L val Exceeds the preset threshold compared to the previous round, it is considered that the current weight allocation strategy is effective, and keep the dynamic weight allocation coefficient w Unchanged; otherwise, perform feedback regulation on the dynamic weight allocation coefficient w :

[0080] For the case where the comprehensive sensitivity S Is relatively high, but the performance does not improve after parameter update, appropriately reduce the dynamic weight allocation coefficient w , that is ;

[0081] For the case where the comprehensive sensitivity S Is relatively low, but the performance does not decrease after parameter update, appropriately increase the dynamic weight allocation coefficient w , that is 。

[0082] Preferably, updating the parameters of the target network regularly includes:

[0083] Determine whether the current step is a multiple of the target network update interval C If so, assign the parameters of the evaluation network to the target network That is 。 。

[0084] Compared with the prior art, a millimeter-wave radar blind spot detection method based on deep reinforcement learning provided by the present invention has the following beneficial effects:

[0085] 1) In terms of signal modulation, chaotic modulation is performed on the millimeter-wave carrier signal using a chaotic sequence, combined with an adaptive modulation coefficient and a channel impulse response, enhancing the environmental adaptability and anti-interference ability of the signal;

[0086] 2) In terms of data processing, the Kalman filter, DBSCAN algorithm, and multi-object tracking MOT technology are used to greatly improve the reliability of data processing and the target tracking accuracy;

[0087] 3) The decision-making algorithm adopts an improved DQN algorithm. By considering multi-dimensional factors and dynamically adjusting the reward coefficient, the design of the reward function is optimized to adapt to complex and changing traffic scenarios and make real-time decisions efficiently; a prioritized experience replay mechanism is introduced during training to preferentially select important experience samples for learning and accelerate the model convergence speed; a gradient descent algorithm based on dynamic weight allocation and feedback regulation is used to update the parameters of the evaluation network, and the parameters of the target network are updated regularly to reduce the Q value estimation deviation during training, make real-time decisions more efficiently, accurately judge dangerous situations, and give the best response strategies. BRIEF DESCRIPTION OF THE DRAWINGS

[0088] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.

[0089] Figure 1 is a flow diagram of the present invention;

[0090] Figure 2 is a flow diagram of the improved DQN algorithm in the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0091] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0092] A millimeter-wave radar blind spot detection method based on deep reinforcement learning, as Figure 1 and Figure 2 shown, S1: Detect the target information in the vehicle blind spot area by transmitting and receiving millimeter-wave signals, specifically including:

[0093] S11: Use a chaotic sequence x n (chaotic signals have characteristics such as non-periodicity, broadband spectrum, and sensitivity to initial conditions) to perform chaotic modulation on the millimeter-wave carrier signal . The millimeter-wave signal after chaotic modulation s ( t ) is:

[0094] ;

[0095] Among them, A c , f c are the amplitude and frequency of the millimeter-wave carrier signal respectively, t is time, k is the modulation coefficient, n is the chaotic sequence x n 's time index, N is the chaotic sequence x n 's total length, rect is the rectangular pulse function, is the sampling period;

[0096] S12: In the actual vehicle driving scenario, to better adapt to the influence of different environmental factors on the propagation of the millimeter-wave signal s ( t ), make the modulation coefficient k adaptively change with events or the environment, denoted as k ( t ). At the same time, considering the multipath fading characteristics of the channel, introduce the channel impulse response h ( t ) and the millimeter-wave signal s ( t)Perform convolution, and the updated millimeter-wave signal after chaotic modulation s’ ( t ) is as follows:

[0097] .

[0098] S2. Filter, cluster, and track the millimeter-wave radar data to obtain the target data in the blind spot area, specifically including:

[0099] S21. Perform Kalman filtering on the millimeter-wave radar data to remove the noise interference in the millimeter-wave radar data;

[0100] S22. Cluster the data points with similar characteristics according to the density connectivity between data points to distinguish different targets;

[0101] S23. Associate the trajectories of the same target in the consecutive-frame millimeter-wave radar data to obtain the target data in the blind spot area.

[0102] Specifically, performing Kalman filtering on the millimeter-wave radar data in S21 to remove the noise interference in the millimeter-wave radar data includes:

[0103] The state equation and observation equation of Kalman filtering are:

[0104] ;

[0105] Among them, X k , X k-1 are the state vectors at times k and k -1 respectively, A is the state transition matrix, U k is the control input vector at time k , B is the input matrix, W k is the process noise vector at time k ;

[0106] Z k is the observation vector at time k , H is the observation matrix, V k is the observation noise vector at time k ;

[0107] In S22, clustering the data points with similar characteristics according to the density connectivity between data points to distinguish different targets includes:

[0108] According to the density connectivity between data points, the DBSCAN algorithm is used to cluster data points with similar features to distinguish different targets;

[0109] In S23, the trajectories of the same target are associated in the millimeter-wave radar data of consecutive frames to obtain the target data in the blind spot area, including:

[0110] The multi-object tracking MOT algorithm is used to associate the trajectories of the same target in the millimeter-wave radar data of consecutive frames to obtain the target data in the blind spot area.

[0111] S3. Based on the target data in the blind spot area, the improved DQN algorithm is used to make real-time decisions to judge whether there is a dangerous situation and give the best response strategy.

[0112] In the technical solution of this application, as Figure 2 shown, the improved DQN algorithm optimizes the design of the reward function by considering multi-dimensional factors and dynamically adjusting the reward coefficient to adapt to complex and changeable traffic scenarios and make real-time decisions efficiently;

[0113] In the training process of the improved DQN algorithm, a prioritized experience replay mechanism is introduced to preferentially select important experience samples for learning to accelerate the model convergence speed; at the same time, a gradient descent algorithm based on dynamic weight allocation and feedback regulation is used to update the parameters of the evaluation network, and the parameters of the target network are updated regularly to reduce the Q value estimation deviation in the training process, make real-time decisions more efficiently, accurately judge dangerous situations and give the best response strategy.

[0114] 1) The improved DQN algorithm optimizes the design of the reward function by considering multi-dimensional factors and dynamically adjusting the reward coefficient, including:

[0115] S311. In the reward function, a speed-distance reward function and an environmental interference reward function are introduced:

[0116] ;

[0117] Among them, is the relative speed to the target, t is the time, is the relative distance to the target, , are the speed change weight coefficient and the distance change weight coefficient respectively, which are used to adjust the influence degree of the speed and distance change rate on the reward. When the target approaches rapidly and the distance is too close, an additional negative reward is given;

[0118] I is the interference intensity in the environment,u is the environmental interference reward coefficient. When the interference intensity is large and the target can still be accurately detected, a positive reward compensation is given;

[0119] S312. In the reward function, introduce reward terms for successfully detecting blind spots, adopting correct avoidance strategies, false alarms, missed alarms, and excessive avoidance. The corresponding reward coefficients for each part are 、 、 、 、 :

[0120] ;

[0121] Among them, 、 、 、 、 are the initial reward values in the reward coefficients for successfully detecting blind spots, adopting correct avoidance strategies, false alarms, missed alarms, and excessive avoidance, respectively. And , , , , , 、 、 、 、 are the growth coefficients in the reward coefficients for successfully detecting blind spots, adopting correct avoidance strategies, false alarms, missed alarms, and excessive avoidance, respectively. n is the number of training steps, N is the total number of training steps;

[0122] S313. The reward function is finally expressed as:

[0123] ;

[0124] Among them, s is the current state, a is the action taken in the current state s , s’ is the next state after taking the action a , S 1、 S 2、 S 3、 S 4、 S 5 are the state identifiers for successfully detecting blind spots, adopting correct avoidance strategies, false alarms, missed alarms, and excessive avoidance, respectively, and their values are all 0 or 1.

[0125] 2) The improved DQN algorithm introduces a prioritized experience replay mechanism during training, preferentially selecting important experience samples for learning to accelerate the model convergence speed, including:

[0126] S321. Let the experience sample at the i th step be: e i :

[0127] ;

[0128] where s i is the state at the i th step, a i is the action taken in the state s i , r i is the reward obtained after taking the action a i , s i+1 is the state at the i +1th step, d i is the flag indicating whether it ends;

[0129] S322. Calculate the TD error e i of the experience sample :

[0130] ;

[0131] where is the value estimate of the evaluation network with parameter s i when taking the action a i in the state Q , is the value estimate of the target network with parameter s i+1 when taking the action a’ in the state Q , is the discount factor, and ;

[0132] S323. Prioritize all experience samples according to the magnitude of the TD error and calculate the priority e i of the experience sample pi :

[0133] ;

[0134] wherein, is a very small integer used to avoid the situation where the priority is zero when ;

[0135] S324. Calculate the sampling probability of the empirical samples e i : P ( i ):

[0136] ;

[0137] wherein, i , j both represent the number of steps, is a parameter controlling the influence degree of the priority, and , when , it is equivalent to random sampling, and when , all empirical samples are sampled completely according to the priority, N is the size of the experience replay buffer;

[0138] S325. Sample a batch of important empirical samples according to the sampling probability of each empirical sample to form an important empirical sample set , wherein n is the size of the important empirical sample set.

[0139] 3) Update the parameters of the evaluation network using a gradient descent algorithm based on dynamic weight allocation and feedback regulation, including:

[0140] S331. Calculate the target value of the important empirical sample in the important empirical sample set :

[0141] ;

[0142] wherein, is the state at the b k th step, is the action taken in the state , is the reward obtained after taking the action , is the state at the b k +1th step, is the flag indicating whether it is the end, For the target network At state Take action a’ At that time Q Value estimation;

[0143] S332. Calculate the evaluation network At state Take action At that time Q Value estimation And calculate the loss function :

[0144] ;

[0145] S333. Update the parameters of the evaluation network using the gradient descent algorithm based on dynamic weight allocation and feedback regulation Of .

[0146] Specifically, in S333, the parameters of the evaluation network are updated using the gradient descent algorithm based on dynamic weight allocation and feedback regulation Of , including:

[0147] S3331. Initialization stage: Initialize the learning rate , dynamic weight allocation coefficient , feedback regulation coefficient , construct the comprehensive sensitivity S To record the loss sensitivity of important experience samples to the parameters , with the initial value set to 0;

[0148] S3332. Gradient calculation stage: Calculate the target value of the important experience sample And the difference between the value estimation Of the evaluation network Of Q Value estimation The square of the difference between them with respect to the parameter Gradient :

[0149] ;

[0150] Among them, Is the value estimation of the evaluation network Of Q Value estimation With respect to the parameter Gradient;

[0151] Calculate the average value of the above gradients for all important experience samples in the important experience sample set To obtain the loss function With respect to the parameter Gradient :

[0152] ;

[0153] S3333. Dynamic Weight Allocation Phase: Calculate the important experience samples for the parameter loss sensitivity , that is, for the parameter make a small perturbation to obtain the change in the loss function and this small perturbation the ratio between them:

[0154] ;

[0155] Calculate the comprehensive sensitivity of all important experience samples in the important experience sample set to the parameter : S :

[0156] ;

[0157] According to the comprehensive sensitivity S update the dynamic weight allocation coefficient w :

[0158] ;

[0159] S3334. Parameter Update Phase: Adjust the learning rate w according to the dynamic weight allocation coefficient , and update the parameter , that is ;

[0160] S3335. Feedback Regulation Phase: After completing one round of parameter update, evaluate the performance of the evaluation network through the average loss L val of the validation set . If the decrease in the average loss of the validation set L val exceeds the preset threshold (0.1) compared with the previous round, it is considered that the current weight allocation strategy is effective, and the dynamic weight allocation coefficient w remains unchanged; otherwise, perform feedback regulation on the dynamic weight allocation coefficient w :

[0161] For the case where the comprehensive sensitivity S is relatively high, but the performance does not improve after parameter update, appropriately reduce the dynamic weight allocation coefficient w , that is ;

[0162] For the comprehensive sensitivity S is low, but the performance does not decline after the parameter update, appropriately increase the dynamic weight allocation coefficient w , that is .

[0163] 4) Regularly update the parameters of the target network, including:

[0164] Judge whether the current step number is a multiple of the target network update interval C , if so, then assign the parameters of the evaluation network to the target network , that is , that is .

[0165] The pseudo-code of the improved DQN algorithm for the above technical solution is shown in Table 1:

[0166] Table 1 Pseudo-code of the improved DQN algorithm

[0167]

[0168] The pseudo-code of the gradient descent algorithm based on dynamic weight allocation and feedback regulation is shown in Table 2:

[0169] Table 2 Pseudo-code of the gradient descent algorithm based on dynamic weight allocation and feedback regulation

[0170]

[0171] S4. Provide driving alerts or complete automatic avoidance actions according to the output results of the improved DQN algorithm.

[0172] In the technical solution of this application, in terms of signal modulation, chaotic sequences are used to perform chaotic modulation on millimeter-wave carrier signals, combined with the adaptive modulation coefficient and the channel impulse response, enhancing the environmental adaptability and anti-interference ability of the signals;

[0173] In terms of data processing, Kalman filtering, DBSCAN algorithm and multi-object tracking MOT technology are adopted, greatly improving the reliability of data processing and the target tracking accuracy;

[0174] The decision-making algorithm adopts the improved DQN algorithm, optimizes the design of the reward function by considering multi-dimensional factors and dynamically adjusting the reward coefficient to adapt to complex and changeable traffic scenarios and make real-time decisions efficiently; in the training process, a prioritized experience replay mechanism is introduced to preferentially select important experience samples for learning to accelerate the model convergence speed; the parameters of the evaluation network are updated using the gradient descent algorithm based on dynamic weight allocation and feedback regulation, and the parameters of the target network are updated regularly to reduce the QValue estimation deviation, make real-time decisions more efficiently, accurately judge dangerous situations and give the best response strategies.

[0175] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than limiting it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements will not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A millimeter-wave radar blind spot detection method based on deep reinforcement learning, characterized in that: It includes the following steps: S1. Detect the target information in the vehicle blind spot area by transmitting and receiving millimeter-wave signals; S2. Filter, cluster, and track the millimeter-wave radar data to obtain the target data in the blind spot area; S3. Based on the target data in the blind spot area, make real-time decisions through an improved DQN algorithm, judge whether there is a dangerous situation, and give the best response strategy; S4. Provide a driving alert or complete an automatic avoidance action according to the output result of the improved DQN algorithm; Among them, the improved DQN algorithm optimizes the design of the reward function by considering multi-dimensional factors and dynamically adjusting the reward coefficient to adapt to complex and changing traffic scenarios and make real-time decisions efficiently; The improved DQN algorithm introduces a prioritized experience replay mechanism during the training process, preferentially selects important experience samples for learning to accelerate the model convergence speed; at the same time, it uses a gradient descent algorithm based on dynamic weight allocation and feedback regulation to update the parameters of the evaluation network and regularly updates the parameters of the target network to reduce the Q-value estimation bias during the training process, make real-time decisions more efficiently, accurately judge dangerous situations, and give the best response strategy; In S1, detecting the target information in the vehicle blind spot area by transmitting and receiving millimeter-wave signals includes: S11. Use the chaotic sequence x n to perform chaotic modulation on the millimeter-wave carrier signal . The millimeter-wave signal s(t) after chaotic modulation is as follows: ; Among them, A c , f c are respectively the amplitude and frequency of the millimeter-wave carrier signal, t is time, k is the modulation coefficient, n is the time index of the chaotic sequence x n , N is the total length of the chaotic sequence x n , rect is the rectangular pulse function, T s is the sampling period; S12. In the actual vehicle driving scenario, to better adapt to the influence of different environmental factors on the propagation of the millimeter-wave signal s(t), let the modulation coefficient k change adaptively with events or the environment, denoted as k(t). At the same time, considering the multipath fading characteristics of the channel, introduce the channel impulse response h(t) to convolve with the millimeter-wave signal s(t). The updated millimeter-wave signal s’(t) after chaotic modulation is: 。 2. The millimeter-wave radar blind spot detection method based on deep reinforcement learning according to claim 1, wherein: In S2, filtering, clustering, and tracking the millimeter-wave radar data to obtain the target data in the blind spot area includes: S21. Perform Kalman filtering on the millimeter-wave radar data to remove the noise interference in the millimeter-wave radar data; S22. Cluster the data points with similar characteristics according to the density connectivity between data points to distinguish different targets; S23. Associate the trajectories of the same target in the continuous-frame millimeter-wave radar data to obtain the target data in the blind spot area.

3. The millimeter-wave radar blind spot detection method based on deep reinforcement learning according to claim 2, characterized in that: In S21, performing Kalman filtering on the millimeter-wave radar data to remove the noise interference in the millimeter-wave radar data includes: The state equation and observation equation of Kalman filtering are: ; where X k and X k-1 are the state vectors at time k and time k-1 respectively, A is the state transition matrix, U k is the control input vector at time k, B is the input matrix, and W k is the process noise vector at time k; Z k is the observation vector at time k, H is the observation matrix, and V k is the observation noise vector at time k; In S22, clustering the data points with similar characteristics according to the density connectivity between data points to distinguish different targets includes: Cluster the data points with similar characteristics according to the density connectivity between data points using the DBSCAN algorithm to distinguish different targets; In S23, associating the trajectories of the same target in the continuous-frame millimeter-wave radar data to obtain the target data in the blind spot area includes: Use the multi-object tracking MOT algorithm to associate the trajectories of the same target in the continuous-frame millimeter-wave radar data to obtain the target data in the blind spot area.

4. The millimeter-wave radar blind spot detection method based on deep reinforcement learning according to claim 1, wherein: The improved DQN algorithm optimizes the design of the reward function by considering multi-dimensional factors and dynamically adjusting the reward coefficient, including: S311. Introduce the speed-distance reward function $R$ v,d and the environmental interference reward function $R$ env : ; Among them, v rel is the relative speed to the target, t is the time, and d rel is the relative distance to the target, , are the speed change weight coefficient and the distance change weight coefficient respectively, which are used to adjust the influence degree of the speed and the distance change rate on the reward. When the target approaches rapidly and the distance is too close, an additional negative reward is given; Let \(I\) be the interference intensity in the environment and \(u\) be the environmental interference reward coefficient. When the interference intensity is large and the target can still be accurately detected, a positive reward compensation is given. S312. In the reward function, introduce reward items for successfully detecting blind spots, adopting correct avoidance strategies, false alarms, missed alarms, and excessive avoidance, and the corresponding reward coefficients for each part are , , , , : ; Among them, , , , , are respectively the initial reward values in the reward coefficients for successfully detecting blind spots, taking correct avoidance strategies, false alarms, missed detections, and over-avoidance reward items, and , , , , , , , , , are respectively the growth coefficients in the reward coefficients for successfully detecting blind spots, taking correct avoidance strategies, false alarms, missed detections, and over-avoidance reward items, where n is the number of training steps and N is the total number of training steps; S313. Reward function Finally expressed as: ; Among them, \(s\) is the current state, \(a\) is the action taken in the current state \(s\), \(s'\) is the next state after taking the action \(a\), and \(S1\), \(S2\), \(S3\), \(S4\), \(S5\) are the state identifiers for successfully detecting blind spots, adopting correct avoidance strategies, false alarms, missed alarms, and over - avoidance reward items respectively, and their values are all 0 or 1.

5. The millimeter-wave radar blind spot detection method based on deep reinforcement learning according to claim 1, characterized in that: The improved DQN algorithm introduces a prioritized experience replay mechanism during training, preferentially selecting important experience samples for learning to accelerate the model convergence speed, including: S321. Let the experience sample \(e\) at the \(i\)-th step be i as follows: ; where s i is the state at the i-th step, a i is the action taken in state s i r i is the reward obtained after taking action a i s i+1 is the state at the (i + 1)-th step, and d i is the flag indicating whether it is the end; S322. Calculate the TD error of the empirical sample e i :​ ; Among them, is the evaluation network with parameters for estimating the Q-value when taking action a in state s i ; i is the target network with parameters for estimating the Q-value when taking action a' in state s i+1 ; is the discount factor, and ;​ S323. Prioritize all experience samples according to the magnitude of the TD error and calculate the priority p i of the experience sample e i : ; wherein, is a very small integer used to avoid the case where the priority is zero; S324. Calculate the sampling probability P(i) of the empirical sample e i : ; where both i and j represent the number of steps, is a parameter for controlling the influence degree of the priority, and when it is equivalent to random sampling, and when it samples all experience samples completely according to the priority, and N is the size of the experience replay buffer; S325. Sample a batch of important experience samples according to the sampling probability P(i) of each experience sample to form an important experience sample set , where n is the size of the important experience sample set.

6. The millimeter-wave radar blind spot detection method based on deep reinforcement learning according to claim 5, characterized in that: The gradient descent algorithm based on dynamic weight allocation and feedback regulation is used to update the parameters of the evaluation network, including: S331. Calculate the important empirical sample set The important empirical samples Target value : ; Among them, is the state of the b k -th step, is the action taken in the state , is the reward obtained after taking the action , is the state of the (b k + 1)-th step, is the flag indicating whether it is the end, is the target network when taking the action a' in the state , the Q-value estimation; S332. Calculate the evaluation network At state take action when estimating the Q value and calculate the loss function : ; S333. Update the parameters of the evaluation network using the gradient descent algorithm based on dynamic weight allocation and feedback regulation of the parameter .

7. The method for blind spot detection of millimeter-wave radar based on deep reinforcement learning according to claim 6, characterized in that: In S333, the gradient descent algorithm based on dynamic weight allocation and feedback regulation is used to update the parameters of the evaluation network parameters , including: S3331. Initialization stage: Initialize the learning rate , the dynamic weight allocation coefficient , the feedback adjustment coefficient , construct the comprehensive sensitivity S to record the loss sensitivity of important experience samples to the parameter , and set the initial value to 0; S3332. Gradient calculation stage: Calculate important experience samples of the target value and the Q-value estimate of the evaluation network The square of the difference between them with respect to the parameter gradient : ; Among them, is the evaluation network Q-value estimation with respect to the parameter gradient; For the important experience sample set calculate the average value of the above gradients of all important experience samples in, and obtain the loss function with respect to the parameter gradient : ; S3333. Dynamic weight allocation phase: Calculate important experience samples For parameter The loss sensitivity That is, for parameter Perform a small perturbation The change in the loss function obtained And this small perturbation The ratio between them: ; Calculate the important experience sample set For all important experience samples in The comprehensive sensitivity S of the parameters ; Update the dynamic weight allocation coefficient \(w\) according to the comprehensive sensitivity \(S\): ; S3334, Parameter Update Phase: Adjust the learning rate according to the dynamic weight allocation coefficient w , and update the parameters , that is ; S3335. Feedback adjustment stage: After completing one round of parameter update, use the average loss L of the validation set val to evaluate the performance of the evaluation network . If the decrease in the average loss L of the validation set val compared to the previous round exceeds the preset threshold, it is considered that the current weight allocation strategy is effective, and the dynamic weight allocation coefficient w remains unchanged; otherwise, perform feedback adjustment on the dynamic weight allocation coefficient w: For the case where the comprehensive sensitivity S is relatively high, but the performance does not improve after parameter update, appropriately reduce the dynamic weight distribution coefficient w, that is ; For the case where the comprehensive sensitivity S is low, but the performance does not degrade after parameter update, appropriately increase the dynamic weight allocation coefficient w, that is .

8. The millimeter-wave radar blind spot detection method based on deep reinforcement learning according to claim 7, characterized in that: Regularly update the parameters of the target network, including: Determine whether the current step is a multiple of the target network update interval C. If so, assign the parameters of the evaluation network to the target network , that is .

Citation Information

Patent Citations

  • Decision planning method for automatic driving vehicle in urban traffic scene based on reinforcement learning

    CN118396034A