Millimeter wave radar blind spot detection method based on deep reinforcement learning
By introducing a millimeter-wave radar method based on deep reinforcement learning into the blind spot detection technology, the existing technology has solved the problem of difficult decision-making in complex traffic scenarios, and efficient and accurate blind spot detection is achieved, meeting the safety needs of intelligent transportation and autonomous driving.
Patent Information
- Application Number
- CN202510570828.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-05-06
AI Technical Summary
The existing blind spot detection technology is difficult to make real-time decisions in a timely and accurate manner in complex and changing traffic scenarios, and the detection accuracy and reliability in harsh environments are reduced, which cannot meet the safety needs of intelligent transportation and autonomous driving.
The millimeter wave radar blind spot detection method based on deep reinforcement learning is adopted to detect target information by transmitting and receiving millimeter wave signals, filtering, clustering and target tracking are performed, real-time decision-making is made using the improved DQN algorithm, and reward coefficients and priority experience playback mechanisms are dynamically adjusted to improve the robustness and adaptability of the system.
Efficient real-time decision-making is achieved in complex traffic scenarios, improving the accuracy and reliability of blind spot detection, enhancing the adaptability and robustness of the system, reducing the risk of traffic accidents, and meeting the safety needs of intelligent traffic and autonomous driving.
Smart Images

Figure CN120103296A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to autonomous driving, and in particular to a millimeter-wave radar blind spot detection method based on deep reinforcement learning. Background Art
[0002] In the field of intelligent transportation and autonomous driving, blind spot detection has always been a key research topic to ensure traffic safety. With the rapid development of intelligent transportation systems and the gradual popularization of autonomous driving technology, the safety issues during vehicle driving have received more and more attention. Reliable blind spot detection technology has a vital impact on improving vehicle driving safety and reducing the incidence of traffic accidents. Especially in complex and changing traffic scenes, accurate blind spot detection can promptly remind drivers of potential dangers and avoid collision accidents caused by visual blind spots, providing strong guarantees for the safe operation of intelligent transportation and autonomous driving.
[0003] Traditional blind spot detection methods mostly rely on simple sensor technology or rule-based algorithms, focusing on some environmental information around the vehicle. However, these methods are easily affected by environmental factors such as weather and lighting. In harsh conditions such as rain, fog, and darkness, their detection accuracy and reliability will drop significantly. At the same time, traditional blind spot detection methods have limited target detection and recognition capabilities in complex traffic scenes, and it is difficult to accurately distinguish different types of targets, which cannot meet the safety requirements of intelligent transportation and autonomous driving.
[0004] At present, although millimeter-wave radar has been used in blind spot detection to a certain extent, most of them are based on its basic distance measurement and speed measurement functions, lacking in-depth mining and effective use of radar data. Moreover, existing blind spot detection systems based on millimeter-wave radar often adopt fixed detection strategies, which are difficult to adapt to different driving scenarios and environmental changes. In addition, when faced with complex traffic environments and dynamically changing targets, the system's robustness and adaptability are insufficient, and it is impossible to make decisions in a timely and accurate manner, resulting in unsatisfactory blind spot detection results, making it impossible for drivers to obtain reliable blind spot monitoring and avoidance information during driving, increasing the risk of traffic accidents, and also limiting the further development and application of autonomous driving technology. Summary of the invention
[0005] In view of the above-mentioned shortcomings of the prior art, the present invention provides a millimeter-wave radar blind spot detection method based on deep reinforcement learning, which can effectively overcome the defect of the prior art that it is difficult to make real-time decisions in a timely and accurate manner in complex and changeable traffic scenarios.
[0006] To achieve the above objectives, the present invention is implemented through the following technical solutions: A millimeter wave radar blind spot detection method based on deep reinforcement learning includes the following steps: S1, detecting target information in the blind spot area of the vehicle by transmitting and receiving millimeter wave signals; S2, filtering, clustering and target tracking the millimeter wave radar data to obtain target data in the blind spot area; S3, based on the target data in the blind spot area, make real-time decisions by improving the DQN algorithm to determine whether there is a dangerous situation and give the best response strategy; S4, providing driving warnings or completing automatic evasive actions according to the output results of the improved DQN algorithm; Among them, the improved DQN algorithm optimizes the reward function design by considering multi-dimensional factors and dynamically adjusting the reward coefficient to adapt to complex and changing traffic scenarios and make real-time decisions efficiently; The improved DQN algorithm introduces a priority experience replay mechanism during the training process, giving priority to important experience samples for learning to speed up the model convergence; at the same time, a gradient descent algorithm based on dynamic weight allocation and feedback adjustment is used to update the parameters of the evaluation network, and the parameters of the target network are regularly updated to reduce the training process. Q The system can reduce the value estimation deviation, make real-time decisions more efficiently, accurately judge dangerous situations and give the best response strategy.
[0007] Preferably, detecting target information in the blind spot area of the vehicle by transmitting and receiving millimeter wave signals in S1 includes: S11. Using chaotic sequence x n For millimeter wave carrier signals Chaotic modulation is performed, and the millimeter wave signal after chaotic modulation s ( t )for: ; in, A c , f c are the amplitude and frequency of the millimeter wave carrier signal, t For time, k is the modulation coefficient, n Chaotic sequence x n The time index, N Chaotic sequence x n The total length of rect is a rectangular pulse function, is the sampling period; S12. In actual vehicle driving scenarios, in order to better adapt to different environmental factors on millimeter wave signals s ( t ) propagation, so that the modulation coefficient kAdaptively change with events or environments, denoted as k ( t ), while considering the multipath fading characteristics of the channel, the channel impulse response is introduced h ( t ) and millimeter wave signals s ( t ) is convolved to update the chaotically modulated millimeter wave signal s’ ( t )for: .
[0008] Preferably, in S2, the millimeter wave radar data is filtered, clustered, and the target is tracked to obtain the target data of the blind spot area, including: S21, performing Kalman filtering on the millimeter-wave radar data to remove noise interference in the millimeter-wave radar data; S22, clustering data points with similar characteristics according to the density connectivity between data points in order to distinguish different targets; S23. Associating the trajectory of the same target in the continuous frame millimeter wave radar data to obtain target data in the blind spot area.
[0009] Preferably, in S21, Kalman filtering is performed on the millimeter wave radar data to remove noise interference in the millimeter wave radar data, including: The state equation and observation equation of Kalman filter are: ; in, X k , X k-1 Separately for the moment k ,time k The state vector is -1, A is the state transfer matrix, U k For the moment k The control input vector, B is the input matrix, W k For the moment k The process noise vector of Z k For the moment k The observation vector, H is the observation matrix, V k For the moment k The observation noise vector of In S22, data points with similar characteristics are clustered according to the density connectivity between data points in order to distinguish different targets, including: According to the density connectivity between data points, the DBSCAN algorithm is used to cluster data points with similar characteristics in order to distinguish different targets; In S23, the trajectory of the same target is associated in the continuous frame millimeter wave radar data to obtain the target data of the blind spot area, including: The multi-target tracking (MOT) algorithm is used to associate the trajectory of the same target in continuous frame millimeter-wave radar data to obtain the target data in the blind spot area.
[0010] Preferably, the improved DQN algorithm optimizes the reward function design by considering multi-dimensional factors and dynamically adjusting the reward coefficient, including: S311. In the reward function, introduce the speed distance reward function and the environmental perturbation reward function : ; in, is the relative speed to the target, t For time, is the relative distance to the target, , They are the speed change weight coefficient and the distance change weight coefficient, which are used to adjust the influence of speed and distance change rate on the reward. When the target approaches quickly and the distance is too close, an additional negative reward is given; I is the interference intensity in the environment, u is the environmental interference reward coefficient. When the interference intensity is large and the target can still be accurately detected, positive reward compensation is given; S312. In the reward function, the reward items for successfully detecting blind spots, taking correct avoidance strategies, false positives, missed negatives, and excessive avoidance are introduced. The corresponding reward coefficients for each part are , , , , : ; in, , , , , are the initial reward values in the reward coefficients for successfully detecting blind spots, taking correct avoidance strategies, false positives, missed negatives, and excessive avoidance rewards, respectively, and , , , , , , , , , They are the growth coefficients of the reward coefficients for successfully detecting blind spots, adopting correct avoidance strategies, false positives, missed negatives, and excessive avoidance rewards. n is the number of training steps, N is the total number of training steps; S313. Reward function The final expression is: ; in, s is the current state, a For the current state s The following actions are taken, s’ To take action a The next state after S 1 , S 2 , S 3 , S 4 , S 5 They are the status indicators of successfully detecting blind spots, taking correct avoidance strategies, false positives, missed negatives, and excessive avoidance rewards, and their values are all 0 or 1.
[0011] Preferably, the improved DQN algorithm introduces a priority experience replay mechanism during the training process, prioritizes important experience samples for learning, and accelerates the model convergence speed, including: S321, set i Step 1 experience sample e i for: ; in, s i For the i The state of the step, a i For the status s i The following actions are taken, r i To take action a i After the reward, s i+1 For the i +1 step status, d i A sign to indicate whether it is finished; S322. Calculate experience samples e i TD error : ; in, The parameter is Evaluation network In Status s i Take action a i of Q Value estimation, The parameter is Target network In Status s i+1 Take action a’ of Q Value estimation, is the discount factor, and ; S323, according to TD error Prioritize all experience samples by size and calculate the experience samples e i Priority p i : ; in, is a small integer to avoid When the priority is zero; S324. Calculate experience samples e i The sampling probability P ( i ): ; in, i , j All represent the number of steps. is a parameter that controls the degree of influence of priority, and ,when is equivalent to random sampling when When all experience samples are sampled completely according to the priority, N The size of the experience replay buffer; S325. Based on the sampling probability of each experience sample , sample a batch of important experience samples to form an important experience sample set ,in n is the size of the important empirical sample set.
[0012] Preferably, the updating of the parameters of the evaluation network using a gradient descent algorithm based on dynamic weight allocation and feedback regulation includes: S331. Calculate important experience sample sets Important experience samples Target value : ; in, For the b k The state of the step, For the status The following actions are taken, To take action After the reward, For the b k +1 step status, It is a sign of whether it is finished. For the target network In Status Take action a’ of Q Value estimation; S332, Calculation Evaluation Network In Status Take action of Q Value Estimation , and calculate the loss function : ; S333, using the gradient descent algorithm based on dynamic weight allocation and feedback adjustment to update the evaluation network Parameters .
[0013] Preferably, in S333, a gradient descent algorithm based on dynamic weight allocation and feedback regulation is used to update the evaluation network. Parameters ,include: S3331, Initialization phase: Initialize learning rate , dynamic weight allocation coefficient , feedback adjustment coefficient , construct comprehensive sensitivity S , used to record important experience sample parameters The loss sensitivity of is set to 0 initially; S3332, Gradient calculation phase: Calculate important experience samples Target value and evaluation network of Q Value Estimation The square of the difference between Gradient : ; in, To evaluate the network of Q Value Estimation About parameters The gradient of Important experience sample set The above gradient calculation average of all important experience samples in the result is the loss function About parameters Gradient : ; S3333, Dynamic weight allocation stage: Calculation of important experience samples Parameters Sensitivity to loss , that is, for the parameter Make a small disturbance The loss function change obtained With this small disturbance The ratio between: ; Calculate important experience sample sets All important empirical samples in the parameters The overall sensitivity S : ; According to the comprehensive sensitivity S Dynamic weight allocation coefficient w To update: ; S3334, parameter update phase: coefficients are allocated according to dynamic weights w Adjusting the learning rate , and update the parameters ,Right now ; S3335, Feedback adjustment stage: After completing a round of parameter updates, the average loss of the validation set is L val Evaluation Network The performance of the validation set is evaluated. L val If the decrease from the previous round exceeds the preset threshold, the current weight allocation strategy is considered effective and the dynamic weight allocation coefficient is maintained. w unchanged; otherwise, the dynamic weight allocation coefficient w To perform feedback adjustment: For the overall sensitivityS If the performance is not improved after the parameter update, the dynamic weight allocation coefficient should be appropriately reduced. w ,Right now ; For the overall sensitivity S If the performance is not reduced after the parameter update, the dynamic weight allocation coefficient should be appropriately increased. w ,Right now .
[0014] Preferably, the periodically updating the parameters of the target network includes: Determine whether the current step number is the target network update interval C If it is a multiple of , the network will be evaluated Parameters Assign to target network ,Right now .
[0015] Compared with the prior art, the millimeter wave radar blind spot detection method based on deep reinforcement learning provided by the present invention has the following beneficial effects: 1) In terms of signal modulation, chaotic sequences are used to perform chaotic modulation on millimeter wave carrier signals, and combined with adaptive modulation coefficients and channel impulse responses, the environmental adaptability and anti-interference ability of the signal are enhanced; 2) In terms of data processing, Kalman filtering, DBSCAN algorithm and multi-target tracking MOT technology are used to greatly improve the reliability of data processing and target tracking accuracy; 3) The decision algorithm adopts the improved DQN algorithm, which optimizes the reward function design by considering multi-dimensional factors and dynamically adjusting the reward coefficient to adapt to complex and changeable traffic scenarios and make real-time decisions efficiently; introduces a priority experience playback mechanism during the training process, prioritizes important experience samples for learning, and accelerates the model convergence speed; uses a gradient descent algorithm based on dynamic weight allocation and feedback adjustment to update the parameters of the evaluation network, and regularly updates the parameters of the target network to reduce the training process. Q The system can reduce the value estimation deviation, make real-time decisions more efficiently, accurately judge dangerous situations and give the best response strategy. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the prior art descriptions are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention, and for ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0017] Figure 1 It is a schematic diagram of the process of the present invention; Figure 2 It is a flowchart of the improved DQN algorithm in the present invention. DETAILED DESCRIPTION
[0018] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0019] A millimeter wave radar blind spot detection method based on deep reinforcement learning, such as Figure 1 and Figure 2 As shown, S1, detecting target information in the blind spot area of the vehicle by transmitting and receiving millimeter wave signals, specifically including: S11. Using chaotic sequence x n (Chaotic signals have characteristics such as non-periodicity, broadband spectrum and sensitivity to initial conditions) Chaotic modulation is performed, and the millimeter wave signal after chaotic modulation s ( t )for: ; in, A c , f c are the amplitude and frequency of the millimeter wave carrier signal, t For time, k is the modulation coefficient, n Chaotic sequence x n The time index, N Chaotic sequence x n The total length of rect is a rectangular pulse function, is the sampling period; S12. In actual vehicle driving scenarios, in order to better adapt to different environmental factors on millimeter wave signals s ( t ) propagation, so that the modulation coefficient k Adaptively change with events or environments, denoted as k ( t ), while considering the multipath fading characteristics of the channel, the channel impulse response is introduced h ( t ) and millimeter wave signalss ( t ) is convolved to update the chaotically modulated millimeter wave signal s’ ( t )for: .
[0020] S2. Filter, cluster and track the millimeter wave radar data to obtain the target data in the blind spot area, including: S21, performing Kalman filtering on the millimeter-wave radar data to remove noise interference in the millimeter-wave radar data; S22, clustering data points with similar characteristics according to the density connectivity between data points in order to distinguish different targets; S23. Associating the trajectory of the same target in the continuous frame millimeter wave radar data to obtain target data in the blind spot area.
[0021] Specifically, in S21, Kalman filtering is performed on the millimeter wave radar data to remove noise interference in the millimeter wave radar data, including: The state equation and observation equation of Kalman filter are: ; in, X k , X k-1 Separately for the moment k ,time k The state vector is -1, A is the state transfer matrix, U k For the moment k The control input vector, B is the input matrix, W k For the moment k The process noise vector of Z k For the moment k The observation vector, H is the observation matrix, V k For the moment k The observation noise vector of In S22, data points with similar characteristics are clustered according to the density connectivity between data points in order to distinguish different targets, including: According to the density connectivity between data points, the DBSCAN algorithm is used to cluster data points with similar characteristics in order to distinguish different targets; In S23, the trajectory of the same target is associated in the continuous frame millimeter wave radar data to obtain the target data of the blind spot area, including: The multi-target tracking (MOT) algorithm is used to associate the trajectory of the same target in continuous frame millimeter-wave radar data to obtain the target data in the blind spot area.
[0022] S3. Based on the target data in the blind spot area, the improved DQN algorithm is used to make real-time decisions to determine whether there is a dangerous situation and give the best response strategy.
[0023] In the technical solution of this application, if Figure 2 As shown in the figure, the improved DQN algorithm optimizes the reward function design by considering multi-dimensional factors and dynamically adjusting the reward coefficient to adapt to complex and changing traffic scenarios and make real-time decisions efficiently; The improved DQN algorithm introduces a priority experience replay mechanism during the training process, giving priority to important experience samples for learning to speed up the model convergence; at the same time, a gradient descent algorithm based on dynamic weight allocation and feedback adjustment is used to update the parameters of the evaluation network, and the parameters of the target network are regularly updated to reduce the training process. Q The system can reduce the value estimation deviation, make real-time decisions more efficiently, accurately judge dangerous situations and give the best response strategy.
[0024] 1) Improve the DQN algorithm by considering multi-dimensional factors and dynamically adjusting the reward coefficient to optimize the reward function design, including: S311. In the reward function, introduce the speed distance reward function and the environmental perturbation reward function : ; in, is the relative speed to the target, t For time, is the relative distance to the target, , They are the speed change weight coefficient and the distance change weight coefficient, which are used to adjust the influence of speed and distance change rate on the reward. When the target approaches quickly and the distance is too close, an additional negative reward is given; I is the interference intensity in the environment, u is the environmental interference reward coefficient. When the interference intensity is large and the target can still be accurately detected, positive reward compensation is given; S312. In the reward function, the reward items for successfully detecting blind spots, taking correct avoidance strategies, false positives, missed negatives, and excessive avoidance are introduced. The corresponding reward coefficients for each part are , , , , : ; in, , , , , are the initial reward values in the reward coefficients for successfully detecting blind spots, taking correct avoidance strategies, false positives, missed negatives, and excessive avoidance rewards, respectively, and , , , , , , , , , They are the growth coefficients of the reward coefficients for successfully detecting blind spots, adopting correct avoidance strategies, false positives, missed negatives, and excessive avoidance rewards. n is the number of training steps, N is the total number of training steps; S313. Reward function The final expression is: ; in, s is the current state, a For the current state s The following actions are taken, s’ To take action a The next state after S 1 , S 2 , S 3 , S 4 , S 5 They are the status indicators of successfully detecting blind spots, taking correct avoidance strategies, false positives, missed negatives, and excessive avoidance rewards, and their values are all 0 or 1.
[0025] 2) Improve the DQN algorithm by introducing a priority experience replay mechanism during training, giving priority to important experience samples for learning and accelerating the model convergence, including: S321, set i Step 1 experience sample e i for: ; in, s i For the iThe state of the step, a i For the status s i The following actions are taken, r i To take action a i After the reward, s i+1 For the i +1 step status, d i A sign to indicate whether it is finished; S322. Calculate experience samples e i TD error : ; in, The parameter is Evaluation network In Status s i Take action a i of Q Value estimation, The parameter is Target network In Status s i+1 Take action a’ of Q Value estimation, is the discount factor, and ; S323, according to TD error Prioritize all experience samples by size and calculate the experience samples e i Priority p i : ; in, is a small integer to avoid When the priority is zero; S324. Calculate experience samples e i The sampling probability P ( i ): ; in, i , j All represent the number of steps. is a parameter that controls the degree of influence of priority, and ,when is equivalent to random sampling when When all experience samples are sampled completely according to the priority, N The size of the experience replay buffer; S325. Based on the sampling probability of each experience sample , sample a batch of important experience samples to form an important experience sample set ,in n is the size of the important empirical sample set.
[0026] 3) Use the gradient descent algorithm based on dynamic weight allocation and feedback adjustment to update the parameters of the evaluation network, including: S331. Calculate important experience sample sets Important experience samples Target value : ; in, For the b k The state of the step, For the status The following actions are taken, To take action After the reward, For the b k +1 step status, It is a sign of whether it is finished. For the target network In Status Take action a’ of Q Value estimation; S332, Calculation Evaluation Network In Status Take action of Q Value Estimation , and calculate the loss function : ; S333, using the gradient descent algorithm based on dynamic weight allocation and feedback adjustment to update the evaluation network Parameters .
[0027] Specifically, in S333, a gradient descent algorithm based on dynamic weight allocation and feedback regulation is used to update the evaluation network. Parameters ,include: S3331, Initialization phase: Initialize learning rate , dynamic weight allocation coefficient , feedback adjustment coefficient , construct comprehensive sensitivity S , used to record important experience sample parameters The loss sensitivity of is set to 0 initially; S3332, Gradient calculation phase: Calculate important experience samples Target value and evaluation network of Q Value Estimation The square of the difference between Gradient : ; in, To evaluate the network of Q Value Estimation About parameters The gradient of Important experience sample set The above gradient calculation average of all important experience samples in the result is the loss function About parameters Gradient : ; S3333, Dynamic weight allocation stage: Calculation of important experience samples Parameters Sensitivity to loss , that is, for the parameter Make a small disturbance The loss function change obtained With this small disturbance The ratio between: ; Calculate important experience sample sets All important empirical samples in the parameters The overall sensitivity S : ; According to the comprehensive sensitivity S Dynamic weight allocation coefficient w To update: ; S3334, parameter update phase: coefficients are allocated according to dynamic weights w Adjusting the learning rate , and update the parameters ,Right now ; S3335, Feedback adjustment stage: After completing a round of parameter updates, the average loss of the validation set is L val Evaluation Network The performance of the validation set is evaluated. L val If the decrease from the previous round exceeds the preset threshold (0.1), the current weight allocation strategy is considered effective and the dynamic weight allocation coefficient is maintained. w unchanged; otherwise, the dynamic weight allocation coefficient w To perform feedback adjustment: For the overall sensitivity S If the performance is not improved after the parameter update, the dynamic weight allocation coefficient should be appropriately reduced. w ,Right now ; For the overall sensitivity S If the performance is not reduced after the parameter update, the dynamic weight allocation coefficient should be appropriately increased. w ,Right now .
[0028] 4) Regularly update the parameters of the target network, including: Determine whether the current step number is the target network update interval C If it is a multiple of , the network will be evaluated Parameters Assign to target network ,Right now .
[0029] The pseudo code of the above technical solution and the improved DQN algorithm is shown in Table 1: Table 1 Pseudo code of the improved DQN algorithm
[0030] The pseudo code of the gradient descent algorithm based on dynamic weight allocation and feedback adjustment is shown in Table 2: Table 2 Pseudocode of the gradient descent algorithm based on dynamic weight allocation and feedback adjustment
[0031] S4. Provide driving warnings or complete automatic avoidance actions based on the output results of the improved DQN algorithm.
[0032] In the technical solution of the present application, in terms of signal modulation, chaotic modulation is performed on the millimeter wave carrier signal using a chaotic sequence, and combined with adaptive modulation coefficients and channel impulse responses, the environmental adaptability and anti-interference ability of the signal are enhanced; In terms of data processing, Kalman filtering, DBSCAN algorithm and multi-target tracking MOT technology are used to greatly improve the reliability of data processing and target tracking accuracy; The decision algorithm adopts the improved DQN algorithm. By considering multi-dimensional factors and dynamically adjusting the reward coefficient, the reward function design is optimized to adapt to complex and changeable traffic scenarios and make real-time decisions efficiently. In the training process, the priority experience playback mechanism is introduced to give priority to important experience samples for learning to speed up the model convergence. The gradient descent algorithm based on dynamic weight allocation and feedback adjustment is used to update the parameters of the evaluation network, and the parameters of the target network are regularly updated to reduce the training process. Q The system can reduce the value estimation deviation, make real-time decisions more efficiently, accurately judge dangerous situations and give the best response strategy.
[0033] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. Such modifications or replacements will not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A millimeter wave radar blind spot detection method based on deep reinforcement learning, characterized in that: The following steps are involved: S1, detecting target information in the blind spot area of the vehicle by transmitting and receiving millimeter wave signals; S2, filtering, clustering and target tracking the millimeter wave radar data to obtain target data in the blind spot area; S3, based on the target data in the blind spot area, make real-time decisions by improving the DQN algorithm to determine whether there is a dangerous situation and give the best response strategy; S4, providing driving warnings or completing automatic evasive actions according to the output results of the improved DQN algorithm; Among them, the improved DQN algorithm optimizes the reward function design by considering multi-dimensional factors and dynamically adjusting the reward coefficient to adapt to complex and changing traffic scenarios and make real-time decisions efficiently; The improved DQN algorithm introduces a priority experience replay mechanism during the training process, giving priority to important experience samples for learning to speed up the model convergence; at the same time, a gradient descent algorithm based on dynamic weight allocation and feedback adjustment is used to update the parameters of the evaluation network, and the parameters of the target network are regularly updated to reduce the training process. Q The system can reduce the value estimation deviation, make real-time decisions more efficiently, accurately judge dangerous situations and give the best response strategy.
2. The millimeter wave radar blind spot detection method based on deep reinforcement learning according to claim 1, characterized in that: S1 detects target information in the vehicle's blind spot area by transmitting and receiving millimeter wave signals, including: S11. Using chaotic sequence x n For millimeter wave carrier signals Chaotic modulation is performed, and the millimeter wave signal after chaotic modulation s ( t )for: ; in, A c , f c are the amplitude and frequency of the millimeter wave carrier signal, t For time, k is the modulation coefficient, n Chaotic sequence x n The time index, N Chaotic sequence x n The total length of rect is a rectangular pulse function, is the sampling period; S12. In actual vehicle driving scenarios, in order to better adapt to different environmental factors on millimeter wave signals s ( t ) propagation, so that the modulation coefficient k Adaptively change with events or environments, denoted as k ( t ), while considering the multipath fading characteristics of the channel, the channel impulse response is introduced h ( t ) and millimeter wave signals s ( t ) is convolved to update the chaotically modulated millimeter wave signal s’ ( t )for: 。 3. The millimeter wave radar blind spot detection method based on deep reinforcement learning according to claim 1, characterized in that: In S2, the millimeter wave radar data is filtered, clustered and tracked to obtain the target data in the blind spot area, including: S21, performing Kalman filtering on the millimeter-wave radar data to remove noise interference in the millimeter-wave radar data; S22, clustering data points with similar characteristics according to the density connectivity between data points in order to distinguish different targets; S23. Associating the trajectory of the same target in the continuous frame millimeter wave radar data to obtain target data in the blind spot area.
4. The millimeter wave radar blind spot detection method based on deep reinforcement learning according to claim 3 is characterized in that: S21 performs Kalman filtering on the millimeter-wave radar data to remove noise interference in the millimeter-wave radar data, including: The state equation and observation equation of Kalman filter are: ; in, X k , X k-1 Separately for the moment k ,time k The state vector is -1, A is the state transfer matrix, U k For the moment k The control input vector, B is the input matrix, W k For the moment k The process noise vector of Z k For the moment k The observation vector, H is the observation matrix, V k For the moment k The observation noise vector of In S22, data points with similar characteristics are clustered according to the density connectivity between data points in order to distinguish different targets, including: According to the density connectivity between data points, the DBSCAN algorithm is used to cluster data points with similar characteristics in order to distinguish different targets; In S23, the trajectory of the same target is associated in the continuous frame millimeter wave radar data to obtain the target data of the blind spot area, including: The multi-target tracking (MOT) algorithm is used to associate the trajectory of the same target in continuous frame millimeter-wave radar data to obtain the target data in the blind spot area.
5. The millimeter wave radar blind spot detection method based on deep reinforcement learning according to claim 1, characterized in that: The improved DQN algorithm optimizes the reward function design by considering multi-dimensional factors and dynamically adjusting the reward coefficient, including: S311. In the reward function, introduce the speed distance reward function and the environmental perturbation reward function : ; in, is the relative speed to the target, t For time, is the relative distance to the target, , They are the speed change weight coefficient and the distance change weight coefficient, which are used to adjust the influence of speed and distance change rate on the reward. When the target approaches quickly and the distance is too close, an additional negative reward is given; I is the interference intensity in the environment, u is the environmental interference reward coefficient. When the interference intensity is large and the target can still be accurately detected, positive reward compensation is given; S312. In the reward function, the reward items for successfully detecting blind spots, taking correct avoidance strategies, false positives, missed negatives, and excessive avoidance are introduced. The corresponding reward coefficients for each part are , , , , : ; in, , , , , are the initial reward values in the reward coefficients for successfully detecting blind spots, taking correct avoidance strategies, false positives, missed negatives, and excessive avoidance rewards, respectively, and , , , , , , , , , They are the growth coefficients of the reward coefficients for successfully detecting blind spots, adopting correct avoidance strategies, false positives, missed negatives, and excessive avoidance rewards. n is the number of training steps, N is the total number of training steps; S313. Reward function The final expression is: ; in, s is the current state, a For the current state s The following actions are taken, s’ To take action a The next state after S 1. S 2. S 3. S 4. S 5 are status indicators of successful blind spot detection, correct avoidance strategy, false alarm, missed alarm, and excessive avoidance reward items, and their values are 0 or 1.
6. The millimeter wave radar blind spot detection method based on deep reinforcement learning according to claim 1, characterized in that: The improved DQN algorithm introduces a priority experience replay mechanism during the training process, prioritizes important experience samples for learning, and accelerates the model convergence speed, including: S321, set i Step 1 experience sample e i for: ; in, s i For the i The state of the step, a i For the status s i The following actions are taken, r i To take action a i After the reward, s i+1 For the i +1 step status, d i A sign to indicate whether it is finished; S322. Calculate experience samples e i TD error : ; in, The parameter is Evaluation network In Status s i Take action a i of Q Value estimation, The parameter is Target network In Status s i+1 Take action a’ of Q Value estimation, is the discount factor, and ; S323, according to TD error Prioritize all experience samples by size and calculate the experience samples e i Priority p i : ; in, is a small integer to avoid When the priority is zero; S324. Calculate experience samples e i The sampling probability P ( i ): ; in, i , j All represent the number of steps. is a parameter that controls the degree of influence of priority, and ,when is equivalent to random sampling when When all experience samples are sampled completely according to the priority, N The size of the experience replay buffer; S325. Based on the sampling probability of each experience sample , sample a batch of important experience samples to form an important experience sample set ,in n is the size of the important empirical sample set.
7. The millimeter wave radar blind spot detection method based on deep reinforcement learning according to claim 6 is characterized in that: The method of updating the parameters of the evaluation network using a gradient descent algorithm based on dynamic weight allocation and feedback regulation includes: S331. Calculate important experience sample sets Important experience samples Target value : ; in, For the b k The state of the step, For the status The following actions are taken, To take action After the reward, For the b k +1 step status, It is a sign of whether it is finished. For the target network In Status Take action a’ of Q Value estimation; S332, Calculation Evaluation Network In Status Take action of Q Value Estimation , and calculate the loss function : ; S333, using the gradient descent algorithm based on dynamic weight allocation and feedback adjustment to update the evaluation network Parameters .
8. The millimeter wave radar blind spot detection method based on deep reinforcement learning according to claim 7, characterized in that: S333 uses a gradient descent algorithm based on dynamic weight allocation and feedback regulation to update the evaluation network Parameters ,include: S3331, Initialization phase: Initialize learning rate , dynamic weight allocation coefficient , feedback adjustment coefficient , construct comprehensive sensitivity S , used to record important experience sample parameters The loss sensitivity of is set to 0 initially; S3332, Gradient calculation phase: Calculate important experience samples Target value and evaluation network of Q Value Estimation The square of the difference between Gradient : ; in, To evaluate the network of Q Value Estimation About parameters The gradient of Important experience sample set The above gradient calculation average of all important experience samples in the result is the loss function About parameters Gradient : ; S3333, Dynamic weight allocation stage: Calculation of important experience samples Parameters Loss sensitivity , that is, for the parameter Make a small disturbance The loss function change obtained With this small disturbance The ratio between: ; Calculate important experience sample sets All important empirical samples in the parameters The overall sensitivity S : ; According to the comprehensive sensitivity S Dynamic weight allocation coefficient w To update: ; S3334, parameter update phase: coefficients are allocated according to dynamic weights w Adjusting the learning rate , and update the parameters ,Right now ; S3335, Feedback adjustment stage: After completing a round of parameter updates, the average loss of the validation set is L val Evaluation Network The performance of the validation set is evaluated. L val If the decrease from the previous round exceeds the preset threshold, the current weight allocation strategy is considered effective and the dynamic weight allocation coefficient is maintained. w unchanged; otherwise, the dynamic weight allocation coefficient w To perform feedback adjustment: For the overall sensitivity S If the performance is not improved after the parameter update, the dynamic weight allocation coefficient should be appropriately reduced. w ,Right now ; For the overall sensitivity S If the performance is not reduced after the parameter update, the dynamic weight allocation coefficient should be appropriately increased. w ,Right now .
9. The millimeter wave radar blind spot detection method based on deep reinforcement learning according to claim 8, characterized in that: The periodic updating of the parameters of the target network includes: Determine whether the current step number is the target network update interval C If it is a multiple of , the network will be evaluated Parameters Assign to target network ,Right now .
Citation Information
Patent Citations
Interference detection shared signal design method based on deep reinforcement learning
CN116738842A
Decision planning method for automatic driving vehicle in urban traffic scene based on reinforcement learning
CN118396034A
Backward Anti-collision driving decision-making method for heavy commercial vehicle
US20230182725A1
Radar system utilizing chaotic coding
US5321409A