An Evolvable Intelligent Single-Mode Airborne Radar Target Tracking Method

By establishing a radar rule base and optimizing radar strategies using a deep reinforcement learning network, the problem of insufficient target tracking accuracy of traditional radar in complex environments is solved, and high-precision target tracking is achieved in various environments.

CN116609754BActive Publication Date: 2026-04-07UNIKINFO TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-04
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Traditional radars struggle to adjust their operating parameters in real time under complex environments, resulting in insufficient accuracy in target detection and tracking. They also lack self-learning and environmental adaptability, failing to meet the demands for accurate, real-time, and robust target tracking in complex environments.

Method used

A radar rule base is established, and combined with a deep reinforcement learning network, target motion information is obtained through noncoherent accumulation, CFAR detection, pulse compression, and Bayesian clutter suppression. The radar strategy is then optimized through the deep reinforcement learning network to form a self-evolving closed-loop system.

Benefits of technology

It achieves high-precision target tracking of radar in complex environments, improves radar's adaptability and target detection accuracy, and meets the real-time tracking requirements in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116609754B_ABST
    Figure CN116609754B_ABST
Patent Text Reader

Abstract

This invention discloses an evolvable intelligent single-mode airborne radar target tracking method, relating to the technical field of intelligent radar. The method includes: establishing a rule base containing radar parameters; selecting the optimal radar tracking strategy from the rule base to illuminate and track the target based on different target actions; preprocessing the radar echo; performing Bayesian clutter suppression on the preprocessed radar echo to obtain basic motion information of the target; adaptively tracking the target; learning and evaluating the tracking data; constructing a deep reinforcement learning network to determine the radar's operating strategy for the next moment; extracting the tracking data as a training set to train the deep reinforcement learning network; using the deep reinforcement learning network to select the target tracking strategy with the maximum reward value from the rule base, forming a closed loop, thereby improving the radar's tracking performance and enabling the deep reinforcement learning network to self-evolve.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of intelligent radar, in particular to an evolvable intelligent single-mode airborne radar target tracking method. BACKGROUND

[0002] The waveform parameters of the signals transmitted and received by the conventional radar are relatively fixed, and the working parameters of the radar cannot be adjusted in real time according to the changes of the target and the environment, so the accuracy and real-time performance of the detection and tracking of the target cannot be guaranteed, and the target cannot be classified and identified according to the characteristics and motion model of the target, the information of the target and the environment cannot be continuously perceived, learned and predicted, and the radar does not have the self-learning and environmental adaptive ability, and cannot meet the requirements of accurate, real-time and robust target tracking and target detection in complex environments.

[0003] The conventional radar engineering applies typical or empirical models for target identification, environment clutter and interference suppression. These commonly used classical models include: Swerling model of target flicker, clutter distribution model, channel attenuation model, etc. Since these models usually have specific use ranges, such as only for sea surface, only for land, etc., their universality is poor in complex environments, such as the sea-land junction.

[0004] Taking a typical clutter suppression model as an example, it is usually assumed that the clutter distribution satisfies certain stationary characteristics, such as plain areas, desert areas, low sea state areas, etc. However, the area detected by the airborne radar is usually complex, and it is difficult to satisfy the stationary characteristics, such as large terrain undulation (mountainous environment), large change in ground cover (different vegetation, water body and road network), existence of sea-land junction, many high-rise buildings, high-voltage towers, etc. These non-uniform clutters will destroy the effectiveness of the clutter suppression model. In addition, the working parameters of the conventional radar, such as signal waveform, bandwidth, antenna polarization, detection threshold, accumulation time, etc., are generally one or more preset values that can be selected, and cannot be dynamically adjusted according to the application environment, thereby severely limiting the environmental adaptability of the radar system.

[0005] Intelligent radar adjusts its operation and processing scheme according to the changes of the environment and the target, so as to realize better target detection performance than the conventional radar. Such radar can learn and develop from its experience, and is an important research direction of the next generation of target detection. SUMMARY

[0006] The purpose of the present application is to provide an evolvable intelligent single-mode airborne radar target tracking method that overcomes the complex external environment and has higher tracking performance.

[0007] The technical solution of the present application is to provide an evolvable intelligent single-mode airborne radar target tracking method, which comprises:

[0008] S1. Establish a rule base containing radar waveform type, waveform parameters, beam direction and scanning mode. Based on the different actions of the tracked target, select the optimal radar tracking strategy from the rule base to illuminate and track the target.

[0009] S2. Preprocess the radar echo by sequentially performing noncoherent accumulation, CFAR detection, and pulse compression to reduce environmental interference.

[0010] S3. Perform Bayesian clutter suppression on the preprocessed radar echo to obtain the basic motion information of the tracked target.

[0011] S4. Adaptive tracking of the target is achieved through a series of processes including initial track extraction, velocity extraction, track association, track filtering and prediction.

[0012] S5. Learn and evaluate the tracking data to obtain parameters for evaluating the tracking effect;

[0013] S6. Construct a deep reinforcement learning network as the radar decision network to determine the radar's operating strategy at the next moment; extract tracking data as a training set to train the deep reinforcement learning network; extract the adaptive tracking results from step S4 and the tracking learning evaluation results from step S5 to update the deep reinforcement learning network; use the deep reinforcement learning network to select the tracking target strategy with the maximum reward value from the rule base in step S1, forming a closed loop to improve the radar's tracking performance, and the deep reinforcement learning network achieves self-evolution.

[0014] In any of the above technical solutions, step S1 further includes:

[0015] An execution memory is established to store the parameters of each selected radar tracking strategy.

[0016] In any of the above technical solutions, the Bayesian clutter suppression in step S3 is further aided by prior information in a preset environmental knowledge base, including a digital map and a clutter map.

[0017] In any of the above technical solutions, further, step S3 performs Bayesian clutter suppression on the data contained in the radar echo, the steps of which include:

[0018] Find under the corresponding constraints Minimum value:

[0019]

[0020] Among them, w k Let R be the optimal filtering weight vector based on prior information at the i-th Doppler cell in the distance cell l to be calculated. kThe clutter covariance matrix based on prior information, s(f i ) is the time-domain steering vector of the i-th Doppler cell to be detected.

[0021] In any of the above technical solutions, step S4 further includes:

[0022] Establish a target knowledge base and write the velocity, direction and position information of the tracked target obtained from each adaptive tracking into the target knowledge base;

[0023] The second and subsequent adaptive tracking will read information recorded in the target knowledge base as knowledge assistance.

[0024] In any of the above technical solutions, the parameter for evaluating the tracking effect in step S5 is further defined as the F-score. The larger the F-score, the better the tracking effect. The evaluation steps include:

[0025] Define accuracy:

[0026]

[0027] Define recall:

[0028]

[0029] F-score:

[0030]

[0031] Among them, G t This indicates the actual location of the target. If the target disappears, then... A t (τ θ ) represents the predicted target location, τ θ The threshold value for a target is represented by θ. t This represents the prediction certainty score at time t. If the score at time t is less than a threshold, i.e., θ... t <τ θ ,So Ω(A t (τ θ ),G t ) represents the intersection of the predicted location and the actual target location; τ Ω The threshold representing precision; N p N represents the sum of the number of frames when the prediction set is not empty. g G represents t The sum of the number of frames that are not empty sets.

[0032] In any of the above technical solutions, further, the tracking data of the deep reinforcement learning network as the training set in step S6 includes data from the execution memory and the target knowledge base.

[0033] In any of the above technical solutions, the loss function for deep reinforcement learning is further defined as follows:

[0034]

[0035] Where, r t It is the reward of the present moment. The weights, s, are updated iteratively by the deep reinforcement network. t It is the target state at the current moment, a t γ is the target behavior; D is the set of all radar decisions, A is the set of all target behaviors, and γ is the discount factor. The previous tracking strategy of the radar, This represents the radar's current tracking strategy.

[0036] The beneficial effects of this invention are:

[0037] The technical solution of this invention establishes a radar rule base that includes radar waveform types and parameters, beams, and scanning methods. Appropriate parameters are extracted from this rule base to change the radar's illumination state, overcoming the shortcomings of traditional radars that cannot effectively cope with complex external environments, resulting in poor target tracking performance. A radar decision-making method based on deep reinforcement learning is constructed, enabling the self-evolution of intelligent radar. Based on prior information such as digital maps and clutter maps, the clutter suppression capability of the Bayesian space-time processing method is improved. Knowledge-assisted Bayesian clutter suppression can be applied to radar echoes to obtain target motion information, overcoming the interference of non-uniform clutter on radar when measuring complex areas with diverse terrains, enabling the radar to achieve high accuracy in various complex environments. Attached Figure Description

[0038] The advantages of the above and additional aspects of the present invention will become apparent and readily understood in the description of the embodiments in conjunction with the following drawings, wherein:

[0039] Figure 1 This is a flowchart of an evolutionary intelligent single-mode airborne radar target tracking method according to an embodiment of the present invention. Detailed Implementation

[0040] To better understand the above-mentioned objectives, features, and advantages of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that, unless otherwise specified, the embodiments of the present invention and the features thereof can be combined with each other.

[0041] In the following description, many specific details are set forth in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and therefore the scope of protection of the invention is not limited to the specific embodiments disclosed below.

[0042] like Figure 1 As shown in the figure, this embodiment provides an evolvable intelligent single-mode airborne radar target tracking method, which includes:

[0043] S1. Establish a rule base containing radar waveform type, waveform parameters, beam direction and scanning mode. Based on the different actions of the tracked target, select the optimal radar tracking strategy from the rule base to illuminate and track the target. Establish an execution memory library to store the parameters of the selected radar tracking strategy each time.

[0044] S2. Preprocess the radar echo by sequentially performing noncoherent accumulation, CFAR detection, and pulse compression to reduce environmental interference.

[0045] Among them, noncoherent accumulation, CFAR detection, and pulse compression are signal processing methods frequently used by radar. Noncoherent accumulation does not need to consider the phase relationship between sampling points and uses the average of multiple samples to improve the signal-to-noise ratio. CFAR detection determines whether a signal exists or not in the presence of noise, and thus determines whether a target is present or absent. Pulse compression shortens the time domain width of the signal by increasing the signal bandwidth, thereby improving the time resolution of the signal.

[0046] S3. Based on prior information such as digital maps and clutter maps in the preset environmental knowledge base, knowledge-assisted Bayesian clutter suppression is applied to the preprocessed radar echo to obtain the basic motion information of the tracked target. The knowledge-assisted Bayesian clutter suppression method is expressed as follows:

[0047]

[0048] Among them, w k Let be the optimal filtering weight vector based on prior information at the i-th Doppler cell in the distance cell l to be calculated. It is w k The conjugate transpose of R k The clutter covariance matrix based on prior information, s(f i ) is the time-domain steering vector of the i-th Doppler cell to be detected, and st refers to the constraint condition.

[0049] The optimal filtering weight vector is the weight vector that minimizes the mean square error between the estimated signal and the true signal output by the receiver under the minimum mean square error criterion. It is also called the optimal solution. By using the optimal filtering weight vector, clutter signals and target signals can be distinguished, thereby achieving target detection and target tracking. The clutter covariance matrix is ​​a matrix used to describe the characteristics of clutter signals received by radar. Each element of this matrix represents the cross-correlation of clutter signals received by different antennas or receiving channels. By performing cross-correlation calculations on the signals, the values ​​of each element of the clutter covariance matrix are obtained, which allows for further processing and analysis of the received signals.

[0050] Specifically, existing radar target tracking methods often employ clutter suppression models that assume certain stationary characteristics in clutter distribution, limiting their accuracy to relatively uniform environments such as plains, deserts, and coastal areas. These models cannot adequately handle non-uniform clutter detected in complex environments, such as land-sea junctions, mountainous regions, and urban areas. The radar target tracking method provided by this invention utilizes prior information from an environmental knowledge base to support a clutter suppression model employing Bayesian spatiotemporal processing. This allows for dynamic adjustment of radar operating parameters based on specific target and environmental conditions, thereby accurately detecting and tracking targets and meeting the requirements for accurate, real-time, and robust target tracking and detection in complex environments.

[0051] S4. Adaptive tracking of the target is performed through a series of processes, including initial track extraction, velocity extraction, track association, track filtering and prediction. A target knowledge base is established, and the velocity, direction and position information of the target obtained from each adaptive tracking is written into the target knowledge base. The information recorded in the target knowledge base will be read as knowledge assistance during the second and subsequent adaptive tracking.

[0052] S5. The tracking data is evaluated using a learning model. The parameter for evaluating the tracking effect is defined as the F-score. A higher F-score indicates better tracking performance. The evaluation steps include:

[0053] Define accuracy:

[0054]

[0055] Define recall:

[0056]

[0057] F-score:

[0058]

[0059] Among them, G t This indicates the actual location of the target. If the target disappears, then... At (τ θ ) represents the predicted target location, τ θ This represents the threshold for being judged as a target, if θ is used. t Let θ represent the prediction certainty score at time t. If the score at time t is less than a threshold, i.e., θ t <τ θ ,So Ω(A t (τ θ ),G t ) represents the intersection of the predicted location and the actual target location; τ Ω The threshold representing precision; N p Let N be the sum of the number of frames when the prediction set is not empty. g For G t The sum of the number of frames that are not empty sets.

[0060] S6. Construct a deep reinforcement learning network as the radar decision network to determine the radar's operating strategy for the next moment; extract execution memory and target knowledge base data as training sets to train the deep reinforcement learning network; extract the adaptive tracking results from step S4 and the tracking learning evaluation results from step S5 to update the deep reinforcement learning network.

[0061] In this embodiment, a Deep Q-Networks (DQN) algorithm is used to determine and optimize radar decisions. The radar-acquired data and radar system state are input into the Q-Network, and the Q-Network outputs the fitted Q-value corresponding to each possible decision action. The Q-value (or action-value function) represents the expected cumulative reward that the agent can obtain under a given state and selected action.

[0062] The Deep Q-Network algorithm continuously optimizes the parameters of the Q-Network through interaction with the environment. At each time step, the system selects an action based on the current state and the Q-Network's estimate. At the same time, the system observes the reward signal and the next state fed back by the environment and stores this information as experience samples.

[0063] The Deep Q-Network (DQN) algorithm randomly selects a batch of data from empirical samples and uses it as training data to train the Q-network. During training, DQN gradually adjusts the output of the Q-network to be closer to the true Q-value by minimizing the mean squared error loss function. In this way, the Q-network gradually learns the cumulative reward relationship between states and actions, thereby guiding the decision-making at the next time step. Through continuous interaction, learning, and optimization, the DQN algorithm enables the radar decision network to gradually improve its predictive capabilities and decision-making strategies.

[0064] The loss function for deep reinforcement learning is:

[0065]

[0066] Where, r t It is the reward of the present moment. The weights, s, are updated iteratively by the deep reinforcement network. t It is the target state at the current moment, a t Let A be the target behavior, D be the set of all radar decisions, γ be the discount factor and γ∈[0,1), and A be the set of all target behaviors. This indicates that the target state is s. t The target behavior is a t At that time, the radar's previous tracking strategy, This indicates the radar's current tracking strategy.

[0067] By using a deep reinforcement learning network to select a strategy with the highest reward value to track the target, a continuously evolving and updating closed loop is formed between radar and target, constantly updating the deep reinforcement learning network according to changes in the target's state to achieve self-evolution.

[0068] In summary, this invention proposes an evolvable intelligent single-mode airborne radar target tracking method, which includes:

[0069] S1. Establish a rule base containing radar waveform type, waveform parameters, beam direction and scanning mode. Based on the different actions of the tracked target, select the optimal radar tracking strategy from the rule base to illuminate and track the target.

[0070] S2. Preprocess the radar echo by sequentially performing noncoherent accumulation, CFAR detection, and pulse compression to reduce environmental interference.

[0071] S3. Perform Bayesian clutter suppression on the preprocessed radar echo to obtain the basic motion information of the tracked target.

[0072] S4. Adaptive tracking of the target is performed through a series of processes, including initial track extraction, velocity extraction, track association, track filtering and prediction.

[0073] S5. Learn and evaluate the tracking data to obtain parameters for evaluating the tracking effect;

[0074] S6. Construct a deep reinforcement learning network as the radar decision network to determine the radar's operating strategy at the next moment; extract tracking data as a training set to train the deep reinforcement learning network; extract the adaptive tracking results from step S4 and the tracking learning evaluation results from step S5 to update the deep reinforcement learning network; use the deep reinforcement learning network to select the tracking target strategy with the maximum reward value from the rule base in step S1, forming a closed loop to improve the radar's tracking performance and also enable the deep reinforcement learning network to achieve self-evolution.

[0075] The steps in this invention can be adjusted, combined, or deleted according to actual needs.

[0076] Although the invention has been disclosed in detail with reference to the accompanying drawings, it should be understood that these descriptions are merely exemplary and not intended to limit the application of the invention. The scope of protection of the invention is defined by the appended claims and may include various modifications, alterations, and equivalents made to the invention without departing from the scope and spirit of the invention.

Claims

1. An evolutionary intelligent single-mode airborne radar target tracking method, characterized in that, The method includes: S1. Establish a rule base containing radar waveform type, waveform parameters, beam direction and scanning mode. Based on the different actions of the tracked target, select the optimal radar tracking strategy from the rule base to illuminate and track the target. S2. Preprocess the radar echo by sequentially performing noncoherent accumulation, CFAR detection, and pulse compression to reduce environmental interference. S3. Perform Bayesian clutter suppression on the preprocessed radar echo to obtain the basic motion information of the tracked target. S4. Adaptive tracking of the target is achieved through a series of processes including initial track extraction, velocity extraction, track association, track filtering and prediction. S5. Learn and evaluate the tracking data to obtain parameters for evaluating the tracking effect; The parameter for evaluating the tracking effect in step S5 is defined as the F-score. A higher F-score indicates a better tracking effect. The evaluation steps include: Define accuracy: Define recall: F-score: Among them, G t This indicates the actual location of the target. If the target disappears, then... A t (τ θ ) represents the predicted target location, τ θ The threshold value for a target is represented by θ. t This represents the prediction certainty score at time t. If the score at time t is less than a threshold, i.e., θ... t <τ θ ,So Ω(A t (τ θ ),G t ) represents the intersection of the predicted location and the actual target location; τ Ω The threshold representing precision; N p N represents the sum of the number of frames when the prediction set is not empty. g G represents t The sum of the number of frames that are not empty sets; S6. Construct a deep reinforcement learning network as the radar decision network to determine the radar's operating strategy at the next moment; extract tracking data as a training set to train the deep reinforcement learning network; extract the adaptive tracking results from step S4 and the tracking learning evaluation results from step S5 to update the deep reinforcement learning network; use the deep reinforcement learning network to select the tracking target strategy with the maximum reward value from the rule base in step S1, forming a closed loop to improve the radar's tracking performance, and the deep reinforcement learning network achieves self-evolution.

2. The evolving intelligent single-mode airborne radar target tracking method as described in claim 1, characterized in that, Step S1 further includes: An execution memory is established, and the parameters of each selected radar tracking strategy are stored in the execution memory.

3. The evolving intelligent single-mode airborne radar target tracking method as described in claim 1, characterized in that, The Bayesian clutter suppression in step S3 is aided by prior information in a preset environmental knowledge base, which includes a digital map and a clutter map.

4. The evolvable intelligent single-mode airborne radar target tracking method as described in claim 3, characterized in that, Step S3 involves Bayesian clutter suppression of the data contained in the radar echo, and includes the following steps: Find under the corresponding constraints Minimum value: Among them, w k Let R be the optimal filtering weight vector based on prior information at the i-th Doppler cell in the distance cell l to be calculated. k The clutter covariance matrix based on prior information, s(f i ) is the time-domain steering vector of the i-th Doppler cell to be detected.

5. The evolvable intelligent single-mode airborne radar target tracking method as described in claim 2, characterized in that, Step S4 further includes: Establish a target knowledge base, and write the speed, direction and position information of the tracked target obtained from each adaptive tracking into the target knowledge base; The second and subsequent adaptive tracking will read information recorded in the target knowledge base as knowledge aids.

6. The evolvable intelligent single-mode airborne radar target tracking method as described in claim 5, characterized in that, In step S6, the tracking data used as the training set for the deep reinforcement learning network includes data from the execution memory and the target knowledge base.

7. The evolving intelligent single-mode airborne radar target tracking method as described in claim 1, characterized in that, The loss function for the deep reinforcement learning is: Where, r t It is the reward of the present moment. The weights, s, are updated iteratively by the deep reinforcement network. t It is the target state at the current moment, a t γ is the target behavior; D is the set of all radar decisions, A is the set of all target behaviors, and γ is the discount factor. The previous tracking strategy of the radar, This represents the radar's current tracking strategy.

Citation Information

Patent Citations

  • Waveform optimization method suitable for RSN in target tracking task

    CN113238219A