A complex mine radar adaptive anti-interference detection method based on reinforcement learning
By combining reinforcement learning and multi-domain feature fusion models with incomplete information game adversarial reinforcement learning algorithms, and dynamically adjusting radar strategies, the problem of detection stability and accuracy under multi-source interference in mine tunneling faces is solved, achieving highly reliable and high-precision target detection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- PEKING UNIV
- Filing Date
- 2026-01-08
- Publication Date
- 2026-05-01
AI Technical Summary
Traditional detection and sensing technologies are easily interfered with in the high dust, high humidity and multipath reflection environment of mine tunneling faces, resulting in increased false alarm rate and poor detection stability, making it difficult to meet the needs of accurate safety monitoring. Existing radar technology has failed to effectively solve the adaptive anti-interference strategy under multi-source interference.
A reinforcement learning-based approach is adopted, which uses multi-domain feature extraction and incomplete information game adversarial reinforcement learning algorithm to construct an adaptive anti-jamming strategy for radar. This strategy includes power spectral density, Doppler time, range-time and range-Doppler analysis, and dynamically adjusts radar waveform parameters and filtering thresholds to suppress multi-source interference.
Improving the reliability and accuracy of radar detection in complex mining environments, significantly enhancing target identification capabilities and system stability, and adapting to multi-source interference from changing environments.
Smart Images

Figure CN121477158B_ABST
Abstract
Description
An Adaptive Anti-jamming Detection Method for Complex Mine Radar Based on Reinforcement Learning Technical Field
[0001] This application relates to the field of mine safety monitoring and intelligent sensing technology, and in particular to a complex mine radar adaptive anti-interference detection method based on reinforcement learning. Background Technology
[0002] Due to the combined effects of tunneling machines cutting through fractured rock, high-pressure spray dust suppression, and confined spaces, mine tunneling faces are consistently exposed to extreme environments of high dust concentration and high humidity. Simultaneously, dense metal support structures and large electromechanical equipment result in severe multipath reflections, and the operation of electrical equipment is accompanied by strong electromagnetic radiation. In such an environment, traditional detection and sensing technologies face significant challenges: laser scanning and lidar are easily obscured by dust scattering, leading to signal attenuation; induced polarization (IP) methods and sonar technologies suffer from false echoes and multipath effects caused by metal supports, resulting in a significant decrease in signal-to-noise ratio. The combined effect of these interferences leads to an increased false alarm rate and poor detection stability at the tunneling end, making it difficult to meet the demands of accurate safety monitoring.
[0003] Currently, in the field of underground roadway detection, existing technologies attempt to use millimeter-wave radar to enhance its ability to penetrate dust and fog. For example, some existing technologies state that "millimeter-wave radar with a speed of 30-300 GHz has a strong ability to penetrate fog, smoke, and dust." However, this technology is limited to static identification or positioning assistance and has not yet systematically solved the adaptive anti-interference strategy of radar in dynamic mine interference environments (dust scattering, multipath reflection, and electromagnetic interference superposition).
[0004] Furthermore, some existing technologies disclose methods for constructing 3D models of mine roadways based on radar point cloud data. While these technologies involve the application of radar in roadways, they focus on spatial structure modeling rather than multi-interference suppression mechanisms in radar signal processing. Therefore, existing technologies still have significant gaps in achieving simultaneous suppression of multi-source interference (dust scattering, multipath reflection, and electromagnetic noise), adaptive selection of optimal waveforms and signal processing strategies, and maintaining long-term stable operation in a mining environment.
[0005] In the construction of intelligent mines, remotely controlled unmanned tunneling is a major requirement, and one of the key foundations of this requirement is the real-time acquisition of high-precision spatiotemporal information of the tunneling head. Therefore, to improve the reliability and accuracy of radar detection in mine roadways, there is an urgent need for a new type of radar anti-interference detection method that can adaptively adjust detection strategies based on environmental changes and suppress multi-source interference (dust, multipath, electromagnetic interference) in real time. This will enable highly reliable and accurate target detection and safety monitoring in complex mine environments. Summary of the Invention
[0006] In view of the above problems, this application proposes an adaptive anti-interference detection method for complex mine radar based on reinforcement learning, which overcomes the shortcomings of the prior art.
[0007] This application provides an adaptive anti-jamming detection method for complex mine radar based on reinforcement learning, including:
[0008] The echo signal is received, and multi-domain features are extracted based on the echo signal;
[0009] Based on the multi-domain features, a reinforcement learning state space is constructed, and a composite reward function based on detection probability, false alarm probability, and signal-to-interference ratio is designed. This composite reward function is used to describe the balance between the radar system's target detection accuracy, false alarm suppression, and interference suppression performance.
[0010] Based on the reinforcement learning state space and the composite reward function, an incomplete information game adversarial reinforcement learning algorithm is introduced to iteratively train the radar policy network. By alternately optimizing the policy parameters of the agent and the opponent, an initial anti-jamming strategy for the radar system under multi-source interference conditions is generated. In the incomplete information game adversarial reinforcement learning algorithm, the radar policy network is used as the agent and the interference source is used as the opponent.
[0011] The initial anti-jamming strategy is applied to the radar system. During the actual detection process of the radar system, the action of the initial anti-jamming strategy is selected and executed according to the current environmental state. The action includes: waveform parameter configuration, filter threshold adjustment and signal processing strategy switching.
[0012] When the performance of the radar system degrades or the interference characteristics change abruptly, the parameters of the radar strategy network are monitored and fine-tuned online based on real-time feedback data, which is pre-processed data obtained from the real-time environment.
[0013] The performance of the radar system after online monitoring and fine-tuning is evaluated based on the aforementioned composite reward function;
[0014] If the performance evaluation result does not meet the preset performance threshold, the weights of the composite reward function and the parameters of the radar policy network are adjusted, and the following steps are executed: an incomplete information game adversarial reinforcement learning algorithm is introduced to iteratively train the adjusted radar policy network. By alternately optimizing the policy parameters of the agent and the antibody, a new anti-jamming strategy for the radar system under multi-source interference conditions and subsequent steps are generated until the optimal anti-jamming strategy is obtained. The performance evaluation result corresponding to the optimal anti-jamming strategy meets the preset performance threshold.
[0015] Optionally, the echo signal is received, and multi-domain features are extracted based on the echo signal, including:
[0016] Using millimeter-wave radar to transmit linear frequency modulated continuous wave signals;
[0017] Millimeter-wave radar is used to receive echo signals that include target echoes, dust scattering, multipath reflections, and electromagnetic interference.
[0018] The echo signal is subjected to fast Fourier transform and time-frequency analysis to extract the multi-domain features, which include: power spectral density domain features, Doppler time domain features, distance-time domain features, and distance-Doppler domain features.
[0019] Optionally, the extracted multi-domain features include:
[0020] The expression for obtaining the linear frequency modulated continuous wave signal transmitted by the radar is as follows:
[0021]
[0022] In the above formula, A represents the amplitude of the transmitted signal; Indicates the starting frequency of the signal; The frequency modulation bandwidth of the signal is represented by T; the frequency modulation period is represented by t; and the time variable is represented by t.
[0023] The echo signal is obtained, and its expression is:
[0024]
[0025] In the above formula, This refers to the echo signal; This represents the reflection coefficient of the i-th target, reflecting the degree of attenuation of the echo signal; The propagation delay of the signal reaching the target is represented by the following formula: ,in Let represent the distance between the i-th target and the radar, and c represent the speed of electromagnetic wave propagation; Represents the amplitude coefficient of the j-th type of interference component; Indicates the propagation delay of the interference signal; This represents system noise, assumed to be Gaussian white noise;
[0026] Performing Fast Fourier Transform and time-frequency analysis on the echo signal yields power spectral density domain characteristics, Doppler time domain characteristics, range-time domain characteristics, and range-Doppler domain characteristics, expressed as follows:
[0027]
[0028] In the above formula, This represents the power spectral density domain characteristics; This represents the Doppler time-domain feature; This represents the distance-time domain feature; This represents the distance Doppler domain feature.
[0029] Optionally, a reinforcement learning state space is constructed based on the multi-domain features, including:
[0030] The multi-domain features are used as state vectors in the reinforcement learning state space, and their expression is:
[0031]
[0032] In the above formula, This represents the environmental state perceived by the agent at time t; The power spectral density domain characteristics at time t are represented; Represents the Doppler time-domain characteristics at time t; Represents the distance-time domain characteristics at time t; This represents the distance-Doppler domain feature at time t.
[0033] Optionally, the expression for the designed composite reward function is:
[0034]
[0035] In the above formula, This represents the agent's immediate reward at time t; It represents the target detection probability, reflecting the radar system's ability to correctly detect targets; It represents the false alarm probability, reflecting the probability that the radar system misjudges noise or interference as a target; Indicates the signal-to-interference ratio; , , These represent the weighting coefficients of each indicator in the reward function, used to balance the contributions of detection performance, false alarm control, and interference suppression.
[0036] Optionally, an adversarial reinforcement learning algorithm based on incomplete information game theory is introduced to iteratively train the radar policy network, including:
[0037] An anti-antibody is introduced during training to simulate different types and evolutionary forms of interference sources, forming an adversarial interaction with the agent. The goal of the anti-antibody is to minimize the detection performance of the radar system by continuously changing the interference strategy to approximate the most unfavorable interference situation in the real environment. The goal of the agent is to maximize the cumulative reward by optimizing the anti-interference strategy to maintain the preset detection capability and robustness under adversarial conditions.
[0038] The radar policy network is trained using the incomplete information game adversarial learning algorithm, and the optimal anti-jamming policy is searched using the policy gradient method.
[0039] Optionally, the expression for the agent's goal is:
[0040]
[0041] In the above formula, This represents the optimal anti-interference strategy; This represents the radar strategy network; T represents the total number of interaction steps; This represents the discount factor, used to measure the importance of future rewards;
[0042] The agent updates the optimal anti-interference strategy using a near-end policy optimization framework, the expression of which is:
[0043]
[0044] In the above formula, Represents the loss function; The parameters represent the radar strategy network; The policy probability ratio is represented by:
[0045]
[0046] In the above formula, Indicates the action to be taken under the current anti-interference strategy. The probability of; This indicates the action taken under the old anti-interference strategy. The probability of; The advantage function measures the advantage of the current action relative to the average policy. This represents the cutoff coefficient, used to prevent the anti-interference strategy from being updated beyond a preset value;
[0047] By alternately optimizing the policy parameters of the agent and the antibody, an adaptive game optimization of the anti-interference strategy under complex interference conditions is carried out.
[0048] Among them, the environmental state transition satisfies:
[0049]
[0050] In the above formula, Represents the actions of the intelligent agent. S represents the action on the antibody. t This indicates an anti-interference strategy.
[0051] Optionally, during the actual detection process of the radar system, the actions of selecting and executing the initial anti-jamming strategy according to the current environmental state include:
[0052] During the actual radar detection process, the radar system adjusts its operation based on the real-time status of the current environment. Select the optimal strategy from the initial anti-interference strategy. And execute its corresponding action, the expression of which is:
[0053]
[0054] In the above formula, This represents the optimal anti-interference strategy. The action representing the optimal decision includes: waveform parameter configuration, filter threshold adjustment, and signal processing strategy switching.
[0055] Optionally, when the performance of the detection radar system degrades or the interference characteristics change abruptly, the parameters of the radar strategy network are monitored and fine-tuned online based on real-time feedback data, including:
[0056] When a performance degradation or sudden change in interference characteristics of the radar system is detected, the parameters of the radar strategy network are updated using an online fine-tuning mechanism based on real-time feedback data. The expression is as follows:
[0057]
[0058] In the above formula, This indicates the updated parameters; Indicates the current parameter; Indicates the learning rate; This indicates the loss function with respect to the parameters. The gradient.
[0059] Optionally, the performance of the detection results of the radar system after online monitoring and fine-tuning is evaluated based on the composite reward function, including:
[0060] The detection probability, the false alarm probability, and the signal-to-interference ratio are calculated respectively.
[0061] The formula for calculating the detection probability is as follows:
[0062]
[0063] In the above formula, Indicates the detection probability; Indicates the number of correctly detected items; Represented as the total;
[0064] The formula for calculating the false alarm probability is:
[0065]
[0066] In the above formula, Indicates the probability of a false alarm; Indicates the number of false alarms; Indicates the number of noise detection events;
[0067] The formula for calculating the signal-to-interference ratio is:
[0068]
[0069] In the above formula, Indicates the signal-to-interference ratio; Indicates the power of the target signal; This is expressed as the power of the interference signal.
[0070] This application proposes a reinforcement learning-based adaptive anti-jamming detection method for complex mine radar, which can effectively cope with multi-source interference factors such as dust scattering, multipath reflection, and electromagnetic radiation in mine roadways. The mine roadway environment is complex, and traditional fixed-parameter or static filtering radar signal processing methods struggle to achieve stable detection in this environment. This application combines multi-domain feature extraction with reinforcement learning decision optimization to construct a multi-domain feature fusion model that includes power spectral density (PSD) analysis, Doppler-time (D-T) analysis, range-time (R-T) analysis, and range-Doppler (R-D) analysis. Furthermore, it utilizes an incomplete information game-theoretic reinforcement learning algorithm to dynamically adjust the radar anti-jamming strategy. This method maintains high detection accuracy and robustness under strong interference conditions, significantly improving the target recognition capability and long-term stability of millimeter-wave radar systems in complex mine environments, demonstrating broad application prospects and high practicality. Attached Figure Description
[0071] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the description of the embodiments taken in conjunction with the following drawings, in which:
[0072] Figure 1 is a flowchart of a complex mine radar adaptive anti-interference detection method based on reinforcement learning in an embodiment of this application;
[0073] Figure 2 is a block diagram of a complex mine radar adaptive anti-jamming detection system based on reinforcement learning in an embodiment of this application. Detailed Implementation
[0074] The embodiments of this application will now be described in detail. Examples of these embodiments are illustrated in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain this application, and should not be construed as limiting this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.
[0075] The adaptive anti-jamming detection method for complex mine radar based on reinforcement learning proposed in this application, as shown in the flowchart in Figure 1, includes the following steps:
[0076] Step 101: Receive the echo signal and extract multi-domain features based on the echo signal.
[0077] First, the echo signal is received, and then multi-domain features are extracted based on the received echo signal. Generally, a millimeter-wave radar can transmit a linear frequency modulated continuous wave (FMCW) signal. The radar system will then receive echo signals containing various information such as target echoes, dust scattering, multipath reflections, and electromagnetic interference. Fast Fourier Transform and time-frequency analysis are performed on these echo signals. Alternatively, other known methods can be used to extract multi-domain features, including: power spectral density (PSD) features, Doppler-time (D–T) features, range-time (R–T) features, and range-Doppler (R–D) features.
[0078] In one embodiment of this application, preferably, the extraction of multi-domain features may include:
[0079] The expression for obtaining the linear frequency modulated continuous wave signal transmitted by the radar is as follows:
[0080]
[0081] In the above formula, A represents the amplitude of the transmitted signal; Indicates the starting frequency of the signal; The frequency modulation bandwidth of the signal is represented by T; the frequency modulation period is represented by t; and the time variable is represented by t.
[0082] The echo signal is obtained by the following expression:
[0083]
[0084] In the above formula, Indicates the echo signal; This represents the reflection coefficient of the i-th target, reflecting the degree of attenuation of the echo signal; The propagation delay of the signal reaching the target is represented by the following formula: ,in Let represent the distance between the i-th target and the radar, and c represent the speed of electromagnetic wave propagation; Represents the amplitude coefficient of the j-th type of interference component; Indicates the propagation delay of the interference signal; This represents system noise, assumed to be Gaussian white noise;
[0085] Fast Fourier Transform and time-frequency analysis were performed on the echo signal to obtain power spectral density domain characteristics, Doppler time domain characteristics, range-time domain characteristics, and range-Doppler domain characteristics, expressed as follows:
[0086]
[0087] In the above formula, Indicates the power spectral density domain characteristics; Indicates the time-domain characteristics of Doppler; Represents distance-time domain features; This represents the distance-Doppler domain characteristics.
[0088] Step 102: Construct a reinforcement learning state space based on multi-domain features, and design a composite reward function based on detection probability, false alarm probability and signal-to-interference ratio. This composite reward function is used to describe the balance between the radar system's target detection accuracy, false alarm suppression and interference suppression performance.
[0089] Based on the aforementioned multi-domain features, a reinforcement learning state space is constructed, and a composite reward function based on detection probability, false alarm probability, and signal-to-interference ratio is designed. This composite reward function is used to describe the balance between the radar system's target detection accuracy, false alarm suppression, and interference suppression performance.
[0090] A preferred approach is to jointly encode the power spectral density domain features, Doppler time domain features, range time domain features, and range-Doppler domain features, encoding them into a state vector. Using these multi-domain features as the state vector in the reinforcement learning state space, its expression can be:
[0091]
[0092] In the above formula, This represents the environmental state perceived by the agent at time t; The power spectral density domain characteristics at time t are represented; Represents the Doppler time-domain characteristics at time t; Represents the distance-time domain characteristics at time t; This represents the distance-Doppler domain feature at time t.
[0093] A composite reward function is designed, with detection probability, false alarm probability, and signal-to-interference ratio as core indicators, to describe the balance between the system's target detection accuracy, false alarm suppression, and interference suppression performance.
[0094] In one embodiment of this application, a preferred design of the composite reward function is expressed as follows:
[0095]
[0096] In the above formula, This represents the agent's immediate reward at time t; It represents the target detection probability, reflecting the radar system's ability to correctly detect targets; It represents the false alarm probability, reflecting the probability that the radar system misjudges noise or interference as a target; Indicates the signal-to-interference ratio; , , These represent the weighting coefficients of each indicator in the reward function, used to balance the contributions of detection performance, false alarm control, and interference suppression.
[0097] To address the practical challenges of millimeter-wave radar detection in complex environments like mine tunnels, traditional methods often focus on a single metric (such as detection probability or false alarm rate), leading to an inability to balance overall performance under complex interference conditions. This application introduces a three-element constraint concept of "detection capability—false alarm suppression—interference suppression" into the design of a composite reward function. By comprehensively incorporating three key performance indicators—target detection probability, false alarm probability, and signal-to-interference ratio (SIR)—into the same reward function, a global trade-off for performance optimization is achieved. Specifically, target detection probability reflects the system's ability to identify real targets and is a direct manifestation of detection performance; false alarm probability measures the risk of the system misjudging noise or interference as targets and is a key indicator of anti-interference robustness; and the SIR describes the enhancement of the target signal relative to the interference energy and is an important quantitative indicator of the radar system's overall anti-interference effectiveness. By weighting and combining these three indicators, the agent can be guided to simultaneously improve detection accuracy, reduce false alarm rate, and enhance anti-interference capability during training, thereby obtaining an adaptive detection strategy with optimal overall performance.
[0098] Step 103: Based on the reinforcement learning state space and composite reward function, an incomplete information game adversarial reinforcement learning algorithm is introduced to iteratively train the radar policy network. By alternately optimizing the policy parameters of the agent and the opponent, the initial anti-jamming strategy of the radar system under multi-source interference conditions is generated. In the incomplete information game adversarial reinforcement learning algorithm, the radar policy network is used as the agent and the interference source is used as the opponent.
[0099] After constructing the reinforcement learning state space and designing the composite reward function, an incomplete information game-theoretic reinforcement learning algorithm is introduced based on the reinforcement learning state space and composite reward function to iteratively train the radar policy network. By alternately optimizing the policy parameters of the agent and the opponent, an initial anti-jamming strategy for the radar system under multi-source interference conditions is generated. In the incomplete information game-theoretic reinforcement learning algorithm, the radar policy network is used as the agent, and the interference source is used as the opponent. Specifically:
[0100] An adversarial antibody can be introduced during training to simulate different types and evolutionary forms of interference sources, forming an adversarial interaction with the agent. The goal of this adversarial antibody is to minimize the detection performance of the radar system by continuously changing the interference strategy to approximate the most unfavorable interference situation in the real environment. The goal of the agent is to maximize the cumulative reward by optimizing the adversarial interference strategy to maintain the preset detection capability and robustness under adversarial conditions. The radar policy network is trained using an incomplete information game adversarial learning algorithm, and the optimal anti-interference strategy is searched using the policy gradient method.
[0101] The expression for the agent's objective is:
[0102]
[0103] In the above formula, This represents the optimal anti-interference strategy; This represents the radar strategy network; T represents the total number of interaction steps. This represents the discount factor, used to measure the importance of future rewards.
[0104] The agent updates the optimal anti-interference strategy using a proximal policy optimization framework, the expression of which is:
[0105]
[0106] In the above formula, Represents the loss function; These represent the parameters of the radar strategy network. The policy probability ratio is represented by:
[0107]
[0108] In the above formula, Indicates the action to be taken under the current anti-interference strategy. The probability of; This indicates the action taken under the old anti-interference strategy. The probability of; The advantage function measures the advantage of the current action relative to the average policy. This represents the truncation coefficient, used to prevent the anti-interference strategy from updating beyond a preset value.
[0109] By alternately optimizing the policy parameters of the agent and the antibody, an adaptive game optimization of the anti-interference strategy under complex interference conditions is carried out.
[0110] Among them, the environmental state transition satisfies:
[0111]
[0112] In the above formula, Represents the actions of the intelligent agent. S represents the action on the antibody. t This indicates an anti-interference strategy.
[0113] The reason for adopting the incomplete information game adversarial reinforcement learning algorithm is that in complex scenarios such as mine tunnels, the behavior of interference sources has strong non-static and adversarial characteristics: on the one hand, interference components such as dust concentration, reflection intensity, and electromagnetic radiation change dynamically with time and space; on the other hand, some interference has "strategic" characteristics, such as the frequency, phase, power, and time structure of electromagnetic interference signals will adaptively adjust according to the radar's working mode. This makes it easy for simple reinforcement learning agents to overfit fixed interference models and have insufficient generalization ability.
[0114] To address the aforementioned challenges, this application creatively introduces an incomplete information game-theoretic adversarial reinforcement learning mechanism: during training, an antibody is introduced to simulate different types and evolutionary modes of interference sources, forming an adversarial interaction with the agent. The antibody aims to minimize radar detection performance by continuously changing its interference strategy to approximate the most unfavorable interference scenarios in the real environment; the agent's goal is to maximize cumulative rewards, maintaining high detection capability and robustness under adversarial conditions through policy optimization. The radar policy network is trained using the incomplete information game-theoretic adversarial learning mechanism, and the optimal policy is searched using the policy gradient method.
[0115] Step 104: Apply the initial anti-jamming strategy to the radar system. During the actual detection process of the radar system, select and execute the actions of the initial anti-jamming strategy according to the current environmental state. These actions include: waveform parameter configuration, filter threshold adjustment, and signal processing strategy switching.
[0116] After obtaining the initial anti-jamming strategy, it needs to be applied to the radar system. During the actual detection process of the radar system, the actions of the initial anti-jamming strategy are selected and executed according to the current environmental state. These actions include: waveform parameter configuration, filter threshold adjustment, and signal processing strategy switching.
[0117] In one embodiment of this application, a preferred method for selecting and executing an initial anti-jamming strategy based on the current environmental state during actual radar system detection includes:
[0118] During actual radar detection, the radar system adjusts its response based on the real-time status of the current environment. Select the optimal strategy from the initial anti-interference strategy. And execute its corresponding action, the expression of which is:
[0119]
[0120] In the above formula, This represents the optimal anti-interference strategy. The action representing the optimal decision includes: waveform parameter configuration, filter threshold adjustment, and signal processing strategy switching.
[0121] Step 105: When the performance of the radar system degrades or the interference characteristics change abruptly, the parameters of the radar strategy network are monitored and fine-tuned online based on real-time feedback data. The real-time feedback data is pre-processed data obtained from the real-time environment.
[0122] The execution of the initial anti-jamming strategy may maintain or improve the performance of the radar system, but it may also degrade the performance of the radar system. Sudden changes in the jamming characteristics may also cause the initial anti-jamming strategy to fail to meet the radar's operational requirements.
[0123] In reinforcement learning-driven adaptive anti-jamming detection of radar, the trained strategy may experience decreased detection accuracy, increased false alarm rate, or weakened interference suppression capability due to dynamic changes in conditions such as dust concentration, electromagnetic interference source power, and multipath reflection characteristics in the mine roadway environment. To ensure the stability and long-term robustness of the radar system in complex dynamic environments, it is necessary to quantitatively evaluate the effectiveness of the current anti-jamming strategy and re-optimize and iteratively update the radar strategy network based on the evaluation results, thus forming a closed-loop self-learning process. This requires online monitoring and fine-tuning of the radar strategy network parameters based on real-time feedback data, where the real-time feedback data is pre-processed data acquired from the real-time environment. Specifically:
[0124] When a performance degradation or sudden change in interference characteristics of the radar system is detected, the parameters of the radar strategy network are updated using an online fine-tuning mechanism based on real-time feedback data. The expression for this update is:
[0125]
[0126] In the above formula, This indicates the updated parameters; Indicates the current parameter; Indicates the learning rate; This indicates the loss function with respect to the parameters. The gradient.
[0127] Step 106: Evaluate the performance of the radar system's detection results after online monitoring and fine-tuning based on the composite reward function.
[0128] After online fine-tuning and updating, it is naturally necessary to perform performance evaluation on the detection results of the radar system after online monitoring and fine-tuning using a composite reward function. In one embodiment of this application, a preferred method for performance evaluation of the detection results of the radar system after online monitoring and fine-tuning based on a composite reward function includes:
[0129] Calculate the detection probability, false alarm probability, and signal-to-interference ratio separately;
[0130] The formula for calculating the detection probability is as follows:
[0131]
[0132] In the above formula, Indicates the detection probability; Indicates the number of correctly detected items; Represented as the total;
[0133] The formula for calculating the false alarm probability is:
[0134]
[0135] In the above formula, Indicates the probability of a false alarm; Indicates the number of false alarms; Indicates the number of noise detection events;
[0136] The formula for calculating the signal-to-interference ratio is:
[0137]
[0138] In the above formula, Indicates the signal-to-interference ratio; Indicates the power of the target signal; This is expressed as the power of the interference signal. Performance evaluation results can be obtained using the above method.
[0139] Step 107: If the performance evaluation result does not meet the preset performance threshold, adjust the weights of the composite reward function and the parameters of the radar policy network, and execute the following steps: introduce an incomplete information game adversarial reinforcement learning algorithm, iteratively train the adjusted radar policy network, and generate a new anti-jamming strategy for the radar system under multi-source interference conditions and subsequent steps by alternately optimizing the policy parameters of the agent and the antibody, until the optimal anti-jamming strategy is obtained, and the performance evaluation result corresponding to the optimal anti-jamming strategy meets the preset performance threshold.
[0140] After obtaining the performance evaluation results, if the performance evaluation results do not meet the preset performance threshold, the weights of the composite reward function and the parameters of the radar policy network are adjusted, and the following steps are executed: an incomplete information game adversarial reinforcement learning algorithm is introduced to iteratively train the adjusted radar policy network. By alternately optimizing the policy parameters of the agent and the antibody, a new anti-jamming strategy for the radar system under multi-source interference conditions and subsequent steps are generated until the optimal anti-jamming strategy is obtained. The performance evaluation results corresponding to the optimal anti-jamming strategy meet the preset performance threshold.
[0141] In other words, the interference suppression effect of radar detection results is evaluated based on performance indicators such as target detection probability, false alarm probability, and signal-to-interference ratio. When the evaluation results do not meet the preset performance threshold, the weights of the reinforcement learning reward function and the parameters of the radar strategy network are automatically adjusted to retrain and optimize the anti-interference strategy, forming a closed-loop self-learning mechanism of "evaluation-optimization-relearning". This ensures the continuous learning, long-term stable operation, and detection accuracy of the radar system under complex mining environments and dynamic changes in multi-source interference.
[0142] Based on the above-mentioned reinforcement learning-based adaptive anti-jamming detection method for complex mine radar, this application also proposes a reinforcement learning-based adaptive anti-jamming detection system for complex mine radar, as shown in the system block diagram in Figure 2, which includes:
[0143] Extraction module 210 is used to receive echo signals and extract multi-domain features based on the echo signals;
[0144] The design module 220 is used to construct a reinforcement learning state space based on the multi-domain features and design a composite reward function based on detection probability, false alarm probability and signal-to-interference ratio. This composite reward function is used to describe the balance between the radar system's target detection accuracy, false alarm suppression and interference suppression performance.
[0145] The game-theoretic pre-training module 230 is used to introduce an incomplete information game-theoretic reinforcement learning algorithm based on the reinforcement learning state space and the composite reward function to iteratively train the radar policy network. By alternately optimizing the policy parameters of the agent and the opponent, the initial anti-jamming strategy of the radar system under multi-source interference conditions is generated. In the incomplete information game-theoretic reinforcement learning algorithm, the radar policy network is used as the agent and the interference source is used as the opponent.
[0146] The execution module 240 is used to apply the initial anti-interference strategy to the radar system. During the actual detection process of the radar system, the initial anti-interference strategy is selected and executed according to the current environmental state. The action includes: waveform parameter configuration, filter threshold adjustment and signal processing strategy switching.
[0147] The online monitoring and fine-tuning module 250 is used to monitor and fine-tune the parameters of the radar strategy network online based on real-time feedback data when the performance of the radar system degrades or the interference characteristics change abruptly. The real-time feedback data is pre-processed data obtained from the real-time environment; that is, the radar strategy network is updated with small-step gradients using a small number of real-time samples to quickly adapt to the new environment.
[0148] Performance evaluation module 260 is used to evaluate the detection results of the radar system after online monitoring and fine-tuning based on the composite reward function;
[0149] The adaptive optimization module 270 is used to adjust the weights of the composite reward function and the parameters of the radar policy network if the performance evaluation result does not meet the preset performance threshold. The steps are as follows: an incomplete information game adversarial reinforcement learning algorithm is introduced to iteratively train the adjusted radar policy network. By alternately optimizing the policy parameters of the agent and the antibody, a new anti-jamming strategy for the radar system under multi-source interference conditions and subsequent steps are generated until the optimal anti-jamming strategy is obtained. The performance evaluation result corresponding to the optimal anti-jamming strategy meets the preset performance threshold.
[0150] Each of the above modules corresponds to the complex mine radar adaptive anti-interference detection method described in steps 101 to 107 above. This can be intuitively understood by combining the above explanation and description, and will not be repeated here.
[0151] In summary, the reinforcement learning-based adaptive anti-jamming detection method for complex mine radar proposed in this application can effectively cope with multi-source interference factors such as dust scattering, multipath reflection, and electromagnetic radiation in mine roadways. The mine roadway environment is complex, and traditional fixed-parameter or static filtering radar signal processing methods struggle to achieve stable detection in this environment. This application combines multi-domain feature extraction with reinforcement learning decision optimization to construct a multi-domain feature fusion model that includes power spectral density (PSD) analysis, Doppler-time (D-T) analysis, range-time (R-T) analysis, and range-Doppler (R-D) analysis. Furthermore, it utilizes an incomplete information game-theoretic reinforcement learning algorithm to dynamically adjust the radar anti-jamming strategy. This method maintains high detection accuracy and robustness under strong interference conditions, significantly improving the target recognition capability and long-term stability of millimeter-wave radar systems in complex mine environments, demonstrating broad application prospects and high practicality.
[0152] Although preferred embodiments of the present application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the embodiments of the present application.
[0153] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal device. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal device that includes said element.
[0154] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims. All of these forms are within the protection scope of this application.
Claims
1. A complex mine radar adaptive anti-jamming detection method based on reinforcement learning, characterized in that, include: Receive the echo signal and extract multi-domain features based on the echo signal; A reinforcement learning state space is constructed based on the aforementioned multi-domain features, and a composite reward function based on detection probability, false alarm probability, and signal-to-interference ratio is designed. This composite reward function describes the balance between target detection accuracy, false alarm suppression, and interference suppression performance of the radar system. Based on the reinforcement learning state space and the composite reward function, an incomplete information game-theoretic reinforcement learning algorithm is introduced to iteratively train the radar policy network. By alternately optimizing the policy parameters of the agent and the opponent, an initial anti-interference strategy for the radar system under multi-source interference conditions is generated. In the incomplete information game-theoretic reinforcement learning algorithm, the radar policy is used as the basis for the initial anti-interference strategy. The network acts as an intelligent agent, with interference sources as its countermeasures. The initial anti-interference strategy is applied to the radar system. During actual radar detection, the system selects and executes actions based on the current environmental state, including waveform parameter configuration, filter threshold adjustment, and signal processing strategy switching. When the radar system's performance degrades or interference characteristics abruptly change, the parameters of the radar strategy network are monitored and fine-tuned online based on real-time feedback data, which is pre-processed data obtained from the real-time environment. The detection results of the radar system after online monitoring and fine-tuning are evaluated based on the composite reward function. The system can evaluate performance; if the performance evaluation result does not meet the preset performance threshold, the weights of the composite reward function and the parameters of the radar policy network are adjusted, and the following steps are executed: An incomplete information game-theoretic reinforcement learning algorithm is introduced to iteratively train the adjusted radar policy network. By alternately optimizing the policy parameters of the agent and the antibody, a new anti-jamming strategy for the radar system under multi-source interference conditions and subsequent steps are generated until the optimal anti-jamming strategy is obtained. The performance evaluation result corresponding to the optimal anti-jamming strategy meets the preset performance threshold. Specifically, an incomplete information game-theoretic reinforcement learning algorithm is introduced to iteratively train the radar policy network. This includes: introducing an anti-interference agent during training to simulate different types and evolutionary forms of interference sources, creating an adversarial interaction with the agent. The anti-interference agent aims to minimize the radar system's detection performance by continuously changing its interference strategy to approximate the most unfavorable interference scenario in the real environment. The agent's objective is to maximize cumulative reward by optimizing its anti-interference strategy to maintain preset detection capabilities and robustness under adversarial conditions. The radar policy network is trained using the incomplete information game-theoretic adversarial learning algorithm, and the optimal anti-interference strategy is searched using the policy gradient method. The agent's objective is expressed as: In the above formula, This represents the optimal anti-interference strategy; This represents the radar strategy network; This represents the immediate reward the agent receives at time t; T represents the total number of interaction steps. The discount factor is used to measure the importance of future rewards; the agent updates the optimal anti-interference strategy using a proximal policy optimization framework, the expression of which is: In the above formula, Represents the loss function; Represents the dominance function; The parameters represent the radar strategy network; The policy probability ratio is represented by: In the above formula, Indicates the action to be taken under the current anti-interference strategy. The probability of; This indicates the action taken under the old anti-interference strategy. The probability of; The advantage function measures the advantage of the current action relative to the average policy. represents the cutoff coefficient, used to prevent the anti-interference strategy update from exceeding a preset value; adaptive game optimization of the anti-interference strategy under complex interference conditions is performed by alternately optimizing the policy parameters of the agent and the antibody; wherein, the environmental state transition satisfies: In the above formula, Represents the actions of the intelligent agent. S represents the action on the antibody. t It indicates the current state of the environment.
2. The adaptive anti-interference detection method for complex mine radar according to claim 1, characterized in that, Receiving echo signals and extracting multi-domain features based on the echo signals includes: transmitting linear frequency modulated continuous wave signals using millimeter-wave radar; receiving echo signals containing target echoes, dust scattering, multipath reflections, and electromagnetic interference using millimeter-wave radar; performing fast Fourier transform and time-frequency analysis on the echo signals to extract the multi-domain features, which include: power spectral density domain features, Doppler time domain features, range time domain features, and range-Doppler domain features.
3. The adaptive anti-interference detection method for complex mine radar according to claim 1, characterized in that, The extracted multi-domain features include: obtaining the radar-transmitted linear frequency modulated continuous wave signal, the expression of which is: In the above formula, A represents the amplitude of the transmitted signal; Indicates the starting frequency of the signal; The frequency modulation bandwidth of the signal is represented by T; the frequency modulation period is represented by t; the echo signal is obtained by the following expression: In the above formula, This refers to the echo signal; This represents the reflection coefficient of the i-th target, reflecting the degree of attenuation of the echo signal; The propagation delay of the signal reaching the target is represented by the following formula: ,in Let represent the distance between the i-th target and the radar, and c represent the speed of electromagnetic wave propagation; Represents the amplitude coefficient of the j-th type of interference component; Indicates the propagation delay of the interference signal; The system noise is represented as Gaussian white noise. Fast Fourier Transform and time-frequency analysis are performed on the echo signal to obtain power spectral density domain characteristics, Doppler time domain characteristics, range-time domain characteristics, and range-Doppler domain characteristics, expressed as: In the above formula, This represents the power spectral density domain characteristics; This represents the Doppler time-domain feature; This represents the distance-time domain feature; This represents the distance Doppler domain feature.
4. The adaptive anti-interference detection method for complex mine radar according to claim 1, characterized in that, Constructing a reinforcement learning state space based on the multi-domain features includes: using the multi-domain features as state vectors of the reinforcement learning state space, the expression of which is: In the above formula, This represents the environmental state perceived by the agent at time t; The power spectral density domain characteristics at time t are represented; Represents the Doppler time-domain characteristics at time t; Represents the distance-time domain characteristics at time t; This represents the distance-Doppler domain feature at time t.
5. The adaptive anti-interference detection method for complex mine radar according to claim 1, characterized in that, The expression for the designed composite reward function is as follows: In the above formula, This represents the agent's immediate reward at time t; It represents the target detection probability, reflecting the radar system's ability to correctly detect targets; It represents the false alarm probability, reflecting the probability that the radar system misjudges noise or interference as a target; Indicates the signal-to-interference ratio; 、 、 These represent the weighting coefficients of each indicator in the reward function, used to balance the contributions of detection performance, false alarm control, and interference suppression.
6. The adaptive anti-interference detection method for complex mine radar according to claim 1, characterized in that, During the actual detection process of the radar system, the actions of selecting and executing the initial anti-jamming strategy based on the current environmental state include: during the actual radar detection process, the radar system selects and executes the initial anti-jamming strategy based on the real-time state of the current environment. Select the optimal strategy from the initial anti-interference strategy. And execute its corresponding action, the expression of which is: In the above formula, This represents the optimal anti-interference strategy. The action representing the optimal decision includes: waveform parameter configuration, filter threshold adjustment, and signal processing strategy switching.
7. The adaptive anti-interference detection method for complex mine radar according to claim 1, characterized in that, When a performance degradation or sudden change in interference characteristics of the radar system is detected, the parameters of the radar strategy network are monitored and fine-tuned online based on real-time feedback data. This includes: when a performance degradation or sudden change in interference characteristics of the radar system is detected, the parameters of the radar strategy network are updated using an online fine-tuning mechanism based on real-time feedback data, the expression of which is: In the above formula, This indicates the updated parameters; Indicates the current parameter; Indicates the learning rate; This indicates the loss function with respect to the parameters. The gradient.
8. The adaptive anti-interference detection method for complex mine radar according to claim 1, characterized in that, The performance evaluation of the detection results of the radar system after online monitoring and fine-tuning is performed based on the composite reward function, including: calculating the detection probability, the false alarm probability, and the signal-to-interference ratio respectively; wherein, the formula for calculating the detection probability is: In the above formula, Indicates the detection probability; Indicates the number of correctly detected items; The total number is represented as the number of false alarms; the formula for calculating the false alarm probability is: In the above formula, Indicates the probability of a false alarm; Indicates the number of false alarms; This represents the number of noise detection events; the formula for calculating the signal-to-interference ratio is: In the above formula, Indicates the signal-to-interference ratio; Indicates the power of the target signal; This is expressed as the power of the interference signal.
Citation Information
Patent Citations
Deep reinforcement learning anti-interference method for frequency agile radar
CN114509732A
Radar and communication integrated unmanned aerial vehicle cooperative multi-target detection method
CN114679729A