Signal processing method, device, electronic device and storage medium
Through deep reinforcement learning model, the interference characteristic information in the radar echo signal is identified, and the anti-interference algorithm strategy is quickly and accurately determined, which solves the problems of low efficiency and low accuracy caused by artificial methods in the prior art, and improves the anti-interference ability of the radar system.
Patent Information
- Application Number
- CN202411457355.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-18
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2044-10-18
AI Technical Summary
In the prior art, the anti-interference strategy is manually determined by means of a manual method, resulting in low interference suppression efficiency and prone to errors, affecting the anti-interference effect.
The deep reinforcement learning model is used to identify the interference characteristic information in the radar echo signal, and the anti-interference algorithm strategy is quickly and accurately determined based on the strategy determination model, and anti-interference processing is carried out.
It improves the efficiency and accuracy of anti-interference processing, and enhances the adaptability and intelligence level of radar systems in complex electromagnetic environments.
Smart Images

Figure CN119414350B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to the field of signal processing technology, and in particular to a signal processing method, device, electronic device, and storage medium. Background Art
[0002] With the evolution of electromagnetic spectrum interference technology becoming more complex, flexible and intelligent, the interference scenarios faced by radar systems have become more complex and changeable. Anti-interference processing of received signals by radar systems has become a very important link.
[0003] In the prior art, an anti-interference strategy is typically determined manually, and then anti-interference processing is performed on the received echo signal using the determined anti-interference strategy. However, during the implementation of the present invention, it was discovered that the prior art suffers from at least the following technical problems: Manually determining the anti-interference strategy results in low interference suppression efficiency and is prone to errors, resulting in poor accuracy of the determined anti-interference strategy, which affects the anti-interference effect. Summary of the Invention
[0004] The embodiments of the present invention provide a signal processing method, device, electronic device and storage medium, which can quickly and accurately determine the current anti-interference algorithm strategy through a strategy determination model, thereby improving the efficiency and accuracy of anti-interference processing.
[0005] According to one aspect of the present invention, there is provided a signal processing method, comprising:
[0006] receiving a current radar echo signal, and identifying current interference feature information of a current interference signal in the current radar echo signal;
[0007] Inputting the current interference feature information into a pre-trained strategy determination model, and obtaining a current anti-interference algorithm strategy for processing the current interference signal based on an output result of the strategy determination model, so as to perform anti-interference processing on the current radar echo signal according to the current anti-interference algorithm strategy;
[0008] Wherein, the strategy determination model includes a deep reinforcement learning model.
[0009] According to another aspect of the present invention, there is provided a signal processing device, comprising:
[0010] A signal receiving module is used to receive the current radar echo signal and identify the current interference feature information of the current interference signal in the current radar echo signal;
[0011] An information input module is used to input the current interference feature information into a pre-trained strategy determination model, and based on the output results of the strategy determination model, obtain the current anti-interference algorithm strategy for processing the current interference signal, so as to perform anti-interference processing on the current radar echo signal according to the current anti-interference algorithm strategy; wherein the strategy determination model includes a deep reinforcement learning model.
[0012] According to another aspect of the present invention, an electronic device is provided, comprising:
[0013] at least one processor; and
[0014] a memory communicatively connected to the at least one processor; wherein,
[0015] The memory stores a computer program that can be executed by the at least one processor. The computer program is executed by the at least one processor to enable the at least one processor to perform the signal processing method described in any embodiment of the present invention.
[0016] According to another aspect of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the signal processing method according to any embodiment of the present invention when executed.
[0017] The technical solution of the embodiment of the present invention receives the current radar echo signal and identifies the current interference feature information of the current interference signal in the current radar echo signal; by inputting the current interference feature information into a pre-trained strategy determination model, based on the output result of the strategy determination model, the current anti-interference algorithm strategy for processing the current interference signal is obtained; wherein, the strategy determination model includes a deep reinforcement learning model; through the strategy determination model, the current anti-interference algorithm strategy can be determined quickly and accurately without human intervention, and the current radar echo signal can be subjected to anti-interference processing according to the current anti-interference algorithm strategy, which is conducive to improving the efficiency and accuracy of anti-interference processing.
[0018] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present invention, nor is it intended to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0020] Figure 1 is a flowchart of a signal processing method provided according to an embodiment of the present invention;
[0021] Figure 2 is a flowchart of another signal processing method provided according to an embodiment of the present invention;
[0022] Figure 3 This is a radar interference suppression training framework diagram based on deep reinforcement learning provided by an embodiment of the present invention;
[0023] Figure 4 is a structural diagram of a signal processing device provided according to an embodiment of the present invention;
[0024] Figure 5 It is a structural diagram of an electronic device for implementing the signal processing method according to an embodiment of the present invention. DETAILED DESCRIPTION
[0025] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.
[0026] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "etc." and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0027] Figure 1 This is a flow chart of a signal processing method according to an embodiment of the present invention. This embodiment is applicable to performing anti-interference processing on received radar echo signals. The method can be performed by a signal processing device, which can be implemented in hardware and / or software.
[0028] like Figure 1 As shown, the method of this embodiment may specifically include:
[0029] S110: Receive a current radar echo signal, and identify current interference feature information of a current interference signal in the current radar echo signal.
[0030] It should be noted that a radar system can transmit a current radar transmit signal into the surrounding environment. Upon encountering a target object in the surrounding environment, the current radar transmit signal is reflected. The radar system then determines the received reflected echo signal as the current radar echo signal. The current radar echo signal includes echo signals from multiple targets, interference signals, and environmental noise. The current interference signature information is the characteristic information of the interference signal contained in the current radar echo signal.
[0031] In a specific implementation, the radar system can receive the current radar echo signal, identify the interference signal in the current radar echo signal, and use the characteristic information of the interference signal in the current radar echo signal as the current interference characteristic information. The received current radar echo signal u(t) can be expressed as:
[0032]
[0033] Among them, N ant Indicates the number of receiving antennas. For example, it may include 8 main antennas and 4 auxiliary antennas, for a total of 12. k,m (t) represents the echo signal of the kth target on the mth antenna. i,m (t) represents the signal of the i-th interference signal on the m-th antenna. m (t) represents the ambient noise on the mth antenna. N tar Indicates the number of targets. N int Indicates the number of interference sources.
[0034] Optionally, identifying the current interference characteristic information of the current interference signal in the current radar echo signal includes: determining the corresponding multiple signal classification spectrum in the current radar echo signal based on a multiple signal classification algorithm; identifying the current interference type and current characteristic parameters of the current interference signal in the current radar echo signal; and determining the multiple signal classification spectrum, the current interference type and the current characteristic parameters as the current interference characteristic information.
[0035] The current interference type includes at least one of a main-sidelobe interference type, a suppression-forward interference type, and a pulse-continuous wave interference type; and the current characteristic parameter includes frequency and / or amplitude. The main-sidelobe interference type can be used to distinguish whether the interference signal is mainlobe interference or sidelobe interference, the suppression-forward interference type is used to distinguish whether the interference signal is suppression interference or forwarding interference, and the pulse-continuous wave interference type is used to distinguish whether the interference signal is pulse interference or continuous interference.
[0036] Specifically, the Multiple Signal Classification (MUSIC) algorithm solves the arrival angle information of each radiation source and outputs a multiple signal classification spectrum, which includes the received signal strength information at the angle corresponding to the radiation source arrival angle information.
[0037] Furthermore, when determining the interference type of the current interference signal, the main lobe and side lobe interference types can be determined first. For the main lobe and side lobe interference types, the difference in directivity of the radar's main and auxiliary beams is utilized, and the interference judgment is performed by setting the main and auxiliary antenna receiving gains of the side lobe arrival signals. Exemplarily, the ratio of the first average power of the auxiliary beam receiving interference to the second average power of the main beam receiving interference is determined. If the ratio is greater than or equal to the preset judgment threshold value, the type of the interference signal is determined to be side lobe interference; if the ratio is less than the preset judgment threshold value, the type of the interference signal is determined to be main lobe interference.
[0038] When determining whether the current interference signal is of the continuous wave interference type, a first signal crossing threshold ratio of the current interference signal can be determined based on the average amplitude of the background noise and the continuous wave interference detection absolute threshold. The first signal crossing threshold ratio is compared with a preset crossing absolute threshold ratio. When the first signal crossing threshold ratio is greater than or equal to the preset crossing absolute threshold ratio, the current interference signal is determined to have continuous wave interference and the interference type may be continuous wave interference. When the first signal crossing threshold ratio is less than the preset crossing absolute threshold ratio, the current interference signal is determined to not have continuous wave interference. When determining whether the current interference signal is of the pulse interference type, background normalization processing can be performed on the current interference signal. Furthermore, based on the pulse detection relative background threshold, a second signal crossing threshold ratio of the current interference signal is determined and compared with a preset pulse interference judgment ratio threshold. If the second signal crossing threshold ratio is greater than or equal to the pulse interference judgment ratio threshold, the current interference signal is determined to have pulse interference and the interference type may be pulse interference. If the second signal crossing threshold ratio is less than the pulse interference judgment threshold, the current interference pulse signal is determined to not have pulse interference.
[0039] Furthermore, forwarding interference manifests as pulses in the pulse pressure output domain. Therefore, detection and classification of dense forwarding interference requires only the same processing as for pulse interference detection at the pulse pressure output. Based on the dense forwarding interference detection relative background threshold, the signal-to-relative threshold ratio is determined. If the signal-to-relative threshold ratio is greater than or equal to the preset relative threshold ratio, dense forwarding interference is determined to be present in the current interfering signal. If the signal-to-relative threshold ratio is less than the preset relative threshold ratio, dense forwarding interference is determined to be absent.
[0040] In this embodiment, the determination result of the current interference type can be reflected in the form of a combination sequence, which includes a logical value corresponding to each interference type, that is, the main lobe interference type, the side lobe interference type, the suppressed interference type, the forwarding interference type, the pulse interference type, and the continuous interference type each correspond to a logical value, and the logical value is used to determine whether the current interference signal has this type of interference. Exemplarily, the logical value includes two values 1 or 0, 1 indicates that the current interference signal has this interference type, and 0 indicates that the current interference signal does not have this interference type. For example, when the logical value corresponding to the main lobe interference type is 1, it means that the current interference signal has main lobe interference; the logical value corresponding to the suppressed interference type is 0, which means that the current interference signal does not have suppressed interference.
[0041] In this embodiment, the multiple signal classification spectrogram, the current interference type, and the current characteristic parameters are determined as the current interference characteristic information, thereby facilitating a more comprehensive understanding of the characteristics of the interference signal. By utilizing the multiple signal classification algorithm, the source and characteristics of the interference signal can be accurately identified, providing richer information for subsequent interference suppression. This helps improve the accuracy of interference identification and the universality of the interference identification method, making it easier to cope with complex and changing electromagnetic environments.
[0042] S120. Input the current interference feature information into a pre-trained strategy determination model, and based on the output result of the strategy determination model, obtain the current anti-interference algorithm strategy for processing the current interference signal, so as to perform anti-interference processing on the current radar echo signal according to the current anti-interference algorithm strategy; wherein the strategy determination model includes a deep reinforcement learning model.
[0043] In a specific implementation, the current interference feature information includes a multiple signal classification spectrum, a current interference type, and current feature parameters. The multiple signal classification spectrum, the current interference type, and the current feature parameters can be respectively input into the trained deep reinforcement learning model. The hidden layer of the deep reinforcement learning model includes a convolutional neural network layer for extracting feature information of the multiple signal classification spectrum, and a fully connected layer for extracting feature information of the current interference type and the current feature parameters. The fully connected layer serves as a feature channel output layer to output the extracted features to the connection layer. The connection layer integrates different types of output feature information and transmits it to subsequent hidden layers for further feature extraction. The final output layer is a fully connected layer, which is used to output the output result of the strategy determination model. The output result includes the current anti-interference algorithm strategy for processing the current interference signal. Exemplarily, when there is pulse interference in the current interference signal, the current anti-interference algorithm strategy is a strategy for processing pulse interference.
[0044] In this embodiment, before inputting the current interference characteristic information into the pre-trained strategy determination model, it also includes: obtaining the current state information of the radar corresponding to the current radar echo signal; inputting the current interference characteristic information into the pre-trained strategy determination model, including: inputting the current state information and the current interference characteristic information into the pre-trained strategy determination model to obtain the next radar transmission parameters based on the output result of the strategy determination model, and in response to the transmission trigger operation of the next radar transmission signal, transmitting the next radar transmission signal according to the next radar transmission parameters.
[0045] The current status information includes at least one of the working mode, transmission carrier frequency, transmission bandwidth, radar antenna elevation angle and sector range.
[0046] To reduce the interference of radar state information during signal transmission on the return signal, the current radar state information corresponding to the current radar return signal can be obtained before the current interference characteristic information is input into the strategy determination model. Both the current state information and the current interference characteristic information are input into the strategy determination model to obtain an output result. The output result may include the current anti-interference algorithm strategy and radar transmission parameter adjustment information. The radar transmission parameter adjustment information can be used to adjust the current state information to obtain the next radar transmission parameters. The next radar transmission parameters are the corresponding radar state information before the next signal transmission. Alternatively, the output result may include the current anti-interference algorithm strategy and the next radar transmission parameters.
[0047] In a specific implementation, the next radar transmission parameter can be determined based on the output result. In a scenario where signals are continuously transmitted, in response to a trigger operation for transmitting the next radar transmission signal, the next transmission signal can be generated based on the adjustment of the next radar transmission parameter and transmitted. Furthermore, the current interference signal is processed using the determined current anti-interference algorithm strategy.
[0048] This embodiment uses both current interference signature information and current state information as input, providing a more comprehensive observation space state for the strategy determination model. This allows the strategy determination model to comprehensively consider both the current interference signature information and the current state information, thereby determining not only the current anti-interference algorithm strategy but also the next radar transmission parameters. This improves the adaptability and flexibility of the interference suppression strategy, enabling it to more effectively cope with complex and changing electromagnetic environments. Furthermore, it helps the radar quickly respond to new interference situations and optimizes radar signal transmission and reception, significantly enhancing the radar system's intelligence and dynamic adaptability.
[0049] The technical solution of the embodiment of the present invention receives the current radar echo signal and identifies the current interference feature information of the current interference signal in the current radar echo signal; by inputting the current interference feature information into a pre-trained strategy determination model, based on the output result of the strategy determination model, the current anti-interference algorithm strategy for processing the current interference signal is obtained; wherein, the strategy determination model includes a deep reinforcement learning model; through the strategy determination model, the current anti-interference algorithm strategy can be determined quickly and accurately without human intervention, and the current radar echo signal can be subjected to anti-interference processing according to the current anti-interference algorithm strategy, which is conducive to improving the efficiency and accuracy of anti-interference processing.
[0050] Figure 2 This is a flow chart of another signal processing method provided according to an embodiment of the present invention. This embodiment inputs sample interference feature information corresponding to sample radar echo signals into a deep reinforcement learning model to be trained before inputting the current interference feature information into a pre-trained strategy model to obtain a strategy determination model. Explanations of terms that are identical or corresponding to those in the above embodiments are omitted here.
[0051] like Figure 2 As shown, the method includes:
[0052] S210. Input the sample interference feature information corresponding to the sample radar echo signal into the deep reinforcement learning model to be trained for action decision-making; perform anti-interference processing on the sample radar echo signal based on the training anti-interference algorithm strategy output by the deep reinforcement learning model to be trained to obtain a processed echo signal; determine the target reward value corresponding to the training anti-interference algorithm strategy based on the processed echo signal and a preset target reward function.
[0053] It should be noted that the strategy determination model can be trained based on a jammer and electromagnetic environment simulation system. A radar simulation system generates sample radar transmit signals and transmits them. The jammer generates interference, causing the radar simulation system to receive the jammed sample radar echo signals.
[0054] In a specific implementation, sample interference signature information corresponding to a sample radar echo signal can be determined. This signature information includes a multiple signal classification spectrum, interference type, and characteristic parameters corresponding to the sample radar echo signal. This signature information can be input into a deep reinforcement learning model to be trained to make action decisions and output a trained anti-interference algorithm strategy corresponding to the sample radar echo signal. Anti-interference processing is then performed on the sample radar echo signal according to the trained anti-interference algorithm strategy, and the resulting signal is used as the processed echo signal.
[0055] In this embodiment, based on the processed echo signal, the sample radar echo signal and the preset target reward function, the target reward value corresponding to the training anti-interference algorithm strategy is determined, including: determining the detection target point information and signal receiving power corresponding to the processed echo signal; based on the detection target point information, the signal receiving power, the actual target point information of the predetermined sample radar echo signal, the interference signal power, the noise signal power and the target reward function, determining the target reward value corresponding to the training anti-interference algorithm strategy.
[0056] Among them, the detected target point information is the position information of the target object predicted by the processed echo signal, and the detected target point information includes the target point predicted trajectory position, target point predicted azimuth, target point predicted pitch angle and target predicted radial velocity. The actual target point information is the position information of the target object actually set when arranging the simulation scene, and the actual target point information includes the target point real trajectory position, target point real azimuth, target point real pitch angle and target real radial velocity. The target reward value is used to reflect the prediction accuracy of the detected target point information. The higher the target reward value, the higher the prediction accuracy of the detected target point information; conversely, the lower the prediction accuracy of the detected target point information. It should be noted that the target reward function can be set to different functions according to different tasks.
[0057] In a specific implementation, the jammer's jamming signal power and noise signal power can be determined based on the jamming signal generation parameters set by the jammer. The target reward value corresponding to the anti-interference algorithm training strategy is determined by detecting target point information, signal reception power, actual target point information of predetermined sample radar echo signals, jamming signal power, noise signal power, and a target reward function. Exemplarily, the target reward function includes at least one of a sparse reward, a dense reward, a shape reward, and a potential reward.
[0058] This embodiment provides an effective method for determining the target reward value, which can quickly and accurately determine the target reward value, thereby improving the accuracy of the model training results.
[0059] S220. Adjust model parameters in the deep reinforcement learning model to be trained based on the target reward value until the training is completed when preset training conditions are met, and determine the deep reinforcement learning model after the training as the policy determination model.
[0060] The preset training condition may be that the loss function converges or the number of training times reaches a preset threshold.
[0061] In a specific implementation, model parameters can be adjusted based on the determined target reward value, and the deep reinforcement learning model can be updated based on the adjusted model parameters, thereby maximizing the target reward value corresponding to the updated deep reinforcement learning model. Furthermore, each time the target reward value is determined, it can be determined whether a preset training condition is currently met, that is, whether the current number of training times has reached a preset number threshold; or whether the loss function corresponding to the trained deep learning model has converged. If the preset number threshold is reached or the loss function has converged, training can be determined to be complete, and the deep reinforcement learning model obtained after training can be determined as the policy determination model.
[0062] This embodiment adjusts training parameters by determining a target reward function, and obtains a policy determination model when preset training conditions are met, thereby facilitating the determination of an accurate and effective policy determination model.
[0063] Furthermore, the model parameters in the deep reinforcement learning model to be trained are adjusted based on the target reward value until the training is terminated when the preset training conditions are met, including: adjusting the current radar transmission parameters based on the training radar scheduling parameters output by the deep reinforcement learning model to be trained, and transmitting the sample transmission signal according to the adjusted current radar transmission parameters; updating the interference feature information of the received radar echo signal corresponding to the sample transmission signal to sample interference feature information, and inputting the updated sample interference feature information into the deep reinforcement learning model after adjusting the model parameters for action decision-making, until the training is terminated when the preset training conditions are met.
[0064] It should be noted that during the training process of the deep reinforcement learning model, it is necessary to continuously send sample transmission signals to obtain radar echo signals. By determining the sample interference characteristics of the radar echo signals, the model parameters of the deep reinforcement learning model are continuously adjusted. Therefore, the transmission parameters of the sample transmission signals need to be continuously adjusted to achieve the training of the deep reinforcement learning model.
[0065] Optionally, the deep reinforcement learning model can also output training radar scheduling parameters, which are the transmission parameters required to indicate the next radar transmission during the training process. The training radar scheduling parameters can be updated to the current radar transmission parameters to transmit a sample transmission signal according to the updated current radar transmission parameters. A radar echo signal corresponding to the sample transmission signal is received, and interference characteristic information of the radar echo signal corresponding to the sample transmission signal is extracted as sample interference characteristic information. The updated sample interference characteristic information is input into the deep reinforcement learning model after adjusting the model parameters to make an action decision, thereby obtaining a trained anti-interference algorithm strategy. The radar echo signal is subjected to anti-interference processing using the trained anti-interference algorithm strategy according to the above steps, and the model parameters of the deep reinforcement learning model are adjusted based on the processed echo signal, thereby completing the training of the deep reinforcement learning model. The current radar transmission parameters include at least one of the radar's operating mode number, the transmission carrier frequency number, the transmission bandwidth, the radar antenna pitch angle, and the sector range.
[0066] This embodiment adjusts the current radar transmission parameters through the training radar scheduling parameters output by the deep reinforcement learning model to continuously adjust the sample transmission signal and quickly and efficiently complete the training of the deep reinforcement learning model.
[0067] S230: Receive a current radar echo signal, and identify current interference feature information of a current interference signal in the current radar echo signal.
[0068] S240. Input the current interference feature information into a pre-trained strategy determination model, and based on the output result of the strategy determination model, obtain the current anti-interference algorithm strategy for processing the current interference signal, so as to perform anti-interference processing on the current radar echo signal according to the current anti-interference algorithm strategy; wherein the strategy determination model includes a deep reinforcement learning model.
[0069] This embodiment uses deep reinforcement learning and artificial neural networks to generate anti-interference strategies without human intervention, thereby improving the real-time performance of interference suppression decisions and improving the efficiency and accuracy of interference suppression.
[0070] The embodiments corresponding to the signal processing method are described in detail above. In order to make those skilled in the art further understand the technical solution of the present method, the training process of the strategy determination model required for signal processing is described in detail below.
[0071] Figure 3 This is a radar interference suppression training framework diagram based on deep reinforcement learning provided by an embodiment of the present invention. Figure 3As shown, the training process for a deep reinforcement learning model includes four parts: electromagnetic environment simulation, radar system simulation, deep reinforcement learning network, and reward function calculation. In a specific implementation, the radar simulation system generates a sample transmit signal and transmits it into a simulation environment containing an interference signal provided by a jammer and a preset target object. After reflection in the simulation environment, the sample transmit signal generates a sample radar echo signal, which is then received by the radar system. Both the sample transmit signal and the sample radar echo signal are transmitted via a channel. Furthermore, interference identification is performed on the sample radar echo signal to obtain sample interference signature information. This sample interference signature information is then sent to the deep reinforcement learning model to be trained, where the agent makes decisions and outputs a trained anti-interference algorithm strategy and trained radar scheduling parameters. The trained anti-interference algorithm strategy processes the sample radar echo signal to obtain a processed echo signal, which is then sent to the radar system for signal processing to obtain detected target point information and signal received power. The target reward value is determined based on the detected target point information, signal received power, actual target point information, interference signal power, and noise signal power. The detected target point information includes the target point predicted trajectory position, target point predicted azimuth, target point predicted pitch angle and target predicted radial velocity. The actual target point information includes the target point true trajectory position, target point true azimuth, target point true pitch angle and target true radial velocity.
[0072] Specifically, the specific process of determining the target reward value includes: 1. Determining the absolute value of the difference between the target point's predicted trajectory position and the target point's actual trajectory position as the point track distance error; 2. Determining the absolute value of the difference between the target point's predicted azimuth and the target point's actual azimuth as the point track azimuth error; 3. Determining the absolute value of the difference between the target point's predicted pitch angle and the target point's actual pitch angle as the point track pitch angle error; 4. Determining the absolute value of the difference between the target's predicted radial velocity and the target's actual radial velocity as the point track radial velocity error. Determine the signal-to-interference-plus-noise ratio (SINR). The calculation formula for the SINR is:
[0073] SINR = P1 / (P2+P3)
[0074] Among them, P1 is the signal receiving power, P2 is the interference signal power, and P3 is the noise signal power.
[0075] 5. Determine the monitoring point distance L between the detection target domain and the real target. The calculation formula of the monitoring point distance L is as follows:
[0076] L=μ1ΔR+μ2Δα+μ3Δθ+μ4Δυ
[0077] Where μ1 is the preset weight coefficient corresponding to the track range error, μ2 is the preset weight coefficient corresponding to the track azimuth error, μ3 is the preset weight coefficient corresponding to the track pitch error, and Δυ is the preset weight coefficient corresponding to the track radial velocity error. ΔR is the track range error, Δα is the track azimuth error, Δθ is the track pitch error, and Δυ is the track radial velocity error.
[0078] 6. By comparing the set distance threshold w with the distance L, we can get the true target capture judgment result logic (L <= w). This formula shows that when L is less than or equal to w, the logic value is 1; when L is greater than w, the logic value is 0. If no target is detected, the target reward value R wd = 0. If the target is detected, the target reward value R is determined based on the logic value, SINR, ΔR, Δα, Δθ and Δυ wd Specifically, R wd is determined as follows:
[0079] R wd =g(+logic(L <w),+SINR,-ΔR,-Δα,-Δθ,-Δυ)
[0080] Among them, "+" indicates positive correlation and "-" indicates negative correlation.
[0081] Furthermore, the model parameters of the deep reinforcement learning model are adjusted based on the determined target reward value. The model parameters in the deep reinforcement learning model are adjusted based on the obtained target reward value. Furthermore, the current radar transmission parameters are adjusted based on the training radar scheduling parameters output from the deep reinforcement learning model, and a radar scheduling instruction is generated so that the radar system performs radar scheduling operations based on the radar scheduling instruction, generates a sample transmission signal corresponding to the current radar transmission parameters, and repeats the above steps of signal transmission, signal reception, interference identification, and determining the target reward value until the training is completed when the preset training conditions are met. The deep reinforcement learning model after training is determined as the policy determination model.
[0082] This embodiment uses a deep reinforcement learning model to output the anti-interference algorithm strategy and radar scheduling instructions together, not only achieving passive interference suppression at the signal processing end, but also adjusting the transmission waveform parameters through active radar scheduling instructions to achieve active anti-interference measures. This approach can more comprehensively consider the dynamic changes of the radar system in a complex electromagnetic environment, and actively respond to interference by adjusting the radar's operating mode and parameters in real time, such as the transmission carrier frequency, transmission bandwidth, and radar antenna pitch angle. In addition, it can more effectively improve the adaptability and flexibility of the radar system when facing variable interference sources, and enhance its anti-interference capability in complex environments.
[0083] This embodiment also determines a target reward value and weights the signal-to-interference-and-noise ratio (SINR) and monitoring point distance, thereby integrating multi-dimensional evaluation indicators. By assigning corresponding weight coefficients to different indicators, the effectiveness of radar detection can be more comprehensively and objectively evaluated. The SINR reflects the relative strength of the signal and noise, while the monitoring point distance measures the proximity between the detection result and the actual target. By weighting these indicators, the accuracy and reliability of radar detection can be more accurately quantified, improving the radar system's adaptability to interference and anti-interference capabilities while also enhancing the real-time feedback and optimization capabilities of detection results.
[0084] Figure 4 1 is a schematic diagram of the structure of a signal processing device provided according to an embodiment of the present invention, which is used to execute the signal processing method provided in any of the above embodiments. The device and the signal processing methods of the above embodiments belong to the same inventive concept. For details not fully described in the embodiments of the signal processing device, please refer to the embodiments of the above signal processing methods. Figure 4 As shown, the device includes:
[0085] The signal receiving module 10 is used to receive the current radar echo signal and identify the current interference feature information of the current interference signal in the current radar echo signal;
[0086] The information input module 11 is used to input the current interference feature information into a pre-trained strategy determination model, and based on the output result of the strategy determination model, obtain the current anti-interference algorithm strategy for processing the current interference signal, so as to perform anti-interference processing on the current radar echo signal according to the current anti-interference algorithm strategy; wherein, the strategy determination model includes a deep reinforcement learning model.
[0087] Based on any optional technical solution in the embodiment of the present invention, optionally, the signal receiving module 10 includes:
[0088] The spectrum determination submodule is used to determine the corresponding multiple signal classification spectrum in the current radar echo signal based on the multiple signal classification algorithm;
[0089] a parameter identification submodule, configured to identify a current interference type and current characteristic parameters of a current interference signal in a current radar echo signal; wherein the current interference type includes at least one of a main-side lobe interference type, a suppression-forwarding interference type, and a pulse-continuous-wave interference type; and the current characteristic parameters include frequency and / or amplitude;
[0090] The current interference characteristic information determination submodule is used to determine the multiple signal classification spectrogram, the current interference type and the current characteristic parameters as the current interference characteristic information.
[0091] Based on any optional technical solution in the embodiments of the present invention, the following may be optionally included:
[0092] An information input module is used to input the sample interference feature information corresponding to the sample radar echo signal into the deep reinforcement learning model to be trained for action decision-making before inputting the current interference feature information into the pre-trained strategy determination model;
[0093] An anti-interference processing module is used to perform anti-interference processing on the sample radar echo signal based on the trained anti-interference algorithm strategy output by the deep reinforcement learning model to be trained, thereby obtaining a processed echo signal;
[0094] A target reward value determination module is used to determine the target reward value corresponding to the training anti-interference algorithm strategy based on the processed echo signal and a preset target reward function;
[0095] The strategy determination model determination module is used to adjust the model parameters in the deep reinforcement learning model to be trained based on the target reward value until the training is completed when the preset training conditions are met, and the deep reinforcement learning model after the training is determined as the strategy determination model.
[0096] Based on any optional technical solution in the embodiments of the present invention, optionally, the strategy determination model determination module includes:
[0097] A sample transmission signal transmission submodule is used to adjust the current radar transmission parameters based on the training radar scheduling parameters output by the deep reinforcement learning model to be trained, and transmit the sample transmission signal according to the adjusted current radar transmission parameters;
[0098] The sample interference feature information update submodule is used to update the interference feature information of the received radar echo signal corresponding to the sample transmission signal to the sample interference feature information, and input the updated sample interference feature information into the deep reinforcement learning model after adjusting the model parameters for action decision-making until the training is completed when the preset training conditions are met.
[0099] Based on any optional technical solution in the embodiments of the present invention, optionally, the target reward value determination module includes:
[0100] The target reward value determination submodule is used to determine the detection target point information and signal reception power corresponding to the processed echo signal; based on the detection target point information, signal reception power, the actual target point information of the predetermined sample radar echo signal, the interference signal power, the noise signal power and the target reward function, the target reward value corresponding to the training anti-interference algorithm strategy is determined.
[0101] Based on any optional technical solution in the embodiments of the present invention, the following may be optionally included:
[0102] A current state information acquisition module is used to obtain the current state information of the radar corresponding to the current radar echo signal before inputting the current interference feature information into the pre-trained strategy determination model;
[0103] The information input module 11 includes:
[0104] The transmission parameter acquisition submodule is used to input the current state information and the current interference characteristic information into the pre-trained strategy determination model to obtain the next radar transmission parameters based on the output results of the strategy determination model, and respond to the transmission trigger operation of the next radar transmission signal to transmit the next radar transmission signal according to the next radar transmission parameters.
[0105] Based on any optional technical solution in the embodiments of the present invention, optionally, the current status information includes at least one of an operating mode, a transmission carrier frequency, a transmission bandwidth, a radar antenna elevation angle, and a sector range.
[0106] The technical solution of the embodiment of the present invention receives the current radar echo signal and identifies the current interference feature information of the current interference signal in the current radar echo signal; by inputting the current interference feature information into a pre-trained strategy determination model, based on the output result of the strategy determination model, the current anti-interference algorithm strategy for processing the current interference signal is obtained; wherein, the strategy determination model includes a deep reinforcement learning model; through the strategy determination model, the current anti-interference algorithm strategy can be determined quickly and accurately without human intervention, and the current radar echo signal can be subjected to anti-interference processing according to the current anti-interference algorithm strategy, which is conducive to improving the efficiency and accuracy of anti-interference processing.
[0107] It is worth noting that in the embodiment of the above-mentioned signal processing device, the various units and modules included are only divided according to functional logic, but are not limited to the above-mentioned division, as long as the corresponding functions can be achieved; in addition, the specific names of the functional units are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of the present invention.
[0108] Figure 5 Schematic diagram of the structure of an electronic device that implements the signal processing method of an embodiment of the present invention. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are only examples and are not intended to limit the implementation of the present invention described and / or required herein.
[0109] like Figure 5 As shown, the electronic device 20 includes at least one processor 21, and a memory connected to the at least one processor 21, such as a read-only memory (ROM) 22, a random access memory (RAM) 23, etc., wherein the memory stores a computer program that can be executed by the at least one processor, and the processor 21 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 22 or the computer program loaded from the storage unit 28 to the random access memory (RAM) 23. Various programs and data required for the operation of the electronic device 20 can also be stored in the RAM 23. The processor 21, ROM 22 and RAM 23 are connected to each other via a bus 24. An input / output (I / O) interface 25 is also connected to the bus 24.
[0110] Multiple components in the electronic device 20 are connected to the I / O interface 25, including an input unit 26, such as a keyboard, a mouse, etc.; an output unit 27, such as various types of displays, speakers, etc.; a storage unit 28, such as a magnetic disk, an optical disk, etc.; and a communication unit 29, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 29 allows the electronic device 20 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0111] The processor 21 may be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the processor 21 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 21 performs the various methods and processes described above, such as the signal processing method.
[0112] In some embodiments, the signal processing method can be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as the storage unit 28. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 20 via the ROM 22 and / or the communication unit 29. When the computer program is loaded into the RAM 23 and executed by the processor 21, one or more steps of the signal processing method described above can be performed. Alternatively, in other embodiments, the processor 21 can be configured to perform the signal processing method in any other appropriate manner (for example, by means of firmware).
[0113] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0114] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer program is executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer program may be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0115] In the context of the present invention, computer-readable storage media can be tangible media that can contain or store a computer program for use with an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. Computer-readable storage media can include but are not limited to electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, computer-readable storage media can be machine-readable signal media. More specific examples of machine-readable storage media can include electrical connections based on one or more lines, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0116] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0117] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0118] A computing system may include clients and servers. The clients and servers are typically remote from each other and typically interact via a communication network. This client-server relationship arises through computer programs running on the respective computers, creating a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host. This server is a hosting product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosting and VPS services.
[0119] This embodiment also provides a computer program product, including a computer program, which, when executed by a processor, implements the signal processing method provided in any embodiment of the present application.
[0120] The computer program product may be implemented by writing computer program code for performing the operations of the present invention in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0121] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in the present invention can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved. This is not limited herein.
[0122] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.
Claims
1. A signal processing method, characterized in that: include: Receive a current radar echo signal and determine a corresponding multiple signal classification spectrum in the current radar echo signal based on a multiple signal classification algorithm; wherein the multiple signal classification spectrum includes received signal strength information at an angle corresponding to the radiation source arrival angle information; Identifying a current interference type and current characteristic parameters of a current interference signal in the current radar echo signal; wherein the current interference type includes at least one of a main-side lobe interference type, a suppressed forwarding interference type, and a pulse continuous wave interference type; and the current characteristic parameters include frequency and / or amplitude; Determining the multiple signal classification spectrogram, the current interference type, and the current characteristic parameter as the current interference characteristic information; Inputting the current interference feature information into a pre-trained strategy determination model, and obtaining a current anti-interference algorithm strategy for processing the current interference signal based on an output result of the strategy determination model, so as to perform anti-interference processing on the current radar echo signal according to the current anti-interference algorithm strategy; Wherein, the strategy determination model includes a deep reinforcement learning model.
2. The method according to claim 1, characterized in that Before inputting the current interference feature information into the pre-trained strategy determination model, the method further includes: The sample interference feature information corresponding to the sample radar echo signal is input into the deep reinforcement learning model to be trained for action decision-making; performing anti-interference processing on the sample radar echo signal based on the training anti-interference algorithm strategy output by the deep reinforcement learning model to be trained to obtain a processed echo signal; Determining a target reward value corresponding to the training anti-interference algorithm strategy based on the processed echo signal and a preset target reward function; The model parameters in the deep reinforcement learning model to be trained are adjusted based on the target reward value until the training is completed when a preset training condition is met, and the deep reinforcement learning model after the training is determined as the strategy determination model.
3. The method according to claim 2, characterized in that The adjusting the model parameters in the deep reinforcement learning model to be trained based on the target reward value until the training ends when a preset training condition is met includes: Adjusting current radar transmission parameters based on the training radar scheduling parameters output by the deep reinforcement learning model to be trained, and transmitting a sample transmission signal according to the adjusted current radar transmission parameters; The interference characteristic information of the received radar echo signal corresponding to the sample transmission signal is updated to the sample interference characteristic information, and the updated sample interference characteristic information is input into the deep reinforcement learning model after adjusting the model parameters for action decision-making until the training is completed when the preset training conditions are met.
4. The method according to claim 2, characterized in that The determining, based on the processed echo signal, the sample radar echo signal, and a preset target reward function, a target reward value corresponding to the training anti-interference algorithm strategy includes: Determining detection target point information and signal receiving power corresponding to the processed echo signal; Based on the detected target point information, the signal reception power, the predetermined actual target point information of the sample radar echo signal, the interference signal power, the noise signal power and the target reward function, the target reward value corresponding to the training anti-interference algorithm strategy is determined.
5. The method according to claim 1, wherein Before inputting the current interference feature information into the pre-trained strategy determination model, the method further includes: Acquiring current status information of the radar corresponding to the current radar echo signal; The inputting the current interference feature information into a pre-trained strategy determination model includes: The current state information and the current interference characteristic information are input into a pre-trained strategy determination model to obtain next radar transmission parameters based on an output result of the strategy determination model, and in response to a transmission trigger operation for the next radar transmission signal, the next radar transmission signal is transmitted according to the next radar transmission parameters.
6. The method according to claim 5, characterized in that The current status information includes at least one of a working mode, a transmission carrier frequency, a transmission bandwidth, a radar antenna elevation angle, and a sector range.
7. A signal processing device, characterized in that: include: a signal receiving module, configured to receive a current radar echo signal and identify current interference feature information of a current interference signal in the current radar echo signal; an information input module, configured to input the current interference feature information into a pre-trained strategy determination model, and based on an output result of the strategy determination model, obtain a current anti-interference algorithm strategy for processing the current interference signal, so as to perform anti-interference processing on the current radar echo signal according to the current anti-interference algorithm strategy; wherein the strategy determination model includes a deep reinforcement learning model; The signal receiving module includes: a spectrum determination submodule, configured to determine a corresponding multiple signal classification spectrum in the current radar echo signal based on a multiple signal classification algorithm; wherein the multiple signal classification spectrum includes received signal strength information at an angle corresponding to the radiation source arrival angle information; a parameter identification submodule, configured to identify a current interference type and current characteristic parameters of a current interference signal in the current radar echo signal; wherein the current interference type includes at least one of a main-side lobe interference type, a suppressed forwarding interference type, and a pulse continuous wave interference type; and the current characteristic parameters include frequency and / or amplitude; The current interference characteristic information determination submodule is used to determine the multiple signal classification spectrum, the current interference type and the current characteristic parameter as the current interference characteristic information.
8. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor. The computer program is executed by the at least one processor to enable the at least one processor to perform the signal processing method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the signal processing method according to any one of claims 1 to 6 when executed.
Citation Information
Patent Citations
Radar anti-interference strategy optimization method based on double-layer Q learning
CN115236607A