A method and device for simulating a target based on reinforcement learning for a monopulse radar
Through a reinforcement learning-based method, the detection radar is used to obtain position data and adjust the parameters of multi-point source signals, which solves the problem of safe landing of single-pulse radar in complex terrain, achieves a precise and safe landing process, reduces the risk of equipment damage and improves operational efficiency.
Patent Information
- Application Number
- CN202510058166.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2045-01-14
AI Technical Summary
In the existing technology, it is difficult for single-pulse radar to land safely under complex terrain conditions, and there is a lack of effective technical solutions.
A reinforcement learning-based method is used to obtain the spatial position data of the single-pulse radar in real time through the detection radar. The deep deterministic policy gradient algorithm is combined to generate a multi-point source signal control strategy, and the signal parameters are dynamically adjusted to enable the single-pulse radar to land safely and accurately in the predetermined area.
It achieved safe and precise landing of the monopulse radar, reduced the risk of equipment damage, simplified the operating process, and improved operational efficiency.
Smart Images

Figure CN119986580B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of signal processing, and in particular relates to a single-pulse radar target simulation method and device based on reinforcement learning. Background Art
[0002] With its exceptional accuracy and reliability, monopulse radar plays a key role in numerous fields, including search and rescue missions and weather monitoring. In particular, during natural disasters and emergencies, monopulse radar systems mounted on drones become crucial for search and rescue missions, playing a vital role in protecting the lives of those affected.
[0003] However, disaster-stricken areas often have irregular and complex terrain, including mountains, canyons, and ravines. The terrain can be steep and rugged, creating numerous uncertainties in the radar system's landing. In these situations, ensuring the safe landing of a monopulse radar system after the search and rescue mission is completed is extremely challenging. To date, there has been a lack of technical solutions for the safe landing of monopulse radars. Summary of the Invention
[0004] To this end, the present invention provides a single-pulse radar target simulation method and device based on reinforcement learning. By performing real-time analysis and processing on the single-pulse radar position data provided by the detection radar, the parameters of the multi-point source signal can be flexibly adjusted, so that the single-pulse radar continuously receives the model target signal until the single-pulse radar can land in the pre-set target area.
[0005] To achieve the above object, the present invention provides the following technical solution: a monopulse radar target simulation method based on reinforcement learning, comprising:
[0006] Call the detection module to obtain the spatial position data of the single pulse radar in real time through the detection radar;
[0007] Calling a control strategy generation module to generate a multi-point source signal control strategy based on the real-time spatial position data provided by the detection module and in combination with a reinforcement learning algorithm;
[0008] The multi-point source signal generation module is called to output the multi-point source signal according to the multi-point source signal control strategy generated by the control strategy generation module, and the output multi-point source signal is used to enable the monopulse radar to receive the simulated target signal.
[0009] As a preferred solution for the monopulse radar target simulation method based on reinforcement learning, in the process of calling the detection module to obtain the spatial position data of the monopulse radar in real time through the detection radar, the pointing angle θ of the monopulse radar after being controlled by the two-point source signal is s for:
[0010]
[0011] Where θ e is the angle between the two signal sources relative to the alignment axis of the monopulse radar; α is the amplitude ratio of the two signal sources; φ is the phase difference between the two signal sources.
[0012] As a preferred solution for the single-pulse radar target simulation method based on reinforcement learning, the control strategy generation module is called to generate a multi-point source signal control strategy based on the real-time spatial position data provided by the detection module and combined with the reinforcement learning algorithm. The reinforcement learning algorithm used is a deep deterministic policy gradient algorithm;
[0013] The reinforcement learning algorithm uses the amplitude ratio and phase difference of the signal source to determine the action of the driven intelligent agent.
[0014] As a preferred solution for the monopulse radar target simulation method based on reinforcement learning, the action of the intelligent agent is defined as: an adjustment strategy for the amplitude ratio in the pitch and azimuth directions of the monopulse radar pointing angle. The adjustment strategy formula is:
[0015] π=[π ε ,π ξ ]
[0016] Where, π ε is the azimuth adjustment strategy, π ξ is the pitch adjustment strategy, π ε and π ξ Any number between -1 and 1.
[0017] As a preferred solution for the monopulse radar target simulation method based on reinforcement learning, the action state of the agent is the three-dimensional position information of the monopulse radar:
[0018] s=[x,y,z]
[0019] Where s is the action state of the agent, and [x, y, z] is the three-dimensional position information of the monopulse radar.
[0020] As a preferred solution for the single-pulse radar target simulation method based on reinforcement learning, the reward r of the agent is:
[0021]
[0022] Wherein, R is the distance between the final landing point of the monopulse radar and the center point of the signal source, R1 and R2 are the shortest and farthest distances of the target area from the center point of the signal source, respectively. The values of R1 and R2 are determined according to the terrain. In the present invention, R1 is 1000m and R2 is 4000m.
[0023] The present invention also provides a single pulse radar target simulation device based on reinforcement learning, comprising:
[0024] The detection module is used to obtain the spatial position data of the single pulse radar in real time through the detection radar;
[0025] A control strategy generation module, configured to generate a multi-point source signal control strategy based on the real-time spatial position data provided by the detection module and in combination with a reinforcement learning algorithm;
[0026] The multi-point source signal generating module is used to output the multi-point source signal according to the multi-point source signal control strategy generated by the control strategy generating module, and use the output multi-point source signal to enable the monopulse radar to receive the simulated target signal.
[0027] As a preferred solution of the single pulse radar simulation target device based on reinforcement learning, in the detection module, the pointing angle θ of the single pulse radar after being controlled by the two point source signals is s for:
[0028]
[0029] Where θ e is the angle between the two signal sources relative to the alignment axis of the monopulse radar; α is the amplitude ratio of the two signal sources; φ is the phase difference between the two signal sources.
[0030] As a preferred solution for the monopulse radar target simulation device based on reinforcement learning, the reinforcement learning algorithm used in the control strategy generation module is a deep deterministic policy gradient algorithm; the reinforcement learning algorithm uses the amplitude ratio and phase difference of the signal source to determine the action of the driven intelligent agent;
[0031] In the control strategy generation module, the action of the intelligent agent is defined as: an adjustment strategy for the amplitude ratio in the pitch and azimuth directions of the monopulse radar pointing angle. The adjustment strategy formula is:
[0032] π=[π ε ,π ξ ]
[0033] Where, π ε is the azimuth adjustment strategy, π ξ is the pitch adjustment strategy, π ε and π ξ Any number between -1 and 1;
[0034] In the control strategy generation module, the action state of the agent is the three-dimensional position information of the monopulse radar:
[0035] s=[x,y,z]
[0036] Where s is the action state of the agent, and [x, y, z] is the three-dimensional position information of the monopulse radar.
[0037] As a preferred solution of the single pulse radar target simulation device based on reinforcement learning, in the multi-point source signal generation module, the reward r of the agent is:
[0038]
[0039] Wherein, R is the distance between the final landing point of the monopulse radar and the center point of the signal source, R1 and R2 are the shortest and farthest distances of the target area from the center point of the signal source, respectively. The values of R1 and R2 are determined according to the terrain. In the present invention, R1 is 1000m and R2 is 4000m.
[0040] The present invention has the following advantages: by utilizing the characteristic that monopulse radars are easily controlled by multi-point source signals, combined with a reinforcement learning algorithm, the parameters of the multi-point source signals are dynamically adjusted through real-time analysis of the monopulse radar position data provided by the detection radar, so that the monopulse radar receives the model target signal. This process will continue until the monopulse radar can accurately and safely land in the predetermined designated area, thereby ensuring the integrity of the equipment and the smooth progress of subsequent tasks. The present invention not only provides a feasible technical path for the safe landing of monopulse radars, simplifies the operational process of radar landing, and improves operational efficiency, but also significantly reduces the risk of damage to the equipment during landing, and also makes an important contribution to technological progress in related fields. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] To more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for the embodiments or the description of the prior art. Obviously, the drawings described below are merely exemplary, and those skilled in the art can, without inventive effort, derive other implementation drawings based on the provided drawings.
[0042] Figure 1 A schematic flow chart of a method for simulating a target using a monopulse radar based on reinforcement learning provided in an embodiment of the present invention;
[0043] Figure 2 Provides a logic diagram of a single-pulse radar target simulation method based on reinforcement learning in an embodiment of the present invention;
[0044] Figure 3 A schematic diagram of a monopulse radar being controlled by two point sources in a monopulse radar target simulation method based on reinforcement learning provided in an embodiment of the present invention;
[0045] Figure 4This is a simulation diagram of a successful landing result of a monopulse radar provided in an embodiment of the present invention;
[0046] Figure 5 This is a schematic diagram of the architecture of a single-pulse radar target simulation device based on reinforcement learning provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0047] The following describes the implementation of the present invention using specific embodiments. Those skilled in the art will readily understand the other advantages and benefits of the present invention from the disclosure herein. Obviously, the embodiments described are only a portion of the present invention, not all of it. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without inventive effort are intended to fall within the scope of protection of the present invention.
[0048] Example 1
[0049] See also Figure 1 and Figure 2 Embodiment 1 of the present invention provides a method for simulating a target using a monopulse radar based on reinforcement learning, comprising the following steps:
[0050] S1. Call the detection module to obtain the spatial position data of the single pulse radar in real time through the detection radar;
[0051] S2. Calling a control strategy generation module to generate a multi-point source signal control strategy based on the real-time spatial position data provided by the detection module and in combination with a reinforcement learning algorithm;
[0052] S3. Calling a multi-point source signal generating module to output a multi-point source signal according to the multi-point source signal control strategy generated by the control strategy generating module, and utilizing the output multi-point source signal to enable the monopulse radar to receive a simulated target signal.
[0053] In this embodiment, in step S1, the detection module utilizes a detection radar to specifically track and accurately obtain the spatial position information of the monopulse radar in real time, thereby ensuring continuous monitoring of the monopulse radar position during the entire landing process.
[0054] In this embodiment, in step S2, the control strategy generation module uses the real-time spatial position data provided by the detection module, combined with a reinforcement learning algorithm, to generate a multi-point source signal control strategy. The core of the multi-point source signal control strategy is to achieve precise control of the monopulse radar landing process to achieve a safe monopulse radar landing effect.
[0055] Among them, the core of the control strategy generation module is the reinforcement learning algorithm. The reinforcement learning algorithm of the present invention is the Deep Deterministic Policy Gradient algorithm (DDPG), which determines the actions, states and rewards of the intelligent agent in the reinforcement learning algorithm according to the purpose of the present invention.
[0056] See also Figure 3 , is a schematic diagram of a monopulse radar controlled by two point sources. After the monopulse radar is controlled by two point source signals, its pointing angle θ s The amplitude ratio α of the two signal sources, the phase difference φ, and the angle θ between the two signal sources relative to the alignment axis of the monopulse radar e Specifically, the pointing angle θ of the monopulse radar after being controlled by the two-point source signal s for:
[0057]
[0058] Where θ e is the angle between the two signal sources relative to the alignment axis of the monopulse radar; α is the amplitude ratio of the two signal sources; φ is the phase difference between the two signal sources.
[0059] As can be seen, the agent's actions driven by the reinforcement learning algorithm in the control strategy generation module should be closely related to the signal source's amplitude ratio α and phase difference φ. To ensure the monopulse radar's safe landing, the present invention focuses on precisely controlling the signal source's amplitude ratio α. Given that the monopulse radar's pointing angle involves two dimensions, pitch and azimuth, the agent's actions are ultimately defined as strategies for adjusting the amplitude ratio in both pitch and azimuth directions—a two-dimensional amplitude ratio adjustment strategy.
[0060] In one possible embodiment, the control strategy generation module generates a multi-point source signal control strategy based on the real-time spatial position data provided by the detection module and in combination with a reinforcement learning algorithm. The reinforcement learning algorithm used is a deep deterministic policy gradient algorithm; the reinforcement learning algorithm uses the amplitude ratio and phase difference of the signal source to determine the action of the driven intelligent agent. The action of the intelligent agent is defined as: the adjustment strategy of the amplitude ratio in the pitch and azimuth directions of the single pulse radar pointing angle. The adjustment strategy formula is:
[0061] π=[π ε ,π ξ ]
[0062] Where, π ε is the azimuth adjustment strategy, π ξ is the pitch adjustment strategy, π ε and π ξ Any number between -1 and 1.
[0063] Specifically, when a monopulse radar receives signals from multiple point sources, its pointing angle will change, causing the radar's spatial state to change. Therefore, it is appropriate to use the monopulse radar's three-dimensional position information (x, y, z) as the agent's state. The agent's action state is the monopulse radar's three-dimensional position information:
[0064] s=[x,y,z]
[0065] Where s is the action state of the agent, and [x, y, z] is the three-dimensional position information of the monopulse radar.
[0066] Specifically, the reward is an interactive means to guide the agent to take the desired action. Since the purpose of the present invention is to ensure that the monopulse radar can safely land in the designated area, the reward r of the agent is:
[0067]
[0068] Where R is the distance between the final landing point of the monopulse radar and the center point of the signal source, R1 and R2 are the pre-set closest and farthest distances between the target area and the center point of the signal source, respectively. The values of R1 and R2 are determined according to the terrain. Without loss of generality, in this embodiment, R1 is 1000 m and R2 is 4000 m.
[0069] In this embodiment, in step S3, the multi-point source signal generation module, based on the precise strategy developed by the control strategy generation module, is responsible for outputting corresponding multi-point source signals, enabling the monopulse radar to receive simulated target signals. These signals are crucial for the safe and accurate landing of the monopulse radar. Working together according to the established strategy, they guide the radar to a smooth landing in the predetermined landing area. In this way, the multi-point source signal generation module ensures the stability and safety of the monopulse radar throughout the landing process, thereby guaranteeing the integrity of the equipment and the smooth execution of subsequent missions.
[0070] See also Figure 4 In this embodiment, the method proposed in the embodiment of the present invention is verified by simulation experiments. The present invention uses a reinforcement learning algorithm to dynamically control the amplitude ratio parameters of the multi-point source signal based on the single pulse radar position information obtained by the detection radar, and ultimately can achieve the safe landing of the single pulse radar and reach the predetermined area ( Figure 4 The success rate of the proposed method is 91.73%.
[0071] In summary, the present invention calls a detection module to acquire the spatial position data of a monopulse radar in real time through a detection radar; calls a control strategy generation module to generate a multi-point source signal control strategy based on the real-time spatial position data provided by the detection module and in combination with a reinforcement learning algorithm; calls the multi-point source signal generation module to output a multi-point source signal based on the multi-point source signal control strategy generated by the control strategy generation module, and utilizes the output multi-point source signal to enable the monopulse radar to receive a simulated target signal. In the process of calling the control strategy generation module to generate the multi-point source signal control strategy based on the real-time spatial position data provided by the detection module and in combination with a reinforcement learning algorithm, the reinforcement learning algorithm used is a deep deterministic policy gradient algorithm; the reinforcement learning algorithm utilizes the amplitude ratio and phase difference of the signal source to determine the action of the driven intelligent agent. The present invention utilizes the characteristic that the monopulse radar is susceptible to multi-point source signal control and, in combination with a reinforcement learning algorithm, dynamically adjusts the parameters of the multi-point source signal by real-time analysis of the monopulse radar position data provided by the detection radar, enabling the monopulse radar to receive the model target signal. This process continues until the monopulse radar can accurately and safely land in the designated area, ensuring the integrity of the equipment and the smooth progress of subsequent missions. This invention not only provides a feasible technical path for the safe landing of monopulse radars, simplifies the radar landing process, and improves operational efficiency, but also significantly reduces the risk of damage to the equipment during landing, making an important contribution to technological advancement in related fields.
[0072] Example 2
[0073] See also Figure 5 Embodiment 2 of the present invention further provides a single-pulse radar target simulation device based on reinforcement learning, comprising:
[0074] The detection module 100 is used to obtain the spatial position data of the monopulse radar in real time through the detection radar;
[0075] A control strategy generation module 200 is configured to generate a multi-point source signal control strategy based on the real-time spatial position data provided by the detection module 100 and in combination with a reinforcement learning algorithm;
[0076] The multi-point source signal generating module 300 is used to output a multi-point source signal according to the multi-point source signal control strategy generated by the control strategy generating module 200, and use the output multi-point source signal to enable the monopulse radar to receive a simulated target signal.
[0077] In this embodiment, in the detection module 100, the pointing angle θ of the monopulse radar after being controlled by the two point source signals is s for:
[0078]
[0079] Where θ e is the angle between the two signal sources relative to the alignment axis of the monopulse radar; α is the amplitude ratio of the two signal sources; φ is the phase difference between the two signal sources.
[0080] In this embodiment, the control strategy generation module 200 uses a deep deterministic policy gradient algorithm as the reinforcement learning algorithm. The reinforcement learning algorithm uses the amplitude ratio and phase difference of the signal source to determine the action of the driven intelligent agent.
[0081] In the control strategy generation module 200, the action of the intelligent agent is defined as: an adjustment strategy for the amplitude ratio in the pitch and azimuth directions of the monopulse radar pointing angle. The adjustment strategy formula is:
[0082] π=[π ε ,π ξ ]
[0083] Where, π ε is the azimuth adjustment strategy, π ξ is the pitch adjustment strategy, π ε and π ξ Any number between -1 and 1;
[0084] In the control strategy generation module 200, the action state of the agent is the three-dimensional position information of the monopulse radar:
[0085] s=[x,y,z]
[0086] Where s is the action state of the agent, and [x, y, z] is the three-dimensional position information of the monopulse radar.
[0087] In this embodiment, in the multi-point source signal generation module 300, the reward r of the agent is:
[0088]
[0089] Wherein, R is the distance between the final landing point of the monopulse radar and the center point of the signal source, R1 and R2 are the shortest and farthest distances of the target area from the center point of the signal source, respectively. The values of R1 and R2 are determined according to the terrain. In the present invention, R1 is 1000m and R2 is 4000m.
[0090] It should be noted that the information interaction, execution process, etc. between the modules of the above-mentioned device are based on the same concept as the method embodiment in Example 1 of the present application, and the technical effects they bring are the same as those of the method embodiment of the present application. For specific contents, please refer to the description in the method embodiment shown above in the present application, and will not be repeated here.
[0091] Example 3
[0092] Embodiment 3 of the present invention provides a non-transitory computer-readable storage medium, in which the program code of a single-pulse radar target simulation method based on reinforcement learning is stored. The program code includes instructions for executing the single-pulse radar target simulation method based on reinforcement learning of embodiment 1 or any possible implementation thereof.
[0093] Computer-readable storage media can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more available media. The available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state drives (SSDs)).
[0094] Example 4
[0095] Embodiment 4 of the present invention provides an electronic device, including: a memory and a processor;
[0096] The processor and the memory communicate with each other via a bus; the memory stores program instructions that can be executed by the processor, and the processor calls the program instructions to execute the single-pulse radar target simulation method based on reinforcement learning of Example 1 or any possible implementation thereof.
[0097] Specifically, the processor can be implemented by hardware or by software. When implemented by hardware, the processor can be a logic circuit, an integrated circuit, etc.; when implemented by software, the processor can be a general-purpose processor, which is implemented by reading software code stored in a memory. The memory can be integrated into the processor or located outside the processor and exist independently.
[0098] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present invention is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, computer, server or data center to another website, computer, server or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode.
[0099] Obviously, those skilled in the art will appreciate that the various modules or steps of the present invention described above can be implemented using a general-purpose computing device, centralized on a single computing device, or distributed across a network of multiple computing devices. Alternatively, they can be implemented using program code executable by a computing device, which can then be stored in a storage device and executed by the computing device. In some cases, the steps shown or described can be performed in a different order than that shown, or can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, the present invention is not limited to any particular combination of hardware and software.
[0100] Although the present invention has been described in detail above using general descriptions and specific embodiments, it will be apparent to those skilled in the art that modifications and improvements may be made thereto. Therefore, such modifications and improvements, without departing from the spirit of the present invention, are intended to be within the scope of protection claimed herein.
Claims
1. A single pulse radar target simulation method based on reinforcement learning, characterized in that: include: Call the detection module to obtain the spatial position data of the single pulse radar in real time through the detection radar; Calling a control strategy generation module to generate a multi-point source signal control strategy based on the real-time spatial position data provided by the detection module and in combination with a reinforcement learning algorithm; Calling a multi-point source signal generation module to output a multi-point source signal according to the multi-point source signal control strategy generated by the control strategy generation module, and using the output multi-point source signal to enable the monopulse radar to receive a simulated target signal; When a monopulse radar receives signals from multiple point sources, its pointing angle will change, causing the spatial state of the radar to change.
2. A method for simulating targets using a single pulse radar based on reinforcement learning according to claim 1, characterized in that: In the process of calling the detection module to obtain the spatial position data of the monopulse radar in real time through the detection radar, the pointing angle θ of the monopulse radar after being controlled by the two-point source signal is s for: Where θ e is the angle between the two signal sources relative to the alignment axis of the monopulse radar; α is the amplitude ratio of the two signal sources; φ is the phase difference between the two signal sources.
3. A method for simulating targets using a single pulse radar based on reinforcement learning according to claim 2, characterized in that: Invoking a control strategy generation module to generate a multi-point source signal control strategy based on the real-time spatial position data provided by the detection module and combining it with a reinforcement learning algorithm, the reinforcement learning algorithm used is a deep deterministic policy gradient algorithm; The reinforcement learning algorithm uses the amplitude ratio and phase difference of the signal source to determine the action of the driven intelligent agent.
4. A method for simulating targets using a single pulse radar based on reinforcement learning according to claim 3, characterized in that: The action of the agent is defined as: the adjustment strategy of the amplitude ratio in the pitch and azimuth directions of the monopulse radar pointing angle. The formula of the adjustment strategy is: π=[π ε ,π ξ ] Where, π ε is the azimuth adjustment strategy, π ξ is the pitch adjustment strategy, π ε and π ξ Any number between -1 and 1.
5. The method for simulating a target using a single pulse radar based on reinforcement learning according to claim 4, wherein: The action state of the agent is the three-dimensional position information of the monopulse radar: s=[x,y,z] Where s is the action state of the agent, and [x, y, z] is the three-dimensional position information of the monopulse radar.
6. A method for simulating targets using a single pulse radar based on reinforcement learning according to claim 5, characterized in that: The agent's reward r is: Where R is the distance between the final landing point of the monopulse radar and the center point of the signal source, R1 and R2 are the shortest and farthest distances between the target area and the center point of the signal source, respectively. The values of R1 and R2 are determined according to the terrain.
7. A single pulse radar target simulation device based on reinforcement learning, characterized in that: include: The detection module is used to obtain the spatial position data of the monopulse radar in real time through the detection radar; A control strategy generation module, configured to generate a multi-point source signal control strategy based on the real-time spatial position data provided by the detection module and in combination with a reinforcement learning algorithm; a multi-point source signal generating module, configured to output a multi-point source signal according to the multi-point source signal control strategy generated by the control strategy generating module, and enable the monopulse radar to receive a simulated target signal using the output multi-point source signal; When a monopulse radar receives signals from multiple point sources, its pointing angle will change, causing the spatial state of the radar to change.
8. A single pulse radar target simulation device based on reinforcement learning according to claim 7, characterized in that: In the detection module, the pointing angle θ of the monopulse radar after being controlled by two point source signals is s for: Where θ e is the angle between the two signal sources relative to the alignment axis of the monopulse radar; α is the amplitude ratio of the two signal sources; φ is the phase difference between the two signal sources.
9. A single pulse radar target simulation device based on reinforcement learning according to claim 8, characterized in that: In the control strategy generation module, the reinforcement learning algorithm used is a deep deterministic policy gradient algorithm; the reinforcement learning algorithm uses the amplitude ratio and phase difference of the signal source to determine the action of the driven intelligent agent; In the control strategy generation module, the action of the intelligent agent is defined as: an adjustment strategy for the amplitude ratio in the pitch and azimuth directions of the monopulse radar pointing angle. The adjustment strategy formula is: π=[π ε ,π ξ ] Where, π ε is the azimuth adjustment strategy, π ξ is the pitch adjustment strategy, π ε and π ξ Any number between -1 and 1; In the control strategy generation module, the action state of the agent is the three-dimensional position information of the monopulse radar: s=[x,y,z] Where s is the action state of the agent, and [x, y, z] is the three-dimensional position information of the monopulse radar.
10. A single pulse radar target simulation device based on reinforcement learning according to claim 9, characterized in that: In the multi-point source signal generation module, the agent's reward r is: Where R is the distance between the final landing point of the monopulse radar and the center point of the signal source, R1 and R2 are the shortest and farthest distances between the target area and the center point of the signal source, respectively. The values of R1 and R2 are determined according to the terrain.
Citation Information
Patent Citations
Towed bait interference alarm method based on monopulse angle measurement
CN118625264A
Method of radar protection against antiradar missile based on use of additional radiation source with a lift-type horn aerial
RU2287168C1