Monopulse radar target simulation method and device based on reinforcement learning
Through a reinforcement learning-based method, the real-time position data of the detection radar and the depth deterministic strategy gradient algorithm are used to generate a multi-point source signal control strategy, so that single-pulse radar can achieve safe landing in complex terrain, solving the problem that single-pulse radar system is difficult to safely land, and improving landing efficiency and equipment safety.
Patent Information
- Application Number
- CN202510058166.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-14
AI Technical Summary
In complex terrain, single-pulse radar systems are difficult to land safely, and there is a lack of effective technical solutions.
Using reinforcement learning-based method, the spatial position data of single-pulse radar is obtained in real time through detection radar, and a multi-point source signal control strategy is generated in combination with the depth deterministic strategy gradient algorithm, so that the single-pulse radar can receive analog target signals and achieve safe landing.
It realizes accurate and safe landing of single-pulse radar, ensures that the equipment is intact, simplifies the operation process of radar landing, and improves operating efficiency.
Smart Images

Figure CN119986580A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of signal processing, and in particular relates to a single pulse radar target simulation method and device based on reinforcement learning. Background Art
[0002] With its excellent accuracy and reliability, monopulse radar plays a key role in many fields, including search and rescue missions and weather monitoring. Especially in the face of natural disasters or emergencies, the monopulse radar system carried by drones becomes a key equipment for performing search and rescue missions, and it plays a vital role in protecting the lives of people affected by disasters.
[0003] However, the terrain of disaster-stricken areas is usually irregular and complex, such as mountains, canyons, gullies, etc. The terrain may be steep and rugged, which brings many uncertainties to the landing of the radar system. In this case, it is very difficult to ensure that the monopulse radar system lands safely on the ground after the search and rescue mission is completed. So far, there is a lack of technical solutions for the safe landing of monopulse radars. Summary of the invention
[0004] To this end, the present invention provides a single-pulse radar target simulation method and device based on reinforcement learning. By performing real-time analysis and processing on the single-pulse radar position data provided by the detection radar, the parameters of the multi-point source signal can be flexibly adjusted so that the single-pulse radar continues to receive the model target signal until the single-pulse radar can land in a pre-set target area.
[0005] In order to achieve the above object, the present invention provides the following technical solution: a single pulse radar target simulation method based on reinforcement learning, comprising:
[0006] Call the detection module to obtain the spatial position data of the single pulse radar in real time through the detection radar;
[0007] Calling a control strategy generation module to generate a multi-point source signal control strategy based on the real-time spatial position data provided by the detection module and in combination with a reinforcement learning algorithm;
[0008] The multi-point source signal generation module is called to output the multi-point source signal according to the multi-point source signal control strategy generated by the control strategy generation module, and the output multi-point source signal is used to enable the single pulse radar to receive the simulated target signal.
[0009] As a preferred solution of the single pulse radar simulation target method based on reinforcement learning, in the process of calling the detection module to obtain the spatial position data of the single pulse radar in real time through the detection radar, the pointing angle θ of the single pulse radar after being controlled by the two-point source signal s for:
[0010]
[0011] In the formula, θ e is the angle between the two signal sources relative to the alignment axis of the monopulse radar; α is the amplitude ratio of the two signal sources; φ is the phase difference between the two signal sources.
[0012] As a preferred solution of the single pulse radar target simulation method based on reinforcement learning, the control strategy generation module is called according to the real-time spatial position data provided by the detection module, and in the process of generating the multi-point source signal control strategy in combination with the reinforcement learning algorithm, the reinforcement learning algorithm used is the deep deterministic policy gradient algorithm;
[0013] The reinforcement learning algorithm uses the amplitude ratio and phase difference of the signal source to determine the action of the driven intelligent agent.
[0014] As a preferred solution of the monopulse radar target simulation method based on reinforcement learning, the action of the intelligent agent is defined as: an adjustment strategy for the amplitude ratio in the pitch and azimuth directions of the monopulse radar pointing angle, and the formula of the adjustment strategy is:
[0015] π=[π ε ,π ξ ]
[0016] In the formula, π ε is the azimuth adjustment strategy, π ξ is the pitch adjustment strategy, π ε and π ξ Any number between -1 and 1.
[0017] As a preferred solution of the monopulse radar target simulation method based on reinforcement learning, the action state of the agent is the three-dimensional position information of the monopulse radar:
[0018] s=[x,y,z]
[0019] Where s is the action state of the agent, and [x, y, z] is the three-dimensional position information of the single pulse radar.
[0020] As a preferred solution of the single pulse radar target simulation method based on reinforcement learning, the reward r of the agent is:
[0021]
[0022] Wherein, R is the distance between the final landing point of the single pulse radar and the center point of the signal source, R1 and R2 are respectively the shortest and farthest distances between the target area and the center point of the signal source set in advance, and the values of R1 and R2 are determined according to the terrain. In the present invention, R1 is 1000m and R2 is 4000m.
[0023] The present invention also provides a single pulse radar target simulation device based on reinforcement learning, comprising:
[0024] A detection module, used to obtain the spatial position data of the single pulse radar in real time through the detection radar;
[0025] A control strategy generation module, used to generate a multi-point source signal control strategy based on the real-time spatial position data provided by the detection module and in combination with a reinforcement learning algorithm;
[0026] The multi-point source signal generating module is used to output the multi-point source signal according to the multi-point source signal control strategy generated by the control strategy generating module, and utilize the output multi-point source signal to enable the single pulse radar to receive the simulated target signal.
[0027] As a preferred solution of the single pulse radar simulation target device based on reinforcement learning, in the detection module, the pointing angle θ of the single pulse radar after being controlled by the two point source signals is s for:
[0028]
[0029] In the formula, θ e is the angle between the two signal sources relative to the alignment axis of the monopulse radar; α is the amplitude ratio of the two signal sources; φ is the phase difference between the two signal sources.
[0030] As a preferred solution of the single pulse radar simulation target device based on reinforcement learning, the reinforcement learning algorithm used in the control strategy generation module is a deep deterministic policy gradient algorithm; the reinforcement learning algorithm uses the amplitude ratio and phase difference of the signal source to determine the action of the driven intelligent body;
[0031] In the control strategy generation module, the action of the intelligent agent is defined as: the adjustment strategy of the amplitude ratio in the pitch and azimuth directions of the single pulse radar pointing angle, and the formula of the adjustment strategy is:
[0032] π=[π ε ,π ξ ]
[0033] In the formula, π ε is the azimuth adjustment strategy, π ξ is the pitch adjustment strategy, π ε and π ξ Any number between -1 and 1;
[0034] In the control strategy generation module, the action state of the agent is the three-dimensional position information of the monopulse radar:
[0035] s=[x,y,z]
[0036] Where s is the action state of the agent, and [x, y, z] is the three-dimensional position information of the single pulse radar.
[0037] As a preferred solution of the single pulse radar simulation target device based on reinforcement learning, in the multi-point source signal generation module, the reward r of the agent is:
[0038]
[0039] Wherein, R is the distance between the final landing point of the single pulse radar and the center point of the signal source, R1 and R2 are respectively the shortest and farthest distances between the target area and the center point of the signal source set in advance, and the values of R1 and R2 are determined according to the terrain. In the present invention, R1 is 1000m and R2 is 4000m.
[0040] The present invention has the following advantages: by utilizing the characteristic that the monopulse radar is easily controlled by multi-point source signals, combined with a reinforcement learning algorithm, the parameters of the multi-point source signals are dynamically adjusted by real-time analysis of the monopulse radar position data provided by the detection radar, so that the monopulse radar receives the model target signal. This process will continue until the monopulse radar can accurately and safely land in the predetermined designated area, thereby ensuring the integrity of the equipment and the smooth progress of subsequent tasks. The present invention not only provides a feasible technical path for the safe landing of the monopulse radar, simplifies the operational process of radar landing, and improves operating efficiency, but also significantly reduces the risk of damage that the equipment may suffer during the landing process, and also makes important contributions to the technological progress in related fields. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] In order to more clearly illustrate the implementation methods of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for the implementation methods or the description of the prior art. Obviously, the drawings in the following description are only exemplary, and for ordinary technicians in this field, other implementation drawings can be derived from the provided drawings without creative work.
[0042] Figure 1 A schematic diagram of a flow chart of a single pulse radar target simulation method based on reinforcement learning provided in an embodiment of the present invention;
[0043] Figure 2 A logical schematic diagram of a single pulse radar target simulation method based on reinforcement learning is provided in an embodiment of the present invention;
[0044] Figure 3 A schematic diagram of a monopulse radar being controlled by two point sources in a monopulse radar target simulation method based on reinforcement learning provided in an embodiment of the present invention;
[0045] Figure 4A simulation diagram of a successful landing result of a monopulse radar provided in an embodiment of the present invention;
[0046] Figure 5 Schematic diagram of the architecture of a single pulse radar target simulation device based on reinforcement learning provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0047] The following is a description of the implementation of the present invention by specific embodiments. People familiar with the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0048] Example 1
[0049] See also Figure 1 and Figure 2 Embodiment 1 of the present invention provides a single pulse radar target simulation method based on reinforcement learning, comprising the following steps:
[0050] S1, calling the detection module to obtain the spatial position data of the single pulse radar in real time through the detection radar;
[0051] S2, calling the control strategy generation module to generate a multi-point source signal control strategy based on the real-time spatial position data provided by the detection module and combining it with a reinforcement learning algorithm;
[0052] S3, calling the multi-point source signal generation module to output the multi-point source signal according to the multi-point source signal control strategy generated by the control strategy generation module, and using the output multi-point source signal to enable the single pulse radar to receive the simulated target signal.
[0053] In this embodiment, in step S1, the detection module uses a detection radar to track and accurately obtain the spatial position information of the single pulse radar in real time, ensuring continuous monitoring of the position of the single pulse radar during the entire landing process.
[0054] In this embodiment, in step S2, the control strategy generation module uses the real-time spatial position data provided by the detection module and combines it with the reinforcement learning algorithm to generate a multi-point source signal control strategy. The core of the multi-point source signal control strategy is to achieve a detailed regulation of the single pulse radar landing process to achieve a single pulse radar safe landing effect.
[0055] Among them, the core of the control strategy generation module is the reinforcement learning algorithm. The reinforcement learning algorithm of the present invention is the Deep Deterministic Policy Gradient (DDPG) algorithm. According to the purpose of the present invention, the action, state and reward of the intelligent agent in the reinforcement learning algorithm are determined.
[0056] See also Figure 3 , is a schematic diagram of a monopulse radar controlled by two point sources. After the monopulse radar is controlled by two point source signals, its pointing angle θ s The amplitude ratio α of the two signal sources, the phase difference φ, and the angle θ between the two signal sources relative to the alignment axis of the monopulse radar e Specifically, the pointing angle θ of the monopulse radar after being controlled by the two-point source signal s for:
[0057]
[0058] In the formula, θ e is the angle between the two signal sources relative to the alignment axis of the monopulse radar; α is the amplitude ratio of the two signal sources; φ is the phase difference between the two signal sources.
[0059] It can be seen that the action of the intelligent agent driven by the reinforcement learning algorithm in the control strategy generation module should be closely related to the amplitude ratio α and phase difference φ of the signal source. In order to ensure that the monopulse radar can land safely, the present invention focuses on the precise control of the amplitude ratio α of the signal source. In view of the fact that the pointing angle of the monopulse radar involves two dimensions, pitch and azimuth, the action of the intelligent agent of the present invention is ultimately defined as an adjustment strategy for the amplitude ratio in the pitch and azimuth directions, that is, a two-dimensional amplitude ratio adjustment strategy.
[0060] In a possible embodiment, the control strategy generation module is called to generate a multi-point source signal control strategy based on the real-time spatial position data provided by the detection module and combined with the reinforcement learning algorithm. The reinforcement learning algorithm used is a deep deterministic policy gradient algorithm; the reinforcement learning algorithm uses the amplitude ratio and phase difference of the signal source to determine the action of the driven intelligent agent. The action of the intelligent agent is defined as: the adjustment strategy of the amplitude ratio in the pitch and azimuth directions of the single pulse radar pointing angle, and the formula of the adjustment strategy is:
[0061] π=[π ε ,π ξ ]
[0062] In the formula, π ε is the azimuth adjustment strategy, π ξ is the pitch adjustment strategy, π ε and π ξ Any number between -1 and 1.
[0063] Specifically, when a monopulse radar receives signals from multiple point sources, its pointing angle will change, which will cause the spatial state of the radar to change. In view of this, it is appropriate to use the three-dimensional position information (x, y, z) of the monopulse radar as the state of the agent. The action state of the agent is the three-dimensional position information of the monopulse radar:
[0064] s=[x,y,z]
[0065] Where s is the action state of the agent, and [x, y, z] is the three-dimensional position information of the single pulse radar.
[0066] Specifically, the reward is an interactive means to guide the agent to take the desired action. Since the purpose of the present invention is to ensure that the monopulse radar can land safely in the designated area, the reward r of the agent is:
[0067]
[0068] Wherein, R is the distance between the final landing point of the single pulse radar and the center point of the signal source, R1 and R2 are respectively the shortest and farthest distances between the target area and the center point of the signal source set in advance, and the values of R1 and R2 are determined according to the terrain without loss of generality. In this embodiment, R1 is 1000m and R2 is 4000m.
[0069] In this embodiment, in step S3, the multi-point source signal generation module is responsible for outputting the corresponding multi-point source signal according to the precise strategy formulated by the control strategy generation module, so that the monopulse radar receives the simulated target signal. These signals are the key to the safe and accurate landing of the monopulse radar. They work together according to the established strategy to guide the radar to land smoothly in the predetermined landing area. In this way, the multi-point source signal generation module ensures the stability and safety of the monopulse radar during the entire landing process, thereby ensuring the integrity of the equipment and the smooth execution of subsequent tasks.
[0070] See also Figure 4 In this embodiment, the method proposed in the embodiment of the present invention is verified by simulation experiments. The present invention dynamically controls the amplitude ratio parameters of the multi-point source signal according to the single pulse radar position information obtained by the detection radar through the reinforcement learning algorithm, and finally achieves the safe landing of the single pulse radar and reaches the predetermined area ( Figure 4 The success rate of the proposed method is 91.73%.
[0071] In summary, the present invention calls the detection module to obtain the spatial position data of the single pulse radar in real time through the detection radar; calls the control strategy generation module to generate a multi-point source signal control strategy based on the real-time spatial position data provided by the detection module and in combination with the reinforcement learning algorithm; calls the multi-point source signal generation module to output a multi-point source signal according to the multi-point source signal control strategy generated by the control strategy generation module, and uses the output multi-point source signal to enable the single pulse radar to receive a simulated target signal. In the process of calling the control strategy generation module to generate a multi-point source signal control strategy based on the real-time spatial position data provided by the detection module and in combination with the reinforcement learning algorithm, the reinforcement learning algorithm used is a deep deterministic policy gradient algorithm; the reinforcement learning algorithm uses the amplitude ratio and phase difference of the signal source to determine the action of the driven intelligent body. The present invention utilizes the characteristic that the single pulse radar is easily controlled by the multi-point source signal, combines the reinforcement learning algorithm, and dynamically adjusts the parameters of the multi-point source signal by real-time analysis of the single pulse radar position data provided by the detection radar, so that the single pulse radar receives the model target signal. This process will continue until the monopulse radar can accurately and safely land in the predetermined designated area, thus ensuring the integrity of the equipment and the smooth progress of subsequent tasks. The present invention not only provides a feasible technical path for the safe landing of the monopulse radar, simplifies the operation process of radar landing, and improves the operation efficiency, but also significantly reduces the risk of damage that the equipment may suffer during the landing process, and also makes an important contribution to the technological progress in related fields.
[0072] Example 2
[0073] See also Figure 5 Embodiment 2 of the present invention further provides a single pulse radar target simulation device based on reinforcement learning, comprising:
[0074] The detection module 100 is used to obtain the spatial position data of the monopulse radar in real time through the detection radar;
[0075] A control strategy generation module 200, for generating a multi-point source signal control strategy based on the real-time spatial position data provided by the detection module 100 and in combination with a reinforcement learning algorithm;
[0076] The multi-point source signal generating module 300 is used to output the multi-point source signal according to the multi-point source signal control strategy generated by the control strategy generating module 200, and use the output multi-point source signal to enable the single pulse radar to receive the simulated target signal.
[0077] In this embodiment, in the detection module 100, the pointing angle θ of the single pulse radar after being controlled by the two point source signals is s for:
[0078]
[0079] In the formula, θ e is the angle between the two signal sources relative to the alignment axis of the monopulse radar; α is the amplitude ratio of the two signal sources; φ is the phase difference between the two signal sources.
[0080] In this embodiment, the reinforcement learning algorithm used in the control strategy generation module 200 is a deep deterministic policy gradient algorithm; the reinforcement learning algorithm uses the amplitude ratio and phase difference of the signal source to determine the action of the driven intelligent agent;
[0081] In the control strategy generation module 200, the action of the intelligent agent is defined as: an adjustment strategy for the amplitude ratio in the pitch and azimuth directions of the monopulse radar pointing angle, and the formula for the adjustment strategy is:
[0082] π=[π ε ,π ξ ]
[0083] In the formula, π ε is the azimuth adjustment strategy, π ξ is the pitch adjustment strategy, π ε and π ξ Any number between -1 and 1;
[0084] In the control strategy generation module 200, the action state of the agent is the three-dimensional position information of the monopulse radar:
[0085] s=[x,y,z]
[0086] Where s is the action state of the agent, and [x, y, z] is the three-dimensional position information of the single pulse radar.
[0087] In this embodiment, in the multi-point source signal generation module 300, the reward r of the agent is:
[0088]
[0089] Wherein, R is the distance between the final landing point of the single pulse radar and the center point of the signal source, R1 and R2 are respectively the shortest and farthest distances between the target area and the center point of the signal source set in advance, and the values of R1 and R2 are determined according to the terrain. In the present invention, R1 is 1000m and R2 is 4000m.
[0090] It should be noted that the information interaction, execution process and other contents between the modules of the above-mentioned device are based on the same concept as the method embodiment in Example 1 of the present application, and the technical effects they bring are the same as those of the method embodiment of the present application. For specific contents, please refer to the description in the method embodiment shown in the previous part of the present application, and will not be repeated here.
[0091] Example 3
[0092] Embodiment 3 of the present invention provides a non-transitory computer-readable storage medium, in which a program code of a single-pulse radar target simulation method based on reinforcement learning is stored, and the program code includes instructions for executing the single-pulse radar target simulation method based on reinforcement learning of embodiment 1 or any possible implementation method thereof.
[0093] The computer-readable storage medium may be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that includes one or more available media. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)).
[0094] Example 4
[0095] Embodiment 4 of the present invention provides an electronic device, including: a memory and a processor;
[0096] The processor and the memory communicate with each other via a bus; the memory stores program instructions that can be executed by the processor, and the processor calls the program instructions to execute the single-pulse radar target simulation method based on reinforcement learning of Example 1 or any possible implementation thereof.
[0097] Specifically, the processor can be implemented by hardware or by software. When implemented by hardware, the processor can be a logic circuit, an integrated circuit, etc.; when implemented by software, the processor can be a general-purpose processor implemented by reading software codes stored in a memory. The memory can be integrated in the processor or can be located outside the processor and exist independently.
[0098] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present invention is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium, for example, the computer instructions can be transmitted from a website site, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website site, computer, server or data center.
[0099] Obviously, those skilled in the art should understand that the above modules or steps of the present invention can be implemented by a general computing device, they can be concentrated on a single computing device, or distributed on a network composed of multiple computing devices, and optionally, they can be implemented by a program code executable by a computing device, so that they can be stored in a storage device and executed by the computing device, and in some cases, the steps shown or described can be executed in a different order than here, or they can be made into individual integrated circuit modules, or multiple modules or steps therein can be made into a single integrated circuit module for implementation. Thus, the present invention is not limited to any specific combination of hardware and software.
[0100] Although the present invention has been described in detail above by general description and specific embodiments, it is obvious to those skilled in the art that some modifications or improvements can be made to the present invention. Therefore, these modifications or improvements made without departing from the spirit of the present invention all belong to the scope of protection claimed by the present invention.
Claims
1. A single pulse radar target simulation method based on reinforcement learning, characterized in that: include: Call the detection module to obtain the spatial position data of the single pulse radar in real time through the detection radar; Calling a control strategy generation module to generate a multi-point source signal control strategy based on the real-time spatial position data provided by the detection module and in combination with a reinforcement learning algorithm; The multi-point source signal generation module is called to output the multi-point source signal according to the multi-point source signal control strategy generated by the control strategy generation module, and the output multi-point source signal is used to enable the single pulse radar to receive the simulated target signal.
2. A single pulse radar target simulation method based on reinforcement learning according to claim 1, characterized in that: In the process of calling the detection module to obtain the spatial position data of the single pulse radar in real time through the detection radar, the pointing angle θ of the single pulse radar after being controlled by the two-point source signal s for: In the formula, θ e is the angle between the two signal sources relative to the alignment axis of the monopulse radar; α is the amplitude ratio of the two signal sources; φ is the phase difference between the two signal sources.
3. A single pulse radar target simulation method based on reinforcement learning according to claim 2, characterized in that: In the process of calling the control strategy generation module to generate the multi-point source signal control strategy according to the real-time spatial position data provided by the detection module and combining the reinforcement learning algorithm, the reinforcement learning algorithm used is a deep deterministic policy gradient algorithm; The reinforcement learning algorithm uses the amplitude ratio and phase difference of the signal source to determine the action of the driven intelligent agent.
4. A single pulse radar target simulation method based on reinforcement learning according to claim 3, characterized in that: The action of the agent is defined as: the adjustment strategy of the amplitude ratio in the pitch and azimuth directions of the monopulse radar pointing angle, and the formula of the adjustment strategy is: π=[π ε ,π ξ ] In the formula, π ε is the azimuth adjustment strategy, π ξ is the pitch adjustment strategy, π ε and π ξ Any number between -1 and 1.
5. A single pulse radar target simulation method based on reinforcement learning according to claim 4, characterized in that: The action state of the agent is the three-dimensional position information of the monopulse radar: s=[x,y,z] Where s is the action state of the agent, and [x, y, z] is the three-dimensional position information of the single pulse radar.
6. A single pulse radar target simulation method based on reinforcement learning according to claim 5, characterized in that: The reward r of the agent is: Where R is the distance between the final landing point of the single pulse radar and the center point of the signal source, R1 and R2 are the shortest and farthest distances between the target area and the center point of the signal source, respectively. The values of R1 and R2 are determined according to the terrain.
7. A single pulse radar target simulation device based on reinforcement learning, characterized in that: include: A detection module, used to obtain the spatial position data of the single pulse radar in real time through the detection radar; A control strategy generation module, used to generate a multi-point source signal control strategy based on the real-time spatial position data provided by the detection module and in combination with a reinforcement learning algorithm; The multi-point source signal generating module is used to output the multi-point source signal according to the multi-point source signal control strategy generated by the control strategy generating module, and utilize the output multi-point source signal to enable the single pulse radar to receive the simulated target signal.
8. A single pulse radar target simulation device based on reinforcement learning according to claim 7, characterized in that: In the detection module, the pointing angle θ of the single pulse radar after being controlled by two point source signals is s for: In the formula, θ e is the angle between the two signal sources relative to the alignment axis of the monopulse radar; α is the amplitude ratio of the two signal sources; φ is the phase difference between the two signal sources.
9. A single pulse radar target simulation device based on reinforcement learning according to claim 8, characterized in that: In the control strategy generation module, the reinforcement learning algorithm used is a deep deterministic policy gradient algorithm; the reinforcement learning algorithm uses the amplitude ratio and phase difference of the signal source to determine the action of the driven intelligent body; In the control strategy generation module, the action of the intelligent agent is defined as: the adjustment strategy of the amplitude ratio in the pitch and azimuth directions of the single pulse radar pointing angle, and the formula of the adjustment strategy is: π=[π ε ,π ξ ] In the formula, π ε is the azimuth adjustment strategy, π ξ is the pitch adjustment strategy, π ε and π ξ Any number between -1 and 1; In the control strategy generation module, the action state of the agent is the three-dimensional position information of the monopulse radar: s=[x,y,z] Where s is the action state of the agent, and [x, y, z] is the three-dimensional position information of the single pulse radar.
10. A single pulse radar target simulation device based on reinforcement learning according to claim 9, characterized in that: In the multi-point source signal generation module, the reward r of the agent is: Where R is the distance between the final landing point of the single pulse radar and the center point of the signal source, R1 and R2 are the shortest and farthest distances between the target area and the center point of the signal source, respectively. The values of R1 and R2 are determined according to the terrain.
Citation Information
Patent Citations
Towed bait interference alarm method based on monopulse angle measurement
CN118625264A
Method of radar protection against antiradar missile based on use of additional radiation source with a lift-type horn aerial
RU2287168C1
Approach radar with array antenna having rows and columns skewed relative to the horizontal
US20040196172A1
Monopulse radar signal processing for rotorcraft brownout aid application
US7633429B1