Onboard radar guidance search decision method and system based on proximal policy optimization

By constructing a near-end policy optimization method for upper-layer and lower-layer policy modules, and combining it with reinforcement learning to train the radar search decision model, the radar guidance search problem in the scenario of cluster targets with beyond-line-of-sight and large airspace distribution is solved, achieving efficient and robust autonomous decision-making and performance optimization.

CN118673400BActive Publication Date: 2026-05-19NORTHWESTERN POLYTECHNICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
NORTHWESTERN POLYTECHNICAL UNIV
Filing Date
2024-05-30
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

In scenarios where clustered targets are distributed over a wide airspace beyond visual range, existing radar-guided search missions are unable to meet the real-time decision-making requirements of highly dynamic operations. Traditional methods seek the optimal solution within a known search space, lacking effective optimization of airspace missions and radar resources.

Method used

An airborne radar-guided search decision-making method based on near-end strategy optimization is constructed, including an upper-level strategy module and a lower-level strategy module. The method is trained using a strategy network and a value network to optimize the azimuth of the search airspace, the radar beam dwell time, and the beam search data rate. Finally, it is combined with reinforcement learning to make autonomous decisions.

Benefits of technology

It achieves efficient, robust, and convergent radar-guided search in scenarios with large airspace and beyond-line-of-sight distribution of clustered targets, enabling rapid and accurate autonomous decision-making and optimizing radar search performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118673400B_ABST
    Figure CN118673400B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of radar resource allocation scheduling, in particular to an airborne radar guided search decision method and system based on proximal policy optimization. A decision model including an upper policy module and a lower policy module is constructed, the upper policy module including a policy network and a value network; the decision model is used to collect radar search decision trajectories, the policy network and the value network of the upper policy module are trained using the radar trajectories, and the policy network parameters and the value network parameters are updated; the current observation state is input into the trained decision model, the azimuth coordinate of the space to be searched is obtained based on the upper policy module, the radar beam residence time and the wave position search data rate are obtained based on the lower policy module, and the decision of the radar guided search task in the cluster target space over-the-horizon distribution scene is realized. The method is more accurate in decision-making.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of radar resource allocation and scheduling technology, specifically to an airborne radar guidance and search decision-making method and system based on near-end strategy optimization. Background Technology

[0002] Currently, in scenarios involving large-scale, beyond-line-of-sight target distributions, airborne radar-guided search missions require multi-dimensional and complex decision-making processes that combine radar search airspace ensemble coverage decisions with search performance optimization algorithms. On one hand, by combining battlefield target and friendly aircraft situational information with corresponding radar search resources, the search orientation and order of each airspace are determined to achieve precise coverage of potential target locations, thereby maximizing search efficiency. On the other hand, based on target guidance information, the radar search performance within the search airspace is optimized. This involves determining sub-airspace partitioning strategies and sub-airspace search beam position arrangement strategies, calculating the beam dwell time and beam position search data rate of corresponding sub-airspaces under different search mission resource loads, and constructing a global radar search performance optimization model under different search mission resource loads.

[0003] The radar airspace guidance search problem in the scenario of clustered targets with large airspace distribution requires the construction of an airspace coverage decision model and the adoption of a high-precision real-time solution algorithm to achieve high-quality online generation of airspace guidance search schemes. Jiang Jianlin, Cheng Kun, Wang Cancan et al. published "Set Coverage Problem Based on Improved Genetic Algorithm" in Mathematics in Practice and Theory, 2012, 42(05):120-126, pointing out an improved genetic algorithm to solve the set coverage problem, mainly focusing on improvements in initial population generation, repair of infeasible solutions, and generalization of multi-point crossover. However, existing research mainly focuses on improving the algorithm solution performance of typical problems, lacks the ability to handle the high dynamic constraints of actual targets, and has less research on constraints such as airspace tasks and radar resources. Regarding the search performance optimization problem after the search airspace of the carrier radar is determined, the traditional radar optimal search model usually needs to combine target guidance information and the performance and occupancy information of the local radar to determine the sub-airspace division and wave position arrangement strategy, and optimize for different search parameters and sub-airspace search resource allocation respectively.

[0004] Therefore, in scenarios involving large-scale, beyond-visual-range airspace distributions of clustered targets, the aforementioned radar-guided search missions require the construction of an online decision-making model covering the search airspace set, and optimization of the search performance within the decision airspace set based on the situational information of both sides and the aircraft's radar parameters. For this multi-dimensional decision-making problem, traditional methods require pre-determining the optimization objective and constraints, and can only find the optimal solution within a known search space, making it difficult to meet the real-time decision-making requirements of future highly dynamic operations. Summary of the Invention

[0005] To address the shortcomings of existing technologies, this invention provides an airborne radar-guided search decision-making method and system based on near-end strategy optimization, which solves the radar-guided search problem in scenarios with large airspace beyond-line-of-sight distribution of clustered targets, and has good robustness and convergence.

[0006] The objective of this invention is achieved through the following technical solutions:

[0007] An airborne radar-guided search decision-making method based on near-end strategy optimization includes:

[0008] S1. Construct a decision model that includes an upper-level strategy module and an lower-level strategy module, wherein the upper-level strategy module includes a strategy network and a value network;

[0009] S2. Collect radar trajectories using the decision model obtained in step S1, and use the radar trajectories to train the policy network and value network of the upper-level policy module, and update the policy network parameters and value network parameters.

[0010] S3. Input the current observation status into the decision model trained in step S2. Based on the upper-level strategy module, obtain the azimuth coordinates of the airspace to be searched. Based on the lower-level strategy module, obtain the radar beam dwell time and beam position search data rate to realize the decision-making of radar-guided search tasks in the scenario of beyond-line-of-sight distribution of cluster targets in the airspace.

[0011] As a further improvement of the present invention, the spatial domain set coverage model satisfies the condition that each target can be covered by at least one spatial domain in the spatial domain set, and the sum of the spatial domain costs of all spatial domains in the spatial domain set is minimized. The calculation expression is as follows:

[0012]

[0013] x j ∈{0,1},j=1,2,...,m.

[0014] In the formula, c j Let $a$ be the cost of the $j$-th spatial domain in spatial domain set $X$, $j$ be a specific spatial domain in spatial domain set $X$, $m$ be the total number of spatial domains in spatial domain set $X$, and $n$ be the total number of targets. When $a$... ij =1 indicates that the i-th target is covered by the j-th spatial domain, a ij = 0 indicates that the i-th target is not covered by the j-th spatial domain, x j =1 means that the j-th spatial domain is contained within the spatial domain set X, x j =0 indicates that the j-th spatial domain is not in the spatial domain set X.

[0015] As a further improvement of the present invention, the expression for calculating the spatial cost is as follows:

[0016]

[0017] Where, n j To guide the start of the search, the number of targets falling into the j-th spatial domain is calculated based on the guidance information. j Let P be the coverage area of ​​the j-th spatial domain. dj / S j As the search progresses, the overall interception probability P represents the probability of capturing all targets per unit area in the j-th spatial domain. dj It is calculated based on the coverage area of ​​the j-th spatial domain and the Gaussian mixture model based on the probability distribution of the constant velocity target position.

[0018] As a further improvement of the present invention, the network structure of the upper-layer strategy module includes a long short-term memory neural network layer, a multi-head attention mechanism unit, and a fully connected neural network layer connected in sequence.

[0019] As a further improvement of the present invention, the radar search parameter optimization model includes a sub-spatial beam dwell time optimization model based on the maximum expected detection range of clustered targets, and a sub-spatial beam position search data rate optimization model based on the maximum average accumulated detection probability of clustered targets. The sub-spatial beam dwell time optimization model is as follows:

[0020]

[0021] In the formula, The weighted expected discovery distance of the target across the entire search space, where N represents the number of subspaces, and α i Ω represents the threat level weighting coefficient for the sub-space domain. i τ is a constant related to the radar system in each sub-space domain. si The beam dwell time (SNR) is denoted as i in sub-spatial domain i. D N represents the echo signal-to-noise ratio at the radar detection range. s V represents the number of wavenumbers in the subspace search. k The value represents the target speed; n represents the number of targets, and w represents the target velocity. k The normalized threat coefficient of the target, satisfying

[0022] The sub-spatial position search data rate optimization model is as follows:

[0023]

[0024] In the formula, p d0 Let t be the probability of radar detecting a target, where t is a radar search frame period. f The k-th wave position within the j-th spatial domain of the inner pair was... Second photo, The search data rate for each wave position corresponding to target i. This represents the probability that the target will appear at this wave position.

[0025] As a further improvement of the present invention, the observation space and action space of the upper-level strategy module are respectively:

[0026]

[0027] In the formula, This serves as the observation space for the upper-level strategy module. This represents the action space of the upper-level strategy module, where U represents the set of targets not found. These represent the target azimuth and pitch coordinate information contained in the guidance information, respectively. The observation space is a variable observation space, where azimuth_center is the azimuth center and pitch_center is the pitch coordinate.

[0028] As a further improvement of the present invention, the observation space and action space of the underlying strategy module are respectively:

[0029]

[0030] In the formula, This serves as the observation space for the underlying strategy module. For the action space of the underlying strategy module, τ si Let i be the beam dwell time in sub-space domain i. Let N be the search data rate for each wave position corresponding to target i, and let N represent the number of sub-space domains. s This indicates the number of wavenumbers for subspace search.

[0031] As a further improvement of the present invention, the training process of the decision model includes:

[0032] The training process of the decision model includes:

[0033] Initialize the policy network, value network, policy network parameters, and value network parameters in the upper-level policy module; initialize the policy network learning rate, value network learning rate, maximum training steps, maximum number of rounds, and experience replay pool.

[0034] Based on the observation state of the current update step, a comprehensive action is obtained based on the upper-layer policy module and the lower-layer policy module. Executing the comprehensive action yields the comprehensive reward function and transition state for that step. The observation state, comprehensive action, reward function, and transition state of that step are combined into a quadruple and stored in the experience replay pool. This step is updated cyclically until the amount of quadruple data in the experience replay pool reaches the set minimum training data amount. Several quadruples are then selected from the experience replay pool and input into the upper-layer policy module. The policy network parameters are continuously updated based on the policy network learning rate and the quadruples. The value network parameters are updated based on the value network learning rate and expected value. After the update is complete, it is determined whether the maximum number of training steps has been reached. If not, the trajectory is re-collected and training continues until the maximum number of training steps is reached, resulting in the trained policy network and value network.

[0035] As a further improvement of the present invention, the round-based comprehensive reward function corresponding to the decision model is:

[0036]

[0037] In the formula, reward1 is the reward function obtained based on the greedy algorithm; reward2 is the process reward function; reward3 is the round search spatial redundancy reward function, which is obtained based on the number of training steps and the overlapping area of ​​the search spatial domain; reward4 is the underlying strategy module optimization effect reward function, which is obtained based on the target weighted expected discovery distance in the search spatial domain and the comprehensive accumulated discovery probability of cluster targets; reward5 is the round task completion reward function, which is obtained based on the maximum round step size during training; and T represents the total number of execution steps in the round. This represents the number of undiscovered targets before each airspace search. This indicates the number of targets not found after the search is completed.

[0038] This invention also provides an airborne radar-guided search decision system based on near-end strategy optimization, which, based on the above-mentioned airborne radar-guided search decision method based on near-end strategy optimization, includes:

[0039] The building module constructs a decision model that includes an upper-level strategy module and an lower-level strategy module. The upper-level strategy module includes a strategy network and a value network.

[0040] The training module collects trajectories through the decision model, uses the trajectories to train the policy network and value network of the upper-level policy module, and updates the parameters of the policy network and value network.

[0041] The testing module inputs the current observation status into the trained decision model to obtain the azimuth coordinates of the airspace to be searched, the radar beam dwell time, and the beam position search data rate, thereby realizing the decision-making of radar-guided search tasks in the scenario of beyond-line-of-sight distribution of cluster targets in the airspace.

[0042] The beneficial effects of this invention are as follows: This invention provides an airborne radar-guided search decision-making method and system based on near-end strategy optimization. It constructs an upper-level strategy module and a lower-level strategy module for guidance information of clustered targets. The upper-level strategy module includes a strategy network and a value network. Based on reinforcement learning training of the upper-level strategy module, the strategy network can obtain the azimuth coordinates of the airspace to be searched. The value network further evaluates the quality of the strategy network's decision-making actions. The lower-level strategy module obtains the radar beam dwell time and beam position search data rate, further optimizing the results obtained by the upper-level strategy module. After training in a reinforcement learning environment, it can quickly obtain accurate autonomous decisions based on the current observation state. The radar-guided search intelligent decision-making method based on near-end strategy optimization designed in this invention can effectively solve the radar-guided search problem in scenarios with large airspace beyond-line-of-sight distribution of clustered targets, and has good robustness and convergence. Attached Figure Description

[0043] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0044] Figure 1 This is a network structure block diagram of the upper-layer strategy module in an embodiment of the present invention;

[0045] Figure 2 This is a schematic diagram of the sub-space domain partitioning results in an embodiment of the present invention;

[0046] Figure 3 This is a schematic diagram of the two-dimensional search spatial domain in an embodiment of the present invention;

[0047] Figure 4 This is a weighted target expected discovery distance curve diagram in an embodiment of the present invention;

[0048] Figure 5 This is a target average cumulative discovery probability curve in an embodiment of the present invention;

[0049] Figure 6 This is a schematic diagram of the airborne radar-guided search decision-making method based on near-end strategy optimization in an embodiment of the present invention. Detailed Implementation

[0050] To make the objectives and technical solutions of this invention clearer and easier to understand, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. The specific embodiments described herein are for illustrative purposes only and are not intended to limit the invention.

[0051] The present invention provides an airborne radar-guided search decision-making method and system based on near-end strategy optimization. The method mainly includes:

[0052] S1. Construct a decision model that includes upper-level strategy modules and lower-level strategy modules. The upper-level strategy modules include a strategy network and a value network.

[0053] S2. Collect radar search decision trajectories using the decision model obtained in step S1, and use the radar trajectories to train the policy network and value network of the upper-level policy module, and update the policy network parameters and value network parameters.

[0054] S3. Input the current observation status into the decision model trained in step S2. Based on the upper-level strategy module, obtain the azimuth coordinates of the airspace to be searched. Based on the lower-level strategy module, obtain the radar beam dwell time and beam position search data rate to realize the decision-making of radar-guided search tasks in the scenario of beyond-line-of-sight distribution of cluster targets in the airspace.

[0055] The technical solution of the present invention will be clearly and completely described below with reference to the accompanying drawings and specific embodiments. The described embodiments are only some embodiments of the present invention, and not all embodiments.

[0056] Example of an airborne radar-guided search decision-making method based on near-end strategy optimization:

[0057] like Figure 6 The airborne radar-guided search decision method based on near-end strategy optimization shown in this embodiment includes the following implementation steps. In this embodiment, n targets and m airspaces are used as an example for specific explanation. The n targets are distributed in the airspace, and the entire airspace is divided into m small airspaces that can be covered by phased array radars, using matrix A = (a ij Let i∈N, j∈M represent the spatial coverage of the target, where N={1,2,...,n} and M={1,2,...,m} represent the sets of the target and the spatial domain of matrix A, respectively. Then, if a ij =1, indicating that the i-th target is covered by the j-th spatial domain; otherwise, if a ij =0, which means that the i-th target is not covered by the j-th spatial domain.

[0058] S1. Construct a decision model including an upper-level strategy module and a lower-level strategy module. The upper-level strategy module includes a strategy network and a value network. In this embodiment, the decision model includes an upper-level strategy module and a lower-level strategy module. The upper-level strategy module uses the current target coverage and guidance information as the observation space, and the azimuth and elevation coordinates of the center of the searched airspace as the action space, to obtain the azimuth coordinates to be searched.

[0059] The upper-level strategy module is constructed based on the spatial domain set coverage model. This model ensures that every objective can be covered by at least one spatial domain in the set, and that the sum of the spatial domain costs in that set is minimized. Specifically, the spatial domain set coverage model aims to find such a set of spatial domains. This ensures that every target in the target set N is covered by at least one spatial domain in the spatial domain set X, and minimizes the sum of the costs of all spatial domains in X. The cost matrix C = (c j ),j∈M, where c j Let c represent the cost of spatial domain j. j >0, The calculation expression is:

[0060]

[0061] x j ∈{0,1},j=1,2,...,m. (3)

[0062] In the formula, c j Let $a$ be the cost of the $j$-th spatial domain in spatial domain set $X$, $j$ be a specific spatial domain in spatial domain set $X$, $m$ be the total number of spatial domains in spatial domain set $X$, and $n$ be the total number of targets. When $a$... ij =1 indicates that the i-th target is covered by the j-th spatial domain, a ij = 0 indicates that the i-th target is not covered by the j-th spatial domain, x j =1 means that the j-th spatial domain is contained within the spatial domain set X, x j =0 indicates that the j-th spatial domain is not in the spatial domain set X.

[0063] In this embodiment, the expression for calculating the cost in the spatial cost matrix is ​​as follows:

[0064]

[0065] Where, n j To guide the start of the search, the number of targets falling into the j-th spatial domain is calculated based on the guidance information. j Let P be the coverage area of ​​the j-th spatial domain. dj / S jLet P be the overall intercept probability of all targets per unit area in the j-th spatial domain as the search progresses. dj The cost is calculated based on the coverage area of ​​the j-th spatial domain and a Gaussian mixture model of the probability distribution of the constant-velocity target position. The purpose of setting the spatial domain cost in conjunction with the Gaussian mixture probability distribution model of the constant-velocity moving target is to accurately evaluate the quality of the spatial domain, which facilitates the calculation of reward2 during training.

[0066] The Gaussian mixture model based on the probability distribution of the position of a constant-velocity target is obtained from the established two-dimensional normal distribution model of the probability distribution of the position of a constant-velocity moving target. Specifically, the probability distribution model of the position of a constant-velocity moving target based on the two-dimensional normal distribution is as follows:

[0067]

[0068] Where σ and μ represent the standard deviations of the position and velocity guidance errors, respectively; X0 represents the true value of the target position at the initial moment. This represents the initial target position guidance information; V represents the actual value of the target velocity. Guidance information indicating the target speed.

[0069] Therefore, the calculation expression based on the Gaussian mixture model of the probability distribution of the constant velocity target position is:

[0070]

[0071] Then, based on the Gaussian mixture model, the comprehensive interception probability P of spatial domain j for all targets is defined. dj for:

[0072]

[0073] Where S j Let j be the area covered by the spatial domain.

[0074] In this embodiment, the network structure of the upper-layer policy module is as follows: Figure 1As shown, the network comprises a Long Short-Term Memory (LSTM) neural network layer, a multi-head attention mechanism unit (i.e., the multi-head attention module shown in Figure 1), and a fully connected neural network layer (i.e., the fully connected layer), connected sequentially. The policy module receives the observation state at its input port and connects to the input of the LSTM layer via an embedding layer. The output data of the LSTM layer is then processed through a dropout layer and normalized before being input to the multi-head attention mechanism unit. The output data of the multi-head attention mechanism unit and the LSTM layer are normalized and then input to the fully connected neural network layer. After reshaping through a reshape layer, the data is input to the second fully connected neural network layer, finally outputting the corresponding action space data. The network structure is fixed, while the network parameters are continuously learned based on the solution process of the spatial domain set coverage model. Combining the LSTM layer, the multi-head attention mechanism, and the fully connected neural network allows for better processing of long sequence observation state data while effectively focusing on target guidance information in different parts of the sequence. LSTM is used to process global observation state sequences of variable length and capture long-term dependencies in the sequence; multi-head attention can help fully connected neural networks better focus on state information at different positions in the input sequence, thereby improving the performance of decision models.

[0075] The bottom-level strategy module uses the observation space and action space of the upper-level strategy module as the observation space and the radar search parameters as the action space to obtain the radar beam dwell time and beam position search data rate. In this embodiment, the purpose of the bottom-level strategy module is to perform sub-space division, beam position arrangement, and radar search parameter optimization on the airspace determined by the upper-level strategy module. The bottom-level strategy module is constructed based on the radar search parameter optimization model. The radar search parameter optimization model includes a sub-space beam dwell time optimization model based on the maximum expected detection distance of clustered targets and a sub-space beam position search data rate optimization model based on the maximum average accumulated detection probability of clustered targets.

[0076] Specifically, the sub-spatial beam dwell time optimization model is as follows:

[0077]

[0078] In the formula, The weighted expected discovery distance of the target across the entire search space, where N represents the number of subspaces, and α i Ω represents the threat level weighting coefficient for the sub-space domain. i τ is a constant related to the radar system in each sub-space domain. si The beam dwell time (SNR) is denoted as i in sub-spatial domain i. D N represents the echo signal-to-noise ratio at the radar detection range. sV represents the number of wavenumbers in the subspace search. k The value represents the target speed; n represents the number of targets, and w represents the target velocity. k The normalized threat coefficient of the target, satisfying

[0079] Among them, the sub-space threat weight coefficient α i The calculation expression is:

[0080]

[0081] Q i Let α represent the set of targets contained in subspace i. min ,α max The minimum threat weight coefficient and the maximum threat weight coefficient are both constants.

[0082] The constants Ω related to the radar system in each sub-space domain i The calculation formula is:

[0083]

[0084] In equation (10), P av For the average transmit power, G ti For the transmit antenna gain, G ri λ is the receiver antenna gain, σ is the radar wavelength, σ is the target RCS, k is the Boltzmann constant, T0 is the receiver noise temperature (290K at room temperature), and F is the receiver noise temperature. n Let L be the receiver noise figure, and L be the radar system loss.

[0085] Echo signal-to-noise ratio (SNR) at radar detection range D This can be obtained from the relationship between radar detection probability and radar false alarm probability, that is:

[0086]

[0087] Where p fa p represents the radar false alarm probability. d The radar detection probability is given by setting the radar false alarm probability p. fa and radar detection probability p d At that time, SNR D It is a constant.

[0088] The sub-space wave position search data rate optimization model is as follows:

[0089]

[0090] In the formula, p d0 Let t be the probability of radar detecting a target, where t is a radar search frame period. fThe k-th wave position within the j-th spatial domain of the inner pair was... Second photo, The search data rate for each wave position corresponding to target i. τ represents the probability that the target will appear at this wave position. si The optimal beam dwell time for sub-space domain i.

[0091] Among them, the optimal beam dwell time τ in sub-spatial domain i si :

[0092]

[0093] Among them, the probability of the target appearing at this wave position. for:

[0094]

[0095] Among them, S jk This represents the coverage area of ​​the k-th wave position within sub-space domain j.

[0096] The search data rate for each wave position corresponding to target i for:

[0097]

[0098] Based on equation (15), after superimposing and normalizing the wavelet search data rates corresponding to each target, the expression for calculating the wavelet search data rate based on the maximum average cumulative discovery probability of cluster targets is as follows:

[0099]

[0100] The underlying strategy module, based on the azimuth of the search airspace and the performance parameters of the local radar from the upper strategy module, combined with target guidance information, divides the search airspace into two sub-airspaces according to a fixed strategy. It also optimizes the beam dwell time and beam position search data rate of the sub-airspaces based on the expected target detection distance and the accumulated target detection probability, in order to achieve better guidance and search performance.

[0101] The airspace search strategy in this embodiment includes an upper-layer reinforcement learning radar search azimuth optimization strategy and a lower-layer fixed radar search parameter optimization strategy.

[0102] This embodiment uses dual observation information, combining the current target coverage and guidance information, as the observation space for the upper-layer strategy module. The observation space of the upper-level strategy module It can be represented as:

[0103]

[0104] Where U represents the set of undiscovered targets. Let the target azimuth and elevation coordinates be represented respectively, as the guidance information includes, then the observation space of the upper-level strategy module... The dimensions will change as the search progresses, i.e., the variable observation space.

[0105] The azimuth and elevation coordinates of the center of the airspace to be searched are used as the action space of the upper-level strategy module.

[0106]

[0107] In the formula, azimuth_center is the azimuth center and pitch_center is the pitch coordinate.

[0108] The observation space and action space of the underlying strategy module are as follows:

[0109]

[0110] In the formula, This serves as the observation space for the underlying strategy module. For the action space of the underlying strategy module, τ si Let i be the beam dwell time in sub-space domain i. Let N be the search data rate for each wave position corresponding to target i, and let N represent the number of sub-space domains. s This indicates the number of wavenumbers for subspace search.

[0111] The observation space and action space of the entire airspace search environment can then be represented as follows:

[0112]

[0113] In this embodiment, the global observation space is used as the input to the decision-making model, and the global action space is used as the output.

[0114] S2. Using the decision model obtained in step S1, radar search decision trajectories are collected. These radar trajectories are then used to train the upper-level policy module's policy network and value network, updating their parameters. In this embodiment, the value network is used to evaluate the merits of the current decision action. During training, the optimal decision action needs to be selected through the value network, and the value network and policy network need to be trained synchronously. The decision model training process in this embodiment is as follows:

[0115] a. Initialize the upper-level strategy module and value network Policy network parameters θ0 and value network parameters φ0; re-initialize hyperparameters: policy network learning rate l actor Value network learning rate lcritic Maximum training steps T max Maximum number of rounds T, current number of steps t, experience replay pool D.

[0116] Among them, the network model parameters θ0 and φ0 represent the weights and biases of the neural network; the network learning rate is used to update the network parameters, and the learning rate is needed to control the update size when using gradient to update the network parameters; the maximum training steps are the number of training iterations; the maximum number of epochs represents the maximum length of a single epoch trajectory collected; and the experience replay pool represents a list or array that stores the collected trajectories.

[0117] b. Under the current observation state, use the policy network in the upper-level policy module. Select an action state. Based on the upper-level observed state. Select the upper-level strategy action. Where the definition implement Obtain the underlying observation state The observation space based on equations (13), (16), and the underlying strategy module Calculate the action space of the underlying strategy module separately, i.e. This leads to the comprehensive action space a. t ,Right now The comprehensive action space is used to obtain the round-based comprehensive reward function r. t and transition state And combine them into a quadruple (s t ,a t ,r t ,s t+1 Store it in the experience replay pool D.

[0118] Execute decision action a t You will receive a reward r t and new state s t+1 s t+1 With s t The meaning is the same, indicating that the state will change after the action is performed; the experience replay pool contains multiple round trajectories, each round containing several steps, i.e., several quadruplets, as follows:

[0119] From state s t Select action a t Receive reward r t and new state s t+1 ;

[0120] Determine if the round termination condition is met. If it is, reset the state and re-execute; otherwise, proceed as follows:

[0121] From state s t+1 Select action a t+1 Receive reward rt+1 and new state s t+2 ;

[0122] Repeat the above process until the number of collected trajectories meets the training requirements. Therefore, the quadruple will change at each step of the action.

[0123] The round-based comprehensive reward function is calculated for each round. In this embodiment, the round-based comprehensive reward function r corresponding to the decision model is... t for:

[0124]

[0125] In the formula, reward1 is the reward function obtained based on the greedy algorithm; reward2 is the process reward function; reward3 is the round search spatial redundancy reward function, which is obtained based on the number of training steps and the overlapping area of ​​the search spatial domain; reward4 is the underlying policy module optimization effect reward function, which is obtained based on the target weighted expected discovery distance in the search spatial domain and the comprehensive accumulation of cluster target discovery probability; reward5 is the round task completion reward function, which is obtained based on the maximum round step size during training; and T represents the total number of execution steps in the round. This represents the number of undiscovered targets before each airspace search. This indicates the number of targets not found after the search is completed.

[0126] The calculation expression for the reward function reward1 obtained based on the greedy algorithm is as follows:

[0127] reward1=α(n greedy -n rl ) (twenty two)

[0128] Where α represents a positive constant, n greedy n represents the number of searches required by the greedy algorithm. rl This indicates the number of searches performed by the reinforcement learning algorithm in this round.

[0129] To prevent reward sparsity from affecting the network's learning speed, the process reward reward2 is defined as:

[0130] reward2=βc ost (n start -n end ) (twenty three)

[0131] Where β represents a positive constant, n start This indicates that no target was found in the current search airspace, n. end c represents the number of undiscovered targets after the search is completed. ostThis indicates the airspace cost for this search.

[0132] To prevent network overfitting and increase network exploration, the round search spatial redundancy reward function reward3 is defined as follows:

[0133]

[0134] Where γ,t0 represents positive constants, S represents the overlapping area of ​​the search space in this round, and t represents the number of training steps.

[0135] To reflect the merits of the upper-level airspace selection strategy, the reward function reward4 for the optimization effect of the lower-level strategy module is defined as follows:

[0136]

[0137] in, and These represent the target weighted expected discovery distance in the entire search airspace, achieved using the method described herein and the uniformly distributed sub-spatial beam dwell time method, respectively. and These represent the cluster target comprehensive accumulation discovery probability based on this method and the uniform beam distribution search data rate method, respectively.

[0138] To prevent the coupling of the above-mentioned optimization metrics, the reward function reward5 for round task completion is defined as follows:

[0139]

[0140] Where max_step represents the maximum round step size, and r0 represents a positive constant.

[0141] c. When the amount of data in the experience replay pool D reaches the set minimum training data amount, samples are taken from the experience replay pool D, and the quadruplets in the samples are input into the upper-level policy module of the decision model for training:

[0142] Update the policy network using the new observation space and corresponding reward function, based on the quadruple s. t ,a t ,s t+1 ,r t The dominance function is calculated to evaluate the quality of higher-level strategies. The expression for the dominance function is as follows:

[0143]

[0144] in, Let s represent the agent's advantage function at time t. t Let a represent the state of the agent at time t. tγ represents the action output by the policy network at time t. rl λ represents the discount factor. GAE Represents the GAE coefficient. Indicates the value network in s t State and φ k The state value output under the parameters.

[0145] The policy network parameter θ0 is updated as follows:

[0146]

[0147] Where ε represents the clipping coefficient of the proximal policy optimization model PPO-clip; clip represents the clipping function; π θ (a t |s t ) represents the probability of the policy network action that needs to be updated, and θ represents its network parameters; θ represents the probability of a policy network action in interaction with the environment. k This represents the policy network parameters used for interaction.

[0148] Update the parameter φ0 of the value network module. The updated value network parameter φ0 is as follows:

[0149]

[0150] Among them, This represents the expected value of the trajectory after time t.

[0151] Repeat step bc until training is complete, and you will obtain the trained decision model.

[0152] S3. Input the current observation status into the decision model trained in step S2. Based on the upper-level strategy module, obtain the azimuth coordinates of the airspace to be searched. Based on the lower-level strategy module, obtain the radar beam dwell time and beam position search data rate to realize the decision-making of radar-guided search tasks in the scenario of beyond-line-of-sight distribution of cluster targets in the airspace.

[0153] This embodiment constructs a spatial domain ensemble coverage model and a radar search parameter optimization model for cluster target guidance information, respectively. Based on these models, upper-layer and lower-layer policy modules are constructed. After training in a reinforcement learning environment, these modules can quickly make accurate autonomous decisions based on the current observation state. The radar guidance search intelligent decision-making method based on near-end policy optimization designed in this embodiment can effectively solve the radar guidance search problem in scenarios with large spatial domain beyond-line-of-sight distribution of cluster targets, and has good robustness and convergence. To verify the effectiveness of this method, this embodiment further illustrates it through the following simulation examples.

[0154] The number of targets in this instance is determined to be n=20, the coordinates are randomly generated, the standard deviation of the guidance information σ=1°, μ=0.3°, and the target parameters such as threat level are set as shown in Table 1:

[0155] Table 1 Target Parameter Table

[0156]

[0157] The radar system constant Ω0 is set to 4.85 × 10⁻⁶. 26 Detection probability p d =0.9; False alarm probability p fa =10 -6 Beamwidth = 5°; Azimuth and elevation search airspace range: rx = ry = 30°, and reward function parameter settings: α = 1.0, β = 1.0, r0 = 20, γ = 1.0 / (rx*ry). Sub-spatial threat weight upper and lower limits α max =0.3,α min =0.1; Subspace partitioning results are as follows Figure 2 As shown.

[0158] The reinforcement learning parameters are set as follows: policy network learning rate l actor =0.001, Value Network Learning Rate l critic =0.001, maximum training steps T max =3×10 5 Maximum number of rounds T = 20, discount factor γ rl =0.99, Experience replay pool size D=10 6 Multi-head attention module parameter settings: embedding layer dimension n_embd = 16, state sequence length ns = 4, number of attention heads I = 4; number of LSTM layers is 1, number of hidden layer neurons is 64; fully connected neuron dimension: 64×32×2.

[0159] The changes in the target parameters after training are as follows: Figures 3-5 As shown in Table 2.

[0160] Table 2 Target Discovery Status

[0161]

[0162]

[0163] Through Table 2, Figures 3-5As shown, the trained upper-layer strategy module can output the effective search airspace azimuth based on target guidance information, meaning that all targets can fall into the searched sub-airspace. The lower-layer strategy module can achieve a higher expected detection range and accumulated detection probability for cluster targets compared to traditional single-target radar search parameter optimization algorithms. The trained agent can realize intelligent decision-making for radar-guided search tasks in scenarios with cluster targets distributed over long distances beyond visual range.

[0164] Example of an airborne radar-guided search decision system based on near-end strategy optimization:

[0165] An airborne radar-guided search decision system based on near-end strategy optimization, comprising:

[0166] The building module constructs a decision model that includes an upper-level strategy module and an lower-level strategy module. The upper-level strategy module includes a strategy network and a value network.

[0167] The training module collects trajectories through the decision model, uses the trajectories to train the policy network and value network of the upper-level policy module, and updates the parameters of the policy network and value network.

[0168] The testing module inputs the current observation status into the trained decision model to obtain the azimuth coordinates of the airspace to be searched, the radar beam dwell time, and the beam position search data rate, thereby realizing the decision-making of radar-guided search tasks in the scenario of beyond-line-of-sight distribution of cluster targets in the airspace.

[0169] The specific implementation principles and methods of each module of the system have been described in detail in the embodiment of the airborne radar-guided search decision method based on near-end strategy optimization, and will not be repeated here.

Claims

1. An airborne radar-guided search decision-making method based on near-end strategy optimization, characterized in that, include: S1. Construct a decision model that includes an upper-level strategy module and an lower-level strategy module, wherein the upper-level strategy module includes a strategy network and a value network; S2. Collect radar trajectories using the decision model obtained in step S1, and use the radar trajectories to train the policy network and value network of the upper-level policy module, and update the policy network parameters and value network parameters. S3. Input the current observation status into the decision model trained in step S2, obtain the azimuth coordinates of the airspace to be searched based on the upper-level strategy module, and obtain the radar beam dwell time and beam position search data rate based on the lower-level strategy module, so as to realize the decision of radar-guided search task in the scenario of over-the-horizon distribution of cluster targets in the airspace. The upper-level strategy module is constructed based on a spatial domain set coverage model. This model ensures that each target can be covered by at least one spatial domain in the spatial domain set, and that the sum of the spatial domain costs in the set is minimized. The calculation expression is as follows: In the formula, For the set of empty domains X The Middle Airspace costs, For the set of empty domains X A certain airspace in, m For the set of empty domains X The total number of airspaces in the region, n For the target total quantity, when Time is represented as the first The first goal was the Airspace coverage, Time indicates the first The first goal was not achieved. Airspace coverage, For the first Airspace is contained in the airspace set X Inside, Indicates the first Airspace not in the airspace set X middle; The underlying strategy module is constructed based on a radar search parameter optimization model. This model includes a sub-spatial beam dwell time optimization model based on the maximum expected detection range of clustered targets, and a sub-spatial beam position search data rate optimization model based on the maximum average accumulated detection probability of clustered targets. The sub-spatial beam dwell time optimization model is as follows: In the formula, The target is weighted and the expected discovery distance is calculated for the entire search space. Indicates the number of subspaces. For the threat level weight coefficient of the sub-space domain, These are constants related to the radar system in each sub-space domain. Represented as a subspace Beam dwell time, This indicates the signal-to-noise ratio of the echo at the radar detection range. Indicates the number of wavenumbers in the subspace search. Indicates the target speed; Indicates the number of targets. The normalized threat coefficient of the target, satisfying ; The sub-spatial position search data rate optimization model is as follows: In the formula, The probability of radar detecting a target, where one radar search frame period is given. Inner pair of sub-space domains The first in the airspace Wave position was performed Second photo, For the goal The corresponding search data rate for each wave position, The probability of the target appearing at this wave position. For subspace Optimal beam dwell time.

2. The airborne radar-guided search decision method based on near-end strategy optimization according to claim 1, characterized in that, The expression for calculating spatial cost is: in, To guide the start of the search, the falling into the first position is calculated based on the guiding information. Number of targets in the airspace For the first The airspace coverage area, As the search progresses, the first The overall probability of intercepting all targets per unit area of ​​airspace, the overall probability of interception According to the The coverage area of ​​the airspace and the Gaussian mixture model based on the probability distribution of the position of a constant-velocity target are calculated.

3. The airborne radar guidance and search decision method based on near-end strategy optimization according to claim 1, characterized in that, The network structure of the upper-level strategy module includes a long short-term memory neural network layer, a multi-head attention mechanism unit, and a fully connected neural network layer connected in sequence.

4. The airborne radar-guided search decision method based on near-end strategy optimization according to claim 1, characterized in that, The observation space and action space of the upper-level strategy module are as follows: In the formula, This serves as the observation space for the upper-level strategy module. This serves as the action space for the upper-level strategy module. This indicates that no target set was found. These represent the target azimuth and elevation coordinates included in the guidance information, respectively. The observation space is a variable observation space. Center of azimuth, These are the pitch angle coordinates.

5. The airborne radar-guided search decision method based on near-end strategy optimization according to claim 4, characterized in that, The observation space and action space of the underlying strategy module are as follows: In the formula, This serves as the observation space for the underlying strategy module. This is the action space for the underlying strategy module. Represented as a subspace Beam dwell time, For the goal The corresponding search data rate for each wave position, Indicates the number of subspaces. This indicates the number of wavenumbers for subspace search.

6. The airborne radar-guided search decision method based on near-end strategy optimization according to claim 1, characterized in that, The training process of the decision model includes: Initialize the policy network, value network, policy network parameters, and value network parameters in the upper-level policy module; initialize the policy network learning rate, value network learning rate, maximum training steps, maximum number of rounds, and experience replay pool. Based on the observation state of the current update step, a comprehensive action is obtained based on the upper-layer policy module and the lower-layer policy module. Executing the comprehensive action yields the comprehensive reward function and transition state for that step. The observation state, comprehensive action, reward function, and transition state of that step are combined into a quadruple and stored in the experience replay pool. This step is updated cyclically until the amount of quadruple data in the experience replay pool reaches the set minimum training data amount. Several quadruples are then selected from the experience replay pool and input into the upper-layer policy module. The policy network parameters are continuously updated based on the policy network learning rate and the quadruples. The value network parameters are updated based on the value network learning rate and expected value. After the update is complete, it is determined whether the maximum number of training steps has been reached. If not, the trajectory is re-collected and training continues until the maximum number of training steps is reached, resulting in the trained policy network and value network.

7. The airborne radar-guided search decision method based on near-end strategy optimization according to claim 6, characterized in that, The round-based comprehensive reward function corresponding to the decision-making model is: In the formula, The reward function is obtained based on a greedy algorithm; For the process reward function; The round-based search spatial redundancy reward function is obtained based on the number of training steps and the overlapping area of ​​the search spatial domain. The reward function for optimizing the underlying strategy module is derived from the weighted expected discovery distance of the target in the search space and the comprehensive accumulated discovery probability of the cluster target. Here is the round-based task completion reward function, which is obtained based on the maximum round step size during training. Indicates the total number of execution steps in a round. This represents the number of undiscovered targets before each airspace search. This indicates the number of targets not found after the search is completed.

8. An airborne radar-guided search and decision system based on near-end strategy optimization, based on the airborne radar-guided search and decision method based on near-end strategy optimization as described in any one of claims 1-7, characterized in that, include: The building module constructs a decision model that includes an upper-level strategy module and an lower-level strategy module. The upper-level strategy module includes a strategy network and a value network. The training module collects trajectories through the decision model, uses the trajectories to train the policy network and value network of the upper-level policy module, and updates the parameters of the policy network and value network. The testing module inputs the current observation status into the trained decision model to obtain the azimuth coordinates of the airspace to be searched, the radar beam dwell time, and the beam position search data rate, thereby realizing the decision-making of radar-guided search tasks in the scenario of beyond-line-of-sight distribution of cluster targets in the airspace.