Path-power joint distribution method and system based on improved hierarchical reinforcement learning

By improving the path-power joint allocation method of hierarchical reinforcement learning, the problem of power and path coupling in interference resource allocation is solved, more efficient interference path planning and power allocation are achieved, and the security and detection probability of radar networking are improved.

CN120669209APending Publication Date: 2025-09-19HARBIN ENG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510826062.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

In multi-base radar collaborative detection, the traditional interference resource allocation method fails to effectively solve the mutual coupling problem between interference power and interference path, making it difficult to achieve the optimal suppression effect and reducing the security of radar networking missions.

Method used

An improved hierarchical reinforcement learning-based path-power joint allocation method is adopted. By constructing an interference countermeasure scenario, the interference resource allocation is decomposed into two sub-tasks: path and power, which are executed by independent decision makers respectively. The action key encoding and advantage experience replay technology are used to optimize the decision-making process, thus achieving joint decision-making of interference path and power.

Benefits of technology

It improves the security and detection probability of the jamming process, reduces the penetration time of high-risk targets, and enhances the security and jamming effect of radar networking.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120669209A_ABST
    Figure CN120669209A_ABST
Patent Text Reader

Abstract

The invention discloses a path-power joint distribution method and system based on improved hierarchical reinforcement learning, and the method comprises the steps: firstly constructing a confrontation scene containing an aircraft cluster (a jammer group + a target aircraft) and multiple radar nodes, and building a frequency modulation noise jamming model and a radar detection model employing the rank 1 criterion for fusion; modeling interference resource allocation as a Markov decision process, and decomposing a path-power joint decision task by adopting a continuous near-end strategy optimization algorithm: an upper-layer path decision maker plans a flight direction according to a fleet-radar distance, and a lower-layer power decision maker generates an interference resource allocation matrix in combination with the flight direction; independently setting a state set, an action set and a reward function of the two decision-making devices; and finally, optimizing a power distribution strategy through action key coding, playing back accelerated training in combination with advantage experience, and finally executing joint interference control. According to the method, the interference power is reasonably distributed while the high-risk target is avoided, the penetration safety coefficient is higher, and the time when the detection probability exceeds the safety threshold is shorter.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of radar electronic countermeasures and intelligent decision-making, and in particular relates to a path-power joint allocation method and system based on improved hierarchical reinforcement learning. Background Art

[0002] In radar countermeasure systems, multi-base radar collaborative detection technology is widely used to offset the vulnerability of individual radars to interference and significantly improve radar detection accuracy. This forces the jammer to continuously improve its jamming capabilities to effectively counter radar networking. Among them, collaborative jamming by multiple companion jammers is an effective countermeasure. However, as the number of jamming devices increases, how to rationally allocate jamming resources to maximize the suppression effect on radar networks is an important issue that needs to be addressed. However, during the collaborative jamming process, the interference power and interference path interact with each other and are highly coupled. Traditional methods often allocate power resources separately, making it difficult to achieve optimal suppression effects, thereby reducing the security of the radar network breakthrough mission. To maximize the suppression of radar networks, a joint allocation method for interference power and interference paths is crucial.

[0003] Currently, multi-domain joint resource allocation methods are often used to address parameter coupling in joint resource allocation. This allocation research primarily relies on the interference equation as its theoretical foundation, focusing on optimizing the allocation of time, space, frequency, and energy domain parameters that influence the interference-to-signal ratio. However, as research deepens, practical applications of these technologies have also exposed some problems. A literature search revealed that the paper "A Collaborative Electronic Jamming Resource Allocation Method Based on Fuzzy Multi-Attributes" uses extremely discretized time, space, frequency, and energy domain parameters and allocates them to different jammers. Joint allocation is indirectly achieved by making decisions based on jammer usage. However, this approach oversimplifies the joint decision-making problem, making it difficult to fine-tune in complex battlefield environments and failing to meet actual combat requirements. Another paper identified established a dynamic combat scenario involving long-range support jamming and multi-aircraft accompanying jamming, and proposed a formation flight trajectory planning scheme to ensure the safety of the accompanying aircraft group. However, in practice, accompanying aircraft groups often take risks to shorten jamming time, increasing the risk of detection. Summary of the Invention

[0004] The purpose of the present invention is to overcome the problem of mutual influence between spatial and energy domain resource allocation caused by the mutual coupling between interference power and interference path during cooperative interference, and propose a path-power joint allocation method and system based on improved hierarchical reinforcement learning.

[0005] The purpose of the present invention is achieved through the following technical solutions:

[0006] A path-power joint allocation method based on improved hierarchical reinforcement learning, the specific steps are as follows:

[0007] Step 1: Construct a jamming countermeasure scenario: A jamming countermeasure scenario is constructed within a pre-set simulation environment, including a jammer swarm and target aircraft, as well as multiple radars. A jamming resource allocation model is established, using frequency-modulated noise as a means of suppressing the jammers. A radar detection model is established, selecting the rank-1 criterion as the radar information fusion rule to calculate the system detection probability.

[0008] Step 2: The jamming resource allocation problem is modeled as a Markov decision process and solved using a continuous proximal strategy optimization algorithm. The joint path-power decision task in the jamming process is decomposed into two subtasks, each executed by an independent decision maker. The upper-level jamming path decider plans the flight direction based on the current distance between the fleet and the radar; the lower-level jamming power decider generates the jamming resource allocation matrix based on the flight direction. At the same time, the state sets, action sets, reward functions, and state transition probabilities of the two decision makers are set for different tasks.

[0009] Step 3: Use action key encoding (AKE) and advantage experience replay (AER) technology to optimize the decision process, and perform joint interference control based on the flight direction of the path decider and the allocation strategy of the power decider.

[0010] Furthermore, the step 1 includes the following steps:

[0011] Step 1.1: Obtain the target aircraft, jammer, and radar at time k. The jamming signal transmitted by jammer m to radar n at time k is as follows:

[0012]

[0013] in, is the amplitude of the noise FM signal, f c is the initial frequency of the FM signal, K FM Used to control the size of the frequency increment, u(t) is white noise, is the initial phase of the signal, and its value is randomly selected with the same probability in [0,2π). At the same time, u(t) is The two are independent of each other; and They represent the interference beam distribution coefficient and interference power distribution coefficient transmitted by jammer m to radar n at time k respectively;

[0014] At time k, the interference resource allocation matrix of the jammer group is established, namely, the interference resource allocation model:

[0015]

[0016] Step 1.3: Build a radar detection model and calculate the radar signal-to-interference ratio using the target echo signal power, receiver thermal noise, and interference signal power:

[0017]

[0018] in, is the target echo signal power, P N is the receiver thermal noise, is the interference signal power, P t is the transmission power of a single radar, G t Radar antenna gain, λ is the wavelength of the radar transmission signal, σ is the reflection area of ​​the target aircraft, is the distance between the target aircraft and radar n at time k, G j represents the energy gain of the jammer on the transmitted signal, λ j is the wavelength of the interference signal, γ j represents the polarization loss, is the distance between jammer m and radar n at time k, It represents the system gain of radar n when echoing signal in the direction of jammer m;

[0019] Step 1.4: Calculate the system detection probability based on the radar information fusion rule based on the rank 1 criterion The detection probability of target n by radar at time k Where V T is the detection threshold.

[0020] Furthermore, the interference signal power for:

[0021]

[0022] described The angle between the main lobe direction and this direction is small The specific relationship is:

[0023]

[0024] Among them, β rad is an empirical constant, θ 0.5 is the beam width of the radar antenna.

[0025] Furthermore, the continuous proximal strategy optimization algorithm in step 2 adopts Beta distribution as the strategy function of the algorithm, and strictly limits the sampling results to the interval [0,1].

[0026] Furthermore, the state sets of the two decision makers are:

[0027]

[0028] in, represents the Euclidean distance between the target aircraft and the radar at time k, θ path is the flight direction of the jammer group, Pd j represents the detection probability of the radar network under interference conditions;

[0029] The action sets of the two decision makers are:

[0030] A=[A path ,A power ]

[0031] Among them, A path =[θ path ] is the flight direction of the drone, with the direction directly in front of the starting position as 0, and the value range is [-π / 2, -π / 2], A power =[u·P] is the interference resource allocation matrix;

[0032] The reward function of the path decision maker is:

[0033]

[0034] Among them, Pd o represents the detection probability of the radar network when there is no interference;

[0035] The reward function of the power decider is:

[0036]

[0037] The state transition probability P of the two decision makers is the dynamic characteristic of the environment.

[0038] Furthermore, in step 3, the action key code AKE is used to encode the interference power allocation strategy. The integer part of the encoded real code represents the target radar number, and the decimal part identifies the power allocation ratio. The advantage experience replay AER technology is used to accelerate training: a common experience set D and an advantage experience set D′ are constructed, and random sampling is performed in the two sets. The advantage experience must meet the following requirements:

[0039]

[0040] Among them, R sum is the cumulative reward value of this round, R avg The average value of historical cumulative rewards;

[0041] Sample training data from D′ according to probability β.

[0042] Furthermore, the action key code AKE is used to number the interference beams of all jammers, and then the power decision maker generates the corresponding interference key. The action space of the power decision maker is:

[0043] A power =[a1,a2,…,a i ],0≤a i ≤L

[0044] Among them, a i It represents the interference power of each beam obtained by decoding according to the AKE rule, and L is the total number of interference beams.

[0045] Furthermore, the β=5%.

[0046] A computer device / apparatus / system comprises a memory, a processor and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of a path-power joint allocation method based on improved hierarchical reinforcement learning.

[0047] The beneficial effects of the present invention are:

[0048] (1) By introducing the advantage replay technology, the present invention can grasp the advantage action more quickly in the early stage of the algorithm, thereby accelerating the convergence process of the algorithm;

[0049] (2) The present invention adopts a joint allocation strategy of hierarchical reinforcement learning, which can reasonably allocate interference power while avoiding high-risk targets. Compared with traditional methods, it has a higher penetration safety factor, and the time when the detection probability exceeds the safety threshold during the entire penetration process is the shortest. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 : Interference process simulation environment.

[0051] Figure 2 : Framework of path-power joint allocation method based on improved hierarchical reinforcement learning.

[0052] Figure 3 : Path planning results of the interference formation in a single round.

[0053] Figure 4 : Interference power allocation result of a single round.

[0054] Figure 5 : The detection probability of the radar network in each round during training.

[0055] Figure 6 : Round average reward curves of different algorithms.

[0056] Figure 7 : Penetration safety factor curves of different algorithms. DETAILED DESCRIPTION

[0057] The present invention will be further described below with reference to the accompanying drawings.

[0058] The present invention provides a path-power joint allocation method and system based on improved hierarchical reinforcement learning. The method solves the problem that a single parameter decision is difficult to achieve the optimal value through joint allocation. The specific steps are as follows:

[0059] Step 1: Construct the jamming adversarial scenario.

[0060] A simulation confrontation environment with a side length of 120km was built. Figure 1 As shown in the figure, there is an aircraft swarm consisting of a jammer and a target aircraft. Four radars are deployed in the swarm's path. The radars coordinate detection to compensate for each other's blind spots, thereby increasing the probability of target detection. The swarm's task is to choose a reasonable flight direction and plan a jamming path that poses a lower threat. At the same time, they rationally allocate limited jamming power during the penetration process, minimizing the radar's detection probability and improving the penetration safety factor.

[0061] Assume that at time k, the target aircraft, jammer, and radar have all been acquired by the electronic detection system in advance. The jamming signal emitted by jammer m to radar n at time k is defined as follows:

[0062]

[0063] in, is the amplitude of the noise FM signal, f c is the initial frequency of the FM signal, K FM Used to control the size of the frequency increment, u(t) is white noise, is the initial phase of the signal; its value is randomly selected with the same probability in [0,2π), and u(t) is The two are independent of each other; and They represent the interference beam distribution coefficient and interference power distribution coefficient emitted by jammer m to radar n at time k respectively.

[0064] The safety factor in the process of penetrating through the radar network is defined as follows:

[0065]

[0066] Among them, T safe With T total They are the time when the detection probability is less than 50% and the total penetration time during the penetration process.

[0067] Since the transmit power of the jammer has a certain upper limit, in order to maximize the use of the interference power, the present invention sets the transmit power of different interference beams as a continuously controllable variable. Based on this, the interference resource allocation matrix defining the jammer group at time k can be expressed as follows:

[0068]

[0069] The radar detects our target aircraft by transmitting pulse signals and receiving target echoes. Assuming that the operating parameters of each radar are the same, the target echo signal power at radar n at time k is as follows:

[0070]

[0071] Among them, P t is the transmission power of a single radar, G t Radar antenna gain, λ is the wavelength of the radar transmission signal, σ is the reflection area of ​​the target aircraft, is the distance between the target aircraft and radar n at time k.

[0072] The radar receiver itself has a certain amount of internal thermal noise, which can generally be defined as follows:

[0073] P N =k B T0B n F n

[0074] where k B is the Boltzmann constant, T0 represents the absolute temperature of the system, B n is the radar receiver bandwidth, F n Used to measure the noise figure of a receiver.

[0075] Therefore, the interference signal power received by radar n from jammer m at time k is as follows:

[0076]

[0077] Among them, G j represents the energy gain of the jammer on the transmitted signal, λ j is the wavelength of the interference signal, γ j represents the polarization loss, is the distance between jammer m and radar n at time k, It represents the system gain of radar n when echoing the signal in the direction of jammer m. Its magnitude is affected by the angle between the main lobe direction and this direction. The specific relationship is shown in the following formula:

[0078]

[0079] Among them, β rad is an empirical constant, θ 0.5 is the beam width of the radar antenna.

[0080] Then, the signal-to-interference ratio of the echo signal received by radar n at time k in the interference environment is as follows:

[0081]

[0082] Different radars can make up for the lack of detection range through information fusion, thereby improving the joint detection probability of the system. Usually, radars use the rank K criterion to perform information fusion, which means that when the number of radar nodes that detect the target in multiple radars exceeds the threshold K, d When the system detects the target, it considers it a false alarm. Otherwise, it considers it a false alarm. However, the interfering party usually cannot know which fusion rule the radars follow. Therefore, the present invention selects the rank 1 criterion with the highest detection probability as the fusion rule. That is, when any radar detects the target, the system considers it a target. At this time, the radar system's detection probability for the target aircraft is expressed as follows:

[0083]

[0084] in, is the detection probability of the target by radar n at time k, which can be expressed as follows:

[0085]

[0086] Among them, V T is the detection threshold. When the constant false alarm detection method is used, P FA is the false alarm probability.

[0087] Step 2: Construct a path-power joint allocation model for improved hierarchical reinforcement learning. The model framework is as follows: Figure 2 As shown in the figure, the joint path-power decision task during the interference process is decomposed into two subtasks: the upper layer is responsible for planning the interference path, and the lower layer is responsible for determining the interference power. These two tasks are each handled by an independent decision maker, and the upper layer's decision action serves as an input to the lower layer's state element, thereby reducing task complexity and avoiding weight conflicts.

[0088] In the continuous proximal strategy optimization algorithm used to solve the interference resource problem, Beta distribution is used as the algorithm's strategy function, and the sampling results are strictly limited to the interval [0,1]. This can avoid action clipping during action selection and effectively improve algorithm performance. This invention mainly focuses on the joint decision-making of interference power and flight path during the interference process for the following reasons:

[0089] According to P N =k B T0B n F n During the jamming process, the main parameter that the jammer can actively influence the detection probability of the other party is the power of the jamming beam. The distance between our aircraft cluster and the radar and and polarization loss γ j Since the polarization loss changes little during the entire penetration process, it can be considered a constant. Therefore, the present invention mainly makes a joint decision on the interference power and flight path during the interference process.

[0090] In the scenario described in step 1, the confrontation between the cognitive jammer and the radar can be simplified as a Markov process, typically represented by the tuple (S, A, P, R). The present invention involves an m×n-dimensional jamming power action and a one-dimensional flight direction action. Both actions are key parameters in the interference equation, and when making joint decisions about them, conflicting action weights are inevitable. Furthermore, since both the action and state settings are continuous values, the state space and action space of the agent are further increased. Furthermore, as the number of agents requiring decision-making increases, the joint decision-making dimension grows exponentially, making the training process very slow.

[0091] The present invention consists of a decision maker, an interference effect evaluator and a network training optimization module. The decision maker is divided into an interference path decider and an interference power decider. At time k in the penetration process, the interference path decider will determine the flight direction based on the distance between the current cluster of our aircraft and the radar, and pass the decision result to the interference power decider; the power decider combines environmental information with the upper-level motion direction to decide the interference resource allocation matrix of the intelligent agent, thereby generating and implementing a path power joint allocation plan. Subsequently, the interference effect evaluator will calculate the radar detection probability after the implementation of the plan and determine the reward value of the plan, and the network training optimization module will train and update the algorithm according to the reward value. Among them, the specific description of the state, action and reward mechanism of each decision maker corresponding to each element in the MDP for different tasks is as follows:

[0092] The state set setting (S):

[0093] The penetration mission requirements indicate that both decision makers need to select the optimal jamming action at different locations within space. The path decision maker focuses on selecting a flight path with a low threat level, while the power decision maker focuses on reducing the probability of detection by the radar network. Therefore, the present invention uses the flight direction of our aircraft cluster and the detection probability of the radar network as auxiliary states for the two decision makers, respectively. The state sets of the two decision makers are represented as follows:

[0094]

[0095] in, represents the Euclidean distance between the target aircraft and the four radars at time k, θ path is the flight direction of our aircraft cluster, Pd j represents the detection probability of the radar network under interference.

[0096] The action set settings (A):

[0097] The two decision makers control the flight direction of the UAV formation and the power allocation matrix of the jammer group respectively. Therefore, the actions of the two decision makers are expressed as follows:

[0098] A=[A path ,A power ]

[0099] Among them, A path =[θ path ] is the flight direction of the drone, with the direction directly in front of the starting position as 0, and the value range is [-π / 2, -π / 2]. power =[u·P] is the interference resource allocation matrix.

[0100] The reward function setting (R):

[0101] The reward function guides the update direction of the agent, so the reward function is also designed through the objectives of different decision makers.

[0102] The decision-making goal of the path decider is to choose a flight path with a lower threat level as much as possible and reach the destination, so its reward function is as follows:

[0103]

[0104] where Pd o represents the detection probability of the radar network when there is no interference.

[0105] The goal of the power decision maker is to reduce the detection probability of the radar network, thereby ensuring the safety of our aircraft cluster. Generally speaking, the radar has a false alarm probability P fa =10 -6 and detection probability Pd o = 0.5 is the maximum range of the standard radar. Therefore, the present invention assumes that the jammer is effective when the jamming-to-signal ratio (JSR) is above 2, that is, the radar detection probability under the Swerling I condition is less than 50%. The power decision maker reward function is:

[0106]

[0107] Guided by the reward function, the intelligent agent chooses a less threatening flight route during the update process, and on this basis interferes with appropriate radars, further reducing the detection probability of the radar network and improving safety during the penetration process.

[0108] The state transition probability (P):

[0109] The state transition probability describes the dynamic characteristics of the environment. The present invention assumes that our aircraft cluster can obtain the radar position of the other party, thereby indirectly obtaining the state transition probability.

[0110] Step 3: Use action key encoding (AKE) and advantage experience replay (AER) technology to optimize the decision process, and perform joint interference control based on the flight direction of the path decider and the allocation strategy of the power decider.

[0111] Action Key Coding: The interference resource allocation problem is essentially a nonlinear mixed-integer programming problem with multiple constraints. If decisions are replicated individually for each element in the interference resource allocation matrix, the action dimension of the power decider will become extremely large, making direct convergence difficult. This is because each jammer's interference beam has an upper limit on the number of transmissions, which limits the number of non-zero elements in each row of the matrix. Based on this characteristic, if decisions are made only on the size of these non-zero elements, the action dimension of the power decider can be effectively reduced. In this case, the jammer only needs to determine the direction and allocated power of each interference beam. Therefore, the present invention performs key coding on the interference beam allocation matrix and the interference power allocation matrix to ensure that the power decider can accurately generate interference keys and thus reasonably define the action space. The encoding is in the form of real numbers, with the integer portion representing the target radar number to which the interference beam is directed, and the decimal portion representing the proportion of the power allocated to the corresponding beam to the jammer's total power.

[0112] By numbering the interference beams of all jammers, the power decider generates the corresponding interference key. On this basis, the action space of the power decider is:

[0113] A power =[a1,a2,...,a i ],0≤a i ≤L

[0114] Among them, a i It represents the interference power of each beam obtained by decoding according to the AKE rule, and L is the total number of interference beams.

[0115] Advantage experience replay: Based on the original experience set D, an additional advantage experience set D′ is added. During the algorithm update process, random sampling will be performed from the experience set D or the advantage experience set D′ to achieve the purpose of multiple utilization of advantage experience and thus accelerate the convergence of the algorithm. Advantage experience is defined here as the number of times the agent completes the given task in this round or the cumulative reward value R in this round. sum Compared with the historical cumulative reward average R avg The following relationship is satisfied:

[0116]

[0117] At the end of a round, if the cumulative reward meets the aforementioned requirements, all MDP tuples from that round are stored in the superior experience set. However, it should be noted that while superior experience sets can accelerate algorithm convergence, they contain far fewer samples than the average experience set. Therefore, excessive sampling from the superior experience set can lead to overfitting of the policy network. Therefore, to control the probability β of sampling from the superior experience set, this paper sets β to 5%.

[0118] The results of allocating interference resources using the method of the present invention are as follows:

[0119] Figure 3 The results of a single round of formation path planning are presented in the form of a heat map. The red triangles represent the four radars, and the shades of the heat map represent the radar detection probability in the absence of interference. The dark blue line and light blue area represent the flight path of the active aircraft cluster and the target destination, respectively. The formation is able to proactively fly to areas with lower detection probabilities, avoiding high-threat areas and ultimately arriving near the target destination, demonstrating that the proposed method can successfully plan interference paths.

[0120] Likewise, Figure 4 The distribution of jamming power during a single penetration round is also displayed in a heat map format. The horizontal axis represents the penetration time, the vertical axis represents the numbers of the four radars, and the color depth indicates the jammer's jamming power distribution to the corresponding radar at that time. The jammer prioritizes jamming power on radars with higher threat levels, thereby increasing the overall detection probability of the radars.

[0121] Figure 5The figure shows the detection probability of the radar network during each round of scenario training. The horizontal axis represents the penetration time per round, the vertical axis represents the number of test rounds, and the color depth indicates the magnitude of the detection probability. As can be seen from the figure, in the early stages of training, due to the agent's insufficient understanding of the environment, the swarm of our aircraft failed to select an effective path-power allocation scheme. Consequently, it was unable to avoid or jam high-risk targets. This resulted in the jammer's jamming signal power reaching the radar being far less than the radar's return power, reducing the jammer's JSR and maintaining a high detection probability throughout the penetration process. As training progressed, the agent gradually mastered the adversarial environment and formed a well-developed strategy network. At this point, the swarm of our aircraft was able to proactively avoid areas with high detection probability and concentrate jamming power on high-risk radars. Consequently, the period of high detection probability during the penetration process was shortened. By the end of training, there were almost no moments of high-risk detection probability throughout the penetration process, demonstrating that the method proposed in this invention can effectively guide path and power allocation in penetration tasks.

[0122] In order to verify the advantages of the present invention, experiments were conducted on different interference resource allocation methods under the same environment, including:

[0123] (1) Path-power joint allocation based on hierarchical reinforcement learning (HRL-PPJA): A hierarchical PPO algorithm is used to perform interference path and power allocation.

[0124] (2) Jamming resource allocation based on PPO (PPO-JRA): The PPO algorithm is used to allocate interference power under a fixed path.

[0125] (3) Fixed jamming resource allocation (Fixed-JRA): An interference resource allocation method based on fixed paths and interference power allocation.

[0126] Since the penetration paths under the guidance of the policy network at different moments of training are different, the time consumed by each test round is different, which makes it difficult to compare the cumulative rewards of rounds at different stages. Therefore, the average reward of the round is used as the indicator to evaluate the performance of the algorithm. Figure 6As shown in the figure, the horizontal axis represents the number of test rounds, and the vertical axis represents the average reward per round. It can be seen that the average reward per round using the hierarchical reinforcement learning algorithm is much higher than that of the other two algorithms. Furthermore, the proposed method introduces advantage replay technology, which allows for faster identification of advantageous actions in the early stages, accelerating the algorithm's convergence process.

[0127] Figure 7 The safety factor convergence curves of different algorithms during the penetration process were compared. The fixed strategy struggled to autonomously adjust the power allocation scheme based on environmental changes, resulting in the lowest safety factor. However, the joint allocation strategy, which utilizes hierarchical reinforcement learning, was able to rationally allocate jamming power while avoiding high-risk targets, resulting in a higher penetration safety factor than traditional methods. This validates the superiority of the joint allocation algorithm in terms of penetration safety.

[0128] By training and optimizing the model, the path-power joint allocation method based on improved hierarchical reinforcement learning can select flight trajectories with lower detection probabilities when interference resources are limited and interfere with high-threat radars as much as possible, thereby improving the safety of our aircraft cluster.

[0129] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.

Claims

1. A path-power joint allocation method based on improved hierarchical reinforcement learning, characterized by: The specific steps are as follows: Step 1: Construct a jamming countermeasure scenario: A jamming countermeasure scenario is constructed within a pre-set simulation environment, including a jammer swarm and target aircraft, as well as multiple radars. A jamming resource allocation model is established, using frequency-modulated noise as a means of suppressing the jammers. A radar detection model is established, selecting the rank-1 criterion as the radar information fusion rule to calculate the system detection probability. Step 2: The jamming resource allocation problem is modeled as a Markov decision process and solved using a continuous proximal strategy optimization algorithm. The joint path-power decision task during the jamming process is decomposed into two subtasks, each executed by an independent decision maker. The upper-level jamming path decider plans the flight direction based on the current distance between the fleet and the radar. The lower-layer interference power decision maker combines the flight direction to generate an interference resource allocation matrix. At the same time, the state set, action set, reward function, and state transition probability of the two decision makers are set for different tasks. Step 3: Use action key encoding (AKE) and advantage experience replay (AER) technology to optimize the decision process, and perform joint interference control based on the flight direction of the path decider and the allocation strategy of the power decider.

2. The path-power joint allocation method based on improved hierarchical reinforcement learning according to claim 1, characterized in that: The step 1 comprises the following steps: Step 1.1: Obtain the target aircraft, jammer, and radar at time k. The jamming signal transmitted by jammer m to radar n at time k is as follows: in, is the amplitude of the noise FM signal, f c is the initial frequency of the FM signal, K FM Used to control the size of the frequency increment, u(t) is white noise, is the initial phase of the signal, and its value is randomly selected with the same probability in [0,2π). At the same time, u(t) is The two are independent of each other; and They represent the interference beam distribution coefficient and interference power distribution coefficient transmitted by jammer m to radar n at time k respectively; At time k, the interference resource allocation matrix of the jammer group is established, namely, the interference resource allocation model: Step 1.3: Build a radar detection model and calculate the radar signal-to-interference ratio using the target echo signal power, receiver thermal noise, and interference signal power: in, is the target echo signal power, P N is the receiver thermal noise, is the interference signal power, P t is the transmission power of a single radar, G t Radar antenna gain, λ is the wavelength of the radar transmission signal, σ is the reflection area of ​​the target aircraft, is the distance between the target aircraft and radar n at time k, G j represents the energy gain of the jammer on the transmitted signal, λ j is the wavelength of the interference signal, γ j represents the polarization loss, is the distance between jammer m and radar n at time k, It represents the system gain of radar n when echoing signal in the direction of jammer m; Step 1.4: Calculate the system detection probability based on the radar information fusion rule based on the rank 1 criterion The detection probability of target n by radar at time k Where V T is the detection threshold.

3. The path-power joint allocation method based on improved hierarchical reinforcement learning according to claim 2, characterized in that: The interference signal power for: described The angle between the main lobe direction and this direction is small The specific relationship is: Among them, β rad is an empirical constant, θ 0.5 is the beam width of the radar antenna.

4. The path-power joint allocation method based on improved hierarchical reinforcement learning according to claim 1, characterized in that: The continuous proximal strategy optimization algorithm described in step 2 adopts Beta distribution as the strategy function of the algorithm, and strictly limits the sampling results to the interval [0,1].

5. The path-power joint allocation method based on improved hierarchical reinforcement learning according to claim 1 or 4, characterized in that: The state sets of the two decision makers are: in, represents the Euclidean distance between the target aircraft and the radar at time k, θ path is the flight direction of the jammer group, Pd j represents the detection probability of the radar network under interference conditions; The action sets of the two decision makers are: A=[A path ,A power ] Among them, A path =[θ path ] is the flight direction of the drone, with the direction directly in front of the starting position as 0, and the value range is [-π / 2, -π / 2], A power =[u·P] is the interference resource allocation matrix; The reward function of the path decision maker is: Among them, Pd o represents the detection probability of the radar network when there is no interference; The reward function of the power decider is: The state transition probability P of the two decision makers is the dynamic characteristic of the environment.

6. The path-power joint allocation method based on improved hierarchical reinforcement learning according to claim 1, characterized in that: The action key code AKE in step 3 is used to encode the interference power allocation strategy, the integer part of the encoded real number represents the target radar number, and the decimal part identifies the power allocation ratio; Adopting the AER technology to accelerate training: constructing a common experience set D and an AER set D′, and randomly sampling from the two sets. The AER set must meet the following requirements: Among them, R sum is the cumulative reward value of this round, R avg The average value of historical cumulative rewards; Sample training data from D′ according to probability β.

7. The path-power joint allocation method based on improved hierarchical reinforcement learning according to claim 6, characterized in that: The action key code AKE is used to number the interference beams of all jammers, and then the power decision maker generates the corresponding interference key. The action space of the power decision maker is: <h2 style=";text-align:left;direction:ltr">A<h2 style=";text-align:left;direction:ltr"> power <h2 style=";text-align:left;direction:ltr"> =[a1,a2,...,a<h2 style=";text-align:left;direction:ltr"> i <h2 style=";text-align:left;direction:ltr"> ],0≤a<h2 style=";text-align:left;direction:ltr"> i <h2 style=";text-align:left;direction:ltr"> ≤L Among them, a i It represents the interference power of each beam obtained by decoding according to the AKE rule, and L is the total number of interference beams.

8. The path-power joint allocation method based on improved hierarchical reinforcement learning according to claim 6, characterized in that: Said β=5%.

9. A computer device / apparatus / system comprising a memory, a processor, and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 8.