Networking multifunctional radar interference decision-making method based on threat level sorting

Through the interference decision-making method based on Viagra level sorting, Markov decision-making process is modeled on the network multifunction radar scenario, and the interference decision-making network is used to optimize the interference strategy, which solves the problem that traditional interference means are difficult to effectively combat the network multifunction radar, and achieves efficient and accurate interference decision-making.

CN120178174APending Publication Date: 2025-06-20XIDIAN UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510277688.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

Traditional interference means are difficult to effectively combat highly integrated and intelligently collaborative networked multi-function radars, especially when the number of jammers and radars is not equal and the interference resources are limited, the high algorithm complexity leads to low decision-making efficiency.

Method used

The interference decision-making method based on threat level sorting is adopted to model the network multi-function radar scenario through the Markov decision-making process, and the interference decision-making network is used to optimize the interference strategy based on the radar threat level sorting and interference style mask matrix.

Benefits of technology

It effectively improves the algorithm learning efficiency, improves decision-making accuracy, and can quickly generate collaborative interference resource scheduling strategies under limited resources, improving the interference capability of networking multi-function radar.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120178174A_ABST
    Figure CN120178174A_ABST
Patent Text Reader

Abstract

The invention discloses a networking multifunctional radar interference decision-making method based on threat level sorting. The method comprises the following steps: A, determining the working mode of each radar of a networking multifunctional radar; b, performing threat level sorting on the radar according to the working mode; c, inputting the interference pattern mask matrix and the one-hot code of the radar with the highest threat level into an interference decision network according to the radar sorting, so that the network outputs an interference pattern; d, updating the interference pattern mask matrix according to the interference pattern selected at the previous moment, and performing interference decision on the radar of the next threat level; repeating the step D until the interference patterns of all radars are obtained; and the interference decision network is used for realizing a networking multifunctional radar formation anti-interference decision based on a Markov decision. According to the method, the algorithm learning efficiency is effectively improved, the decision-making precision is improved, and the cooperative interference resource scheduling strategy can be quickly and effectively generated under the conditions that the number of jammers is not equal to that of radars and interference resources are limited.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of radar, and particularly relates to a networking multi-functional radar interference decision-making method based on threat level ranking. Background Art

[0002] With the rapid progress of information technology, networking multi-functional radars have gradually been highly regarded for their excellent multi-task processing, waveform diversification, and powerful networked collaborative detection capabilities. Such radar systems make full use of the data of each single radar and conduct fusion and sharing, thereby greatly improving their overall performance. Through efficient information sharing and collaborative operations, the detection accuracy has been greatly improved. Traditional interference means often have difficulty effectively implementing interference when facing such highly integrated and intelligent collaborative radar networks. Therefore, the interference decision-making problem for networking multi-functional radars has become the research focus. How to design a fast-response, intelligent decision-making, and efficient execution interference strategy for the characteristics of networking multi-functional radars in a complex and changeable electromagnetic environment has become the key to determining the success or failure of interference operations.

[0003] The rapid development of artificial intelligence technology has provided a new solution for the interference decision-making of networking multi-functional radars. By introducing advanced machine learning and deep reinforcement learning algorithms, the interference system continuously learns and optimizes the interference strategy, thereby achieving effective interference against networking multi-functional radars and self-protection under threat detection. With the continuous progress of technology and the in-depth expansion of applications, the interference decision-making problem of networking multi-functional radars will be solved more comprehensively and deeply. However, the interference action space of networking multi-functional radars increases exponentially, which will bring huge pressure to the calculation of traditional reinforcement learning algorithms. The algorithm complexity increases exponentially, resulting in slow or difficult convergence of the algorithm, making it difficult to quickly and effectively achieve collaborative interference decision-making in the case of unequal numbers of jammers and radars and limited interference resources. Summary of the Invention

[0004] In order to solve the above problems existing in the prior art, the present invention provides a networking multi-functional radar interference decision-making method based on threat level ranking.

[0005] The technical problems to be solved by the present invention are realized through the following technical solutions:

[0006] A networking multi-functional radar interference decision-making method based on threat level ranking includes:

[0007] A. Determine the working modes of N radars of the networking multi-functional radar; each radar has E working modes, N>1; E>1;

[0008] B. Rank the threat levels of the N radars according to the working modes of each radar to obtain a radar ranking;

[0009] C. Sort by radar, and input the interference pattern mask matrix and the one-hot encoding of the radar with the highest threat level into the interference decision-making network, so that the interference decision-making network outputs the interference pattern for this radar;

[0010] D. Update the interference pattern mask matrix according to the selected interference pattern at the previous moment, and make an interference decision on the radar of the next threat level; Repeat step D until the interference patterns of all radars are obtained;

[0011] Among them, the interference decision-making network is used to implement the networked multi-functional radar formation penetration interference decision-making based on Markov decision-making; the interference pattern mask matrix is used to characterize whether the ML interference patterns that M jammers can generate are available; M is the number of jammers in the interference formation, and each jammer has L interference patterns, M>1, L>1; the one-hot encoding is used to encode the working mode of the radar.

[0012] Optionally, the interference decision-making network is obtained through online training based on multiple empirical samples;

[0013] The empirical samples include: input state, output action, action reward, transition state, and end signal; among them, the input state is the mask matrix and one-hot encoding input to the interference decision-making network during training, the output action is the interference pattern output by the interference decision-making network during training for the input state, and the action reward is the reward obtained by taking the output action in the input state; the transition state is the new state migrated to after taking the output action, and the end signal is used to indicate whether the penetration fails.

[0014] Optionally, the action reward is obtained in the following way:

[0015] If the penetration task is completed, the action reward is set to 100;

[0016] If the threat level of the radar decreases, the action reward is set to 1;

[0017] If the threat level of the radar increases, the action reward is set to -1;

[0018] If the threat level of the radar rises to the highest, the action reward is set to -100.

[0019] Optionally, when the threat level of the radar rises to the highest, the end signal is set to 1, indicating that the penetration fails.

[0020] Optionally, the method further includes: constructing a new empirical sample according to the interference pattern mask matrix, one-hot encoding input to the interference decision-making network in step C, and the interference pattern output by the interference decision-making network corresponding to it, and storing it in the empirical pool; the empirical pool is used to store empirical samples.

[0021] Optionally, the interference decision network is a DQN network; the DQN network includes an online network and a target network. The parameters of the online network are represented by θ, and the parameters of the target network are represented by θ - . The parameters of the target network are periodically copied from the online network.

[0022] Optionally, the online network includes a first fully connected layer, a first ReLU activation layer, a second fully connected layer, a second ReLU activation layer, a third fully connected layer, and a fourth fully connected layer connected in cascade in sequence; wherein, the first fully connected layer includes ML + E nodes, and the fourth fully connected layer includes ML nodes; the target network has the same structure as the online network.

[0023] Optionally, during the process of training the interference decision network, a loss function is used to calculate the loss value, and the parameters of the online network are updated according to the loss value;

[0024] The loss function is:

[0025]

[0026] where L MSE represents the loss value, N H represents the number of experience samples, r represents the action reward in the experience sample, γ represents the discount factor, Q(s t , a t ; θ) represents the Q value of the online network when the input state is s t and the output action is a t , Q(s t+1 , a t+1 ; θ - ) represents the Q value of the target network when the input state is s t+1 and the output action is a t+1 , and t represents time.

[0027] Optionally, the method further includes: using the ML interference patterns of the M jammers to perform interference training on the networked multi-functional radar.

[0028] Optionally, the method is applied to an electronic device.

[0029] The multi-functional radar jamming decision-making method based on threat level ranking provided by the present invention models the scenario of a jammer formation breaking through a multi-functional radar network, establishes the jamming decision-making process as a Markov decision-making process, ranks the threat levels of each radar in the multi-functional radar network and then makes jamming decisions in sequence, effectively reducing the dimension of the action space, encodes the radar working modes using one-hot encoding, introduces a jamming pattern mask matrix to constrain the jamming patterns. Finally, the optimal jamming strategy is obtained using the jamming decision-making network. Compared with existing methods, the jamming decision-making method proposed by the present invention effectively improves the learning efficiency of the algorithm and the decision-making accuracy, and can quickly and effectively generate a cooperative jamming resource scheduling strategy when the number of jammers and radars is not equal and the jamming resources are limited. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 FIG. is a flowchart of a multi-functional radar jamming decision-making method based on threat level ranking provided by the present invention;

[0031] Figure 2 FIG. is a jamming decision-making framework diagram adopted by the present invention;

[0032] Figure 3 FIG. is the DQN network structure adopted in the present invention;

[0033] Figure 4 FIG. exemplarily shows the process of online training a jamming decision-making network of the present invention;

[0034] Figure 5 FIG. exemplarily shows the radar working mode transition relationship in the jamming decision-making process of an experiment of the present invention;

[0035] Figure 6 FIG. exemplarily shows the breakthrough distance curve graph in an experiment of the present invention;

[0036] Figure 7 FIG. exemplarily shows the reward curve graph in an experiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0037] The following further describes the present invention in detail with reference to specific embodiments, but the embodiments of the present invention are not limited thereto.

[0038] In order to quickly and effectively generate a cooperative jamming resource scheduling strategy when the number of jammers and radars is not equal and the jamming resources are limited, an embodiment of the present invention provides a multi-functional radar jamming decision-making method based on threat level ranking. This method uses a Markov decision-making process to model the problem of a jammer formation breaking through a multi-functional radar network, establishes the jamming decision-making process as a Markov decision-making process, and uses a jamming decision-making network to obtain the optimal jamming strategy.

[0039] Specifically, the Markov Decision Process (MDP) is a mathematical model used to describe the decision-making of an agent in an environment with Markov properties. It is described by a five-tuple {S, A, P, R, D}, where S is the state space, A is the action space, and P is the state transition probability, that is, the agent is in the current state s t ∈S takes action a t ∈A, the environment state is transferred to s t+1 The probability p(s t+1 |s t ,a t ), t represents time. R is the reward function, that is, the agent is in state s t Take action a t , the environment state is transferred to s t+1 The reward r(s) obtained by the agent t+1 |s t ,a t ), D is the end signal, D = 0 means not ended, D = 1 means ended. The goal of MDP is to find a strategy π, which is a mapping from state to action, which tells the agent which action should be taken in each state to maximize its long-term accumulated reward.

[0040] In the penetration network multifunctional radar scenario, there are N radars and M jammers. Each jammer can generate K beams, so a total of MK beams can be generated. Each jammer has L interference patterns, so a total of ML interference patterns, and each jammer can only generate one interference pattern at the same time, N>1, M>1, L>1, K≥1. The present invention uses a Markov decision process to model the above penetration network multifunctional radar scenario.

[0041] Specifically, the state space S is defined as Characterize the working mode of radar i at time t, and N represents the number of radars. Each radar has 6 working modes, Mod1 to Mod6, where Mod1 has the lowest threat level and Mod6 has the highest threat level. Define the action space A as Indicates that radar i is detected by the fth th There are jamming patterns, and the jamming patterns are {jam1, jam2, …, jam ML}, respectively, noise amplitude modulation interference, noise frequency modulation interference, dense false target interference, comb spectrum interference, slice forwarding interference, distance deception interference, speed deception interference, etc. The state transition probability P is a dynamic characteristic of the environment, and the interference decision network does not directly output or learn it. The reward function R is used to evaluate the interference effectiveness after taking a specific interference. In this invention, it is defined according to the change of radar threat level before and after the interference, and the action reward r is issued. tThe function of the reward function is realized in this way. Specifically, the action reward is obtained in the following way: if the penetration mission is completed after taking the action, the action reward is set to 100; if the threat level of the radar is reduced after taking the action, the action reward is set to 1; if the threat level of the radar is increased after taking the action, the action reward is set to -1; if the threat level of the radar is raised to the highest level after taking the action, the action reward is set to -100. When the working mode of the radar is transferred to Mod6, it means that the penetration fails, and the end signal D = 1, otherwise the end signal D = 0.

[0042] In the present invention, the interference decision goal is to train the neural network (referred to as the interference decision network in the present invention) so that it can output the optimal interference strategy that it can currently output under a given input state. The interference decision network with such capability can generate the interference strategy π in the present invention.

[0043] like Figure 1 As shown, the networked multi-function radar interference decision-making method based on threat level sorting provided by an embodiment of the present invention includes the following steps:

[0044] A. Determine the working modes of the N radars of the networked multifunctional radar; each radar has E working modes, N>1, E>1.

[0045] Specifically, the jammer determines the status information of each radar by receiving signals, and then collects the determined status information to the execution subject of the method of the present invention, which is an electronic device. The electronic device can be a central control jammer in the formation where the jammer is located, or it can be an interference control platform that all jammers communicate with.

[0046] B. Sort the threat levels of N radars according to their working modes to obtain radar ranking.

[0047] Specifically, radars with higher threat levels are placed in front, and radars with lower threat levels are placed in the back.

[0048] C. According to the radar sorting, the interference pattern mask matrix and the unique hot encoding of the radar with the highest current threat level are input into the interference decision network, so that the interference decision network outputs the interference pattern for the radar.

[0049] D. Update the interference pattern mask matrix according to the interference pattern selected at the last moment, and make interference decisions for the radars of the next threat level; repeat step D until the interference patterns of all radars are obtained.

[0050] Among them, the interference pattern mask matrix is used to characterize whether the ML interference patterns that M jammers can generate are available, so as to constrain the interference patterns by using the interference pattern mask matrix. For example, in the interference pattern mask matrix, 1 can be used to indicate that the interference pattern is available, and 0 can be used to indicate that the interference pattern is unavailable. For example, assuming there are three interference patterns, interference pattern 1 and interference pattern 3 are available, and interference pattern 2 is unavailable, then the mask matrix can be expressed as [1, 0, 1].

[0051] One-hot encoding is used to encode the working modes of the radar. For example, the one-hot encoding of working mode Mod3 can be expressed as [0, 0, 1, 0, 0, 0].

[0052] The interference decision network is used for the interference decision of the multi-functional radar formation penetration based on Markov decision-making. The interference decision network is obtained through online training based on multiple experience samples. Each experience sample includes: input state, output action, action reward, transition state, and end signal; among them, the input state is the mask matrix and one-hot encoding input to the interference decision network during training, the output action is the interference pattern output by the interference decision network during training for the input state, the action reward is the reward obtained by taking the output action in the input state; the transition state is the new state migrated to after taking the output action, and the end signal is used to indicate whether the penetration fails. In order to make the layout of the specification clearer, the interference decision network and its online training process will be illustrated by examples in the following text.

[0053] In step C, since each jammer can only generate one interference pattern at the same time, it cannot be guaranteed that each radar is effectively jammed when the number of radars and jammers is not equal. Therefore, the interference pattern is preferentially decided for the radar with a high threat level.

[0054] For example, referring to Figure 2 , the working mode and interference pattern mask matrix of the radar with the highest threat level are used as the input of the interference decision network ( Figure 2 DQN in) to obtain the sub-interference action Update the interference pattern mask matrix according to the selected interference action, and select the radar with the next threat level for interference decision to obtain the sub-interference action Repeat this process until the interference patterns of all radars are obtained. Figure 2 denoted by the set J in t is represented.

[0055] Among them, the interference pattern mask matrix is updated according to the selected interference actions, which can be understood by referring to the following example: Suppose there are 3 jammers, each jammer can generate two beams, and each jammer has 3 interference patterns, so there are a total of 9 interference patterns. At the first step of decision-making, all interference patterns can be selected, and the interference pattern mask matrix is [1, 1, 1, 1, 1, 1, 1, 1, 1]. If the result of the first decision is interference pattern 1, since a jammer can only adopt one interference pattern at the same time, if jammer 1 has adopted interference pattern 1, it cannot adopt interference pattern 2 and interference pattern 3 anymore. Then the interference pattern mask matrix is updated to [1, 0, 0, 1, 1, 1, 1, 1, 1].

[0056] After obtaining the interference patterns of all radars, the interference patterns can be used to conduct interference training on the networked multi-functional radar, so as to improve the task skills of our interference formation, reduce the risk of being detected and intercepted by the opponent's radar, greatly enhance the task security, ensure the smooth progress, and at the same time help our networked multi-functional radar iterate and optimize to improve its anti-interference ability.

[0057] The interference decision-making method for networked multi-functional radar based on threat level ranking provided by the present invention models the scenario of the interference formation breaking through the networked multi-functional radar, establishes the interference decision-making process as a Markov decision-making process, ranks the threat levels of each radar in the networked multi-functional radar and then conducts interference decision-making in turn, effectively reducing the dimension of the action space, encoding the radar working mode using one-hot encoding, and introducing an interference pattern mask matrix to constrain the interference patterns. Finally, the optimal interference strategy is obtained using the interference decision-making network. Compared with the existing methods, the interference decision-making method proposed by the present invention effectively improves the algorithm learning efficiency and the decision-making accuracy, and can quickly and effectively generate a cooperative interference resource scheduling strategy under the condition that the number of jammers and radars is not equal and the interference resources are limited.

[0058] In a preferred implementation manner, the method proposed by the present invention may further include: constructing a new experience sample according to the interference pattern mask matrix, one-hot encoding input to the interference decision-making network in step C, and the interference pattern corresponding to the output of the interference decision-making network, and storing it in the experience pool; this experience pool is used to store experience samples.

[0059] It can be understood that when step C is executed, the interference decision-making network has undergone preliminary training. Therefore, the experience sample constructed according to the interference pattern mask matrix, one-hot encoding input to the interference decision-making network, and the interference pattern corresponding to the output of the interference decision-making network is a valid experience sample and can expand the experience pool.

[0060] The following gives a detailed example of the interference decision-making network and its online training process.

[0061] In the embodiments of the present invention, it is preferably to use a DQN network as the interference decision network. The actual output of the DQN network is the Q value of each action. The action with the largest Q value is defined as the output action of the interference decision network, that is, the interference decision result.

[0062] It can be understood that the DQN network includes an online network and a target network. The parameters of the online network are represented by θ, and the parameters of the target network are represented by θ - '. The parameters of the target network are periodically copied from the online network. The role of the target network is to provide a stable Q value, thereby improving the training stability.

[0063] Among them, the online network and the target network have the same structure. As Figure 3 shown, it includes a first fully connected layer, a first ReLU activation layer, a second fully connected layer, a second ReLU activation layer, a third fully connected layer, and a fourth fully connected layer connected in series in sequence; among them, the first fully connected layer includes ML + E nodes, and the fourth fully connected layer includes ML nodes.

[0064] Figure 4 An exemplary process of online training the interference decision network is shown, including:

[0065] (1) Initialize parameters: Set the total number of training games N round , the maximum number of games in a single round N max , initialize the interference decision network, initialize the experience pool H, and the sampling number N of experience samples H ;

[0066] (2) Initialize the environmental state: Initialize the number of radars N, the number of jammers M, etc.

[0067] (3) Sense the radar state: Reconnaissance and sense the radar state S t ;

[0068] (4) Determine whether the penetration mission fails: Determine whether there is Mod6 ∈ S t . If so, the penetration fails, and return to step (2) to start a new round of training. Otherwise, enter step (5);

[0069] (5) Threat level sorting: Sort the radars according to the threat levels of each radar;

[0070] (6) Interference decision: Concatenate the interference pattern mask matrix and the one-hot encoding of the radar with the highest current threat level to form the input state of the interference decision network, and generate the output action through forward propagation, that is, output the interference pattern

[0071] (7) Implement interference: Use the interference pattern output by the interference decision network to interfere with the radar;

[0072] (8) Update radar status: After interference, detect the new radar status Issue action reward r t , and store it in the experience pool;

[0073] (9) Update network parameters: Using the sampling method of prioritized experience replay, sample N H experience samples from the experience pool H for training and updating the online network. Calculate the loss value using the loss function and update the parameters of the online network according to the loss value. Specifically, backpropagate the loss value to update the parameters θ of the online network, and copy the parameters θ of the online network to the parameters θ of the target network every once in a while. - .

[0074] Among them, the loss function is as follows:

[0075]

[0076] In this loss function, L MSE represents the loss value, N H represents the number of experience samples, r represents the action reward in the experience sample, γ represents the discount factor, Q(s t ,a t ; θ) represents the Q value of the online network when the input state is s t , and the output action is a t . Q(s t+1 ,a t+1 ; θ - ) represents the Q value of the target network when the input state is s t+1 , and the output action is a t+1 .

[0077] (10) Determine whether the penetration mission is completed: Each interference decision made by the interference decision network is regarded as a game. For a single penetration mission, a round of games including multiple interference decisions will be carried out. Determine whether the penetration is successful by judging whether the number of single-round games for the current penetration mission does not exceed N max ; if it does not exceed N max , the penetration is not successful, and return to the step of "detecting radar status" to continue completing the penetration mission; if it exceeds N max , the penetration is successful.

[0078] (11) Determine whether the maximum number of training rounds is reached: When the penetration is successful, determine whether the total number of training games reaches N round . If it reaches, end the training of the interference decision network, otherwise return to the step of initializing the environmental state.

[0079] Based on Figure 4For the process shown, the entire confrontation process in the present invention can be described as follows: determining the status information of each radar by intercepting signals, then adopting a certain jamming pattern according to the jamming strategy, and after being jammed, the radar status transfers to a new state and the change in the radar threat level is analyzed to obtain a reward. During the continuous confrontation process, the jammer continuously updates the jamming strategy through the reward to obtain the maximum reward, realizing the adaptive jamming of the networked multi-functional radar.

[0080] The effects of the present invention can be further illustrated by the following simulation experiments.

[0081] Suppose there are 3 jammers and 6 radars. Each jammer can generate two beams, and each jammer has 3 jamming patterns. Each jammer can only generate one jamming pattern at the same time. Then there are a total of 6 beams and 9 jamming patterns, which are {jam1, jam2, …, jam ML}, and there are 6 radar operating modes, which are {Mod1, Mod2, …, Mod6}, where Mod1 has the lowest threat level and Mod6 has the highest threat level. Figure 5 Exemplarily shows the relationship of radar operating mode transfer in the jamming decision process. When the penetration distance of the jamming formation reaches 40 km, the penetration is successful, and when the radar operating mode transfers to s6, the penetration fails.

[0082] To verify the performance of the jamming decision algorithm based on threat level ranking (Threat LevelsRanking-based Jamming Decision, TLRJD) proposed in the present invention, it is compared with the jamming decision algorithm based on DQN in the same scenario. Here, the jamming decision algorithm based on DQN is a method for jamming decision of networked multi-functional radars that uses DQN for jamming decision, does not rank the radars, and does not use a jamming pattern mask matrix to constrain the jamming patterns. The comparison results are as Figure 6 and Figure 7 shown. Figure 6 Shown is the penetration distance of each round of training. As the number of training rounds increases, the penetration distance of TLRJD and DQN in each round is increasing. The penetration distances of the TLRJD and DQN algorithms both converge to 40 km, indicating that both algorithms can learn effective strategies to enable the jamming formation to complete the penetration task. However, the TLRJD algorithm can converge around 3900 rounds, while DQN converges around 6000 rounds, proving that the proposed algorithm can effectively improve the learning efficiency. Figure 7 Shown is the cumulative reward of each round of training. When the penetration is successful, a reward of 100 is obtained, and when the penetration fails, a reward of -100 is obtained. The reward of the TLRJD algorithm converges to 98, while the reward of the DQN algorithm converges to 91, proving that the proposed algorithm can effectively improve the decision-making accuracy. It can be seen that the method proposed in the present invention can quickly and effectively generate jamming strategies in the scenario of a jamming formation penetrating a networked multi-functional radar.

[0083] In summary, in the method proposed by the present invention, a multi-functional radar scenario for establishing a jammer formation penetration network is established, and the jammer decision-making process is established as a Markov decision-making process. Based on the deep reinforcement learning algorithm DQN, the threat levels of each radar in the networked multi-functional radar are sorted and then jammer decisions are made in sequence, effectively reducing the dimensionality of the action space. The operating modes of the radars are encoded using a one-hot matrix, and a mask matrix is introduced to constrain the jammer patterns, improving the learning efficiency and decision-making accuracy of the jammer strategy and reducing the trial-and-error cost.

[0084] The method provided by the embodiments of the present invention can be applied to electronic devices. Specifically, the electronic device can be: a desktop computer, a portable computer, a server, an aircraft, etc. There is no limitation here, and any electronic device that can implement the present invention belongs to the protection scope of the present invention.

[0085] It should be noted that the terms "first", "second", etc. are used to distinguish similar objects and do not necessarily have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present invention.

[0086] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, those skilled in the art can combine and combine the different embodiments or examples described in this specification.

[0087] Although the present invention has been described in connection with various embodiments herein, however, in the process of implementing the claimed invention, those skilled in the art can understand and implement other variations of the disclosed embodiments by viewing the accompanying drawings and the disclosed content. In the description of the present invention, the term "including" does not exclude other components or steps, the term "a" or "one" does not exclude a plurality of cases, and the meaning of "a plurality" is two or more, unless otherwise specifically defined. In addition, certain measures are described in different embodiments, but this does not mean that these measures cannot be combined to produce good results.

[0088] The above content is a further detailed description of the present invention in combination with specific preferred embodiments. It cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, several simple deductions or substitutions can still be made, and all should be regarded as belonging to the protection scope of the present invention.

Claims

1. A networked multi-function radar jamming decision-making method based on threat level ranking, characterized in that: include: A. Determine the working modes of N radars of the networked multifunctional radar; each radar has E working modes, N>1; E>1; B. Sort the N radars by threat level according to the working mode of each radar to obtain a radar ranking; C. According to the radar sorting, the interference pattern mask matrix and the unique hot code of the radar with the highest threat level are input into the interference decision network, so that the interference decision network outputs the interference pattern for the radar; D. Update the interference pattern mask matrix according to the interference pattern selected at the last moment, and make interference decisions for the radars of the next threat level; repeat step D until the interference patterns of all radars are obtained; Among them, the interference decision network is used to realize the networked multi-function radar formation penetration interference decision based on Markov decision; the interference pattern mask matrix is ​​used to characterize whether the ML interference patterns that can be generated by M jammers are available; M is the number of jammers in the interference formation, each jammer has L interference patterns, M>1, L>1; the one-hot encoding is used to encode the working mode of the radar.

2. The networked multi-function radar jamming decision-making method based on threat level ranking according to claim 1 is characterized in that: The interference decision network is obtained by online training based on multiple experience samples; The experience sample includes: input state, output action, action reward, transition state and end signal; wherein, the input state is the mask matrix and unique hot encoding input to the interference decision network in training, the output action is the interference pattern output by the interference decision network in training corresponding to the input state, the action reward is the reward obtained by taking the output action under the input state; the transition state is the new state to which the output action is transferred after taking the output action, and the end signal is used to indicate whether the penetration fails.

3. The networked multi-function radar jamming decision-making method based on threat level ranking according to claim 2 is characterized in that: The action reward is obtained in the following way: If the breakthrough mission is completed, the action reward is set to 100; If the threat level of the radar is reduced, the action reward is set to 1; If the threat level on the radar is raised, the action reward is set to -1; If the threat level on the radar is raised to maximum, the action reward is set to -100.

4. The networked multi-function radar jamming decision-making method based on threat level ranking according to claim 3 is characterized in that: When the radar threat level reaches the highest level, the end signal is set to 1, indicating that the penetration fails.

5. The networked multi-function radar jamming decision-making method based on threat level ranking according to claim 2 is characterized in that: The method also includes: constructing new experience samples and storing them in an experience pool according to the interference pattern mask matrix and one-hot encoding input to the interference decision network in step C, and the interference pattern corresponding to the output of the interference decision network; the experience pool is used to store experience samples.

6. The networked multi-function radar jamming decision-making method based on threat level ranking according to claim 2 is characterized in that: The interference decision network is a DQN network; the DQN network includes an online network and a target network, the parameters of the online network are represented by θ, and the parameters of the target network are represented by θ - It means that the parameters of the target network are copied from the online network periodically.

7. The networked multi-function radar jamming decision-making method based on threat level ranking according to claim 6 is characterized in that: The online network includes a first fully connected layer, a first ReLU activation layer, a second fully connected layer, a second ReLU activation layer, a third fully connected layer and a fourth fully connected layer, which are cascaded in sequence; wherein the first fully connected layer includes ML+E nodes, and the fourth fully connected layer includes ML nodes; the target network has the same structure as the online network.

8. The networked multi-function radar jamming decision-making method based on threat level ranking according to claim 6 is characterized in that: In the process of training the interference decision network, a loss value is calculated using a loss function, and a parameter of the online network is updated according to the loss value; The loss function is: Among them, L MSE Represents the loss value, N H represents the number of experience samples, r represents the action reward in the experience sample, γ represents the discount factor, Q(s t ,a t ; θ) represents the online network when the input state is s t , output action is a t Q value, Q(s t+1 ,a t+1 θ - ) indicates that the target network is in the input state s t+1 , output action is a t+1 is the Q value at time t represents time.

9. The networked multi-function radar jamming decision-making method based on threat level ranking according to claim 1 is characterized in that: The method further includes: using the M L jamming patterns of the M jammers to perform jamming training on the networked multifunctional radar.

10. The networked multifunctional radar jamming decision-making method based on threat level ranking according to any one of claims 1 to 9, characterized in that: Used in electronic equipment.

Citation Information

Cited By

  • Radar interference pattern decision-making method and device based on interference efficiency evaluation

    CN121232123A