Disturbance decision-making method based on iterative updating of composite factor disturbance report table
By combining the composite elements of interference methods and interference power, an interference return report is constructed and iterative update method and hierarchical analysis method are applied, which solves the problems of poor application and great limitations of the existing technology in complex electromagnetic interference environments, and achieves more effective interference decisions.
Patent Information
- Application Number
- CN202411647005.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-18
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2044-11-18
AI Technical Summary
The existing interference decision-making methods are poorly applied in complex electromagnetic interference environments and have great limitations.
By extracting the two elements with the greatest influence, the interference method and the interference power, it is combined into composite elements, and an interference return report is constructed, combining iterative update method and hierarchical analysis method to optimize the interference decision-making process.
In a complex electromagnetic interference environment, the application and effectiveness of interference decisions are improved, the limitations of template matching method and reinforcement learning method in this environment are overcome, and more diverse and practical interference strategies are realized.
Smart Images

Figure CN119577374B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of radar communication technology, and further relates to an interference decision method based on iteratively updating a composite element interference report table in the field of electronic countermeasure technology. The present invention can be used to generate an interference strategy for a corresponding jammer for a radar. Background Art
[0002] The optimal choice of jammer jamming strategy is the main content of jamming decision-making. It is necessary to correspond the radar state to the specific jamming strategy one by one. Therefore, building a scientific and comprehensive jamming strategy and a reasonable evaluation method is the basis for optimizing the jamming strategy. The selected indicators need to be as relevant as possible to the main influencing factors of the jamming, fit the actual application, and reduce the influence of the invisible factors. The corresponding evaluation method should also consider the factors as comprehensively as possible to maximize the evaluation of the given jamming strategy. At present, there are two main methods for jammer jamming decision-making: the first is a method based on template matching, which is a traditional jamming decision-making method. The core of the template matching-based method lies in sufficient prior knowledge, so the quality of the prior knowledge base directly determines the timeliness and accuracy of such jamming decision-making methods. The second is a method based on reinforcement learning, which is a relatively popular jamming decision-making method.
[0003] Zhu Bakun et al. published a method for interference decision-making based on template matching in their paper "A Review of Radar Interference Decision-making Technology Based on Reinforcement Learning" (Electro-Optics and Control, 2022, 29(4): 52-58, 111.). The implementation steps are as follows: the first step is to build an interference knowledge base, which stores the frequency parameters of different threat object signal samples and the corresponding optimal interference patterns; the second step is to intercept the threat object echo signal in the environment, and pre-process the echo signal to obtain its frequency parameters; the third step is to extract the frequency parameters of the echo signal and the threat object signal samples stored in the interference library, and calculate the distance between the frequency parameters of all threat object signal samples in the interference library and the echo signal frequency parameters; the third step is to select the threat object signal sample with the smallest distance, and output the best interference pattern matched at the end to complete the template matching process. The disadvantage of this method is that it depends on the accuracy of the mapping relationship between the signal and the matcher. When facing complex electromagnetic interference environments and new radar anti-interference strategies, it will be impossible to match effective strategies, and there is a problem of poor applicability of decision results.
[0004] Zhang Baikai et al. published a Q-learning-based interference decision method in their paper "Multifunctional Radar Cognitive Jamming Decision Method Based on Q-Learning" (Telecommunication Technology, 2020, 60(2): 129-136). The implementation steps of this method are as follows: the first step is to determine the working state of the radar and give the state change relationship of the radar under different jamming styles adopted by the jammer; the second step is to define the benefit value of the radar state transition under different jamming styles adopted by the jammer; the third step is to use the Q-learning algorithm for learning and updating, and end the learning when the Q value converges and the radar reaches the search state. This method realizes the optimal jamming style selection of the jammer under different radar states, and considers the influence of the jamming style on the radar state transition reward in the jamming decision. However, the method still has the disadvantage that when defining the radar state change relationship and the benefit value of the radar state transition, it is only classified by the different jamming styles. When the types of jamming resource allocation increase and the radar anti-jamming measures are complicated, the corresponding relationship between the jamming style and the radar state transition reward becomes vague. In actual application scenarios, there is a problem of large limitations in the decision results. Summary of the invention
[0005] The purpose of the present invention is to address the problems existing in the above-mentioned prior art and to provide an interference decision method based on iterative updating of a composite factor interference report table, aiming to solve the problems that the existing interference decision methods have poor applicability in complex electromagnetic interference environments and the large limitations of the existing decision methods.
[0006] The technical idea for achieving the purpose of the present invention is that the present invention extracts the two elements that have the greatest impact on the interference strategy, namely, the interference mode and the interference power, and combines these two elements into a composite element, and sets 8 different combinations of composite elements. According to the different combinations of composite elements, an interference report table is constructed to achieve a one-to-one correspondence between the radar state and the interference strategy. The corresponding radar state transition probability table and state transition reward table are set for the 8 combinations of composite elements respectively. According to the contents of these tables, the interference decision background information of the iterative update method is improved, so that under different radar states, the interference strategy of the radar that can be affected by the jammer and the state that the radar can transfer are more diverse and closer to the complex electromagnetic interference environment, thereby solving the problem of large limitations of the reinforcement learning method in the prior art. The present invention uses an iterative update method for the interference decision-making process. In the process of obtaining the optimal interference strategy under a certain radar state, the interference report table is continuously iteratively updated, and when the content in the interference report table no longer changes, the interference strategy corresponding to the maximum interference report under each radar state in the interference report table is found, and the hierarchical analysis method is used to evaluate the effect of the interference strategy and finally output the best interference plan. The effectiveness of the obtained interference strategy can be verified at the first time, so that the interference effect achieved in a complex electromagnetic interference environment is more advantageous, thereby solving the problem of poor applicability of the template matching method adopted by the existing interference decision-making method in a complex electromagnetic interference environment.
[0007] The specific steps for achieving the purpose of the present invention are as follows: Step 1, combining the two elements of interference mode and interference power into a composite element to generate an interference strategy;
[0008] Step 2, generating a radar feature state transition probability table and a state transition reward table;
[0009] Step 3, construct a composite factor interference return table composed of the benefit values of different interference strategies adopted by the jammer under different radar working conditions;
[0010] Step 4, iteratively update the composite element interference report table and output the optimal interference strategy of the jammer corresponding to each working state of the radar.
[0011] Compared with the prior art, the present invention has the following advantages:
[0012] First, the present invention extracts the two factors that have the greatest impact on the interference strategy, namely the interference mode and the interference power, and combines these two factors into a composite factor. According to different combinations of the composite factors, the state transition probability matrix and the reward matrix are set one by one. According to the setting of these matrix contents, the interference decision background information of the iterative update method is improved, which overcomes the defects of the reinforcement learning method in the prior art. The present invention makes the interference strategy of the radar being affected by the jammer and the working state that the radar can transfer more diverse under different radar working states, and is closer to the complex electromagnetic interference environment.
[0013] Second, the present invention constructs an interference report table that corresponds one-to-one between radar states and interference strategies. The content of the interference report table gives the benefit values of different interference strategies adopted by the jammer under different radar working states, which reduces the running time of the algorithm, so that when the present invention finally outputs the interference strategy, the interference strategy with the highest benefit value can be directly selected.
[0014] Third, in the process of obtaining the optimal interference strategy under a certain radar working state, the interference report table is continuously iteratively updated, and when the content in the interference report table no longer changes, the hierarchical analysis method is used to evaluate the effect of the interference strategy. Compared with the traditional template matching method, it overcomes the defect of poor applicability in complex electromagnetic interference environments, so that the interference decision method generated by the present invention achieves more advantageous interference effects in actual application scenarios and has better applicability in complex electromagnetic interference environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 It is a flow chart of the present invention. DETAILED DESCRIPTION
[0016] The following is combined with Figure 1 , the specific implementation steps of the embodiment of the present invention are further described in detail.
[0017] Step 1: Design radar working status and interference strategy.
[0018] In the embodiment of the present invention, a jamming strategy generation process for point-to-point jamming of a multifunctional radar is considered. In this process, there is a functional radar and a jammer. In this process, the multifunctional radar has four working states, denoted as S(s 1 ,s 2 ,s 3 ,s 4 ), where s 1 Search for the target, s 2 For target confirmation, 3 For target tracking, s 4 For target identification.
[0019] In the embodiment of the present invention, the maximum jamming power of the jammer is 4kw, and it has two jamming modes: suppression jamming and deception jamming. The maximum jamming power of 4kw is discretized into 4 values of 1kw, 2kw, 3kw, and 4kw at intervals of 1kw. In the actual jammer and radar confrontation scenario, the parameter control of different jamming strategies is relatively complex. Considering the engineering application and hardware feasibility, in the embodiment of the present invention, the combination of the jamming style and the above-mentioned discrete jamming power is selected as the jamming strategy, which is recorded as A(a 1 ,a 2 ,…,a 8 ), where a 1 、a 2 、a 3 、a 4 They represent suppression interference 1, suppression interference 2, suppression interference 3, and suppression interference 4, respectively, indicating that under suppression interference, the interference power is 1kw, 2kw, 3kw, and 4kw respectively. 5 、a 6 、a 7 、a 8 They represent deception interference 1, deception interference 2, deception interference 3, and deception interference 4 respectively, indicating that under deception interference, the interference powers are 1kw, 2kw, 3kw, and 4kw respectively.
[0020] Step 2: Design the radar feature state transition probability table and state transition reward table.
[0021] When the jammer adopts the relevant jamming strategy to counter the radar, the radar has a certain probability of changing its working state. At this time, the jammer will be given a certain reward. The decision-making process of the jammer is to continuously maximize the benefit value. Therefore, it is necessary to define the radar characteristic state transition probability table and the state transition reward table. Since the state transition probability and the transfer reward are related to the working state of the radar and the jamming strategy it is subjected to, the radar characteristic state transition probability table and the state transition reward table are determined for the given 8 jamming strategies. In the embodiment of the present invention, refer to the radar characteristic state transition probability table and the state transition reward table proposed in the paper "Joint Optimization of Jamming Type Selection and Power Control for Countering Multifunction Radar Based on Deep Reinforcement Learning" (IEEE Transactions on Aerospace and Electronic Systems 2023, 59 (4): 4651-4665.) published by Panzes et al., when there is only suppression jamming style or deception jamming style, the characteristic state transition probability table of the radar is shown in Tables 1 and 2, and the state transition reward table is shown in Tables 3 and 4. Among them, the radar's working states include search, confirmation, identification and tracking, which correspond to the radar working states in the embodiment of the present invention. The values in the characteristic state transition probability table represent the probability values of the radar when the working state is transferred; the values in the state transition reward table represent the reward values of the radar when the working state is transferred.
[0022] Table 1 Radar characteristic state transition probability table under suppression interference
[0023]
[0024] Table 2 Radar characteristic state transition probability table under deception interference
[0025]
[0026] Table 3. State transfer reward table under suppression interference
[0027]
[0028] Table 4. State transfer reward table under deception interference
[0029]
[0030] When constructing the radar characteristic state transition probability table and state transition reward table based on the interference style and power composite elements, the embodiment of the present invention refers to Table 1 and Table 2 and uses them as the radar characteristic state transition probability table of suppression interference 4 and deception interference 4. The radar characteristic state transition probability table of suppression interference 3 is modified on the basis of Table 1. Considering that suppression interference is used more in the search stage, the state transition probability in the search state is adjusted, and the corresponding probability of search to search is reduced to 0.8, and the corresponding probability of search to confirmation is increased to 0.2, and the rest remains unchanged. Suppression interference 2 is modified on the radar characteristic state transition probability table of suppression interference 3, and the corresponding probability of search to search is reduced to 0.7, and the corresponding probability of search to confirmation is increased to 0.3, and the rest remains unchanged; Suppression interference 1 is modified on the radar characteristic state transition probability table of suppression interference 2, and the corresponding probability of search to search is reduced to 0.6, and the corresponding probability of search to confirmation is increased to 0.4, and the rest remains unchanged. The radar characteristic state transition probability table of deception jammer 3 is modified on the basis of Table 2. Considering that deception jammers are used more in the identification and tracking stages, the state transition probabilities in the identification and tracking states are adjusted. The corresponding probability of identification to search is reduced to 0.55, the corresponding probability of identification to tracking is increased to 0.27, and the corresponding probability of identification to identification is increased to 0.18. The state transition probability in the tracking state is kept the same as the identification state. The radar characteristic state transition probability table of deception jammer 2 is modified on the basis of deception jammer 3. The corresponding probability of identification to search is reduced to 0.5, the corresponding probability of identification to tracking is increased to 0.3, and the corresponding probability of identification to identification is increased to 0.2. The state transition probability in the tracking state is kept the same as the identification state. The radar characteristic state transition probability table of deception jammer 1 is modified on the basis of deception jammer 2. The corresponding probability of identification to search is reduced to 0.45, the corresponding probability of identification to tracking is increased to 0.33, and the corresponding probability of identification to identification is increased to 0.22. The state transition probability in the tracking state is kept the same as the identification state. It can be seen from Table 3 and Table 4 that the state transition rewards of suppression interference and deception interference are different only in the search state. Therefore, when setting the state transition rewards of suppression interference 4 and deception interference 4, the values in Table 3 and Table 4 are directly used. For suppression interference 1, suppression interference 2, and suppression interference 3, the state transition reward is not related to the power setting, but only to the state change, so it can be set equal to suppression interference 4. The state transition reward setting of deception interference 1, deception interference 2, and deception interference 3 is equal to deception interference 4.
[0031] Table 5 Radar characteristic state transition probability table under suppression interference 1
[0032]
[0033] Table 6 State transition reward table under suppression interference 1
[0034]
[0035] Step 3: construct a composite factor interference return table consisting of the benefit values of different interference strategies adopted by the jammer under different radar working conditions.
[0036] The embodiment of the present invention creates a 4*6 matrix to represent the composite factor interference report table. The rows of the composite factor interference report table are composed of radar working states, namely target search, target confirmation, target tracking and target identification. The columns of the interference report table are composed of 8 interference strategies. The content of the interference report table represents the benefit value corresponding to different interference strategies adopted by the jammer.
[0037] The benefit value of the jammer adopting different jamming strategies is obtained by the following formula:
[0038] Q(s i ,a j )←Q(s i ,a j )+β[r(s i ,a j )+γmaxQ(s i ′,a j ′)-Q(s i ,a j )]
[0039] Among them, Q(·) represents the benefit value obtained by the jammer using the jamming strategy, s i 、s i ′ respectively represents the current working state of the radar and the next working state of the radar, i=1,2,3,4, a j 、a j ′ represents the jammer’s current jamming strategy and the next jamming strategy, respectively. j=1,2,…,8. β represents the learning rate, which ranges from 0 to 1. γ represents the discount rate, which ranges from 0 to 1. r(·) represents the immediate benefit value obtained by the jammer according to the state transition reward table. max represents the maximum value operation, and ← represents the assignment operation.
[0040] Step 4, iteratively update the composite factor interference report table.
[0041] Step 4.1, initialize the contents of the composite factor interference reward table to 0, set the parameters of the iterative update algorithm, including the learning rate β, discount factor γ, exploration rate ε and number of iterations num_episodes. Among them, the learning rate β determines the weight of new information relative to old information when updating the strategy or value. The discount factor γ is used to measure the importance of future rewards relative to current rewards, which determines the importance of the agent to future rewards when making decisions. The values of the learning rate β and the discount factor γ are both real numbers between 0 and 1. The number of iterations num_episodes refers to the number of cycles of the algorithm, that is, the number of times the jammer interacts with the radar and updates its strategy or value. Here, an exploration rate ε is also required, which means that the jammer exploits with a probability of ε and explores with a probability of 1-ε under each radar operation. Assuming that the jammer explores at a certain time, the jammer will randomly select a jamming strategy from all possible jamming strategies in this state and execute it and record the reward value. If a certain behavior is exploitation, the jammer will select the jamming strategy with the largest benefit value from all possible jamming strategies based on the experience gained during the exploration process and execute it. The reward values of different strategies are continuously updated during the algorithm iteration process, and form a composite factor interference reward table when the algorithm ends the iteration. According to the composite factor interference reward table, the optimal interference strategy under a certain radar working state is obtained.
[0042] Step 4.2, in the current iteration, set the radar's working state to target search, which means that the radar's working state is transferred from target search. Generate a random number and compare the random number with the exploration rate ε. If ε is greater than the random number, the exploration process is carried out, and a jamming strategy is randomly selected to execute. The next working state and the corresponding benefit value are obtained according to the state transition probability and reward. If ε is less than the random number, the utilization process is carried out, and the jamming strategy with the largest benefit value is selected and executed according to the results of the previous exploration.
[0043] Step 4.3, after selecting the jamming strategy, combine the current radar working state and the state transfer matrix to obtain the next working state of the radar, and obtain the reward obtained by the jammer during the change of the radar working state according to the state reward matrix. Use the same benefit formula as step 3 to recalculate the benefit value, find the maximum benefit value in the composite factor interference return table corresponding to the next working state of the current radar, and update the benefit value of the jammer taking the jamming strategy under the current radar working state.
[0044] Step 4.4, determine whether the current working state of the radar is the target search state; if so, execute step 4.5, otherwise, execute step 4.2.
[0045] Step 4.5, if the content change of the interference report table is less than the threshold value 0.01, the interference report table has stabilized, and the interference strategy of the jammer corresponding to each working state of the radar is selected as the optimal interference strategy according to the interference report table.
[0046] In order to verify the interference effect of the present invention, a judgment matrix A is constructed according to the four indicators of interference effect evaluation, namely, the overlap degree between the interference spectrum and the target signal spectrum, the size of the radiation power, the selection of the interference opportunity and the determination of the interference pattern. The judgment matrix A is a matrix used in the hierarchical analysis method to compare and evaluate the relative importance between different indicators. In the interference effect evaluation, the judgment matrix A is used to compare and evaluate the relative importance between the four indicators of the overlap degree between the interference spectrum and the target signal spectrum, the size of the radiation power, the selection of the interference opportunity and the determination of the interference pattern.
[0047] In the embodiment of the present invention, the magnitude of the radiation power and the determination of the interference pattern are more important than the other two factors, so the judgment matrix A is defined as follows:
[0048]
[0049] The element A of the judgment matrix A ij Indicates the importance ratio of the ith indicator to the jth indicator. Normally, the diagonal elements of the matrix are 1, indicating that the importance of each indicator relative to itself is 1. Except for the elements on the diagonal, the remaining elements of the judgment matrix A are inversely related on both sides of the diagonal elements. In this judgment matrix, the overlap degree between the interference spectrum and the target signal spectrum is the same as the importance of selecting the interference timing, the magnitude of the radiation power is the same as the importance of determining the interference pattern, and the magnitude of the radiation power (or the determination of the interference pattern) is twice as important as the overlap degree between the interference spectrum and the target signal spectrum (or the selection of the interference timing).
[0050] The consistency of the matrix A is determined by calculating its eigenvalues and eigenvectors. If the consistency test is passed, the weight solution is continued.
[0051] The embodiment of the present invention uses the following weight solution method: normalize each row of the judgment matrix A, and then calculate the average value of each row as the weight. The obtained weight is [0.170, 0.330, 0.170, 0.330].
[0052] Define an indicator vector (a, b, c, d). The elements in the indicator vector can only take values of 0 or 1. 0 represents that this parameter is not adopted, and 1 represents that this parameter is adopted. The interference decision method of the present invention simultaneously completes the interference style and interference power selection. Therefore, the values of the second and fourth elements in the indicator vector are 1, and the other two values are 0. At this time, the indicator vector is (0, 1, 0, 1). It is multiplied by the weight vector w to obtain an interference benefit value of 0.833.
Claims
1. An interference decision method based on iterative updating of composite factor interference report table, characterized in that: Generate the interference strategy of the composite factor, construct the composite factor interference report table, and iteratively update the composite factor interference report table; the specific steps of the decision-making method include the following: Step 1, combining the interference mode and interference power into a composite element to generate an interference strategy; Step 2, generating a radar feature state transition probability table and a state transition reward table; Step 3, construct a composite factor interference return table composed of the benefit values of different interference strategies adopted by the jammer under different radar working conditions; The rows of the composite factor interference report table are composed of radar working states, namely target search, target confirmation, target tracking and target identification, and the columns of the interference report table are composed of 8 interference strategies; the content of the interference report table represents the benefit value corresponding to the adoption of this interference strategy under this radar working state; Step 4, iteratively update the composite element interference report table, and output the optimal interference strategy of the jammer corresponding to each working state of the radar; The steps of iteratively updating the composite factor interference report table are as follows: The first step is to set the algorithm parameters, including learning rate β, discount factor γ, exploration rate ε, and number of iterations num_episodes, and initialize the contents of the interference reward value table to 0; The second step is to set the radar's working state to target search in the current iteration, which means that the radar's working state is transferred from target search. A random number is generated and compared with the exploration rate ε. If ε is greater than the random number, the exploration process is carried out and a jamming strategy is randomly selected to execute. The next working state and the corresponding benefit value are obtained according to the state transition probability and reward. If ε is less than the random number, the utilization process is carried out, and the interference strategy with the largest benefit value is selected and executed according to the results of the previous exploration; The third step is to select the jamming strategy, combine the current radar working state and the state transfer matrix to get the next working state of the radar, and get the reward obtained by the jammer in the process of radar working state change according to the state reward matrix, recalculate the benefit value using the benefit formula, find the maximum benefit value in the composite factor jamming return table corresponding to the next working state of the current radar, and update the benefit value of the jammer taking the jamming strategy in the current radar working state; The fourth step is to determine whether the current working state of the radar is a target search state; if so, execute the fifth step, otherwise, execute the second step; In the fifth step, if the content change of the interference report value table is less than the threshold value 0.01, the interference report value table has stabilized, and the interference strategy of the jammer corresponding to each working state of the radar is selected as the optimal interference strategy according to the interference report value table.
2. The interference decision method based on iterative updating of composite factor interference reporting table according to claim 1 is characterized in that: In step 1, the two elements of interference mode and interference power are combined into a composite element to generate an interference strategy, which means that the maximum interference power of the jammer is discretized into four interference power values at equal intervals; four suppression interference strategies after the suppression interference and interference power are combined are generated by combining the two elements corresponding to each other in a one-to-one manner; four deception interference strategies after the deception interference and interference power are combined are generated by combining the two elements corresponding to each other in a one-to-one manner.
3. The interference decision method based on iterative updating of composite factor interference reporting table according to claim 2 is characterized in that: The radar characteristic state transition probability table in step 2 refers to the probability of the radar being transferred to the next working state according to the current working state of the radar, the next working state of the radar after being interfered, and different working states of the radar. Each radar characteristic state transition probability table includes 4 working states of the radar. After one-to-one correspondence with 4 suppression interference strategies and 4 deception interference strategies, 8 radar characteristic state transition probability tables are obtained. Among them, Since suppression jamming is used more frequently in the search phase, only the state transition probability under the search state is adjusted. In the radar characteristic state transition probability table corresponding to each suppression jamming, the corresponding probabilities of search to search are decreased one by one, and the corresponding probabilities of search to confirmation are increased one by one. Since suppression interference is used more frequently in the identification and tracking stages, the state transition probabilities in the identification and tracking states are adjusted. In the radar feature state transition probability table corresponding to the deception interference, the corresponding probabilities of identification to search are decreased one by one, the corresponding probabilities of identification to tracking are increased one by one, and the corresponding probabilities of identification to identification are increased one by one. The state transition probability in the tracking state is adjusted to the same as that in the identification state.
4. The interference decision method based on iterative updating of composite factor interference reporting table according to claim 3 is characterized in that: There are 8 state transition reward tables in step 2, each of which includes 4 working states of the radar. After one-to-one correspondence with the 4 suppression jamming strategies and the 4 deception jamming strategies, 8 state transition reward tables are obtained, among which: The state transition reward table corresponding to one of the four types of suppression interference is directly obtained from the state transition reward table under the suppression interference; since three of the four types of suppression interference state transition rewards are irrelevant to the power setting and only related to the change of the radar working state, the state transition rewards of the remaining three types can be set equal to the other one; One state transition reward table corresponding to the four types of deceptive interference is directly obtained from the content of the state transition reward table under deceptive interference; the state transition rewards corresponding to the other three types of deceptive interference have nothing to do with the power setting, but are only related to the change in the radar working state, so the state transition rewards of the other three types can be set equal to that of another type.
5. The interference decision method based on iterative updating of composite factor interference reporting table according to claim 1 is characterized in that: The benefit value of the jammer adopting different jamming strategies in step 3 is obtained by the following formula: Q(s i ,a j )←Q(s i ,a j )+β[r(s i ,a j )+γmaxQ(s i ′,a j ′)-Q(s i ,a j )] Among them, Q(·) represents the benefit value obtained by the jammer using the jamming strategy, s i 、s i ′ respectively represents the current working state of the radar and the next working state of the radar, i=1,2,3,4, a j 、a j ′ represents the jammer’s current jamming strategy and the next jamming strategy, respectively. j=1,2,…,8. β represents the learning rate, which ranges from 0 to 1. γ represents the discount rate, which ranges from 0 to 1. r(·) represents the immediate benefit value obtained by the jammer according to the state transition reward table. max represents the maximum value operation, and ← represents the assignment operation.
Citation Information
Patent Citations
Radar interference strategy determining method, device, computer device and storage medium
CN109828245A
DQN radar interference decision-making method and device based on priority important sampling fusion
CN114814741A