Cooperative sequential interference decision-making method for cognitive communication countermeasure system
By constructing a collaborative sequential interference decision-making method in the cognitive communication adversarial system, using QMIX algorithm and graph neural network, the problem of mutual influence of agent decision-making results in multi-agent systems is solved, and the global optimal interference effect and the communication party's anti-interference ability are achieved.
Patent Information
- Application Number
- CN202510376496.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-06-27
AI Technical Summary
The existing cognitive communication anti-interference decision-making algorithm fails to effectively consider the task level of multi-agent system, resulting in the agents being unable to fully cooperate with each other in the decision-making process, affecting the overall effect of the decision-making results.
A collaborative sequential interference decision-making method for cognitive communication adversarial systems is proposed. By constructing cognitive communication adversarial scenarios, a model of communication and interference parties is built, and a QMIX algorithm and graph neural network are used to realize information fusion and coordinated decision-making between interfering agents.
This method can improve the ability of each agent of the interfering party to perceive and adapt to the team status when the environment is not completely known, achieve the optimal interference effect in the globally, and enhance the anti-interference ability of the communication party.
Smart Images

Figure CN120224243A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a collaborative sequential interference decision-making method, belonging to the field of communication technologies. Background Art
[0002] Electronic information technology has been widely applied to various weaponry and strategic operations in modern warfare. The electronic warfare capabilities of both sides have become the key to determining the outcome of a war. Cognitive communication countermeasure is an important research area in cognitive electronic warfare, which is divided into three aspects: communication reconnaissance, communication jamming, and communication anti-jamming. As the core issue of communication jamming, the main task of intelligent jamming decision-making is to transmit appropriate jamming signals based on the information obtained from reconnaissance to jam the enemy's communication equipment, and continuously adjust the jamming strategy according to the evaluation results. For a cognitive communication countermeasure system with a large scale, complex tasks, and incomplete individual observations, how to coordinate the jamming devices of the jamming party to work together, improve the ability of the jamming devices to perceive and adapt to the team state during the task execution, and obtain the optimal global jamming strategy of the jamming party is the difficulty and focus of this research direction.
[0003] With the significant improvement in computer computing power, researchers have combined artificial intelligence with intelligent jamming decision-making methods. AI decision-making based on reinforcement learning maximizes the expected long-term reward by spontaneously interacting with the environment, collecting data, and exploring new strategies. Zhang Baikai et al. applied the Q-learning algorithm to radar cognitive jamming decision-making, established an adversarial model, and studied the influence of various parameters of the reinforcement learning algorithm and the model transition probability on the decision-making performance. The development of deep learning has also promoted the development of reinforcement learning. DeepMind combined the two and proposed deep reinforcement learning, which was applied to solve the Go decision-making problem. The deep Q-network algorithm replaced the value function update method of the Q-learning table, better solving the model problem with a large number of states. In fact, a communication jamming system involves multiple objects to complete a certain task, and reinforcement learning has been further introduced into the multi-agent system, that is, multi-agent reinforcement learning. This type of method is mainly divided into two categories: the centralized value function method and the decomposed value function method. The MADDPG algorithm is a classic centralized value function method, which combines local information and global information in the environment, and uses the centralized value function to train the strategy of each agent to achieve the goal of global optimality. The QMIX algorithm proposed by Rashid et al. is a decomposed value function method. The value function of the agent is represented as a set of decomposable functions, each of which acts on the global state and local actions, and introduces the nonlinearity and decomposability of the network structure learning function, making this method applicable to a wider range of tasks. However, the above methods do not consider the problem that in the task level of the multi-agent system, agents need to cooperate with each other and the decision results of each agent will affect each other. Summary of the Invention
[0004] The present invention aims to solve the problem that the existing cognitive communication countermeasure interference decision-making algorithm does not consider that in a multi-agent system at the task level, agents need to cooperate with each other and the results of their respective decisions will affect each other. Therefore, a collaborative sequential interference decision-making method for a cognitive communication countermeasure system is proposed.
[0005] The technical solution adopted by the present invention to solve the above problems is as follows: The steps of the present invention include:
[0006] Step 1: Construct a cognitive communication countermeasure scenario, that is, establish the opposing parties as the communication side and the interference side: Set the communication side to consist of one base station node and five terminal nodes to form five communication links, and set the interference side to consist of four interference nodes to form an interference network;
[0007] Step 2: Construct an anti-interference node model for the communication side: In the reinforcement learning framework, set the communication parameters and status information at the current moment as the input of the communication side model, and output the communication parameters of each communication link to counteract interference;
[0008] Step 3: Construct an interference node model for the interference side, build an information fusion network and train it, and integrate the information of the interference side team: Construct the interference side scenario in the form of a graph, and fuse the local information of each interference node based on the graph neural network to obtain the team characteristics;
[0009] Step 4: Iteratively optimize the interference side decision parameters based on the QMIX algorithm, and perform collaborative sequential interference on the communication side: The QMIX algorithm model introduces an information fusion network, and the interference nodes output the interference side decision parameters according to the globally fused state information.
[0010] Further, in Step 1, all nodes of the interference side are equipped with communication modules, communicate with the remaining nodes to form an interference network to perform collaborative interference on the communication side, and destroy the normal communication ability of the communication side;
[0011] The communication side uses the TCP / IP protocol as the routing protocol and changes the communication parameters after being interfered with. The optional communication parameters of each communication link include:
[0012]
[0013] In formula (1), f represents the optional communication frequencies including f1, f2, and f3, c represents the optional modulation methods including BPSK, QPSK, and EPSK, which are respectively represented as c1, c2, and c3, and p represents the optional transmission power as 20%, 40%, 60%, 80%, and 100% of the maximum transmission power, which are respectively represented as p1, p2, and p3.
[0014] Further, after the communication party is interfered with in step 2, each communication link adjusts its communication parameters according to the communication parameters and status information at the current moment based on the deep Q - neural network algorithm, so as to avoid and counter the interference signal and ensure the stable communication of the communication party system.
[0015] Further, in step 3, an information fusion network is built and trained. The specific steps of integrating the information of the interfering party team include:
[0016] Step 301: Construct the data of the cognitive communication anti - interference party into a graph form, that is, G=(ν,ε), where ν represents the set of nodes, that is, each interference node of the interfering party, and ε represents the set of edges, that is, the communication relationship between each interference node of the interfering party; the interference node calculates the Euclidean distance d(a,b) between itself and other interference nodes, that is The interference node establishes an edge relationship with the interference nodes with the first K Euclidean distances to establish a K - nearest neighbor graph;
[0017] Step 302: The interference node obtains the set N i of K - nearest neighbor neighbor information, and calculates the similarity coefficient e ij between the interference node and the neighbor neighbor nodes, that is represents the importance degree of the neighbor neighbor node feature to the interference node, where W represents the trainable weight matrix, represents the interference node feature, represents the neighbor neighbor interference node feature;
[0018] Step 303: Use the Softmax function to normalize the similarity coefficient e ij to obtain the attention coefficient a ij between adjacent agents, that is
[0019] Step 304: Based on the attention coefficient a ij weight - sum the features of the neighbor neighbor nodes, calculate the output feature of the interference node that fuses the neighbor neighbor node information, and perform a pooling operation to obtain the team feature O all .
[0020] Further, in step 4, based on the QMIX algorithm model, an information fusion network is introduced to construct a cognitive communication anti - interference party interference node model, and the implementation steps of outputting the interference party decision parameters to perform cooperative sequential interference on the communication party include:
[0021] Step 401: Establish the agent evaluation network and the hybrid evaluation network to initialize the parameters θ agent_net of the main network and θ mixing_net , establish the agent target network and the hybrid target network with the same structure to initialize the target network parameters θa ′ gent_net With θ′ mixing_net , initialize the information fusion network, and initialize an experience replay pool D with a capacity of M to store the transition information <s t , a t , s t+1 , r t >. Set the target network update frequency P and the total number of iteration rounds T;
[0022] Step 402: Obtain the environmental state information s t , the information O of each interfering node j , the reward value R, and the interfering node selects an effective action where s t = [f t , c t , p t , that is, all the state information of the interfering nodes in the adversarial environment includes the frequency, modulation pattern, and power of the communication link of the communication party, that is, the information of the interfering nodes actually detected by each interfering node, that is, the algorithm can balance the relationship between the overall complete suppression of the interfering party and the utilization rate of interference resources;
[0023] Step 403: Construct the adjacency matrix A t Obtain the information of the interfering party team That is The interfering node fuses the information of the interfering party team with the information of each interfering node to obtain the observation information O as the input of the agent network;
[0024] Step 404: Each interfering node in the agent network learns the Q value of all interfering nodes selecting effective actions and uses the hidden state of the GRU recurrent layer as the next input;
[0025] Step 405: Each interfering node selects an interference decision action u j according to the interfering node information O t j to obtain the joint action
[0026] Step 406: The interfering nodes of the interfering party execute the joint action u t to obtain the new environmental state s t+1 and the reward value R;
[0027] Step 407: Store in the experience replay pool, and store the transition information <s t , a t , s t+1 , r t > into the experience replay pool D, where the transition information details include:
[0028]
[0029] Step 408: Update the interference party environment state and the optional effective actions of the interference nodes, i.e., s t = s t+1 , Repeat steps 404 to 408 until the experience replay pool is fully stored;
[0030] Step 409: Train the evaluation network. Randomly extract a batch of transfer information <s t , a t , s t+1 , r t > from the same positions corresponding to different cycles in the experience replay pool, update the parameters of the evaluation network, and repeat steps 402 to 409 until the evaluation network converges;
[0031] Step 410: Update the target network by copying the parameters of the evaluation network to the target network;
[0032] Step 411: Utilize the QMIX algorithm model of the introduced information fusion network after training. The interference nodes of the interference party fuse the obtained global state information according to the currently detected communication link state information, and the interference nodes make decisions on the globally optimal joint interference parameters to perform cooperative sequential interference on the communication signals of the communication party.
[0033] The beneficial effects of the present invention are as follows:
[0034] 1. The present invention is directed to a cognitive communication countermeasure interference party system, which enables each interference node to sense and adapt to the team state during execution, correlate with each other under the condition of incomplete environmental knowledge, and then cooperate to complete tasks, obtaining the best global interference effect;
[0035] 2. The communication party system of the present invention is set to communicate through the TCP / IP protocol and adopts the deep Q neural network algorithm with anti-interference ability as the anti-interference method. The present invention has a good global interference effect for this type of cognitive communication system;
[0036] 3. The present invention introduces a graph neural network into the QMIX algorithm. The information fusion ability of the graph neural network enables each agent in the interference party system to make full use of team information and improves the ability of the agent to sense and adapt to the team state during execution. Brief Description of the Drawings
[0037] Figure 1 is the flowchart of the present invention;
[0038] Figure 2 is the schematic diagram of the scenario of the cognitive communication countermeasure system;
[0039] Figure 3 It is the implementation flowchart of the deep Q - neural network algorithm for the communication - side system;
[0040] Figure 4 It is the schematic diagram of the QMIX algorithm introducing the information fusion network;
[0041] Figure 5 The implementation flowchart of the collaborative sequential interference decision - making method for the jammer - side system. Specific implementation manners
[0042] Specific implementation manner one: As Figures 1 to 5 shown, a collaborative sequential interference decision - making method for a cognitive communication countermeasure system, the specific steps include:
[0043] Step 1, construct a cognitive communication countermeasure scenario, that is, establish the two adversarial parties as the communication side and the jammer side: set the communication side to consist of one base - station node and five terminal nodes to form five communication links, and set the jammer side to consist of four jammer nodes to form a jammer network;
[0044] Step 2, construct an anti - jamming node model for the communication side: in the reinforcement - learning framework, set the communication parameters and state information at the current moment as the input of the communication - side model, and output the communication parameters of each communication link to counteract jamming;
[0045] Step 3, construct a jamming node model for the jammer side, build an information fusion network and train it, and integrate the information of the jammer - side team: construct the jammer - side scenario in the form of a graph, and based on the graph neural network, fuse the local information of each jammer node to obtain the team feature;
[0046] Step 4, based on the QMIX algorithm, iteratively optimize the decision - making parameters of the jammer side to perform collaborative sequential interference on the communication side: the QMIX algorithm model introduces an information fusion network, and the jammer nodes output the decision - making parameters of the jammer side according to the globally - fused state information.
[0047] Specific implementation manner two: As Figures 1 to 5 shown, in Step 1, all nodes of the jammer side are equipped with communication modules, communicate with the remaining nodes to form a jammer network to perform collaborative jamming on the communication side, and destroy the normal communication ability of the communication side;
[0048] The communication side uses the TCP / IP protocol as the routing protocol, and changes the communication parameters after being subjected to communication jamming. The optional communication parameters of each communication link include:
[0049]
[0050] In Formula (1), f represents the optional communication frequencies including f1, f2, and f3, c represents the optional modulation methods including BPSK, QPSK, and EPSK, which are respectively represented as c1, c2, and c3, and p represents the optional transmission powers of 20%, 40%, 60%, 80%, and 100% of the maximum transmission power, which are respectively represented as p1, p2, p3.
[0051] Among them, the interfering party node model includes four interfering nodes. All nodes are equipped with communication modules and can communicate with other nodes to form an interference network. Therefore, the interfering party can obtain the communication parameters of the communicating party based on the cooperative sequential signal detection and recognition method, perform cooperative interference on the communicating party, and disrupt the normal communication ability of the communicating party. The present invention assumes that after the communicating party is subjected to communication interference, the communication link will change the above communication parameters to avoid and counter the interference signal and ensure the stable communication of the communicating party system.
[0052] Specific Embodiment 3: As Figures 1 to 3 shown, in Step 2, constructing the anti-interference node model of the communicating party, the specific implementation steps include:
[0053] Step 201: Initialize the parameters θ of the main network by establishing an evaluation network, initialize the target network parameters θ′ of the target network with the same structure, and establish an experience replay pool with an initial capacity of D to store the transition information <s, a, r, s′>;
[0054] Step 202: Randomly select the initial state of the communicating party's communication node. Use the ε-greedy strategy to select the action a with the largest corresponding Q value from the evaluation network with a probability of ε t = [f t , c t , p t , and randomly select an action a from the action space with a probability of 1 - ε t = [f t , c t , p t , and set ε = 0.9;
[0055] Step 203: The communicating party's communication node executes the selected action a t = [f t , c t , p t , observes the reward value r t and the new state s t+1 , and stores the current transition information <s t , a t , r t , s t ′> into the experience replay pool, where the reward value is the bit error rate, that is The state space includes interference frequency, interference modulation method, and interference power, that is
[0056]
[0057] Step 204: Update the state of the communication node of the communication party, that is, s t = s t+1 , update the attenuation function ε, that is, ε = 0.1 + (0.9 - 0.1) × e -0.01 ;
[0058] Step 205: Repeat Step 202, Step 203, and Step 204 until all are stored in the experience replay pool;
[0059] Step 206: Train the evaluation network, randomly extract a batch of transition information [s j , a j , r j , s j+1 from the experience replay pool, calculate the TD target y j , that is, y j = r j + γ max a′ Q(s j+1 , a′; θ), and calculate the difference value loss using the TD target y j and the predicted value of the evaluation network, that is, loss = (y j - Q(s j , a j ; θ)) 2 , update the parameters of the evaluation network, and repeat Step 202, Step 203, Step 204, Step 205, and Step 206 until the evaluation network converges;
[0060] Step 207: Update the target network, copy the parameters of the evaluation network to the target network every certain number of steps;
[0061] Step 208: Using the trained deep Q - neural network algorithm, the communication node of the communication party adjusts the communication parameters according to the current state to counter the interference signal.
[0062] Specific Embodiment 4: As Figures 1 to 5 shown, in Step 3, build an information fusion network and train it. The specific steps of integrating the information of the interference party team include:
[0063] Step 301: Construct the data of the cognitive communication anti - interference party in the form of a graph, that is, G = (ν, ε), where ν represents the set of nodes, that is, each interference node of the interference party, and ε represents the set of edges, that is, the communication relationship between each interference node of the interference party; the interference node calculates the Euclidean distance d(a, b) between it and other interference nodes, that is The interfering node establishes edge relationships with the interfering nodes of the first K Euclidean distances to establish a K-nearest neighbor graph;
[0064] Step 302: The interfering node obtains the set N of K-nearest neighbor neighbor information i , and calculates the similarity coefficient e between the interfering node and the nearest neighbor node ij , that is represents the importance of the nearest neighbor node feature to the interfering node, where W represents a trainable weight matrix, represents the interfering node feature, represents the nearest neighbor interfering node feature;
[0065] Step 303: Use the Softmax function to normalize the similarity coefficient e ij to obtain the attention coefficient a between adjacent agents ij , that is
[0066] Step 304: Based on the attention coefficient a ij weight-sum the features of the nearest neighbor nodes, calculate the output feature of the interfering node that fuses the nearest neighbor node information and perform a pooling operation to obtain the team feature O that fuses the local information of each interfering node of the interfering party all .
[0067] Specific implementation method five: As Figures 1 to 5 shown, the implementation steps of introducing an information fusion network based on the QMIX algorithm model to construct a cognitive communication countermeasure interfering node model and outputting the interfering party decision parameters to perform cooperative sequential interference on the communication party in step 4 include:
[0068] Step 401: Initialize the parameters θ of the main network of the agent evaluation network and the mixed evaluation network agent_net and θ mixing_net , initialize the target network parameters θ of the agent target network and the mixed target network with the same structure a ′ gent_net and θ′ mixing_net , initialize the information fusion network, establish an experience replay pool D with an initial capacity of M for storing transfer information <s t ,a t ,s t+1 ,r t >, set the target network update frequency P, and set the total number of iteration rounds T;
[0069] Step 402: Obtain the environmental state information s t , the information O of each interfering node j , the reward value R, and the interfering node selects an effective action where st = [f t , c t , p t , that is, all interference node state information in the adversarial environment includes the frequency, modulation pattern, and power of the communication link of the communication party. , that is, the interference node information actually detected by each interference node. , that is, the algorithm can balance the relationship between the overall complete suppression of the interfering party and the utilization rate of interference resources.
[0070] Step 403: Construct the adjacency matrix A t Obtain the interfering party team information , that is The interference node fuses the interfering party team information with each interference node information to obtain the observation information O as the input of the agent network.
[0071] Step 404: Each interference node in the agent network learns the Q values of all effective actions selected by the interference nodes and uses the hidden state of the GRU recurrent layer as the next input.
[0072] Step 405: Each interference node selects the interference decision action u according to the interference node information O j and its own Q value to obtain the joint action t j
[0073] Step 406: The interference nodes of the interfering party execute the joint action u t to obtain the new environmental state s t+1 and the reward value R.
[0074] Step 407: Store in the experience replay pool and store the transition information >s t , a t , s t+1 , r t > in the experience replay pool D, where the transition information details include:
[0075]
[0076] Step 408: Update the environmental state of the interfering party and the optional effective actions of the interference nodes, that is, s t = s t+1 , Repeat steps 404 to 408 until the experience replay pool is fully stored.
[0077] Step 409: Train the evaluation network, randomly extract a batch of transition information <s t , a t , s t+1,r t , update the parameters of the evaluation network, and repeat steps 402 to 409 until the evaluation network converges;
[0078] Step 410: Update the target network and copy the parameters of the evaluation network to the target network;
[0079] Step 411: Using the QMIX algorithm model of the information fusion network introduced after training, the interfering party's interfering nodes fuse the globally observed communication link state information according to the currently detected communication link state information, and the interfering nodes make decisions on the globally optimal joint interference parameters to perform cooperative sequential interference on the communication signals of the communicating party.
[0080] The above is only a preferred embodiment of the present invention, and does not impose any form of limitation on the present invention. Although the present invention has been disclosed above with a preferred embodiment, it is not intended to limit the present invention. Any person skilled in the art can make some changes or modifications to the equivalent embodiments by using the above-disclosed technical content within the scope of the technical solution of the present invention. However, as long as it does not depart from the content of the technical solution of the present invention and is based on the technical essence of the present invention, any simple modification, equivalent replacement, and improvement made to the above embodiments still fall within the protection scope of the technical solution of the present invention.
Claims
1. A collaborative sequential interference decision method for cognitive communication countermeasure system, characterized in that: The specific steps include: Step 1: Construct a cognitive communication confrontation scenario, that is, establish the two parties of the confrontation as a communication party and an interference party: set the communication party to consist of five communication links consisting of one base station node and five terminal nodes, and set the interference party to consist of four interference nodes to form an interference network; Step 2: Construct the communication party anti-interference node model: In the reinforcement learning framework, set the current communication parameters and state information as the input of the communication party model, and output the communication parameters of each communication link to counter interference; Step 3: Construct the interference node model of the interferer, build and train the information fusion network, and integrate the interference team information: construct the interference scene in the form of a graph, fuse the local information of each interference node based on the graph neural network, and obtain the team characteristics; Step 4: Iteratively optimize the interference party's decision parameters based on the QMIX algorithm, and perform collaborative sequential interference on the communicating party: The QMIX algorithm model introduces an information fusion network, and the interference node outputs the interference party's decision parameters based on the fused global state information.
2. According to claim 1, a collaborative sequential interference decision method for a cognitive communication countermeasure system is characterized in that: In step 1, all nodes of the interfering party are equipped with communication modules, and communicate with other nodes to form an interference network to perform coordinated interference on the communicating party, thereby destroying the normal communication capability of the communicating party; The communication parties use TCP / IP as the routing protocol and change the communication parameters after being disturbed by the communication. The optional communication parameters of each communication link include: In formula (1), f indicates that the optional communication frequencies include f1, f2, and f3, c indicates that the optional modulation modes include BPSK, QPSK, and EPSK, which are represented by c1, c2, and c3 respectively, and p indicates that the optional transmission power is 20%, 40%, 60%, 80%, and 100% of the maximum transmission power, which are represented by p1, p2, and p3 respectively.
3. According to claim 1, a collaborative sequential interference decision method for a cognitive communication countermeasure system is characterized in that: In step 2, when the communication party is subject to communication interference, each communication link adjusts the communication parameters according to the communication parameters and status information at the current moment based on the deep Q neural network algorithm, thereby avoiding and counteracting the interference signal and ensuring stable communication of the communication party system.
4. According to claim 1, a collaborative sequential interference decision method for a cognitive communication countermeasure system is characterized in that: It is characterized in that In step 3, the information fusion network is built and trained to integrate the jammer team information. The specific steps include: Step 301: construct cognitive communication countermeasure interference party data into a graph form, i.e., G = (ν, ε), where ν represents a node set, i.e., each interference node of the interference party, and ε represents an edge set, i.e., the communication relationship between each interference node of the interference party; the interference node calculates the Euclidean distance d(a, b) between itself and other interference nodes, i.e., The interference node establishes an edge relationship with the interference nodes with the first K Euclidean distances to establish a K-nearest neighbor graph; Step 302: The interfering node obtains a set of K-neighbor information N i , and calculate the similarity coefficient e between the interference node and its neighboring nodes ij ,Right now Indicates the importance of the neighbor node feature to the interfering node, where W represents a trainable weight matrix, represents the interference node characteristics, Indicates the characteristics of neighbor interference nodes; Step 303: Use the Softmax function to convert the similarity coefficient e ij Normalization obtains the attention coefficient a between adjacent agents ij ,Right now Step 304: Based on the attention coefficient a ij The weighted sum of the features of the nearest neighbor nodes is used to calculate the output features of the interference node's fused nearest neighbor node information. The pooling operation obtains the team feature O that integrates the local information of each interference node of the interference party all .
5. According to claim 1, a collaborative sequential interference decision method for a cognitive communication countermeasure system is characterized in that: In step 4, the information fusion network is introduced based on the QMIX algorithm model to build a cognitive communication anti-interference node model, and the interference party decision parameters are output to perform coordinated sequential interference on the communication party. The implementation steps include: Step 401: Establish the agent evaluation network and the hybrid evaluation network to initialize the parameters θ of the main network agent_net With θ mixing_net , establish the same structure of the agent target network and the hybrid target network to initialize the target network parameters θ a ' gent_net With θ m ' ixing_net , establish information fusion network initialization, establish an experience replay pool D with an initial capacity of M for storing transfer information t ,a t ,s t+1 ,r t >, set the target network update frequency P, and set the total number of iterations T; Step 402: Obtain environmental status information s t , each interference node information O j , reward value R, interference node selects effective action where s t =[f t ,c t ,p t ] That is, the state information of all interference nodes in the confrontation environment includes the frequency, modulation style, and power of the communication link of the communication party. That is, the interference node information actually detected by each interference node, That is, the algorithm can balance the relationship between the complete suppression of the jammer as a whole and the utilization of jamming resources; Step 403: Construct adjacency matrix A t Get the interference team information Right now The interference node fuses the interference team information with the information of each interference node to obtain the observation information O as the input of the intelligent agent network; Step 404: Each interference node in the agent network learns that all interference nodes select effective actions The Q value of the GRU loop layer is used as the next input. Step 405: Each interfering node generates an interference signal according to the interference node information O. j Interference decision action u with its own Q value selection t j , get the joint action Step 406: The interfering node of the interfering party performs a joint action u t , get the new environment state s t+1 and reward value R; Step 407: Store the experience replay pool and transfer the information t ,a t ,s t+1 ,r t >Stored in the experience replay pool D, where the transfer information includes: Step 408: Update the interfering party's environment state and the interfering node's optional valid actions, i.e., s t =s t+1 , Repeat steps 404 to 408 until the experience replay pool is fully stored; Step 409: Train the evaluation network and randomly extract a batch of transfer information corresponding to the same position in different periods in the experience replay pool. t ,a t ,s t+1 ,r t >, update the parameters of the evaluation network, and repeat steps 402 to 409 until the evaluation network converges; Step 410: Update the target network and copy the parameters of the evaluation network to the target network; Step 411, using the trained QMIX algorithm model that introduces the information fusion network, the interference node of the interference party obtains the global state information based on the currently detected communication link state information, and the interference node determines the global optimal joint interference parameters to perform coordinated sequential interference on the communication signal of the communication party.