Anti-interference waveform decision-making method based on multi-agent reinforcement learning

By deploying a multi-agent reinforcement learning-based anti-interference waveform decision-making method in the communication system, each communication node in the ad hoc network communication system can autonomously make action decisions, quickly adapt to cognitive interference, and optimize the anti-interference capability and transmission efficiency of the communication system.

CN121751204APending Publication Date: 2026-03-27BEIJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-17
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing anti-interference technologies are unable to quickly adapt to changes in cognitive interference, resulting in insufficient transmission capabilities of communication systems in complex electromagnetic environments.

Method used

An anti-interference waveform decision-making method based on multi-agent reinforcement learning is adopted. The communication channel is grouped and an independent agent is deployed at each communication node. Waveform decisions are made through local observation information, and the actions of the communication nodes are optimized by a multi-agent reinforcement learning model to adapt to changes in cognitive interference.

Benefits of technology

It enables the communication system to adapt quickly and transmit efficiently and reliably in complex electromagnetic environments, reduces collision interference within the system, and optimizes the anti-interference capability of multi-target communication.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121751204A_ABST
    Figure CN121751204A_ABST
Patent Text Reader

Abstract

The invention provides an anti-interference waveform decision-making method based on multi-agent reinforcement learning, and relates to the technical field of communication, and each communication node in an ad hoc network communication system can autonomously perform action decision-making so as to quickly adapt to the change of cognitive interference. According to the method, communication channels in a system are divided into a specified number of groups, and local observation information of each communication node in a current time slot is set to comprise state information of a target channel group corresponding to the communication node and feedback information after the communication node executes a previous time slot waveform decision. In this way, it tends to select a channel that does not interfere with other nodes for communication to mitigate collision interference within the system. In view of positive correlation between the reward of the intelligent agent and the communication rate of the communication node and negative correlation between the reward of the intelligent agent and the power overhead and the waveform switching overhead, the waveform decision parameters output by the model can realize multi-target communication anti-interference capability optimization, and meanwhile, the efficient and reliable transmission capability is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of communication, in particular to an anti-interference waveform decision method based on multi-agent reinforcement learning. BACKGROUND

[0002] With the rapid development of wireless communication technology, wireless communication systems have been widely used in various communication scenarios. However, due to the inherent propagation characteristics of wireless channels, wireless communication systems are vulnerable to various security threats, especially malicious jamming attacks. Unlike noise, malicious jamming has stronger targeting, better adaptability and higher destructiveness. Specifically, some jammers can easily detect the characteristics of the legitimate signal and launch adaptive jamming attacks using this information. Therefore, in order to cope with the influence of malicious jamming and noise, researchers have been committed to extensive research on anti-jamming methods to improve the robustness of communication.

[0003] Currently, anti-jamming systems usually use differential frequency hopping, multi-sequence frequency hopping, variable speed frequency hopping and other technologies. These traditional anti-jamming technologies can greatly reduce and avoid the influence of interference under general interference conditions, but as interference technology continues to advance, interference strategies are increasingly intelligent and complex. Cognitive jamming technology can quickly and intelligently generate interference strategies by sensing changes in the electromagnetic environment, while existing adaptive frequency hopping technology can delete jammed frequency points to achieve dynamic frequency hopping sequence adjustment by sensing changes in the electromagnetic environment. However, regional control nodes need to collect changes in the electromagnetic environment state in a larger local range to periodically issue frequency hopping pattern sequences, which is difficult to quickly adapt to changes in cognitive jamming. SUMMARY

[0004] The purpose of the present application is to provide an anti-interference waveform decision method based on multi-agent reinforcement learning, which can autonomously make action decisions to quickly adapt to changes in cognitive jamming, thereby optimizing the anti-interference ability of multi-target communication while ensuring efficient and reliable transmission capability.

[0005] In a first aspect, the present application provides an anti-interference waveform decision method based on multi-agent reinforcement learning, comprising: dividing the communication channels in a self-organizing network communication system into a specified number of groups to obtain a channel grouping result; obtaining the local observation information of each communication node in the self-organizing network communication system at a current time slot; wherein the local observation information includes state information of a target channel group corresponding to the communication node and feedback information of the communication node after executing waveform decision in a previous time slot; the target channel group includes a channel group to which the communication channel in the previous time slot waveform decision belongs and a random channel group; the state information includes the spectrum energy information of each channel at the current time slot and the occupation state at the previous time slot; the feedback information includes the communication rate and the block error rate; processing the local observation information of the target communication node by using a target waveform decision network model to obtain the waveform decision parameters of the target communication node at the current time slot; wherein the target communication node represents any communication node in the self-organizing network communication system; each communication node in the self-organizing network communication system is deployed one-to-one with each agent in the trained target multi-agent reinforcement learning model, and each agent has an independent waveform decision network model; the state of the agent is the local observation information of the corresponding communication node, the action of the agent is the waveform decision parameter of the corresponding communication node, the reward of the agent is positively correlated with the communication rate of the communication node, and is negatively correlated with the power consumption and the waveform switching overhead; the waveform decision parameter includes the channel frequency, the transmission power and the number of coded blocks in the data packet to be sent.

[0006] In an optional embodiment, further comprising: obtaining an experience replay pool; wherein each set of experience in the experience replay pool includes: the system state of the previous time period, the state set of all agents of the previous time period, the action set and the reward set, the system state of the next time period and the state set of all agents of the next time period; based on the experience replay mechanism and the DDQN network architecture, the initial multi-agent reinforcement learning model and the initial hybrid network model are jointly trained until a preset training number is reached to obtain the target multi-agent reinforcement learning model and the target hybrid network model; wherein the hybrid network model is composed of an attention network and a hyperparameter network.

[0007] In an optional embodiment, during the joint training process, the individual action value function of each agent is represented as: ; wherein, represents the state of the i-th agent at the k-th time period, represents the action of the i-th agent at the k-th time period, represents the individual action value function of the i-th agent at the k-th time period, represents the expectation, represents the long-term reward of the i-th agent, , Indicates reward discount factor of Power of 1 This represents the reward of the i-th agent in the k-th time interval. This indicates the total number of time periods included within the task execution cycle.

[0008] In an optional implementation, the hybrid network model is used to calculate the overall action value function for all agents; the overall action value function is expressed as: ;in, This represents the set of states of all agents in the k-th time interval. This represents the set of actions taken by all agents in the k-th time interval. , This represents the network parameters of the hybrid network model. The network parameters of the waveform decision network model for the i-th agent are represented. This represents the overall action value function of all agents in the k-th time period. This represents the weight vector output by the attention network. This represents the i-th weight. This represents the total number of intelligent agents. A vector representing the value functions of the individual actions of all agents. This represents the weight matrix output by the hyperparameter network. This represents the weight vector output by the hyperparameter network. This represents the bias vector output by the hyperparameter network. This represents the bias scalar output by the hyperparameter network; , .

[0009] In an optional implementation, the loss function used for the joint training is expressed as: ;in, Indicates the batch size of the sampling experience. Indicates from the experience replay pool Mid-sampling, This represents the system state during the k-th time period. This represents the set of states of all agents in the k-th time interval. This represents the set of actions taken by all agents in the k-th time interval. This represents the set of rewards for all agents in the k-th time interval. This represents the system state in the (k+1)th time period. This represents the set of states of all agents in the (k+1)th time interval. This represents the Q-value output by the target network in the DDQN network architecture. , This represents the sum of rewards for all agents in the k-th time interval. This represents the set of actions jointly selected by all agents in the (k+1)th time interval that maximizes the overall action value. This represents the network parameters of the target network in the k-th time period. This indicates that the target network has certain network parameters. Under these conditions, all agents are based on a set of states. Select Action The resulting overall action value function.

[0010] In an optional implementation, the agent's reward is represented as: ;in, This represents the reward of the i-th agent in the k-th time interval. This represents the communication rate of the i-th agent in the k-th time period. This represents the waveform switching cost of the i-th agent in the k-th time period. , This represents the number of coded blocks in the data packet to be sent by the i-th agent in the k-th time period. This represents the channel frequency of the i-th agent in the k-th time period. This represents the transmission power of the i-th agent in the k-th time period. This indicates the preset weighting coefficient.

[0011] In an optional implementation, the communication rate is calculated as follows: ;in, This indicates the number of bits in each coded block. This represents the collision flag of the i-th agent in the k-th time interval. , This represents a conditional function, if... If the condition is met, then ,otherwise , This represents the total number of intelligent agents; , This represents the block error rate of the i-th agent in the k-th time period. This indicates the preset error rate threshold.

[0012] Secondly, the present invention provides an anti-interference waveform decision-making device based on multi-agent reinforcement learning, comprising: a grouping module, used to divide the communication channels in an ad hoc network communication system into a specified number of groups to obtain channel grouping results; an acquisition module, used to acquire local observation information of each communication node in the ad hoc network communication system in the current time slot; wherein, the local observation information includes: the state information of the target channel group corresponding to the communication node and the feedback information of the communication node after performing waveform decision in the previous time slot; the target channel group includes: the channel group to which the communication channel in the previous time slot waveform decision belongs and a random channel group; the state information includes: the spectral energy information of each channel in the current time slot and the occupancy status in the previous time slot; the feedback information includes: communication rate and block error rate; and a processing module. A block is used to process the local observation information of the target communication node using the target waveform decision network model to obtain the waveform decision parameters of the target communication node in the current time slot; wherein, the target communication node represents any communication node in the ad hoc network communication system; each communication node in the ad hoc network communication system is deployed one-to-one with each agent in the trained target multi-agent reinforcement learning model, and each agent has an independent waveform decision network model; the state of the agent is the local observation information of the corresponding communication node, the action of the agent is the waveform decision parameters of the corresponding communication node, the reward of the agent is positively correlated with the communication rate of the communication node, and negatively correlated with power overhead and waveform switching overhead; the waveform decision parameters include: channel frequency, transmit power, and the number of coded blocks in the data packet to be transmitted.

[0013] Thirdly, the present invention provides an electronic device, including a memory and a processor, wherein the memory stores a computer program that can run on the processor, and the processor executes the computer program to implement the anti-interference waveform decision method based on multi-agent reinforcement learning as described in any of the foregoing embodiments.

[0014] Fourthly, the present invention provides a computer-readable storage medium storing computer instructions, which, when executed by a processor, implement the anti-interference waveform decision method based on multi-agent reinforcement learning as described in any of the foregoing embodiments.

[0015] This invention provides an anti-interference waveform decision-making method based on multi-agent reinforcement learning. Each communication node in the ad hoc network communication system is deployed with a corresponding agent, and each agent has an independent waveform decision-making network model. Therefore, each communication node can autonomously make action decisions to quickly adapt to changes in cognitive interference. This method divides the communication channels within the ad hoc network communication system into a specified number of groups and sets the local observation information of each communication node in the current time slot to include: the state information of the target channel group corresponding to the communication node and the feedback information of the communication node after executing the waveform decision in the previous time slot. Therefore, when making action selection, the communication node will automatically tend to choose the channel that will not interfere with other nodes, thereby reducing collision interference within the system. Given that during the training of the multi-agent reinforcement learning model, the agent's reward is positively correlated with the communication rate of the communication node and negatively correlated with power overhead and waveform switching overhead, after the multi-agent reinforcement learning model training is completed, each communication node uses an independent neural network to process its local observation information. The resulting waveform decision parameters can optimize the anti-interference capability of multi-target communication while ensuring efficient and reliable transmission capabilities. Attached Figure Description

[0016] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0017] Figure 1 A flowchart illustrating an anti-interference waveform decision-making method based on multi-agent reinforcement learning, provided as an embodiment of the present invention; Figure 2 This is a schematic diagram of an anti-interference self-organizing network environment provided by an embodiment of the present invention; Figure 3 A schematic diagram of a multi-agent reinforcement learning anti-interference algorithm framework provided in an embodiment of the present invention; Figure 4 This is a schematic diagram of a hybrid network model provided in an embodiment of the present invention; Figure 5 A functional block diagram of an anti-interference waveform decision-making device based on multi-agent reinforcement learning provided in an embodiment of the present invention; Figure 6 This is a schematic diagram of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0018] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.

[0019] Therefore, the following detailed description of the embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.

[0020] The following detailed description of some embodiments of the present invention is provided in conjunction with the accompanying drawings. Unless otherwise specified, the following embodiments and features can be combined with each other.

[0021] Example 1 Figure 1 A flowchart of an anti-interference waveform decision-making method based on multi-agent reinforcement learning provided in this embodiment of the invention is shown below. Figure 1 As shown, the method specifically includes the following steps: Step S102: Divide the communication channels in the ad hoc network communication system into a specified number of groups to obtain the channel grouping results.

[0022] Since a self-organizing network communication system includes multiple communication nodes, communication collisions and self-interference between nodes are prone to occur. To mitigate and avoid such interference, embodiments of the present invention address this issue by... The shared available communication channels are divided into The number of groups is not specifically limited in this embodiment of the invention; users need to configure it according to their actual situation. It should be noted that a larger number of groups results in a lower probability of collisions, but correspondingly, the flexibility of node frequency hopping will also decrease. This is because more groups mean fewer communication channels within each group, thus limiting the number of channel choices for the node.

[0023] Step S104: Obtain the local observation information of each communication node in the current time slot in the ad hoc network communication system.

[0024] The local observation information includes: the status information of the target channel group corresponding to the communication node and the feedback information of the communication node after executing the waveform decision of the previous time slot; the target channel group includes: the channel group to which the communication channel in the waveform decision of the previous time slot belongs and a random channel group; the status information includes: the spectral energy information of each channel in the current time slot and the occupancy status in the previous time slot; the feedback information includes: the communication rate and the block error rate.

[0025] This invention describes the problem of intelligent anti-interference decision-making for multiple communication nodes as a partially observable Markov decision process. Each communication node acts as an agent to make distributed decisions. The node makes waveform parameter decisions by taking the collected information as input through local observation. Since it is only related to the current network state and not to the historical state, it conforms to the Markov decision process.

[0026] Specifically, based on the data content of the local observation information described above, each communication node acquires two aspects of local observation. First, it obtains the state information of two channel groups (i.e., the target channel group). These two channel groups are: the channel group selected in the previous time slot, and a random channel group from the remaining channel groups. The state information includes not only the spectral energy information of each channel in the target channel group, but also the occupancy status of each channel in the previous time slot, including idle and occupied states. Clearly, this observability helps avoid channel collisions in the next decision. Based on this observation setting, after the agent network training converges, each communication node will tend to choose channels that will not interfere with other nodes, thereby reducing collision interference within the system.

[0027] Since the feedback information (communication rate and block error rate) of the previous time slot can reflect the communication performance after the waveform decision in the previous time slot, another aspect of local observation is: if the communication performance is poor, the feedback information after the waveform decision in the previous time slot tends to change the waveform decision in the current time slot, but if the communication performance is good, the waveform decision in the previous time slot tends to be maintained.

[0028] Figure 2 This is a schematic diagram of an anti-interference ad hoc network environment. In an ad hoc network communication system, signaling between communication nodes is conducted through a control channel (i.e., Figure 2 The data is transmitted via the control link in the system, and then transmitted through the communication channel (i.e., the control link in the system). Figure 2 The transmission link in the channel is used for communication data transmission. White noise exists in the channel, and the power of the noise is expressed as... ,in, Indicates channel Average noise power, Values ​​range from 1 to Furthermore, the receiver takes into account large-scale fading and free-space fading present in the environment when receiving signals.

[0029] In addition, in ad hoc network communication environments there are There are 10 jammer nodes, of which 10 are jammers. The transmit power is expressed as The jammer selects one or more channels to apply jamming signals at the beginning of each time slot. The jamming patterns are a mixture of narrowband jamming, wideband jamming, frequency sweeping jamming, and repeater-tracking jamming. The power of the jamming signal for each channel depends on its jamming bandwidth; a larger bandwidth dilutes the power of the jamming signal. Therefore, the jammer can... For the channel The interference power is expressed as: ,in, This represents the ratio of interference bandwidth to channel bandwidth. Due to the existence of sensing-based jamming methods, jammers possess the ability to sense the power of the transmitted signal and apply repeating tracking jamming.

[0030] In this embodiment of the invention, there exists a self-organizing network communication environment. There are several communication nodes, each with transmit and receive capabilities. At the start of a time slot, each node selects the waveform parameters for that time using an algorithm and propagates parameter signaling through a preset control channel. The control channel employs a high-reliability transmission method (1 / 2 polar coding + BPSK modulation) to ensure normal communication. Upon receiving the signaling, the receiver adjusts its receiving parameters to prepare for reception. After receiving and demodulating the data, the receiver uses CRC checksums to count the erroneous code blocks generated during the reception. The BLER (Block Error Rate) is calculated by determining the ratio of erroneous code blocks to the total number of coded blocks. The BLER and data communication rate are then fed back to the transmitting communication node to provide performance metrics for that node's next time slot decision.

[0031] It is known that different coding modulation rates and SINR conditions result in different BLERs. A preset block error rate threshold is set. When the BLER exceeds this threshold, the data packet is considered a demodulation failure and is discarded. In this embodiment of the invention, the SINR formula for the receiver of the communication node is: ,in, Represents a node In the channel The ratio of signal power to interference noise is used to measure communication quality. Represents a node The transmission power, This indicates the gain of the transmitting antenna. Indicates the receiving antenna gain. This refers to free space fading during the transmission of communication signals. Indicates channel Average noise power, Indicates jammer For the channel Interference power, Indicates the gain of jamming transmission. This indicates the time slot overlap loss caused by incomplete time slot alignment between interference signals and communication signals. This indicates the free-space fading of the interference signal.

[0032] Step S106: The local observation information of the target communication node is processed using the target waveform decision network model to obtain the waveform decision parameters of the target communication node in the current time slot.

[0033] In this context, the target communication node represents any communication node in the ad hoc network communication system; each communication node in the ad hoc network communication system is deployed one-to-one with each agent in the trained target multi-agent reinforcement learning model, and each agent has an independent waveform decision network model; the state of the agent is the local observation information of the corresponding communication node, the action of the agent is the waveform decision parameter of the corresponding communication node, the reward of the agent is positively correlated with the communication rate of the communication node, and negatively correlated with the power overhead and waveform switching overhead; the waveform decision parameters include: channel frequency, transmit power, and the number of coded blocks in the data packet to be transmitted.

[0034] As described above, in this embodiment of the invention, each communication node in the ad hoc network communication system participates in distributed anti-interference waveform decision-making as an intelligent agent. A deep reinforcement learning intelligent anti-interference waveform decision-making module (i.e., the target waveform decision-making network model mentioned above) is deployed on the node. Each node maintains its own network model. Each time slot makes waveform decisions based on its local observation information and sends the waveform decision parameters to the receiver through the signaling channel to complete synchronization before receiving.

[0035] In this embodiment of the invention, the communication node The waveform decision parameters are expressed as ,in, Represents communication node The channel frequency, Represents communication node The transmission power, Represents communication node The number of encoded blocks in the data packet to be sent. , Indicates communication channel Each communication channel corresponds to a specific channel frequency; selecting a channel frequency is equivalent to selecting a communication channel. , Indicates the lower limit of transmission power. Indicates the upper limit of transmission power. , This indicates the lower limit of the number of coded blocks in a data packet. This indicates the upper limit of the number of coded blocks in a data packet.

[0036] In this embodiment of the invention, the channel frequency during the action of the intelligent agent The action space depends on the target channel group observed by the corresponding communication node. The channels that can be selected in the current time slot are only the channel group of the previous time slot and a channel in a random channel group. The following cases are discussed: 1) If a node successfully communicates in the previous time slot, the probability of collision when the node continues to choose the channel in the previous time slot will be lower (because other nodes can observe that there are nodes communicating in the channel group, and will try to avoid the node to avoid collision).

[0037] 2) If a node experiences a collision in the previous time slot, causing communication failure, then in the next time slot, the node will tend to choose a different channel combination (the number of nodes using the channel combination can be observed) to avoid the collision.

[0038] Based on the above settings, the node has the ability to adaptively adjust the channel selection after being interfered with, and will eventually converge to a collision-free state after a certain period of time.

[0039] The model continuously updates network parameters through centralized learning and distributed execution until the algorithm converges. In this embodiment of the invention, the goal of the multi-agent deep reinforcement learning method is to find a joint optimal policy that enables the agents to select the optimal joint action according to the policy to obtain the maximum cumulative reward. In this embodiment of the invention, the agent's reward is positively correlated with the communication rate of the communication nodes and negatively correlated with power overhead and waveform switching overhead.

[0040] To quantify the overhead caused by communication node switching parameters and transmit power, power overhead and waveform switching overhead are used as communication overhead indicators. The power overhead of a communication node is set to its transmit power, and the waveform switching overhead is determined by the following formula: ,in, and This represents two adjacent time slots. Based on this formula, it can be seen that if the first... Time slot communication node The selected channel frequency and the number of coded blocks in the data packet to be transmitted are related to the first... If all time slots are selected identically, then the waveform switching overhead... The value is 0; if either the channel frequency or the number of coding blocks changes, the waveform switching overhead... The value is 1.

[0041] This invention provides an anti-interference waveform decision-making method based on multi-agent reinforcement learning. Each communication node in the ad hoc network communication system is deployed with a corresponding agent, and each agent has an independent waveform decision-making network model. Therefore, each communication node can autonomously make action decisions to quickly adapt to changes in cognitive interference. This method divides the communication channels within the ad hoc network communication system into a specified number of groups and sets the local observation information of each communication node in the current time slot to include: the state information of the target channel group corresponding to the communication node and the feedback information of the communication node after executing the waveform decision in the previous time slot. Therefore, when making action selection, the communication node will automatically tend to choose the channel that will not interfere with other nodes, thereby reducing collision interference within the system. Given that during the training of the multi-agent reinforcement learning model, the agent's reward is positively correlated with the communication rate of the communication node and negatively correlated with power overhead and waveform switching overhead, after the multi-agent reinforcement learning model training is completed, each communication node uses an independent neural network to process its local observation information. The resulting waveform decision parameters can optimize the anti-interference capability of multi-target communication while ensuring efficient and reliable transmission capabilities.

[0042] In one alternative implementation, the method further includes the following steps: Step S201: Obtain the experience replay pool; wherein, each set of experience in the experience replay pool includes: the system state of the previous time period. The set of states of all agents in the previous time period Action set and reward set The system status for the next time period and the set of states of all agents in the next time period .

[0043] The set of all agent states for any given time period (also known as a time slot or moment) is the system state. Based on the definition of agent state, the system state under time slot k includes: the spectral energy information of all communication channels under time slot k, the occupancy status of all communication channels in time slot k-1, and the set of feedback information after each agent executes the action selected in time slot k-1.

[0044] Step S202: Based on the experience replay mechanism and DDQN network architecture, the initial multi-agent reinforcement learning model and the initial hybrid network model are jointly trained until the preset number of training times are reached, so as to obtain the target multi-agent reinforcement learning model and the target hybrid network model; wherein, the hybrid network model consists of an attention network and a hyperparameter network.

[0045] Figure 3A schematic diagram of a multi-agent reinforcement learning anti-interference algorithm framework, as shown below. Figure 3 As shown, this embodiment of the invention employs a centralized training and distributed execution multi-agent reinforcement learning framework to solve optimization problems. Specifically, the agent nodes are trained centrally, which can use global state information to update network parameters. After training converges, the model is deployed on the nodes, and in application, the individual uses local observations to make its own actions.

[0046] For each communication node, there exists its own decision network for calculating the action value function. In this embodiment of the invention, a DRQN ​​network is used to calculate the individual action value function, that is, the future benefit that the agent can obtain by taking this action. Specifically, during the joint training process, the individual action value function of each agent is expressed as follows: .

[0047] in, This represents the state of the i-th agent in the k-th time interval. This represents the action of the i-th agent in the k-th time interval. Let represent the individual action value function of the i-th agent in the k-th time period. Expressing expectations, This represents the long-term reward of the i-th agent, that is, the reward value that the agent can obtain by taking the current action in the current and future time periods. , Indicates reward discount factor of Power of 1 The reward discount factor is used to determine the importance of future rewards. This represents the reward of the i-th agent in the k-th time interval. This indicates the total number of time periods included within the task execution cycle.

[0048] This invention employs an experience replay mechanism for model training; that is, during the exploratory learning phase, the agent stores a set of experiences. The centralized network learns and updates itself by randomly sampling experience from an experience pool. To improve the efficiency of experience acquisition, during the centralized learning phase, nodes adopt... The greedy strategy is used to select actions. , Represents intelligent agents The set of available actions in each state. The purpose of this strategy is to improve the ability to explore random actions and avoid the algorithm getting stuck in suboptimal solutions.

[0049] Figure 4This is a schematic diagram of a hybrid network model provided in an embodiment of the present invention. The hybrid network model is used to calculate the overall action value function of all agents, and is used to optimize the overall system strategy. Figure 4 As shown, the hybrid network model includes an attention network and a hyperparameter network. It adaptively generates weighted hybrid action values ​​of individuals based on the current environmental state. The overall action value function is expressed as: .

[0050] in, This represents the set of states of all agents in the k-th time interval. This represents the set of actions taken by all agents in the k-th time interval. , Represents the network parameters of the hybrid network model. The network parameters of the waveform decision network model for the i-th agent are represented. This represents the overall action value function of all agents in the k-th time period. This represents the weight vector output by the attention network. This represents the i-th weight. This represents the total number of intelligent agents. A vector representing the value functions of the individual actions of all agents. This represents the weight matrix output by the hyperparameter network. This represents the weight vector output by the hyperparameter network. This represents the bias vector output by the hyperparameter network. This represents the bias scalar output of the hyperparameter network.

[0051] The i-th weight in the weight vector output by the attention network ,in, This represents the attention score calculated by learning the transformations of the local Q-value and the global state. , and These represent the query vector and key vector in the attention network, respectively. They are obtained through a linear transformation of the neural network and are generated during the calculation of the attention score, used for characterization. and The relationship between them , , , , , All of these are learnable parameters of the neural network. This indicates the size of the attention head space dimension.

[0052] To ensure the monotonicity of individual action value and global action value, embodiments of the present invention are set as follows: , Monotonicity constraints ensure that the optimization directions of the global Q-value (i.e., the overall action value function) and the individual Q-value (i.e., the individual action value function) are consistent. This value decomposition-based multi-agent reinforcement learning method can effectively address dynamic interference and policy coordination problems in multi-node anti-interference environments.

[0053] In this embodiment of the invention, the training process of jointly training the multi-agent reinforcement learning model and the hybrid network model is divided into two parts: experience exploration and collection, and network update. During experience exploration, the agents collect data through... The greedy strategy interacts with the environment to obtain rewards and the state and observations of the next time slot, and stores them in the experience replay pool. In the middle, the experience replay pool is a pool of size... The queue follows the first-in, first-out principle. During the experience learning phase, the centralized control center samples a set of experiences for learning, and the neural network updates its parameters by minimizing the TD error (Temporal Difference).

[0054] In one alternative implementation, the loss function used for joint training is expressed as: ;in, This represents the batch size of the sampling experience. Gradient descent is used to update the network parameters, as shown below: , This represents the learning rate in reinforcement learning.

[0055] Indicates from the experience replay pool Mid-sampling, This represents the system state during the k-th time period. This represents the set of states of all agents in the k-th time interval. This represents the set of actions taken by all agents in the k-th time interval. This represents the set of rewards for all agents in the k-th time interval. This represents the system state in the (k+1)th time period. This represents the set of states of all agents in the (k+1)th time period.

[0056] In addition, to overcome the bootstrapping caused by overestimation of the value function in reinforcement learning, the algorithm adopts a target network design to improve the stability of learning. That is, the DDQN network architecture, in which the target network is a copy of the policy network, but its parameters are updated at a lower frequency. Specifically, the target network parameters are set to be updated every F steps to ensure that the target moves slowly, thereby reducing the instability of function approximation in bootstrapping updates. This represents the Q-value output by the target network in the DDQN network architecture. , This represents the sum of rewards for all agents in the k-th time interval. This represents the set of actions jointly selected by all agents in the (k+1)th time interval that maximizes the overall action value. This represents the network parameters of the target network in the k-th time period. Indicates the target network in network parameters Under these conditions, all agents are based on a set of states. Select Action The resulting overall action value function.

[0057] If the model training reaches the preset number of training iterations, the model parameters are considered to have converged, resulting in a target multi-agent reinforcement learning model and a target hybrid network model. Then, each agent in the trained target multi-agent reinforcement learning model is deployed to its corresponding communication node, allowing each node's policy model to independently make optimal waveform decisions based on local observations.

[0058] In this embodiment of the invention, the reward for the intelligent agent is represented as follows: ;in, This represents the reward of the i-th agent in the k-th time interval. This represents the communication rate of the i-th agent in the k-th time period. This represents the waveform switching cost of the i-th agent in the k-th time period. , This represents the number of coded blocks in the data packet to be sent by the i-th agent in the k-th time period. This represents the channel frequency of the i-th agent in the k-th time period. This represents the transmission power of the i-th agent in the k-th time period. This indicates the preset weighting coefficient.

[0059] To calculate the long-term reward for all agents, we simply need to sum the rewards for each agent across all time periods. As shown in the expression for the reward function above, this function optimizes the communication metrics of multiple objectives through weighting, enabling flexible adjustments to communication rate, reliability, and overhead, thus ensuring the robustness and efficiency of the system.

[0060] In one alternative implementation, the communication rate is calculated as follows: ;in, This represents the number of bits in each coded block. In a shared channel environment, multiple nodes may transmit simultaneously on the same channel, leading to potential collisions and ultimately preventing data packet transmission. Therefore, the network employs a listening mechanism similar to Carrier Sense Multiple Access (CSMA) to enable nodes to perform channel sensing to determine if the channel is idle and then select an action. During time period k, the following is used: This represents the collision flag of the i-th agent in the k-th time interval. , This represents a conditional function, if... If the condition is met, then ,otherwise , This represents the total number of intelligent agents; , This represents the block error rate of the i-th agent in the k-th time period. This indicates the preset block error rate threshold. When... When the preset block error rate threshold is exceeded, the entire data packet is considered a transmission failure. .

[0061] In summary, addressing the challenges of complex interference patterns and the difficulty in countering interference events in complex electromagnetic environments, as well as the shortcomings of existing anti-interference methods, this invention provides an anti-interference waveform decision-making method based on multi-agent deep reinforcement learning. Compared with existing methods, the method proposed in this invention has the following advantages: 1. The multi-agent reinforcement learning-based intelligent communication anti-interference waveform decision method proposed in this invention combines the number of coding blocks, frequency hopping points, and power decisions, and optimizes the strategy through artificial intelligence algorithms. Compared with traditional methods, it is more adaptable to the reliable communication requirements in complex dynamic electromagnetic environments.

[0062] 2. In this invention, each communication node is deployed as an intelligent agent, and each node has an independent neural network for anti-interference waveform decision-making, thus constructing a multi-agent network anti-interference communication system.

[0063] 3. This invention takes into account multiple communication transmission indicators in the implementation of the anti-interference algorithm, optimizes the anti-interference capability of multi-target communication, and ensures efficient and reliable transmission capability.

[0064] 4. This invention adopts a centralized training and distributed execution architecture—during training, the centralized network can optimize the training of individual networks through global information, while during execution, the nodes make action decisions only through their own networks. This method can enhance the collaborative performance of nodes and improve the scalability of the node network.

[0065] Example 2 This invention also provides an anti-interference waveform decision-making device based on multi-agent reinforcement learning. This device is mainly used to execute the anti-interference waveform decision-making method based on multi-agent reinforcement learning provided in Embodiment 1 above. The following is a detailed description of the anti-interference waveform decision-making device based on multi-agent reinforcement learning provided in this invention.

[0066] Figure 5 A functional block diagram of an anti-interference waveform decision-making device based on multi-agent reinforcement learning provided in an embodiment of the present invention is shown below. Figure 5 As shown, the device mainly includes: a grouping module 10, an acquisition module 20, and a processing module 30, wherein: The grouping module 10 is used to divide the communication channel in the ad hoc network communication system into a specified number of groups to obtain the channel grouping result.

[0067] The acquisition module 20 is used to acquire local observation information of each communication node in the current time slot in the ad hoc network communication system. The local observation information includes: the status information of the target channel group corresponding to the communication node and the feedback information of the communication node after executing the waveform decision of the previous time slot. The target channel group includes: the channel group to which the communication channel in the waveform decision of the previous time slot belongs and a random channel group. The status information includes: the spectral energy information of each channel in the current time slot and the occupancy status in the previous time slot. The feedback information includes: the communication rate and the block error rate.

[0068] The processing module 30 is used to process the local observation information of the target communication node using the target waveform decision network model to obtain the waveform decision parameters of the target communication node in the current time slot. The target communication node represents any communication node in the ad hoc network communication system. Each communication node in the ad hoc network communication system is deployed one-to-one with each agent in the trained target multi-agent reinforcement learning model, and each agent has an independent waveform decision network model. The state of the agent is the local observation information of the corresponding communication node, and the action of the agent is the waveform decision parameter of the corresponding communication node. The reward of the agent is positively correlated with the communication rate of the communication node and negatively correlated with power overhead and waveform switching overhead. The waveform decision parameters include: channel frequency, transmit power, and the number of coded blocks in the data packet to be transmitted.

[0069] This invention provides an anti-interference waveform decision-making device based on multi-agent reinforcement learning. Each communication node in the ad hoc network communication system is deployed with a corresponding agent, and each agent has an independent waveform decision-making network model. Therefore, each communication node can autonomously make action decisions to quickly adapt to changes in cognitive interference. The device divides the communication channels within the ad hoc network communication system into a specified number of groups and sets the local observation information of each communication node in the current time slot to include: the state information of the target channel group corresponding to the communication node and the feedback information of the communication node after executing the waveform decision in the previous time slot. Therefore, when making action selection, the communication node will automatically tend to choose the channel that will not interfere with other nodes, thereby reducing collision interference within the system. Given that during the training of the multi-agent reinforcement learning model, the agent's reward is positively correlated with the communication rate of the communication node and negatively correlated with power overhead and waveform switching overhead, after the multi-agent reinforcement learning model training is completed, each communication node uses an independent neural network to process its local observation information. The resulting waveform decision parameters can optimize the anti-interference capability of multi-target communication while ensuring efficient and reliable transmission capabilities.

[0070] Optionally, the device is also used for: Obtain the experience replay pool; each set of experience in the experience replay pool includes: the system state of the previous time period, the state set, action set and reward set of all agents in the previous time period, the system state of the next time period and the state set of all agents in the next time period.

[0071] Based on the experience replay mechanism and DDQN network architecture, the initial multi-agent reinforcement learning model and the initial hybrid network model are jointly trained until the preset number of training iterations are reached, resulting in the target multi-agent reinforcement learning model and the target hybrid network model; wherein, the hybrid network model consists of an attention network and a hyperparameter network.

[0072] Optionally, during joint training, the individual action value function of each agent is represented as: ;in, This represents the state of the i-th agent in the k-th time interval. This represents the action of the i-th agent in the k-th time interval. Let represent the individual action value function of the i-th agent in the k-th time period. Expressing expectations, Let represent the long-term reward of the i-th agent. , Indicates reward discount factor of Power of 1 This represents the reward of the i-th agent in the k-th time interval. This indicates the total number of time periods included within the task execution cycle.

[0073] Optionally, a hybrid network model is used to calculate the overall action value function for all agents; the overall action value function is expressed as: ;in, This represents the set of states of all agents in the k-th time interval. This represents the set of actions taken by all agents in the k-th time interval. , Represents the network parameters of the hybrid network model. The network parameters of the waveform decision network model for the i-th agent are represented. This represents the overall action value function of all agents in the k-th time period. This represents the weight vector output by the attention network. This represents the i-th weight. This represents the total number of intelligent agents. A vector representing the value functions of the individual actions of all agents. This represents the weight matrix output by the hyperparameter network. This represents the weight vector output by the hyperparameter network. This represents the bias vector output by the hyperparameter network. This represents the bias scalar value output by the hyperparameter network. , .

[0074] Optionally, the loss function used for joint training is expressed as: ;in, Indicates the batch size of the sampling experience. Indicates from the experience replay pool Mid-sampling, This represents the system state during the k-th time period. This represents the set of states of all agents in the k-th time interval. This represents the set of actions taken by all agents in the k-th time interval. This represents the set of rewards for all agents in the k-th time interval. This represents the system state in the (k+1)th time period. This represents the set of states of all agents in the (k+1)th time interval. This represents the Q-value output by the target network in the DDQN network architecture. , This represents the sum of rewards for all agents in the k-th time interval. This represents the set of actions jointly selected by all agents in the (k+1)th time interval that maximizes the overall action value. This represents the network parameters of the target network in the k-th time period. Indicates the target network in network parameters Under these conditions, all agents are based on a set of states. Select Action The resulting overall action value function.

[0075] Optionally, the agent's reward is represented as: ;in, This represents the reward of the i-th agent in the k-th time interval. This represents the communication rate of the i-th agent in the k-th time period. This represents the waveform switching cost of the i-th agent in the k-th time period. , This represents the number of coded blocks in the data packet to be sent by the i-th agent in the k-th time period. This represents the channel frequency of the i-th agent in the k-th time period. This represents the transmission power of the i-th agent in the k-th time period. This indicates the preset weighting coefficient.

[0076] Optionally, the formula for communication rate is: ;in, This indicates the number of bits in each coded block. This represents the collision flag of the i-th agent in the k-th time interval. , This represents a conditional function, if... If the condition is met, then ,otherwise , This represents the total number of intelligent agents; , This represents the block error rate of the i-th agent in the k-th time period. This indicates the preset error rate threshold.

[0077] Example 3 See Figure 6 This invention provides an electronic device, which includes a processor 60, a memory 61, a bus 62, and a communication interface 63. The processor 60, the communication interface 63, and the memory 61 are connected via the bus 62. The processor 60 is used to execute executable modules, such as computer programs, stored in the memory 61.

[0078] The memory 61 may include high-speed random access memory (RAM) or non-volatile memory, such as at least one disk storage device. Communication between this system network element and at least one other network element is achieved through at least one communication interface 63 (which can be wired or wireless), such as the Internet, wide area network, local area network, metropolitan area network, etc.

[0079] Bus 62 can be an ISA bus, PCI bus, or EISA bus, etc. The bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 6 The symbol is represented by a single double-headed arrow, but this does not mean that there is only one bus or one type of bus.

[0080] The memory 61 is used to store programs. After receiving an execution instruction, the processor 60 executes the program. The method executed by the apparatus defined by the process disclosed in any of the foregoing embodiments of the present invention can be applied to the processor 60 or implemented by the processor 60.

[0081] Processor 60 may be an integrated circuit chip with signal processing capabilities. In implementation, each step of the above method can be completed by the integrated logic circuitry in the hardware of processor 60 or by instructions in software form. Processor 60 can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this invention. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this invention can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software modules can reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The storage medium is located in memory 61. Processor 60 reads the information in memory 61 and, in conjunction with its hardware, completes the steps of the above method.

[0082] The computer program product of the anti-interference waveform decision method based on multi-agent reinforcement learning provided in this embodiment of the invention includes a computer-readable storage medium storing non-volatile program code executable by a processor. The instructions included in the program code can be used to execute the methods described in the preceding method embodiments. For specific implementation, please refer to the method embodiments, which will not be repeated here.

[0083] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0084] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a processor-executable, non-volatile, computer-readable storage medium. Based on this understanding, the technical solution of this invention, essentially, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0085] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.

[0086] In the description of this invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings, or the orientation or positional relationship commonly used when the product of this invention is in use. They are only for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of this invention. In addition, the terms "first," "second," "third," etc., are only used to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0087] Furthermore, terms such as "horizontal," "vertical," and "sag" do not imply that components must be absolutely horizontal or suspended, but rather that they can be slightly tilted. For example, "horizontal" simply means that its direction is more horizontal relative to "vertical," and does not mean that the structure must be completely horizontal, but can be slightly tilted.

[0088] In the description of this invention, it should also be noted that, unless otherwise explicitly specified and limited, the terms "set," "install," "connect," and "link" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.

[0089] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for anti-interference waveform decision-making based on multi-agent reinforcement learning, characterized in that, include: Divide the communication channels within the ad hoc network communication system into a specified number of groups to obtain the channel grouping results; The local observation information of each communication node in the ad hoc network communication system in the current time slot is obtained; wherein, the local observation information includes: the status information of the target channel group corresponding to the communication node and the feedback information of the communication node after executing the waveform decision of the previous time slot; the target channel group includes: the channel group to which the communication channel in the waveform decision of the previous time slot belongs and a random channel group; the status information includes: the spectral energy information of each channel in the current time slot and the occupancy status in the previous time slot; the feedback information includes: communication rate and block error rate; The local observation information of the target communication node is processed using a target waveform decision network model to obtain the waveform decision parameters of the target communication node in the current time slot. The target communication node represents any communication node in the ad hoc network communication system. Each communication node in the ad hoc network communication system is deployed one-to-one with each agent in the trained target multi-agent reinforcement learning model, and each agent has an independent waveform decision network model. The state of the agent is the local observation information of the corresponding communication node, and the action of the agent is the waveform decision parameter of the corresponding communication node. The reward of the agent is positively correlated with the communication rate of the communication node and negatively correlated with power overhead and waveform switching overhead. The waveform decision parameters include: channel frequency, transmit power, and the number of coded blocks in the data packet to be transmitted.

2. The anti-interference waveform decision-making method based on multi-agent reinforcement learning according to claim 1, characterized in that, Also includes: Obtain an experience replay pool; wherein, each set of experience in the experience replay pool includes: the system state of the previous time period, the state set, action set and reward set of all agents in the previous time period, the system state of the next time period and the state set of all agents in the next time period; Based on the experience replay mechanism and DDQN network architecture, the initial multi-agent reinforcement learning model and the initial hybrid network model are jointly trained until a preset number of training iterations are reached, thereby obtaining the target multi-agent reinforcement learning model and the target hybrid network model; wherein, the hybrid network model consists of an attention network and a hyperparameter network.

3. The anti-interference waveform decision-making method based on multi-agent reinforcement learning according to claim 2, characterized in that, During joint training, the individual action value function of each agent is expressed as: ; in, This represents the state of the i-th agent in the k-th time interval. This represents the action of the i-th agent in the k-th time interval. Let represent the individual action value function of the i-th agent in the k-th time period. Expressing expectations, Let represent the long-term reward of the i-th agent. , Indicates reward discount factor of Power of 1 This represents the reward of the i-th agent in the k-th time interval. This indicates the total number of time periods included within the task execution cycle.

4. The anti-interference waveform decision-making method based on multi-agent reinforcement learning according to claim 3, characterized in that, The hybrid network model is used to calculate the overall action value function of all agents; The overall action value function is expressed as: ; in, This represents the set of states of all agents in the k-th time interval. This represents the set of actions taken by all agents in the k-th time interval. , This represents the network parameters of the hybrid network model. The network parameters of the waveform decision network model for the i-th agent are represented. This represents the overall action value function of all agents in the k-th time period. This represents the weight vector output by the attention network. This represents the i-th weight. This represents the total number of intelligent agents. A vector representing the value functions of the individual actions of all agents. This represents the weight matrix output by the hyperparameter network. This represents the weight vector output by the hyperparameter network. This represents the bias vector output by the hyperparameter network. This represents the bias scalar output by the hyperparameter network; , .

5. The anti-interference waveform decision-making method based on multi-agent reinforcement learning according to claim 4, characterized in that, The loss function used in the joint training is expressed as follows: ; in, Indicates the batch size of the sampling experience. Indicates from the experience replay pool Mid-sampling, This represents the system state during the k-th time period. This represents the set of states of all agents in the k-th time interval. This represents the set of actions taken by all agents in the k-th time interval. This represents the set of rewards for all agents in the k-th time interval. This represents the system state in the (k+1)th time period. This represents the set of states of all agents in the (k+1)th time interval. This represents the Q-value output by the target network in the DDQN network architecture. , This represents the sum of rewards for all agents in the k-th time interval. This represents the set of actions jointly selected by all agents in the (k+1)th time interval that maximizes the overall action value. This represents the network parameters of the target network in the k-th time period. This indicates that the target network has certain network parameters. Under these conditions, all agents are based on a set of states. Select Action The resulting overall action value function.

6. The anti-interference waveform decision-making method based on multi-agent reinforcement learning according to claim 1, characterized in that, The reward for the agent is represented as: ; in, This represents the reward of the i-th agent in the k-th time interval. This represents the communication rate of the i-th agent in the k-th time period. This represents the waveform switching cost of the i-th agent in the k-th time period. , This represents the number of coded blocks in the data packet to be sent by the i-th agent in the k-th time period. This represents the channel frequency of the i-th agent in the k-th time period. This represents the transmission power of the i-th agent in the k-th time period. This indicates the preset weighting coefficient.

7. The anti-interference waveform decision-making method based on multi-agent reinforcement learning according to claim 6, characterized in that, The formula for the communication rate is: ; in, This indicates the number of bits in each coded block. This represents the collision flag of the i-th agent in the k-th time interval. , This represents a conditional function, if... If the condition is met, then ,otherwise , This represents the total number of intelligent agents; , This represents the block error rate of the i-th agent in the k-th time period. This indicates the preset error rate threshold.

8. A multi-agent reinforcement learning-based anti-interference waveform decision-making device, characterized in that, include: The grouping module is used to divide the communication channel in the ad hoc network communication system into a specified number of groups to obtain the channel grouping result; The acquisition module is used to acquire local observation information of each communication node in the current time slot of the ad hoc network communication system; wherein, the local observation information includes: the status information of the target channel group corresponding to the communication node and the feedback information of the communication node after executing the waveform decision of the previous time slot; the target channel group includes: the channel group to which the communication channel in the waveform decision of the previous time slot belongs and a random channel group; the status information includes: the spectral energy information of each channel in the current time slot and the occupancy status in the previous time slot; the feedback information includes: communication rate and block error rate; The processing module is used to process the local observation information of the target communication node using the target waveform decision network model to obtain the waveform decision parameters of the target communication node in the current time slot. The target communication node represents any communication node in the ad hoc network communication system. Each communication node in the ad hoc network communication system is deployed one-to-one with each agent in the trained target multi-agent reinforcement learning model, and each agent has an independent waveform decision network model. The state of the agent is the local observation information of the corresponding communication node, and the action of the agent is the waveform decision parameter of the corresponding communication node. The reward of the agent is positively correlated with the communication rate of the communication node and negatively correlated with power overhead and waveform switching overhead. The waveform decision parameters include: channel frequency, transmit power, and the number of coded blocks in the data packet to be transmitted.

9. An electronic device comprising a memory and a processor, wherein the memory stores a computer program executable on the processor, characterized in that, When the processor executes the computer program, it implements the anti-interference waveform decision method based on multi-agent reinforcement learning as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions, which, when executed by a processor, implement the anti-interference waveform decision-making method based on multi-agent reinforcement learning as described in any one of claims 1 to 7.