Spectrum resource allocation method based on multi-agent coordination mechanism

CN117858239BActive Publication Date: 2026-09-11XIDIAN UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310783483.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2023-02-28
Filing Date
2023-06-29
Publication Date
2026-09-11
Estimated Expiration
2043-06-29

AI Technical Summary

Technical Problem

[0005]本发明的目的是提供一种基于多智能体协同机制的频谱资源分配方法,解决了目前频谱资源分配时存在的面对动态复杂环境的适应能力较弱,面对干扰时抗干扰能力较弱有待进一步提高的问题

Benefits of technology

[0055] The beneficial effects of this invention are that the spectrum resource allocation method based on the multi-agent cooperative mechanism is designed with a performance evaluation function for wireless systems with multiple frequency-using devices. By designing corresponding performance evaluation functions and performance thresholds for various frequency-using devices, the performance indicators of different types of devices are transformed from different dimensions into a unified measure of effectiveness, which facilitates the performance evaluation of wireless systems with multiple frequency-using devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117858239B_ABST
    Figure CN117858239B_ABST
Patent Text Reader

Abstract

The application discloses a spectrum resource allocation method based on a multi-agent cooperation mechanism, and specifically comprises the following steps: step 1, initializing environment information; step 2, acquiring position information between frequency using nodes and spectrum situation information at the frequency using nodes; step 3, making a spectrum allocation scheme decision; step 4, executing the spectrum allocation scheme and performing performance evaluation; and step 5, collecting historical experience information and performing experience information shunting training based on a cooperation mechanism to realize spectrum resource allocation for a dynamic environment. The method can timely respond to a complex dynamic environment and has strong anti-interference capability. Meanwhile, multi-agent common decision making can reduce the calculation load of each agent and the required information demand, improve robustness, and has stronger environmental adaptability. Finally, through the cooperation training of each agent, the negative influence of an unstable environment on training is reduced, and the performance of the multi-agent system is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of spectrum resource allocation technology, specifically relating to a spectrum resource allocation method based on a multi-agent cooperative mechanism. Background Technology

[0002] With the development of wireless communication technology, wireless applications based on spectrum resources are becoming increasingly widespread, and the types and numbers of frequency-using devices are constantly increasing. However, spectrum resources are a finite and non-renewable resource. The contradiction between the ever-increasing demand for frequency and the limited spectrum resources makes efficient spectrum resource allocation algorithms an important method to resolve this contradiction. Existing spectrum resource allocation algorithms include traditional algorithms based on segmentation or frequency tables, game theory-based spectrum resource algorithms, graph theory-based spectrum resource algorithms, centralized decision-making algorithms based on artificial intelligence, and distributed decision-making algorithms based on multi-agent models.

[0003] The increasing number of frequency-using devices and the expanding application scope have led to a more complex electromagnetic environment, resulting in frequent incidents of malicious electromagnetic interference. Traditional algorithms struggle to adaptively make decisions regarding spectrum resources in such a dynamic and complex electromagnetic environment. In centralized decision-making algorithms based on artificial intelligence, the decision center needs to acquire global state information from each node and efficiently complete high-load decision-making computations. This places high demands on the system's basic communication links and the computational capabilities of the decision center, which is difficult to meet under harsh environmental conditions. In distributed decision-making algorithms based on multi-agent models, each agent can only acquire local state information and engage in partial information exchange. The reward feedback received by each agent is influenced not only by its own local state information and its own decisions but also by the decisions of other agents. Therefore, most distributed decision-making algorithms based on multi-agent models often suffer from environmental instability, making convergence difficult and prone to getting stuck in local optima. Furthermore, when facing a wide variety of frequency-using devices, each type often has different performance indicators. How to objectively and comprehensively evaluate a system containing multiple frequency-using devices is also a problem that needs to be solved.

[0004] Existing technologies include spectrum resource allocation algorithms based on game theory, graph theory, centralized artificial intelligence, and distributed multi-agent systems. Game theory and graph theory-based algorithms share a common drawback: weak adaptability to dynamic and complex environments. They struggle to make timely decisions when the electromagnetic environment changes, resulting in weak anti-interference capabilities. Centralized artificial intelligence-based algorithms have two drawbacks: first, they require high preconditions; the central node needs to acquire the spectrum situation information of the entire environment, which is difficult to meet in harsh electromagnetic environments. Second, they place extremely high demands on the central node's computing power. Since the central node must allocate spectrum resources to all frequency-using devices using the electromagnetic information of the entire environment, the computational load is enormous, placing significant demands on the central node's hardware performance. Distributed multi-agent algorithms suffer from environmental instability; for any given agent, its decision-making performance depends not only on its own decisions and the environment but also on the decisions of other agents. This leads to a situation where, when other agents make different decisions, agents may exhibit different behavioral outcomes even when making the same decision in the same environment, making the training of multi-agent systems more difficult. Therefore, in summary, current spectrum resource allocation suffers from weak adaptability to dynamic and complex environments and weak anti-interference capabilities, requiring further improvement. Summary of the Invention

[0005] The purpose of this invention is to provide a spectrum resource allocation method based on a multi-agent cooperative mechanism, which solves the problems of weak adaptability to dynamic and complex environments and weak anti-interference ability in the face of interference in current spectrum resource allocation, which need to be further improved.

[0006] The technical solution adopted in this invention is:

[0007] The spectrum resource allocation method based on the multi-agent cooperative mechanism is carried out according to the following steps:

[0008] Step 1: Initialize the environmental information;

[0009] Step 2: Obtain the location information between frequency-using nodes and the spectrum situation information at the frequency-using nodes;

[0010] Step 3: Make a decision on the spectrum allocation scheme;

[0011] Step 4: Implement the spectrum allocation scheme and perform performance evaluation;

[0012] Step 5: Collect historical experience information and conduct diversion training based on the experience information according to the collaborative mechanism to realize spectrum resource allocation in dynamic environment.

[0013] The invention is further characterized by;

[0014] In step 1, the neural network at each frequency node is initialized, as shown in the following formula (1):

[0015]

[0016] Where ω represents different initialization parameters in the neural network, setting the performance threshold thr of the radar equipment. r The performance threshold of communication equipment. c The input variable of a neural network is represented by the state, and the effect of the neural network on the data is represented by the function F.

[0017] In step 2, each frequency-using node communicates the location information of the communication center. With node location information Calculate the distance information between each frequency-using node and the communication equipment using the following formula (2):

[0018]

[0019] Channel scanning is performed at each frequency-using node using power sensing technology to obtain interference I on different channels at different frequency-using nodes. i,f,t ;

[0020] The spectrum situation information at the frequency-using node is composed of electromagnetic environmental noise, the spectrum situation generated by the frequency use of numerous nodes, and interference generated by jammers; that is, the following formula (3):

[0021]

[0022] Where I f,i,t This represents the interference experienced by frequency-using device i on channel f in time slot t. This represents the self-disturbance generated by each node. N represents the interference caused by the jammer. f,i,t Representing noise in the electromagnetic environment, the spectral situation information acquired by device i in time slot t is shown in the following formula (4):

[0023]

[0024] in This represents the spectral status information of communication center cen1 in channel f1 during time slot t.

[0025] Step 3 specifically involves:

[0026] Channel interference information at the frequency-using node and the distance information between that node and the communication center are used as the state input of the neural network. The state is used as input to the neural network, and the expected value of the node decision scheme is obtained through the forward propagation of the neural network. For Q, there are two ways to choose one of the decision schemes as the frequency and power usage schemes for the frequency-using node:

[0027] The first approach is to select the decision scheme with the largest q value for execution. The second approach is to use the ε-greedy algorithm for selection, selecting the behavior with the largest expected value with a probability of 1-ε as the strategy, and selecting a random behavior with a probability of ε, as shown in the following formula (5):

[0028]

[0029] Where 'a' represents a frequency utilization scheme for a frequency utilization node, π(a|s) represents the probability of selecting frequency utilization scheme 'a' when the neural network input state is 's', and num... a This indicates the total number of spectrum allocation schemes.

[0030] Step 4 is as follows:

[0031] Based on the frequency usage plans decided by each frequency-using node, the power and frequency of each frequency-using device are adjusted, and the radar equipment performance threshold is determined accordingly. r With the performance threshold of communication equipment thr c Evaluate the performance of frequency usage schemes;

[0032] Different thresholds are set for the two devices as reference standards for performance estimation, as shown in the following formula (6):

[0033]

[0034] When work i,t A value of 1 indicates that device i is actively working in time slot t; otherwise, it is not actively working. i,t The performance of node i in time slot t is represented by thr type(i) The threshold value represents the category to which frequency-using device i belongs, and type(i) represents the category of node i, as shown in the following formula (7):

[0035]

[0036] Among them, Set c Set represents a set of communication devices. r Indicates a collection of radar equipment;

[0037] Regarding the performance of frequency-using device i in time slot t, Per i,t It is necessary to classify and discuss them according to different equipment types, as shown in the following formula (8):

[0038]

[0039] The objective function of the system is then... The performance evaluation function for radar equipment is as follows (9);

[0040]

[0041] Among them, P i,t Let σ represent the power used by device i in time slot t, G represent the antenna gain, σ represent the radar cross section, R represent the distance between the radar and the target, and λ represent the radar wavelength. The minimum received power for radar monitoring is obtained by summing the minimum signal-to-noise ratio threshold for radar monitoring with interference and noise.

[0042] Set a distance threshold for the radar equipment as a performance threshold: thr r =R thr ;

[0043] The effective node evaluation function of the radar equipment is as follows (10);

[0044]

[0045] The performance evaluation function for communication equipment is designed as follows (11):

[0046]

[0047] Among them, C i,cen,t SINR represents the channel capacity of the communication channel between communication device i and communication center cen in time slot t. i,cen,t This represents the signal-to-interference-plus-noise ratio from frequency-using device i to communication center cen in time slot t;

[0048] The performance threshold for communication equipment is set as the minimum channel capacity C required for the communication equipment to complete the communication task. thr , i.e., thr c =C thr Then the effective node evaluation function of the communication device is as follows (12):

[0049]

[0050] The performance evaluation of the spectrum allocation scheme for all frequency-using nodes in time slot t is as follows:

[0051] Step 5 specifically involves:

[0052] Partial information from steps 2 to 4 is collected into the experience pool of each agent. The adjacent historical information is analyzed, judged and divided through a collaborative mechanism. For the experience of using the same channel to increase power and obtain lower individual performance, under the collaborative mechanism, the historical experience is classified into secondary experience, priority experience and regular experience, as shown in the following formula (13):

[0053]

[0054] The neural network in the frequency-consuming node is trained using a deep double-Q network algorithm, thereby modifying various parameters ω in the neural network. m historical information sequences are extracted from historical experience information, and the mean squared error loss function is applied. Calculate the loss value, Q(S) j A j ,ω) represents the state S under the current agent parameter ω calculation. j Adopting behavior A j The expected performance evaluation is obtained, where y j =R j +γQ'(S' j ,arg max a' Q(S' j When training an agent using gradient replay based on the loss value, the learning rate based on the data with priority historical experience information is lr. L The learning rate for data based on conventional historical experience information is lr. M The learning rate for data based on secondary historical experience information is lr. S , where lr L >lr M >lr S .

[0055] The beneficial effects of this invention are that the spectrum resource allocation method based on the multi-agent cooperative mechanism is designed with a performance evaluation function for wireless systems with multiple frequency-using devices. By designing corresponding performance evaluation functions and performance thresholds for various frequency-using devices, the performance indicators of different types of devices are transformed from different dimensions into a unified measure of effectiveness, which facilitates the performance evaluation of wireless systems with multiple frequency-using devices.

[0056] This paper designs a historical experience diversion mechanism based on multi-agent collaboration. By designing a collaborative approach among multiple agents, a diversion module is designed for experience pool replay. Different priorities and learning rates are applied to historical experience information of different priorities to reduce the impact of environmental instability on agent training and improve system reliability.

[0057] This invention designs an experience pool that collects both overall and local benefits. In existing deep reinforcement learning schemes, the experience pool usually only contains overall or local benefits. This invention achieves the goal of balancing the local and overall benefits among agents by replaying both overall and local benefits together and applying them to a multi-agent collaborative mechanism.

[0058] The method of this invention can respond promptly to complex and dynamic environments and has strong anti-interference capabilities. Simultaneously, utilizing multi-agent collaborative decision-making reduces the computational load and information requirements of each agent, improving robustness and enhancing environmental adaptability. Finally, through collaborative training among the agents, the negative impact of environmental instability on training is mitigated, resulting in a certain improvement in the performance of the multi-agent system. Attached Figure Description

[0059] Figure 1 This is a flowchart of the spectrum resource allocation method based on the multi-agent cooperative mechanism of the present invention;

[0060] Figure 2 This is a flowchart illustrating part of the observable Markov decision-making process in the spectrum resource allocation method based on a multi-agent cooperative mechanism of the present invention.

[0061] Figure 3 This is a schematic diagram of the spectrum resource allocation process based on multi-agent collaboration in the spectrum resource allocation method based on multi-agent collaboration mechanism of the present invention.

[0062] Figure 4 This is a schematic diagram illustrating the application scenario of the spectrum resource allocation method in the spectrum resource allocation method based on the multi-agent cooperative mechanism of the present invention;

[0063] Figure 5 This is a schematic diagram of the resource allocation method in the spectrum resource allocation method based on the multi-agent cooperative mechanism of the present invention. Detailed Implementation

[0064] The spectrum resource allocation method based on a multi-agent cooperative mechanism of the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0065] Example 1;

[0066] like Figure 1 As shown, step 1: Initialize the environmental information;

[0067] Initialize the number of frequency-consuming nodes (num) node Number of communication centers (num) center Number of malicious interference nodes (num) e Number of shared channels (num) f Number of optional power ratings for radar equipment Number of optional power options for communication equipment

[0068] Node location information Communication center location information

[0069] The location of the malicious interference node center is Initialize channel space Radar power set

[0070] Communication power set The neural network at each frequency node is initialized as shown in the following formula (1):

[0071]

[0072] Where ω represents different initialization parameters in the neural network, setting the performance threshold thr of the radar equipment. r The performance threshold of communication equipment. c .

[0073] Step 2: Obtain the location information between frequency-using nodes and the spectrum situation information at the frequency-using nodes.

[0074] Each frequency-using node communicates with the communication center for location information. With node location information Calculate the distance information between each frequency-using node and the communication equipment using the following formula (2):

[0075]

[0076] Channel scanning is performed at each frequency-using node using power sensing technology to obtain interference I on different channels at different frequency-using nodes. i,f,t .

[0077] The spectral situation information at the frequency-using node is composed of electromagnetic environmental noise, the spectral situation generated by the frequency use of numerous nodes, and interference generated by jammers. That is, as shown in formula (3):

[0078]

[0079] Where I f,i,t This represents the interference experienced by frequency-using device i on channel f in time slot t. This represents the self-disturbance generated by each node. N represents the interference caused by the jammer. i,f,t This represents noise in the electromagnetic environment. The spectral situation information acquired by device i in time slot t is shown in formula (4) below:

[0080]

[0081] in, This represents the spectral status information of communication center cen1 in channel f1 during time slot t.

[0082] Step 3: Make a decision on the spectrum allocation scheme;

[0083] Channel interference information at the frequency-using node and the distance information between that node and the communication center are used as the state input of the neural network. Using the state as input to the neural network, the expected value of the node's decision scheme is obtained through the forward propagation of the neural network:

[0084]

[0085] Where q i Indicates the decision option a i Given the expected value of Q, one decision scheme is selected as the frequency and power usage scheme for the frequency-using node. In this invention, there are two ways to select the decision scheme. The first is to select the decision scheme with the largest q value for execution. The second is to use the ε-greedy algorithm for selection, selecting the behavior with the largest expected value with a probability of 1-ε as the strategy, and selecting random behavior with a probability of ε, as shown in the following formula (5):

[0086]

[0087] Where 'a' represents a frequency utilization scheme for a frequency utilization node, π(a|s) represents the probability of selecting frequency utilization scheme 'a' when the neural network input state is 's', and num... a This indicates the total number of spectrum allocation schemes.

[0088] Step 4: Implement the spectrum allocation scheme and perform performance evaluation;

[0089] Based on the frequency usage plans decided by each frequency-using node, the power and frequency of each frequency-using device are adjusted, and the radar equipment performance threshold is determined accordingly. r With the performance threshold of communication equipment thr c Evaluate the performance of the frequency usage scheme.

[0090] Since the system contains two frequency-using devices with different functions, namely radar and communication, different thresholds are set for the two devices to unify the units of measurement and serve as a reference standard for performance estimation, as shown in the following formula (6):

[0091]

[0092] When work i,t A value of 1 indicates that device i is actively working in time slot t; otherwise, it is not actively working. i,tThe performance of node i in time slot t is represented by thr type(i) The threshold value represents the category to which frequency-using device i belongs, and type(i) represents the category of node i, as shown in the following formula (7):

[0093]

[0094] Among them, Set c Set represents a set of communication devices. r This refers to a collection of radar equipment.

[0095] Regarding the performance of frequency-using device i in time slot t, Per i,t It is necessary to classify and discuss them according to different equipment types, as shown in the following formula (8):

[0096]

[0097] The objective function of the system is then... The performance evaluation function for radar equipment is as follows (9);

[0098]

[0099] Among them, P i,t Let σ represent the power used by device i in time slot t, G represent the antenna gain, σ represent the radar cross section, R represent the distance between the radar and the target, and λ represent the radar wavelength. This represents the minimum received power for radar monitoring, measured by the minimum signal-to-noise ratio threshold for radar monitoring. With interference and noise and I rad The result is obtained through calculation, that is:

[0100]

[0101] Set a distance threshold for the radar equipment as a performance threshold:

[0102] thr r =R thr

[0103] The effective node evaluation function of the radar equipment is as follows (10);

[0104]

[0105] The performance evaluation function for communication equipment is designed as follows (11):

[0106]

[0107] Among them, C i,cen,tSINR represents the channel capacity of the communication channel between communication device i and communication center cen in time slot t. i,cen,t This represents the ratio of signal to interference plus noise from frequency-using device i to communication center cen in time slot t.

[0108] The performance threshold for communication equipment is set as the minimum channel capacity C required for the communication equipment to complete the communication task. thr ,Right now:

[0109] thr c =C thr ;

[0110] The effective node evaluation function of the communication device is as follows (12):

[0111]

[0112] The performance evaluation of the spectrum allocation scheme for all frequency-using nodes in time slot t is as follows:

[0113] Step 5: Collect historical experience information and conduct training by diverting the experience information based on the collaborative mechanism.

[0114] Partial information from steps 2 to 4 is collected into the experience pool of each agent. The collected information includes (s, a, s', R, r), such as... Figure 3 As shown.

[0115] Where s represents the input information of the frequency-sharing node neural network, a represents the spectrum allocation scheme selected in step 3, s' represents the environmental information isomorphic to s after the spectrum allocation scheme is executed, and R represents the performance evaluation of all node spectrum allocation schemes obtained in step 4. r represents the performance evaluation result of the frequency-using node itself.

[0116] Example 2;

[0117] Furthermore, the historical information from neighboring regions is analyzed, judged, and categorized through a collaborative mechanism. Experiences that achieve lower individual performance by increasing power using the same channel are considered secondary experiences; experiences that achieve higher performance scores by reducing power using the same channel are considered primary experiences; experiences that achieve lower overall performance using the same power but different channels are considered secondary experiences; experiences that achieve higher overall performance using the same power but different channels are considered primary experiences; and other experiences categorized as conventional experiences are classified as secondary experiences, primary experiences, and conventional experiences under the collaborative mechanism. This is illustrated in the following formula (13):

[0118]

[0119] The neural network in the frequency node is trained by Deep Reinforcement Learning with Double Q-learning (DDQN) algorithm, thereby modifying various parameters ω in the neural network. m historical information sequences are extracted from historical experience information, and the mean squared error loss function is calculated according to the following formula (14):

[0120]

[0121] Calculate the loss value, Q(S) j A j ,ω) represents the state S under the current agent parameter ω calculation. j Adopting behavior A j The expected performance is as follows (15):

[0122] y j =R j +γQ'(S' j ,arg max a' Q(S' j ,a,ω),ω) (15);

[0123] When training the agent using gradient replay based on the loss value, the learning rate based on the data with priority historical experience information is lr. L The learning rate for data based on conventional historical experience information is lr. M The learning rate for data based on secondary historical experience information is lr. S , where lr L >lr M >lr S .

[0124] Example 3;

[0125] like Figure 2 As shown, the theoretical basis of the spectrum resource allocation method based on a multi-agent cooperative mechanism in this invention is to train agents using a multi-agent reinforcement learning algorithm, and then use the trained agents to make decisions on spectrum resource allocation. The basic model of multi-agent reinforcement learning can be represented by a partially observable Markov decision process, which can be represented by tuples. <N,S space A space ,R,T,γ,O> represents, where N represents the set of agents; S represents the global state space of the environment;

[0126] A represents the joint action space of the agent, where the action vector is a = [a1, a2, ..., a...]. i ,...]∈A space , where a iLet R represent the independent actions of agent i, and let R be the current state-action pair (s, a1, a2, a3, ..., a...). i The global reward function under (s, ...); T represents the state transition function of the environment T(s'|s,A)∈(0,1); γ is the discount factor; O is the observation state vector of the agent at each time step [o1,o2,...,o...). i ,...].

[0127] In reinforcement learning, an agent is trained through its interaction with the environment. The basic principle is that if the reward R from the environment is positive after the agent performs an action 'a', the agent's tendency to perform that action 'a' in the future will increase. If the policy is represented as a conditional probability distribution, this means that the probability of taking that action increases under the same conditions. Conversely, if the reward R from the environment is negative, the agent's tendency to take that action will decrease.

[0128] like Figure 4 As shown, after the spectrum resource allocation method based on the multi-agent cooperative mechanism of this invention is completed, the algorithm effect is tested using simulation data. Specifically:

[0129] Consider a scenario where multiple frequency-using nodes perform node functions. Each frequency-using node includes a communication device and a radar device. The communication device performs communication tasks with the communication center, and the radar device performs monitoring tasks. Meanwhile, malicious interference nodes exist in the environment, maliciously interfering with the frequency-using nodes. Due to various factors, such as changes in environmental climate, wireless interference from external devices, and self-interference from the frequency-using nodes within the system, complex and dynamic interference exists in the channel. Let the number of frequency-using nodes in the scenario be num. node The number of frequency-using devices is num. equ =2*num node The number of communication centers is num center The number of malicious interference devices is num e Let num be the number of channels shared by radar equipment, communication equipment, and malicious jamming equipment. f The center frequency of each channel is expressed as Inter-channel interference is considered only in terms of co-channel interference, not inter-channel interference. Let the usable power of the radar equipment be the radar power set. in This represents the amount of power available to the radar equipment. Let the power available to the communication equipment be the communication power set. in This represents the amount of power that the communication equipment can use. Therefore, the number of spectrum resource allocation schemes for a frequency-using node is the number of combinations of the frequencies and power used by the communication equipment and radar equipment. In this invention, it is assumed that the communication equipment and radar equipment cannot use the same frequency, and the number of combinations is... For ease of description, let's assume... Figure 4 The bottom left corner of the map is the origin of the two-dimensional coordinate system. The location of a frequency-using node is represented by coordinates (x, y), where x represents the horizontal coordinate of the frequency-using device, and y represents the vertical coordinate, x ∈ [0, X], y ∈ [0, Y]. X represents the maximum value of the horizontal coordinate, and Y represents the maximum value of the vertical coordinate. The representation of communication and radar equipment within the frequency-using node is the same as that of the frequency-using node itself. Let the locations of the frequency-using node be represented as follows: The location of the communication center is The location of the malicious interference node center is

[0130] In an example of the spectrum resource allocation method based on the multi-agent cooperative mechanism of this invention, during a specific time slot where a multi-agent decision-making model with a training count K of 16000 times makes decisions about the environment, the decision result is that the radar at node one uses ch6 and the communication equipment uses ch1; the radar at node two uses ch5 and the communication equipment uses ch1, and so on. Ultimately, out of all 40 frequency-using devices, 34 devices reach the effective operating threshold. Partial spectrum resource allocation results of the scheme are shown below. Figure 5 Even in this dynamic and complex interference environment, the algorithm performs well.

[0131] When the environment changes, repeating steps 2, 3, 4, and 5 can achieve spectrum resource allocation for the dynamic environment and further enhance the decision-making ability of frequency-using nodes on spectrum allocation schemes.

[0132] This invention presents a spectrum resource allocation method based on a multi-agent cooperative mechanism, which reduces the impact of environmental instability on multi-agent training and improves the reliability of training results. By unifying the playback of local and overall gains, and simultaneously coordinating the local and overall gains of each frequency-using device in the system, the overall performance of the system is enhanced. The system evaluation function designed for the system model lays a numerical foundation for the system's performance analysis and reinforcement learning training.

[0133] This invention presents a spectrum resource allocation method based on a multi-agent cooperative mechanism, which can respond promptly to complex dynamic environments and has strong anti-interference capabilities. Simultaneously, utilizing multi-agent collaborative decision-making reduces the computational load and information requirements of each agent, improving robustness and enhancing environmental adaptability. Finally, through collaborative training among the agents, the negative impact of environmental instability on training is mitigated, resulting in improved performance of the multi-agent system and demonstrating significant practical value.

Claims

1. A spectrum resource allocation method based on a multi-agent coordination mechanism, characterized in that, The specific steps are as follows: Step 1: Initialize the environmental information; Step 2: Obtain the location information between frequency-using nodes and the spectrum situation information at the frequency-using nodes; Step 3: Make a decision on the spectrum allocation scheme; specifically: Channel interference information at the frequency-using node and the distance information between that node and the communication center are used as the state input of the neural network. The state is used as the input to the neural network, and the expected value of the node decision scheme is obtained through the forward propagation of the neural network. ,against There are two ways to choose one of the decision schemes as the frequency and power usage schemes for the frequency-using node: The first option is to choose. The decision with the highest value is executed; the second is to use... The algorithm is selected to... The strategy is to choose the action with the highest expected value based on probability. The probability is used to select random behavior, as shown in the following formula (5): (5); in For a specific frequency usage scheme of an application frequency node, This indicates that the input state of the neural network is Frequency selection scheme The probability of execution Indicates the total number of spectrum allocation schemes; Step 4: Implement the spectrum allocation scheme and perform performance evaluation; specifically: Based on the frequency usage plans decided by each frequency-using node, the power and frequency of each frequency-using device are adjusted, and the performance thresholds of the radar equipment are considered. Performance threshold of communication equipment Evaluate the performance of frequency usage schemes; Different thresholds are set for the two devices as a reference standard for performance estimation, as shown in the following formula (6): (6); when A value of 1 indicates that the device In the time slot If the work is valid, it is considered valid; otherwise, it is considered invalid. Represents a node In the time slot Performance, Indicates frequency-using equipment The threshold value set for the category to which it belongs. Represents a node The categories are as follows, as shown in formula (7): (7); in, Represents a set of communication devices. Indicates a collection of radar equipment; Regarding frequency-consuming devices In the time slot Performance in It is necessary to discuss this in terms of different equipment types, as shown in the following formula (8): (8): The objective function of the system is then... The performance evaluation function for radar equipment is as follows (9). (9); in, Indicates equipment In the time slot The power used Indicates antenna gain. Radar cross section, Indicates the distance between the radar and the target. Indicates the radar wave wavelength. The minimum received power for radar monitoring is obtained by summing the minimum signal-to-noise ratio threshold for radar monitoring with interference and noise. ; Set a distance threshold for the radar equipment as a performance threshold: ; The effective node evaluation function of the radar equipment is as follows (10); (10); The performance evaluation function for communication equipment is designed as follows (11): (11); in, Indicates communication equipment With the communication center In the time slot The channel capacity of the communication channel in the middle, Indicates time slot Medium frequency equipment To the communications center The signal-to-interference-plus-noise ratio; The performance threshold for communication equipment is set as the minimum channel capacity required for the equipment to complete the communication task. ,Right now Then the effective node evaluation function of the communication device is as follows (12): (12); Then all frequency-using nodes in the time slot The performance evaluation of the spectrum allocation scheme in the middle is as follows ; Step 5: Collect historical experience information and conduct diversion training based on the experience information according to the collaborative mechanism to realize the allocation of spectrum resources in dynamic environments.

2. The spectrum resource allocation method based on a multi-agent cooperative mechanism according to claim 1, characterized in that, In step 1, the neural network at each frequency node is initialized, as shown in the following formula (1): (1); in, Different initialization parameters in the neural network are used to set performance thresholds for radar equipment. Communication equipment performance thresholds.

3. The spectrum resource allocation method based on a multi-agent cooperative mechanism according to claim 1, characterized in that, In step 2, each frequency-using node communicates the location information of the communication center. With node location information Calculate the distance information between each frequency-using node and the communication equipment using the following formula (2): (2); Channel scanning is performed at each frequency-using node using power sensing technology to obtain interference on different channels at different frequency-using nodes. ; The spectrum situation information at the frequency usage node is composed of electromagnetic environmental noise, the spectrum situation generated by the frequency usage of numerous nodes, and interference generated by jammers; that is, the following formula (3): (3); in Indicates frequency-using equipment In the time slot In the channel Interference received above, This represents the self-disturbance generated by each node. This indicates interference caused by a jammer. Representing noise in the electromagnetic environment, for equipment... exist The spectrum situation information obtained is shown in the following formula (4): (4); in, Indicates communication center In the time slot Mid-channel Spectrum situation information.

4. The spectrum resource allocation method based on a multi-agent cooperative mechanism according to claim 1, characterized in that, Step 5 specifically involves: The information from steps 2 to 4 is collected into the experience pool of each agent. The adjacent historical information is analyzed, judged and divided through the collaborative mechanism. For the experience of using the same channel to increase power to obtain lower individual performance, under the collaborative mechanism, the historical experience is classified into secondary experience, priority experience and regular experience, as shown in the following formula (13): (13); The neural network in the frequency-consuming nodes is trained using a deep double-Q network algorithm, thereby modifying various parameters in the neural network. Extracting from historical experience information A historical information sequence, based on the mean squared error loss function Calculate the loss value. Indicates the current agent parameters Calculate the state Adopting behavior The expected performance evaluation results are as follows: When training an agent using gradient replay based on the loss value, the learning rate based on the data with priority historical experience information is... The learning rate of data based on conventional historical experience information is The learning rate of data based on secondary historical experience information is ,in .