A Multi-Agent Coverage Path Planning Method Based on Frequency Band Sharing

By employing a multi-agent coverage path planning method based on frequency band sharing and using reinforcement learning algorithms to train the decision network, the communication problem of frequency band monopoly mechanism in high-density agent clusters is solved, achieving stable communication and task continuity in spectrum-constrained scenarios.

CN122496827APending Publication Date: 2026-07-31TIANFU JIANGXI LAB
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
TIANFU JIANGXI LAB
Filing Date
2026-05-21
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

In scenarios with high-density agent clusters and limited spectrum resources, the frequency band exclusive mechanism cannot support simultaneous communication by all agents, resulting in decreased communication quality and impact on task execution continuity.

Method used

A multi-agent coverage path planning method based on frequency band sharing is adopted. The multi-agent decision network is trained by reinforcement learning algorithm. By combining communication quality reward and frequency band switching penalty, the agents are guided to learn distributed coordination strategy autonomously under power domain reuse, so as to achieve stable communication among multiple agents.

Benefits of technology

Under conditions of limited spectrum resources, all agents can maintain communication links through power domain multiplexing, avoiding communication interruptions and task stagnation, thereby improving the system's coverage efficiency and task execution continuity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122496827A_ABST
    Figure CN122496827A_ABST
Patent Text Reader

Abstract

This invention discloses a multi-agent coverage path planning method based on frequency band sharing, belonging to the field of multi-agent cooperative control technology. The method designs a composite reward function including a communication quality reward and a frequency band switching penalty. A reward is given when communication is successful after an agent performs power domain multiplexing; when communication is interrupted, differentiated penalties are applied based on the cause of the interruption, with penalties imposed on agents that change their access frequency band. A reinforcement learning algorithm is used to train a multi-agent decision network, enabling each agent to autonomously learn a frequency band sharing coordination strategy in distributed decision-making. After training, a trained feature extraction network and a corresponding policy network are deployed on each agent. Each agent inputs its local observation vector into the feature extraction network to obtain a hidden state, and then the policy network outputs a joint action containing movement and frequency band selection actions. This effectively avoids communication interruptions and ensures task continuity in high-density clusters and spectrum-constrained scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of multi-agent cooperative control technology, specifically to a multi-agent coverage path planning method based on frequency band sharing. Background Technology

[0002] In multi-agent collaborative coverage tasks, agent swarms (such as drone formations and robot swarms) need to autonomously plan their movement trajectories to efficiently traverse the designated target area in the shortest time or with the lowest energy consumption. Through reasonable path planning, agents can effectively avoid geographical obstacles, expand the physical search range, and thus significantly improve the system's detection efficiency.

[0003] Efficient collaboration among multiple agents relies heavily on stable and reliable wireless communication links. In traditional methods, agents typically access available frequency bands randomly. Due to a lack of coordination among agents, multiple agents may compete for the same available frequency band, resulting in co-channel interference, which degrades communication quality or even causes communication interruption, severely impacting overall coverage efficiency and task execution continuity.

[0004] Although the frequency band exclusivity mechanism can effectively solve the problem of co-frequency conflict, in scenarios with high-density agent clusters and scarce spectrum resources, the available frequency band resources are far less than the number of agents. As a result, the frequency band exclusivity mechanism cannot allocate an independent frequency band for each agent, causing some agents to be unable to obtain communication opportunities, which also affects the overall coverage efficiency and task execution continuity. Summary of the Invention

[0005] The technical problem to be solved by this invention is that in scenarios with high-density agent clusters and scarce spectrum resources, the frequency band exclusive mechanism cannot support simultaneous communication of all agents. The purpose is to provide a multi-agent coverage path planning method based on frequency band sharing, which solves the above-mentioned problem.

[0006] This invention is achieved through the following technical solution:

[0007] In a first aspect, the present invention provides a multi-agent coverage path planning method based on frequency band sharing, comprising:

[0008] Obtain the local observation vectors of all agents within the target task area at the same time.

[0009] A multi-agent decision network is trained using reinforcement learning algorithms based on the local observation vectors of each agent at the same time. This multi-agent decision network includes a feature extraction network and a policy network for each agent. The training process employs a composite reward function that includes a communication quality reward term and a frequency band switching penalty term. The communication quality reward term is used to reward agents who successfully communicate after performing power domain reuse under a frequency band sharing mechanism, and to impose differentiated penalties based on the cause of communication interruption. The frequency band switching penalty term penalizes agents for changing their access frequency band.

[0010] Deploy the trained feature extraction network and the corresponding trained policy network on each agent;

[0011] Each agent inputs its local observation vector at the current moment into the trained temporal feature extraction network to obtain the agent's hidden state, and inputs the hidden state into the corresponding trained policy network to output the agent's joint action at the next moment; the joint action includes a movement action and a frequency band selection action.

[0012] Optionally, the multi-agent decision network further includes two critic networks with identical structures but different initialization parameters; the step of training the multi-agent decision network using a reinforcement learning algorithm based on the local observation vectors of each agent at the same time includes:

[0013] The local observation vectors of all agents at time t are input into the feature extraction network to obtain the hidden state of each agent at time t. The hidden state of each agent at time t is then input into its corresponding policy network to obtain the joint action of each agent at time t.

[0014] Based on the frequency band selection actions of each agent at time t, determine the set of agents that can access each frequency band;

[0015] The water-filling algorithm is used to allocate the total transmit power of the base station to each frequency band, and power allocation is performed on the intelligent agent set accessing the same frequency band;

[0016] Each agent performs power domain multiplexing in non-orthogonal multiple access technology according to the allocated power, and calculates the composite reward value at time t according to the composite reward function;

[0017] Based on the composite reward value and the action values ​​output by the two commentator networks at adjacent time points, a mean squared error loss function is constructed, and the parameters of the two commentator networks are updated respectively through gradient descent.

[0018] Optionally, the method of using a water-filling algorithm to allocate the total transmit power of the base station to each frequency band includes:

[0019] The maximum value among the channel gains of all agents accessing the same frequency band is taken as the equivalent channel gain of that frequency band.

[0020] Calculate the system capacity based on the equivalent channel gain of each frequency band;

[0021] With the goal of maximizing system capacity, and constrained by the fact that the sum of the power allocated to each frequency band does not exceed the total transmit power of the base station, and the power allocated to each frequency band is not lower than the threshold power, a Lagrangian function is constructed.

[0022] Based on the Lagrange function and KKT conditions, the power allocated to each frequency band is derived, and the specific calculation formula is as follows:

[0023]

[0024] in, The power allocated to frequency band n; N represents the total transmit power of the base station; N is the total number of frequency bands. This represents the power of additive white Gaussian noise. Let n be the equivalent channel gain for frequency band n; Let k be the equivalent channel gain for frequency band k. Let n be the total interference power from external interference sources received by frequency band n; This represents the total interference power from external sources experienced by frequency band k.

[0025] Optionally, the power allocation for the set of intelligent agents accessing the same frequency band includes:

[0026] Calculate the channel gain of each agent in a set of agents accessing the same frequency band;

[0027] Calculate the power allocation coefficient for each agent based on the channel gain of each agent; agents with smaller channel gains are allocated larger power allocation coefficients.

[0028] The power of each agent is obtained by multiplying its power allocation coefficient by the power allocated to the same frequency band.

[0029] Optionally, the formula for calculating the power allocation coefficient is as follows:

[0030]

[0031] in, The power allocation coefficient for the j-th agent in access frequency band n; The channel characteristics of the j-th agent in access frequency band n; The number of agents accessing frequency band n; As the attenuation factor, .

[0032] Optionally, after updating the parameters of the two critic networks separately using gradient descent, the method further includes:

[0033] Based on the counterfactual advantage function, the parameters of the feature extraction network and each policy network are updated respectively by the policy gradient theorem; the counterfactual advantage function is used to evaluate the marginal contribution of the current action performed by the current agent to the global benefit when the actions of other agents are fixed.

[0034] When the value of the counterfactual advantage function is greater than zero, the output probability of the current action is increased through gradient update;

[0035] When the value of the counterfactual advantage function is less than zero, the output probability of the current action is reduced through gradient updates.

[0036] Optionally, the expression for the counterfactual advantage function is as follows:

[0037]

[0038] in, Represents the counterfactual advantage function; Let the action space of agent u be defined. This represents any action taken by agent u as it traverses the action space. The hidden state of agent u at time t, as output by the feature extraction network; This represents the action probability distribution output by the policy network corresponding to agent u; This indicates that agent u is in the hidden state. Selecting joint actions The probability of; Indicates the actions of other intelligent agents; This indicates that when the actions of other intelligent agents are fixed as At that time, the current action of the current agent u is replaced with .

[0039] Optionally, the provision that differential penalties are applied based on the cause of communication interruption includes:

[0040] When the signal-to-interference-plus-noise ratio (SIR) of an agent after power domain multiplexing under the frequency band sharing mechanism is lower than the preset SIR threshold, the agent's communication is determined to be interrupted, and the following operations are performed based on the cause of the interruption:

[0041] If the power range of adjacent channel features is less than the minimum power range that satisfies the decoding requirements, then a heavy penalty is applied;

[0042] If the power range of adjacent channel features is greater than or equal to the minimum power range, but the signal-to-interference-plus-noise ratio (SIR) of the agent is still lower than the preset SIR threshold after power allocation, then a standard penalty is applied.

[0043] Wherein, the absolute value of the standard penalty is less than the absolute value of the severe penalty.

[0044] Optionally, the formula for calculating the signal-to-interference-plus-noise ratio is as follows:

[0045]

[0046] in, The signal-to-interference-plus-noise ratio (SIR) of agent u accessing frequency band n; Agents assigned to access frequency band n by base stations The transmission power; Other intelligent agents allocated access frequency band n by the base station The transmission power; Let be the channel gain from the base station to the agent u in the access frequency band n; The channel gain from the interference source o to the agent u; This represents the set of interference sources operating in frequency band n; The transmission power of the interference source o; This represents the power of additive white Gaussian noise.

[0047] Optionally, the composite reward function further includes a coverage reward and a collision penalty; the coverage reward is used to reward the agent for moving from a covered area to an uncovered target area; the collision penalty is used to penalize the agent for colliding with geographical obstacles or going beyond the target task area.

[0048] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0049] This application provides a multi-agent coverage path planning method based on frequency band sharing. This method employs power domain reuse under a frequency band sharing mechanism, allowing multiple agents to transmit simultaneously on the same frequency band through power differences, thus overcoming the limitation imposed by the number of frequency bands on the agent scale. A composite reward function is designed, including a communication quality reward term and a frequency band switching penalty term. The communication quality reward term provides rewards or differentiated penalties based on the communication results after power domain reuse, guiding agents to actively create channel conditions conducive to successful communication. The frequency band switching penalty term is used to suppress unnecessary frequency hopping behavior, guiding agents to prioritize resolving resource contention through power domain reuse and reducing signaling overhead. Based on this, a multi-agent decision network is trained using reinforcement learning algorithms, enabling each agent to autonomously learn distributed coordination strategies through trial and error in continuous interaction with the environment. Therefore, even in scenarios with high-density agent clusters and limited spectrum resources, all agents can still maintain communication links simultaneously through power domain reuse, avoiding communication interruptions and task stagnation caused by frequency band monopolies. Attached Figure Description

[0050] To more clearly illustrate the technical solutions of the exemplary embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly described below. It should be understood that the following drawings only show some embodiments of the present invention and should not be considered as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort. In the drawings:

[0051] Figure 1 A flowchart illustrating the multi-agent coverage path planning method based on frequency band sharing provided in this application embodiment;

[0052] Figure 2 A network workflow diagram for the training and execution phases provided in the embodiments of this application. Detailed Implementation

[0053] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the embodiments and accompanying drawings. The illustrative embodiments and descriptions of the present invention are only used to explain the present invention and are not intended to limit the present invention.

[0054] Please refer to Figure 1 The following is a flowchart illustrating the multi-agent coverage path planning method based on frequency band sharing provided in this application embodiment. Figure 1 The multi-agent coverage path planning method based on frequency band sharing is introduced.

[0055] S1. Obtain the local observation vectors of all agents within the target task area at the same time.

[0056] In the specific implementation process, the target task area is divided into The target area is a two-dimensional discrete grid map. Each grid cell in the map has one of the following coverage states: obstacle area, uncovered target area, or covered area. The target area includes M mobile intelligent agents. An intelligent agent is an unmanned mobile platform with autonomous perception, decision-making, communication, and mobility capabilities, including but not limited to drones and robots.

[0057] Each agent u can acquire the local observation vector at the same moment through its onboard sensors, and its expression is as follows:

[0058]

[0059] in, Let be the local observation vector of agent u at time t; Let be the position vector of agent u at time t; Let be the local coverage matrix of agent u at time t, which is used to characterize the coverage state of each grid within the perception radius of agent u's current position; the coverage state is an obstacle area, an uncovered target area, or a covered area. Let be the local spectrum matrix of agent u at time t, used to measure the signal-to-interference-plus-noise ratio (SINR) of the available frequency bands at the current location for agent u to perceive.

[0060] After obtaining the local observation vectors of all agents at time t, the global state of the target task region at time t can be constructed, and its expression is as follows:

[0061]

[0062] in, Let be the global state at time t; Let be the global coverage matrix at time t, used to characterize the coverage status of each grid within the target task area; The global spectrum matrix at time t is used to characterize the signal-to-interference-plus-noise ratio (SIR) of each grid in each available frequency band within the target mission area. Let be the position vectors of M agents at time t.

[0063] S2. Based on the local observation vectors of each agent at the same time, a reinforcement learning algorithm is used to train a multi-agent decision network.

[0064] In practical implementation, the multi-agent decision network also includes a feature extraction network, a policy network for each agent, and a dual-critic network, which consists of two critic networks with identical structures but different initialization parameters. Reinforcement learning algorithms refer to agents learning the optimal policy through trial and error by continuously interacting with the environment. During training, each agent executes a joint action based on the current observed state. After receiving the joint actions of all agents, the environment provides a corresponding reward signal according to a pre-defined composite reward function and updates its state to proceed to the next time step. The agents continuously update the parameters of the multi-agent decision network with the goal of maximizing the accumulated composite reward, eventually converging to obtain the trained multi-agent decision network.

[0065] S3. Deploy the trained feature extraction network and the corresponding trained policy network on each agent.

[0066] After the multi-agent decision network is trained, a trained feature extraction network and trained policy networks for each of the M agents are obtained. Since all agents share the same feature extraction network, and the parameters of this network are fixed after training, it is only necessary to copy it M times and deploy it on each of the M agents. At the same time, the trained policy network for each agent is deployed on each of the M agents.

[0067] S4. Each agent inputs its local observation vector at the current time into the trained temporal feature extraction network to obtain the agent's hidden state, and inputs the hidden state into the corresponding trained policy network to output the agent's joint action at the next time step.

[0068] After network deployment, each agent, during actual task execution, first acquires its own local observation vector at the current moment through its onboard sensors. Each agent then inputs this local observation vector into a trained feature extraction network. This network extracts temporal features from the current and historical observations, outputting the agent's hidden state at the current moment. Subsequently, each agent inputs its hidden state into its corresponding trained policy network. The policy network, an agent-specific decision network, outputs a probability distribution of joint actions based on the hidden state. The joint action at the current moment is obtained by sampling or maximizing this probability distribution.

[0069] The combined actions include movement actions and frequency band selection actions. Movement actions control the agent's displacement direction within the grid map, and can involve moving up, down, left, right, or hovering. Frequency band selection actions select a frequency band from the available frequency band set for communication access.

[0070] Each agent independently executes the decision-making process described above at each time step, without waiting for the decision results of other agents or engaging in additional interactions with the environment. Throughout the online execution process, agents do not need to communicate in real time; they can make independent decisions solely based on their local observations and locally stored hidden states, achieving fully distributed online real-time execution.

[0071] In one possible embodiment, step S2 includes:

[0072] The local observation vectors of all agents at time t are input into the feature extraction network to obtain the hidden state of each agent at time t. The hidden states of each agent at time t are then input into their respective policy networks to obtain the joint action of each agent at time t. Based on the frequency band selection actions of each agent at time t, the set of agents accessing each frequency band is determined. A water-filling algorithm is used to allocate the total transmit power of the base station to each frequency band, and power allocation is performed on the set of agents accessing the same frequency band. Each agent performs power domain multiplexing in non-orthogonal multiple access technology according to the allocated power, and the composite reward value at time t is calculated based on the composite reward function. Based on the composite reward value and the action values ​​output by the two critic networks at adjacent times, a mean squared error loss function is constructed, and the parameters of the two critic networks are updated using gradient descent.

[0073] In the specific implementation, the feature extraction network includes multiple fully connected layers and a Long Short-Term Memory (LSTM) network. First, the local observation vectors of all agents at time t are used to construct the network input vector. Secondly, After feature extraction and dimensionality reduction through multiple fully connected layers, spatial feature vectors are obtained. These spatial feature vectors are then input into an LSTM to obtain the hidden states of each agent at time t. Then, consider the hidden states of each agent at time t. Input the policy network corresponding to each agent to obtain the joint action of all agents at time t. .

[0074] agent u in The joint action of time is ;in, For movement actions; Select an action for the frequency band. After all agents execute the joint action, update the global state at time t+1. The global coverage matrix includes time t+1. The global spectrum matrix at time t+1 The position vectors of M agents at time t+1 .

[0075] Based on the frequency band selection actions of each agent at time t, the set of agents accessing each frequency band is determined. A water-filling algorithm is used to allocate the total transmit power of the base station to each frequency band, and power allocation is performed on the sets of agents accessing the same frequency band. Each agent performs power domain multiplexing in Non-Orthogonal Multiple Access (NOMA) according to the allocated power. This means that multiple agents transmit on the same frequency band by allocating different transmit powers, and the communication mechanism separates the signals at the receiving end using Successive Interference Cancellation (SIC) technology. Subsequently, the composite reward value at time t is calculated based on the composite reward function. This composite reward value is used for subsequent policy network parameter updates.

[0076] The composite reward function includes a communication quality reward item. Frequency band switching penalty item Coverage of reward items and collision penalties The expression for the compound reward function R is as follows:

[0077]

[0078] (1) Communication quality award items If the agent successfully communicates after performing power domain multiplexing under the frequency band sharing mechanism, it will be rewarded; if communication is interrupted, a differentiated penalty will be imposed based on the reason for the interruption.

[0079] Specifically, when agent u performs power domain multiplexing under the frequency band sharing mechanism, the signal-to-interference-plus-noise ratio (SIR) is higher than the preset SIR threshold. If the communication by agent u is deemed successful, agent u is rewarded. If the signal-to-interference-plus-noise ratio (SIR) of agent u after power domain multiplexing under the frequency band sharing mechanism is lower than a preset SIR threshold... When this occurs, it is determined that the communication of the intelligent agent u is interrupted, and the following operations are performed based on the reason for the interruption:

[0080] If the power range of adjacent channel features is less than the minimum power range that satisfies the decoding requirements, then a heavy penalty is applied;

[0081] If the power range of adjacent channel characteristics is greater than or equal to the minimum power range, but the signal-to-interference-plus-noise ratio of the agent is still lower than the preset signal-to-interference-plus-noise ratio threshold after power allocation due to limited total transmit power of the base station, excessive frequency band load, or poor weak user channel conditions, then a standard penalty is applied.

[0082] The expression for the communication quality reward item is as follows:

[0083]

[0084] in, This is a communication quality award item; For intelligent agents accessing frequency band n Signal-to-interference-plus-noise ratio; The preset signal-to-interference-plus-noise ratio (SINR) threshold is used; The power range of adjacent channel characteristics; To meet the minimum power range required for decoding, noise-canceling decoding cannot be completed if the power range is below this value; This indicates an overload state, meaning that the number of agents with a signal-to-interference-plus-noise ratio (SIR) lower than the preset SIR threshold is greater than the preset overload threshold. A reward given for successful communication; This indicates severe punishment; Indicates standard punishment; .

[0085] In this embodiment, two different causes of communication interruption are distinguished in the communication quality reward item: one is the failure of serial interference cancellation decoding due to insufficient power difference between adjacent channel characteristics; the other is communication interruption due to overload. A severe penalty is applied to the former, and a standard penalty is applied to the latter. The agent learns to identify the difference between the severe penalty and the standard penalty, and actively adjusts its movement trajectory and frequency band selection to create sufficient channel gain differences between agents accessing the same frequency band, meeting the decoding requirements for serial interference cancellation. Furthermore, through the standard penalty for overload, the agent perceives the system capacity limit and automatically migrates to a more resource-rich frequency band or geographical area, avoiding systemic communication paralysis caused by local overload.

[0086] (2) Frequency band switching penalty Used to punish intelligent agents for changing the access frequency band.

[0087] Specifically, when agent u performs the frequency band selection action at time t, if the frequency band selection action at time t is different from the frequency band selection action at time t-1, then a fixed negative penalty is imposed on agent u. (like This avoids frequent frequency switching between agents and ensures stable utilization of frequency band resources.

[0088] (3) Coverage of reward items Used to reward agents for moving from a covered area to an uncovered target area.

[0089] Specifically, after all agents execute the movement action at time t, if any agent moves to an uncovered target area, all agents are given a positive reward.

[0090]

[0091] in, The reward for covering the new area is set to 2. Indicates the current grid To the center point of the grid map The distance; This represents the maximum distance from a point within the raster map to the center point.

[0092] Coverage rewards implement an "edge-first" exploration strategy through a distance-weighted mechanism: the agent receives a greater reward for covering uncovered grids farther from the center of the map, thus guiding the cluster to prioritize exploring the edges of the region and avoid ineffective wandering in the central area.

[0093] (4) Collision penalty items Used to punish agents for colliding with geographical obstacles or going beyond the target task area.

[0094] Specifically, when agent u performs a movement action at time t, if it moves into an obstacle area or exceeds the target task area, a collision or boundary violation is determined, and a fixed negative penalty is imposed on agent u. (like This constrains the legal movement of intelligent agents in geospace and ensures the security of task execution.

[0095] In one possible embodiment, the signal-to-interference-plus-noise ratio (SINR) is calculated as follows:

[0096]

[0097] in, The signal-to-interference-plus-noise ratio (SIR) of agent u accessing frequency band n; Agents assigned to access frequency band n by base stations The transmission power; Other intelligent agents allocated access frequency band n by the base station The transmission power; Let be the channel gain from the base station to the agent u in the access frequency band n; The channel gain from the interference source o to the agent u; This represents the set of interference sources operating in frequency band n; The transmission power of the interference source o; This represents the power of additive white Gaussian noise.

[0098] In this embodiment, the above-mentioned signal-to-interference-plus-noise ratio (SINR) calculation formula incorporates three sources simultaneously: co-channel interference, external interference, and additive white Gaussian noise. This comprehensively depicts the real interference environment encountered by the agent when performing power domain multiplexing under the frequency band sharing mechanism. As a result, the communication quality reward calculated based on this formula can accurately distinguish between three states: successful communication, failure to decode serial interference cancellation due to insufficient power level difference, and overload interruption. This provides a reliable physical layer basis for the differentiated penalty mechanism.

[0099] In one possible embodiment, the step of allocating the total transmit power of the base station to each frequency band using a water-filling algorithm includes:

[0100] The maximum value of the channel gain among all agents accessing the same frequency band is taken as the equivalent channel gain of that frequency band; the system capacity is calculated based on the equivalent channel gain of each frequency band; with the goal of maximizing the system capacity, and constrained by the sum of the power allocated to each frequency band not exceeding the total transmit power of the base station and the power allocated to each frequency band not being lower than the threshold power, a Lagrangian function is constructed; based on the Lagrangian function and the KKT conditions, the power allocated to each frequency band is derived.

[0101] In practical implementation, the expression for the Lagrange function is as follows:

[0102]

[0103] Where max represents the maximum value function; B is the bandwidth; The power allocated to frequency band n; Let n be the equivalent channel gain for frequency band n; This represents the power of additive white Gaussian noise. Let n be the total interference power from external interference sources received by frequency band n; N represents the total transmit power of the base station; N is the number of frequency bands connected to the intelligent agent. To ensure that all agents within frequency band n meet the minimum signal-to-interference-plus-noise ratio requirement, a threshold power is set. The power allocated to frequency band n.

[0104] Based on the Karush-Kuhn-Tucker (KKT) conditions, the transmit power of the frequency band is derived. The calculation formula is as follows:

[0105]

[0106] in, The power allocated to frequency band n; N represents the total transmit power of the base station; N is the total number of frequency bands. This represents the power of additive white Gaussian noise. Let n be the equivalent channel gain for frequency band n; Let k be the equivalent channel gain for frequency band k. Let n be the total interference power from external interference sources received by frequency band n; This represents the total interference power from external sources experienced by frequency band k.

[0107] Frequency band transmission power The calculation formula embodies the core idea of ​​the water-filling algorithm: the power allocated to each frequency band is equal to the common water level minus the noise-interference-gain ratio of that frequency band. This mechanism improves system throughput while ensuring minimum communication quality for agents within each frequency band through threshold power constraints.

[0108] In one possible embodiment, the step of power allocation for a set of smart agents accessing the same frequency band includes:

[0109] Calculate the channel gain of each agent in the set of agents accessing the same frequency band; calculate the power allocation coefficient of each agent based on its channel gain; allocate a larger power allocation coefficient to agents with smaller channel gains; multiply the power allocation coefficient of each agent by the power allocated to the same frequency band to obtain the power of each agent.

[0110] In practical implementation, the formula for calculating the power allocation factor is as follows:

[0111]

[0112] in, The power allocation coefficient for the j-th agent in access frequency band n; The channel characteristics of the j-th agent in access frequency band n; The number of agents accessing frequency band n; As the attenuation factor, .

[0113] The power allocated to each agent is as follows:

[0114]

[0115] in, The power allocated to agent n.

[0116] In this embodiment, by adhering to the core principle of allocating a larger power allocation coefficient to agents with smaller channel gain, weak users with poor channel conditions (such as agents far from the base station or located in areas obstructed by obstacles) receive higher transmit power compensation. This overcomes path loss and large-scale fading, enabling their signal-to-interference-plus-noise ratio (SNR) to reach the demodulation threshold. This avoids prolonged communication interruptions for weak users due to insufficient power, significantly improving system fairness and service quality for edge users, and is particularly effective in scenarios with large agent cluster coverage or complex geographical environments. Combined with the water-filling algorithm, this forms a two-level collaborative optimization architecture, achieving highly efficient and fair multi-agent frequency band sharing in scenarios with scarce frequency resources.

[0117] In one possible embodiment, the mean squared error loss function is expressed as follows:

[0118]

[0119] in, Let the mean squared error loss function be used. Represents the network of the j-th critic; The action value output by the j-th commentator network; Let be the global state of all agents at time t; Let E[ ] represent the joint action of all agents at time t; E[ ] represents the expectation operation.

[0120] Let the target value at time t be calculated using the following formula:

[0121]

[0122] in, Let be the composite reward value at time t; Discount factor; The global state of all agents at time t+1; For the joint action of all agents at time t+1; The index for the commentator network; min denotes the minimum value function.

[0123] Traditional deep reinforcement learning methods use the maximum Q-value of the same network as the target, which can easily lead to systematic overestimation bias. In this embodiment, a dual critic network is used, and the minimum value of the two outputs is taken to calculate the target value. Since it is unlikely that both critic networks will overestimate the probability of the same action at the same time, taking the minimum value can effectively offset the overestimation bias, making the Q-value estimation more accurate.

[0124] In one possible embodiment, the counterfactual advantage function is used to evaluate the marginal contribution of the current action performed by the current agent to the global gain when the actions of other agents are fixed; the expression for the counterfactual advantage function is as follows:

[0125]

[0126] in, Represents the counterfactual advantage function; Let the action space of agent u be defined. This represents any action taken by agent u as it traverses the action space. Let t be the hidden state of agent u at time t, as output by the feature extraction network. This represents the action probability distribution output by the policy network corresponding to agent u; This indicates that agent u is in the hidden state. Selecting joint actions The probability distribution; Indicates the actions of other intelligent agents; This indicates that when the actions of other intelligent agents are fixed as At that time, the current action of the current agent u is replaced with .

[0127] After updating the parameters of the two critic networks using gradient descent, the parameters of the feature extraction network and each policy network are updated using the policy gradient theorem based on the counterfactual advantage function. The specific update rules are as follows:

[0128] When the counterfactual advantage function is greater than zero, it indicates that the current action is better than the average level, and the output probability of the current action is increased through gradient updates; when the counterfactual advantage function is less than zero, it indicates that the current action is worse than the average level, and the output probability of the current action is decreased through gradient updates.

[0129] The formula for calculating the policy gradient of the policy network is as follows:

[0130]

[0131] in, Represents the policy gradient; Let u be the objective function of the agent. Represents the parameters of the policy network Calculate the partial derivative to determine the direction of parameter updates. This represents the action probability distribution output by the policy network corresponding to agent u; This indicates that agent u is in the hidden state. Selecting joint actions The probability distribution of is calculated using the following formula:

[0132]

[0133] in, Let t be the hidden state output by the LSTM network at time t; Let be the weight matrix of the policy network for agent u. Let be the bias vector of the policy network corresponding to agent u; This is the normalization function.

[0134] Considering that in multi-agent systems, the global reward of a team is difficult to directly attribute to the specific actions of a single agent, this application introduces a counterfactual advantage function. By fixing the actions of other agents and counterfactually changing the assumption of the current agent's action, the marginal contribution of each agent's current action to the global reward is accurately evaluated. This effectively alleviates the reputation allocation problem in multi-agent collaboration and improves the efficiency and final performance of multi-agent collaborative learning.

[0135] Please refer to Figure 2 This is a network workflow diagram for the training and execution phases provided in this application embodiment. During the training phase, the feature extraction network and policy network are updated using a dual critic network. During the execution phase, each agent inputs its local observation vector at the current time into the trained feature extraction network to obtain the agent's hidden state, and then inputs the hidden state into the corresponding trained policy network to output the agent's movement action and frequency band selection action at the current time.

[0136] In summary, this application provides a multi-agent coverage path planning method based on frequency band sharing. First, channel gain is calculated, and a water-filling algorithm is used for global power optimization allocation between frequency bands. Then, fractional-order power allocation is further implemented for multiple agents accessing the same frequency band to achieve non-orthogonal multiplexing of the power domain, mitigating co-channel interference conflicts among multiple agents and providing a broader and more stable communication boundary for coverage path selection. Next, the signal-to-interference-plus-noise ratio is calculated based on the allocated power to determine communication success. A composite reward value is calculated based on a composite function, and a counterfactual baseline is introduced for individual reputation allocation. Finally, gradient updates of parameters are performed through a dual-commenter network and a policy network. This process deeply embeds the power domain multiplexing mechanism into the feedback loop of reinforcement learning, guiding the agent cluster to autonomously learn and achieve coverage path planning in scenarios with high-density agent clusters and limited spectrum resources, ensuring the logical integrity and execution continuity of the coverage task.

[0137] Based on the same inventive concept, this application also provides a computer device, which includes a processor, a memory, and a computer program stored in the memory. The computer program is executed by the processor to implement the aforementioned multi-agent coverage path planning method based on frequency band sharing.

[0138] Based on the same inventive concept, this application also provides a computer storage medium storing a computer program, which is executed by a processor to implement the aforementioned multi-agent coverage path planning method based on frequency band sharing.

[0139] In some embodiments, the computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; or it may be a device including one or any combination of the above-mentioned memories. The computer may be a variety of computing devices, including smart terminals and servers.

[0140] In some embodiments, executable instructions may take the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a standalone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.

[0141] As an example, executable instructions may, but do not necessarily, correspond to files in the file system. They may be stored as part of a file that holds other programs or data, for example, in one or more scripts in a Hyper Text Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple collaborative files (e.g., a file that stores one or more modules, subroutines, or code sections).

[0142] As an example, executable instructions can be deployed to execute on a single computing device, or on multiple computing devices located in one location, or on multiple computing devices distributed across multiple locations and interconnected via a communication network.

[0143] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or system that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or system. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or system that includes that element.

[0144] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0145] The above specific embodiments further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A multi-agent coverage path planning method based on frequency band sharing, characterized in that, include: Obtain the local observation vectors of all agents within the target task area at the same time. A multi-agent decision network is trained using reinforcement learning algorithms based on the local observation vectors of each agent at the same time. This multi-agent decision network includes a feature extraction network and a policy network for each agent. The training process employs a composite reward function that includes a communication quality reward term and a frequency band switching penalty term. The communication quality reward term is used to reward agents who successfully communicate after performing power domain reuse under a frequency band sharing mechanism, and to impose differentiated penalties based on the cause of communication interruption. The frequency band switching penalty term penalizes agents for changing their access frequency band. Deploy the trained feature extraction network and the corresponding trained policy network on each agent; Each agent inputs its local observation vector at the current moment into the trained temporal feature extraction network to obtain the agent's hidden state, and inputs the hidden state into the corresponding trained policy network to output the agent's joint action at the next moment; the joint action includes a movement action and a frequency band selection action.

2. The multi-agent coverage path planning method based on frequency band sharing according to claim 1, characterized in that, The multi-agent decision-making network further includes two critic networks with identical structures but different initialization parameters; the training of the multi-agent decision-making network using a reinforcement learning algorithm based on the local observation vectors of each agent at the same time includes: The local observation vectors of all agents at time t are input into the feature extraction network to obtain the hidden state of each agent at time t. The hidden state of each agent at time t is then input into its corresponding policy network to obtain the joint action of each agent at time t. Based on the frequency band selection actions of each agent at time t, determine the set of agents that can access each frequency band; The water-filling algorithm is used to allocate the total transmit power of the base station to each frequency band, and power allocation is performed on the intelligent agent set accessing the same frequency band; Each agent performs power domain multiplexing in non-orthogonal multiple access technology according to the allocated power, and calculates the composite reward value at time t according to the composite reward function; Based on the composite reward value and the action values ​​output by the two commentator networks at adjacent time points, a mean squared error loss function is constructed, and the parameters of the two commentator networks are updated respectively through gradient descent.

3. The multi-agent coverage path planning method based on frequency band sharing according to claim 2, characterized in that, The method of allocating the total transmit power of the base station to each frequency band using a water-filling algorithm includes: The maximum value among the channel gains of all agents accessing the same frequency band is taken as the equivalent channel gain of that frequency band. Calculate the system capacity based on the equivalent channel gain of each frequency band; With the goal of maximizing system capacity, and constrained by the fact that the sum of the power allocated to each frequency band does not exceed the total transmit power of the base station, and the power allocated to each frequency band is not lower than the threshold power, a Lagrangian function is constructed. Based on the Lagrange function and KKT conditions, the power allocated to each frequency band is derived, and the specific calculation formula is as follows: ; in, The power allocated to frequency band n; N represents the total transmit power of the base station; N is the total number of frequency bands. This represents the power of additive white Gaussian noise. Let n be the equivalent channel gain for frequency band n; Let k be the equivalent channel gain for frequency band k. Let n be the total interference power from external interference sources received by frequency band n; This represents the total interference power from external sources experienced by frequency band k.

4. The multi-agent coverage path planning method based on frequency band sharing according to claim 2, characterized in that, The power allocation for a set of intelligent agents accessing the same frequency band includes: Calculate the channel gain of each agent in a set of agents accessing the same frequency band; Calculate the power allocation coefficient for each agent based on the channel gain of each agent; agents with smaller channel gains are allocated larger power allocation coefficients. The power of each agent is obtained by multiplying its power allocation coefficient by the power allocated to the same frequency band.

5. The multi-agent coverage path planning method based on frequency band sharing according to claim 4, characterized in that, The formula for calculating the power allocation coefficient is as follows: ; in, The power allocation coefficient for the j-th agent in access frequency band n; The channel characteristics of the j-th agent in access frequency band n; The number of agents accessing frequency band n; As the attenuation factor, .

6. The multi-agent coverage path planning method based on frequency band sharing according to claim 2, characterized in that, After updating the parameters of the two critic networks respectively using gradient descent, the method further includes: Based on the counterfactual advantage function, the parameters of the feature extraction network and each policy network are updated respectively by the policy gradient theorem; the counterfactual advantage function is used to evaluate the marginal contribution of the current action performed by the current agent to the global benefit when the actions of other agents are fixed. When the value of the counterfactual advantage function is greater than zero, the output probability of the current action is increased through gradient update; When the value of the counterfactual advantage function is less than zero, the output probability of the current action is reduced through gradient updates.

7. The multi-agent coverage path planning method based on frequency band sharing according to claim 6, characterized in that, The expression for the counterfactual advantage function is as follows: ; in, Represents the counterfactual advantage function; Let the action space of agent u be defined. This represents any action taken by agent u as it traverses the action space. The hidden state of agent u at time t, as output by the feature extraction network; This represents the action probability distribution output by the policy network corresponding to agent u; This indicates that agent u is in the hidden state. Selecting joint actions The probability of; Indicates the actions of other intelligent agents; This indicates that when the actions of other intelligent agents are fixed as At that time, the current action of the current agent u is replaced with .

8. The multi-agent coverage path planning method based on frequency band sharing according to claim 1, characterized in that, If communication is interrupted, differentiated penalties will be applied based on the cause of the interruption, including: When the signal-to-interference-plus-noise ratio (SIR) of an agent after power domain multiplexing under the frequency band sharing mechanism is lower than the preset SIR threshold, the agent's communication is determined to be interrupted, and the following operations are performed based on the cause of the interruption: If the power range of adjacent channel features is less than the minimum power range that satisfies the decoding requirements, then a heavy penalty is applied; If the power range of adjacent channel features is greater than or equal to the minimum power range, but the signal-to-interference-plus-noise ratio (SIR) of the agent is still lower than the preset SIR threshold after power allocation, then a standard penalty is applied. Wherein, the absolute value of the standard penalty is less than the absolute value of the severe penalty.

9. A multi-agent coverage path planning method based on frequency band sharing according to claim 8, characterized in that, The formula for calculating the signal-to-interference-plus-noise ratio is as follows: ; in, The signal-to-interference-plus-noise ratio (SIR) of agent u accessing frequency band n; Agents assigned to access frequency band n by base stations The transmission power; Other intelligent agents allocated access frequency band n by the base station The transmission power; Let be the channel gain from the base station to the agent u in the access frequency band n; The channel gain from the interference source o to the agent u; This represents the set of interference sources operating in frequency band n; The transmission power of the interference source o; This represents the power of additive white Gaussian noise.

10. A multi-agent coverage path planning method based on frequency band sharing according to claim 1, characterized in that, The composite reward function further includes a coverage reward and a collision penalty; the coverage reward is used to reward the agent for moving from a covered area to an uncovered target area; the collision penalty is used to penalize the agent for colliding with geographical obstacles or going beyond the target task area.