A Distributed Location Method and System for Ultra-Shortwave Radiation Sources
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-29
- Publication Date
- 2026-08-14
AI Technical Summary
然而,现有多智能体强化学习算法虽然在多智能体协同方面取得一定进展,但在异构无人机群的电磁频谱监测定位中仍面临计算复杂度高、通信开销大、难以平衡异构智能体差异化目标等问题,导致其可扩展性和效率受限,难以满足实际应用中对定位精度和效率的严苛要求
(1)本发明将无人机群明确划分为功能异构的控制节点和监测节点,并通过强化学习分别为其配置差异化的决策策略,实现了任务的合理分工,使控制节点能专注于集群管理与通信优化,而监测节点能专注于信号感知,避免了同构系统下决策目标冲突导致的效率内耗,实现了对超短波辐射源的高精度定位和大范围监测,显著提升了定位成功率并降低了定位误差。不仅有效解决了单一无人机平台在续航、载荷和性能上的局限,以及多无人机系统在复杂任务中存在的决策混乱、通信拥堵和可扩展性差的问题;而且,即使在无人机载荷、续航和通信能力受限条件下,通过异构集群的智能协同,也能实现对超短波辐射源的高精度、高效率分布式定位。
Smart Images

Figure CN121741628B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of ultra-shortwave radiation source localization technology, and in particular to a distributed localization method and system for ultra-shortwave radiation sources based on heterogeneous multi-agent deep learning. Background Technology
[0002] Electromagnetic spectrum resources are becoming increasingly prevalent in daily applications. With the continuous increase in frequency-using devices, the electromagnetic environment is becoming increasingly complex, making the monitoring of interference sources in the electromagnetic environment particularly important and a key task in maintaining order in the electromagnetic space. Ultra-short waves (UHF), as one of the commonly used frequency bands, are widely used in broadcasting, communications, aviation navigation, and emergency command. Their signals have characteristics such as long propagation distance, strong diffraction ability, and relatively low susceptibility to terrain and building obstruction. However, these characteristics also make UHF radiation sources prone to widespread interference in complex electromagnetic environments, making rapid location difficult. Especially in complex geographical environments such as cities and mountainous areas, traditional ground-based monitoring methods are limited by field of view and deployment location, making it difficult to achieve efficient and accurate location of radiation sources.
[0003] In recent years, with the rapid development of unmanned aerial vehicle (UAV) technology, its advantages in mobility, payload diversity, and deployment flexibility have become increasingly prominent. UAVs can carry spectrum monitoring equipment to achieve aerial mobile monitoring, effectively compensating for the shortcomings of ground monitoring stations in terms of spatial coverage and mobility. In particular, UAVs can quickly reach areas inaccessible to personnel, obtaining a superior monitoring perspective from the air and significantly improving the detection probability and positioning accuracy of ultra-shortwave radiation sources. Furthermore, multi-UAV collaborative operations can further expand the monitoring range, achieving multi-angle and multi-level spectrum perception of target areas, thereby enabling rapid identification and continuous tracking of interference sources in complex electromagnetic environments. However, single UAV platforms still have limitations in terms of endurance, payload capacity, and communication distance, making it difficult to independently complete large-scale, long-term monitoring tasks. Therefore, how to utilize UAV swarms for collaborative positioning of ultra-shortwave radiation sources remains a challenge for electromagnetic monitoring in complex electromagnetic environments.
[0004] Currently, electromagnetic spectrum monitoring of heterogeneous UAVs mainly relies on clustering algorithms, along with heuristic and reinforcement learning algorithms for perception and detection. Compared to traditional clustering and heuristic algorithms, data-driven methods such as reinforcement learning are widely used due to their stronger generalization ability and online adaptability. However, while existing multi-agent reinforcement learning algorithms have made some progress in multi-agent collaboration, they still face challenges in electromagnetic spectrum monitoring and positioning of heterogeneous UAV swarms, including high computational complexity, large communication overhead, and difficulty in balancing the differentiated objectives of heterogeneous agents. This limits their scalability and efficiency, making it difficult to meet the stringent requirements for positioning accuracy and efficiency in practical applications.
[0005] Therefore, there is an urgent need for a distributed positioning method and system for ultra-shortwave radiation sources that can achieve rapid and high-precision distributed positioning of ultra-shortwave radiation sources in complex and dynamic electromagnetic environments. Summary of the Invention
[0006] To address the aforementioned problems, this invention provides a distributed positioning method and system for ultra-shortwave radiation sources, which solves the problem of how to achieve high-precision and high-efficiency distributed positioning of ultra-shortwave radiation sources through intelligent collaboration of heterogeneous clusters under conditions of limited UAV payload, endurance, and communication capabilities.
[0007] On one hand, the present invention provides a distributed location method for ultra-shortwave radiation sources, the method comprising: Construct a heterogeneous drone swarm that includes control nodes and monitoring nodes; Based on a multi-agent reinforcement learning algorithm, differentiated decision-making strategies are configured for the control node and the monitoring node respectively, so as to guide the heterogeneous UAV swarm to conduct cooperative search within the mission area; By synchronously collecting signal characteristic parameters of the radiation source using multiple monitoring nodes, aggregating the signal characteristic parameters from multiple spatially distributed monitoring nodes, and calculating the location result of the radiation source.
[0008] On the other hand, the present invention also provides a distributed positioning system for ultra-shortwave radiation sources, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of any of the methods described above.
[0009] In summary, this invention provides a distributed positioning method and system for ultra-shortwave radiation sources, which achieves the following beneficial effects compared with existing technologies: (1) This invention clearly divides the UAV swarm into functionally heterogeneous control nodes and monitoring nodes, and configures differentiated decision-making strategies for them through reinforcement learning, thereby achieving a reasonable division of tasks. This allows control nodes to focus on swarm management and communication optimization, while monitoring nodes can focus on signal perception, avoiding the efficiency friction caused by conflicting decision-making objectives in homogeneous systems. This enables high-precision positioning and large-scale monitoring of ultra-shortwave radiation sources, significantly improving the positioning success rate and reducing positioning errors. It not only effectively solves the limitations of a single UAV platform in terms of endurance, payload, and performance, as well as the problems of decision confusion, communication congestion, and poor scalability in complex tasks of multi-UAV systems; but also, even under conditions where UAV payload, endurance, and communication capabilities are limited, high-precision and high-efficiency distributed positioning of ultra-shortwave radiation sources can be achieved through intelligent collaboration of heterogeneous swarms.
[0010] (2) The present invention adopts the minimum gap positioning algorithm, which uses the spatial distribution of multiple monitoring nodes to accurately locate the radiation source, and performs joint optimization estimation of the difference in received signal strength and the angle of arrival of the signal; effectively overcomes the disadvantage of the single parameter positioning method being susceptible to environmental interference, and can still achieve accurate positioning even in high noise environments; realizes the autonomous coordination and intelligent decision-making of UAV swarms in dynamic environments, and effectively improves the positioning capability and monitoring efficiency of ultra-shortwave radiation sources.
[0011] (3) The present invention establishes a star communication topology and cluster information aggregation processing, which simplifies the fully interconnected or complex routing communication between monitoring nodes into point-to-point communication with control nodes, further reducing the uplink burden with ground base stations, which can significantly reduce the amount of communication interaction and improve system efficiency.
[0012] (4) The present invention makes decisions and dynamically adjusts deployment based on a multi-agent reinforcement learning model. It can adjust its strategy online through reinforcement learning based on real-time positioning results, signal quality changes and even its own energy state, so as to achieve continuous tracking and positioning of mobile radiation sources and give UAV swarms the ability to make autonomous decisions and continuously optimize in dynamic and complex environments. Attached Figure Description
[0013] To more clearly illustrate the technical solutions in this invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. The accompanying drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0014] Figure 1 This is a schematic diagram of the method steps of a distributed positioning method and system for ultra-shortwave radiation sources provided by the present invention; Figure 2 This is a schematic diagram of a heterogeneous UAV swarm, which is a distributed positioning method and system for ultra-shortwave radiation sources provided by the present invention. Figure 3 This is a schematic diagram illustrating the positional relationship between the control node and the monitoring node in a distributed positioning method and system for ultra-shortwave radiation sources provided by this invention, indicating the establishment of intra-cluster association links. Figure 4 This is a schematic diagram of the HMUDRL structural framework of a distributed positioning method and system for ultra-shortwave radiation sources provided by the present invention. Figure 5 This is a schematic diagram of the CH and CM positioning process trajectory of a distributed positioning method and system for ultra-shortwave radiation sources provided by this invention. Figure 1 ; Figure 6 This is a schematic diagram of the CH and CM positioning process trajectory of a distributed positioning method and system for ultra-shortwave radiation sources provided by this invention. Figure 2 ; Figure 7 This is a schematic diagram of the CH and CM positioning process trajectory of a distributed positioning method and system for ultra-shortwave radiation sources provided by this invention. Figure 3 ; Figure 8 This is a schematic diagram of the CH and CM positioning process trajectory of a distributed positioning method and system for ultra-shortwave radiation sources provided by this invention. Figure 4 ; Figure 9 This is a schematic diagram comparing the total reward moving average variation curves of the HMUDRL of the present invention with those of four baseline algorithms, according to the distributed positioning method and system for ultra-shortwave radiation sources provided by the present invention. Figure 10 This is a box plot comparison diagram of the HMUDRL of the present invention and four baseline algorithms for a distributed positioning method and system for ultra-shortwave radiation sources provided by the present invention. Figure 11 This is a schematic diagram comparing the average number of information interactions within different structures of a UAV swarm, based on the distributed positioning method and system for ultra-shortwave radiation sources provided by this invention. Detailed Implementation
[0015] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings and embodiments. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.
[0016] It should be noted that, in the description of the embodiments of the present invention, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a method, step, or apparatus that includes a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to the method, step, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of additional identical elements in the method, step, or apparatus that includes the element.
[0017] To address the challenge of achieving high-precision, high-efficiency distributed localization of ultra-shortwave radiation sources through intelligent collaboration of heterogeneous swarms under conditions of limited payload, endurance, and communication capabilities for unmanned aerial vehicles (UAVs), this invention provides a distributed localization method and system for ultra-shortwave radiation sources. In complex electromagnetic environments, this method utilizes a UAV swarm for efficient and high-precision localization of ultra-shortwave radiation sources. The core of this system is the construction of a heterogeneous multi-UAV deep reinforcement learning (HMUDRL) model, driving a collaborative system composed of two types of UAV nodes with different functions. On one hand, by establishing a large-scale distributed heterogeneous UAV swarm collaborative perception and localization architecture, the efficiency of UAV swarm collaborative localization is improved. On the other hand, by integrating a branch network architecture through a multi-agent reinforcement learning model, computational and information transmission efficiency is improved. Simultaneously, collaborative calculations using the signal power strength difference and signal angle of arrival of monitoring nodes enhance localization accuracy. Furthermore, a reward mechanism effectively balances monitoring coverage, localization accuracy, and monitoring efficiency, achieving high-precision localization and large-scale monitoring of ultra-shortwave radiation sources, significantly improving the localization success rate and reducing localization errors.
[0018] On the one hand, such as Figure 1 As shown, the method specifically includes: S100: Construct a heterogeneous drone swarm that includes control nodes and monitoring nodes.
[0019] The heterogeneous UAV swarm constructed in this invention includes at least one swarm control node and multiple swarm monitoring nodes. The swarm control node (CH) acts as the brain, responsible for swarm management, information aggregation, and communication with ground base stations; the swarm monitoring nodes (CMs) act as sensors, carrying highly sensitive monitoring equipment, and are responsible for signal perception. These nodes automatically establish and maintain a hierarchical star-shaped communication network by periodically broadcasting Hello messages, in which each CM is associated with a CH, forming multiple cooperative sub-swarms.
[0020] It should be noted that control nodes and monitoring nodes can be any mobile or fixed platform with mobility, communication capabilities, and / or signal monitoring capabilities. Specifically, nodes can be one or more combinations of drones, tethered balloons, manned aircraft, ground mobile vehicles, fixed ground stations, and surface vessels. The key to building a heterogeneous cluster lies in achieving collaboration between nodes with different functions (control and monitoring), rather than being limited to a specific platform type.
[0021] As an example, a heterogeneous UAV swarm including control nodes and monitoring nodes is constructed, specifically including: setting UAVs with data processing and communication capabilities as control nodes, serving as cluster heads; setting UAVs with monitoring capabilities as monitoring nodes, serving as cluster members; and establishing a star communication topology centered on the control nodes, wherein each monitoring node is associated with only one control node.
[0022] For example, such as Figure 2 As shown, a heterogeneous UAV swarm is established to perform spectrum monitoring tasks. To improve the monitoring efficiency of the UAVs, they are divided into several sub-swarms, including a central control base station, multiple control UAVs (CHs), and multiple monitoring UAVs (CMs). The central control base station establishes a stable data control link with the CHs for overall control and data integration of the heterogeneous UAV swarm. The CH, as the cluster head, has strong data processing and communication capabilities, managing the CMs within its sub-swarm, controlling the CMs within its sub-swarm, and aggregating the monitoring data of the CMs within the sub-swarm and uploading it to the central control base station. The CMs, as cluster members, are equipped with high-sensitivity monitoring equipment and have strong monitoring capabilities, used for signal sensing and direction finding of radiation sources within the mission area. Through collaborative monitoring by multiple CMs, a larger monitoring coverage area is achieved, and the data is transmitted back to the corresponding CH.
[0023] This heterogeneous drone swarm solves the limitations of drones' own payload weight, type, and power supply duration, as well as the limitations of a single drone platform in simultaneously achieving long working hours, large monitoring range, and good transmission performance.
[0024] However, due to the overall movement of the drone swarm, the relative positions of the CMs may change dynamically, which poses a significant challenge to drone swarm clustering. Therefore, preferably, during the monitoring process, the CMs are also used for autonomous clustering, and the CH can adaptively select the CMs and find the optimal positions.
[0025] As one example, the control node and the monitoring node establish and maintain intra-cluster association links by periodically broadcasting Hello messages. The Hello message contains the sending node identifier, association status, transmission carrier frequency, and signal power information.
[0026] In other words, within the monitoring task area, multiple CHs and multiple CMs move independently within the task area, and each CM is associated with only one CH. To prevent a high probability of collisions in areas with high CM density, the number of CMs in each subgroup is limited to a certain limit. If the number of CMs associated with CH is less than This indicates that the number of CMs belonging to a CH is not saturated, and the CH and CMs communicate in a dedicated intra-cluster association channel at fixed intervals. Broadcast Hello messages; when a CM is associated with a CH, it directly returns a Hello message; when a CM does not belong to any subgroup, it will send Hello messages to the surrounding CHs to request association in order to actively join one of the CHs. Only unsaturated CHs can receive new association requests.
[0027] Based on the received Hello message, CH and CM can establish and update the association list; if the first... The CH and the first Each CM is in each other's association list, and each is in... If a Hello message is received within the period, it indicates that a link has been established between CH and CM; if they are in a state of mutual inertia... If no Hello message is received within the period, it indicates that the link between CH and CM has broken.
[0028] It should be noted that, although the larger It can adapt well to temporary link interruptions caused by the loss of Hello packets, thus stabilizing the cluster structure; however, it has a larger... This will delay the network's response to topology changes caused by drone movement. Therefore, The value depends on specific factors such as the number of drones performing the mission and network setup requirements.
[0029] Based on the transmission carrier frequency and signal power information in the Hello message, the relative positions between the CH and CM that establish the association link can be estimated by utilizing Doppler frequency shift and signal power changes. By adjusting the position of the CH, the cluster structure can be kept stable.
[0030] like Figure 3 As shown, if the shaded area centered at CH has a coverage radius of... The broadcast area, when CM1 is far from CH Less than When CM1 receives the broadcast Hello message sent by CH, it establishes a link with CH; when CM2 is far from CH... Greater than At this time, CM2 cannot receive broadcast messages sent by CH and cannot establish a link with CH. When CH and CM are in the same task area... If the area of the task region conforms to a random independent distribution, then the task region is... The number of CH is The number of CMs is Considering spatial reachability, as an example, the probability that CM falls within the CH coverage circle is: .
[0031] If each CH can be associated at most There are currently CMs. There are CM associations, and the average load per CH is ;when When CH is associated with CM, it indicates that CH is in an unsaturated state. As an example, when CH is associated with CM, the unsaturation probability of CH is: .
[0032] In addition, continuous A cluster association link can only be established if Hello messages are received in every cycle. Therefore, as an example, continuous... The probability of receiving a Hello message in every cycle is: ; in, This represents the probability that the CH successfully receives a Hello message in each cycle. It depends on channel quality, the Doppler effect, and power attenuation. Specifically, It is a distance function, which is: ;in, For reliable communication radius (slightly less than) ), This is the path loss index.
[0033] It should be noted that for a CM to successfully join a CH, the following conditions must be met: at least one CH must exist within the CM's communication range, and the CHs must be in an unsaturated state and consecutive. Only when a CM receives a Hello message in every cycle can it choose to join the nearest unsaturated CH or the CH with the strongest signal. The probability of successfully establishing an intra-cluster association link is then: .
[0034] As an example, within the task area, the probability that a CM and a CH successfully establish a stable intra-cluster association link is: ; in, This represents the probability of successfully receiving a Hello message in a single cycle within the coverage area. The number of consecutive successful cycles required to establish intra-cluster association links.
[0035] To further describe the kinematic model of the UAV node, this invention explores a method for locating electromagnetic spectrum radiation sources by examining changes in velocity and direction. Specifically, it first establishes a position formula, and then provides a set of velocities and directions to update the UAV's position information.
[0036] For ease of calculation and analysis, the UAV's motion is represented on a two-dimensional plane, and the UAV's position and velocity state can be expressed as: ; in, Indicates drone At any moment Position and velocity state, Indicates drone At any moment coordinates Indicates drone At any moment Velocity components along the coordinate axes.
[0037] Acceleration of drones in motion It can be: ;in, Indicates at time drones In a real system, the change in directional velocity is controlled by the flight control system at fixed time steps. Update; Furthermore, the updated state of the drone's position and velocity can be represented as: .
[0038] It should be noted that in order to ensure that the drone flies within a controllable and safe range and avoids going out of control and affecting the safety of other aircraft, each drone includes speed constraints, acceleration constraints, and position constraints.
[0039] The speed constraint is: ; in, For the combined speed of the drone, and These are the minimum (usually 0) and maximum flight speeds of the drone, respectively.
[0040] The acceleration constraint is: ; in, This represents the maximum acceleration of the drone.
[0041] The positional constraints are: .
[0042] Within the monitoring mission area, the CM (Center for Monitoring and Sensing) performs the monitoring and sensing of the VHF radiation source. The air-to-ground channel gain follows the Nakagami-m distributed fading environment model; therefore, the small-scale fading amplitude between the CM and the radiation source is relatively small. The probability density function (PDF) can be expressed as: ; in, For small-scale fading amplitude The value; The average power of the fading signal, i.e. , Expressing expectations; For gamma function, ; This is a shape parameter used to quantify the severity of small-scale fading.
[0043] For ultra-shortwave radiation sources, monitoring and sensing can be conducted via CM (Comprehensive Detection and Sensing), employing a logarithmic distance path loss model from the radiation source. Path loss to CM for: ; in, As a radiation source The Euclidean distance between CM and CM; For reference distance Path loss (dB) at the location. , , where is the wavelength. The speed of light; This is the path loss index. The standard deviation is A zero-mean Gaussian random variable.
[0044] When multiple monitoring centers (CMs) are simultaneously monitoring a radiation source, the precise location of the CMs and their ability to receive radiation sources can be utilized. Determining the radiation source by the difference in signal power intensity Position coordinates and path loss index That is, for the first Each CM receives a signal power of: ; in, As a radiation source Transmission power; As a radiation source Antenna gain; For the first The receiving antenna gain per CM; As a radiation source To the Path loss per CM.
[0045] When the first CM is selected as the reference UAV, the calculation of the first... The signal strength difference between the first CM and the first CM for: .
[0046] based on and The perceptual model is obtained, namely: .
[0047] Furthermore, if the receiving antenna gain of CM is the same, then: .
[0048] For radiation sources The location requires monitoring by multiple CMs, the first one... Each CM at any time To the radiation source The distance through the first Measured signal strength of radiation source in CM That is to say, according to The exponential relationship between CM and distance allows us to obtain the relationship between CM and the radiation source. The distance. Specifically, the distance formula is: Furthermore, due to the limited monitoring range of CM, measurements are only taken within... Valid for a period of time. This represents the maximum monitoring coverage area of CM.
[0049] As an example, in The set of valid CM measurements at any given time is: ;in, A threshold for effectively monitoring signal strength.
[0050] It should be noted that the three-point location method requires at least three spatially distributed CMs to determine the radiation source. The position, therefore, only when That is, at least three CMs simultaneously obtain effective [conditions / conditions]. Only then can successful monitoring and positioning be achieved.
[0051] CM obtains radiation sources through a direction-finding system. Angle of arrival of signal By combining its own location information, a nonlinear observation equation is constructed to obtain the radiation source. coordinates Using the Cramer-Rao lower bound (CRLB) as a theoretical performance benchmark, it describes the performance of any unbiased estimator for the location of a radiation source given an observation model and noise statistics. The smaller the estimated error covariance matrix trace can reach, the smaller the CRLB value indicates that the distribution configuration of CM is more conducive to high-precision positioning.
[0052] The communication path loss of the intra-cluster association link between CH and CM can be expressed as: ;in, Let CH be the distance between CH and CM.
[0053] The SNR of the CH signal received by CM is: ;in, The transmit power of CH; Channel gain; The noise power spectral density; This refers to the communication bandwidth.
[0054] According to Shannon's theorem, the maximum error-free communication rate between CH and CM is... for: .
[0055] The total throughput of all CMs belonging to the CH subgroup is: .
[0056] It should be noted that if the heterogeneous drone swarm includes CH and Each CM divides the drone into Each subgroup contains one CH and multiple CMs, which can be represented by sets. Indicates the first All nodes within a subgroup, of which Then the set for: ; Indicates the first The number of CMs within a subgroup.
[0057] Intra-group communication delay It can be represented as: ; in, The signal transmission delay depends on the distance between the CH and its associated CM. To handle latency, including signal encoding, decoding, and other processing time; Queuing delay depends on network load and queue length; This represents the average number of retransmissions. , Indicates that the given SNR is The probability of a data packet being successfully transmitted under certain conditions can be derived from the bit error rate (BRE). The single transmission delay depends on the data packet size and communication bandwidth.
[0058] Link status between CH and its associated CM It can be represented as: .
[0059] As an example, the total communication latency of a heterogeneous drone swarm This is the weighted average of the communication delays across all subgroups, which is: ; in, For the first The communication weight between CH and CM within each subgroup.
[0060] S200: Based on a multi-agent reinforcement learning model, it configures differentiated decision-making strategies for control nodes and monitoring nodes respectively to guide heterogeneous UAV swarms to conduct collaborative searches within the mission area.
[0061] It should be noted that in order to solve multi-objective nonlinear optimization problems, commonly used methods in recent years include nonlinear scalarization, evolutionary constraint methods, individual adaptive evolutionary algorithms, and interactive evolutionary algorithms. These methods have good optimization effects when dealing with specific problems, but when the environment changes dynamically or the problem dimension increases, these methods need to be searched again and cannot quickly solve the problem by utilizing previous experience. Their ability to generalize and transfer knowledge is not as good as that of reinforcement learning algorithms.
[0062] Therefore, in order to achieve efficient and high-precision positioning of radiation sources through the collaborative work of heterogeneous UAV swarms, this invention proposes a distributed heterogeneous multi-agent deep reinforcement learning algorithm (HMUDRL), which is applied to heterogeneous UAV swarms. Through the autonomous decision-making of CH and CM, the monitoring and positioning of ultra-shortwave radiation sources can be achieved, which can not only reduce the complexity of the structure, but also improve the monitoring efficiency.
[0063] In other words, after establishing a heterogeneous UAV swarm, it is not only necessary to optimize the positioning accuracy, positioning speed, and number of information exchanges between each CH, but also to establish a distributed reinforcement learning model to achieve independent decision-making while cooperating in positioning among CMs, thus solving the problem of detecting and locating ultra-shortwave radiation sources.
[0064] Specifically, driven by the HMUDRL algorithm, the CH and CM begin to search collaboratively: on the one hand, each CM continuously scans the VHF band during its movement, measuring and recording the received signal strength (RSS) and angle of arrival (AOA) from the radiation source; on the other hand, the CH dynamically adjusts its own position based on the analysis of its policy network, and indirectly guides its CM to move towards areas with stronger signals, thereby narrowing the search range.
[0065] It's important to note that HMUDRL employs a two-layer agent architecture, proposing two separate algorithms for the different roles of the CH (Challenger) and CM (Center) to serve CHs and CMs respectively. While sharing the same environment, each has its own independent observation space, action space, and reward function. The agent's actions can be structured using an actor-critic network: the actor network selects actions based on the current state (i.e., executes the policy), while the critic network predicts future outcomes based on the observed state, evaluating the quality of the actions. This two-layer agent architecture avoids the performance bottlenecks caused by traditional multi-agent reinforcement learning using a single policy to handle heterogeneous agents.
[0066] As one embodiment, differentiated decision-making strategies are configured for control nodes and monitoring nodes respectively, specifically including: configuring a first strategy network for control nodes, whose observation space includes at least one of the following: their own coordinate position, position offset from the cluster centroid, distance to the estimated radiation source, spatial dispersion of their respective monitoring nodes, and average position offset between control nodes; configuring a second strategy network for monitoring nodes, whose observation space includes at least one of the following: their own coordinate position, distance from their respective control nodes, received signal strength, change in received signal strength, and angle of arrival of the signal.
[0067] In other words, the first strategy network serves the CHs (Chains of Communications), observing global information and deciding how to move to optimize cluster structure and maintain high-quality communication links. The second strategy network serves the CMs (Communication Centers), observing local information and deciding how to move to better detect and approach radiation sources, thereby optimizing signal sensing quality and positioning accuracy. The reward function of the first strategy network includes at least one of the following: location confidence reward, coverage quality reward, and exploration diversity reward. The reward function of the second strategy network includes at least one of the following: signal strength gain reward, distance preservation reward, and effective signal angle of arrival reward.
[0068] Preferably, the first policy network is trained using the HMUPPO algorithm, and its reward function is a weighted value of location confidence reward, coverage quality reward, and exploration diversity reward; the second policy network is trained using the DDMPPO algorithm, and its reward function is a weighted value of signal strength gain reward, distance preservation reward, and signal angle of arrival effectiveness reward.
[0069] like Figure 4As shown, HMUDRL is a distributed framework comprising the HMUPPO and DDMPPO algorithms. The HMUPPO algorithm controls the cluster selection and group movement of the CHs (Chains Control Centers), while the DDMPPO algorithm guides the CMs (Machine Control Centers) to complete monitoring and search tasks. The CHs, employing the HMUPPO algorithm, can share global information and leverage their computational and information processing capabilities. They can communicate with the control center and share the critic network. The CMs, using the DDMPPO algorithm, focus on monitoring and locating VHF / UHF radiation sources. Through heterogeneous architecture and collaborative interaction with the environment, they provide stable policy updates when dealing with uncertainties in radiation source location and complex multi-UAV coordination. In actual execution, each CM only needs to communicate with its assigned CH, performing monitoring tasks independently and without needing to communicate with other CMs, thus saving computational and communication resource overhead. This distributed architecture allows both types of nodes to be trained independently, enhancing the scalability and robustness of HMUDRL.
[0070] In the scenario of searching for ultra-shortwave radiation sources by heterogeneous UAV swarms, the goal is to find an optimal strategy through a distributed architecture to guide communication centers (CMs) in effectively locating the radiation source. To address the complex challenges faced in ultra-shortwave radiation source search, this invention proposes a variant of the near-end optimization strategy algorithm, the HMUPPO algorithm, which is designed specifically for the control link of communication centers (CHs). During the movement of the UAV swarm, the CHs autonomously associate with the CMs, seeking the most advantageous transmission and control positions, maintaining high-quality communication links, low communication latency, and sufficient load.
[0071] like Figure 3 As shown, CHs periodically send Hello messages to attempt to establish association links with nearby CMs. CHs can obtain routing information within the cluster and assess communication latency. The observation space of the first policy network. It can be represented as: ; represents the observation space of the first policy network. Includes multiple observed variables: the coordinate position of CH Position offset relative to all CH centroids CH to estimated source distance Spatial dispersion of the CM Average positional offset between CH .
[0072] CHs observe their environment and then change their movement accordingly. For example, it can be specifically divided into staying stationary, moving north, south, east, or west, denoted as... Only one action is performed at a time, therefore It is a multi-category random variable whose probability distribution is determined by parameters. Decision; among which, parameters satisfy: .
[0073] In order to observe the space The state information is incorporated into specific actions, leading to the proposal of an Actor Network. This network takes the observation space and the good / bad ratings obtained from a Critic Network as input, and the unnormalized score of each action as input. As the output, the action probability is then obtained through the softmax function: ;in, This is a neural network function with an output dimension of 5. For learnable parameters, Ultimately The probability of choosing an action.
[0074] Based on the observable state variables of the CH, such as the AOA estimate and RSS observation reported by the CM, and the relative geometric positional relationship between the CH and CM, a reward function is constructed to guide the CH to learn the optimal clustering control and movement strategy, enabling the CH to achieve optimal clustering control and movement at each time step. Obtain scalar reward The reward function is defined as a location-based confidence reward. Quality Coverage Rewards Explore diversity rewards The weighted value.
[0075] It should be noted that location-based reliability rewards Confidence level of radiation source estimation based on the intersection of signal arrival angles reported by CMs A reward is assigned. Specifically, the confidence level is determined based on both the diversity of perspectives and the number of effective AOAs. This reflects the strength of the geometric configuration in supporting monitoring and positioning. Confidence level. The calculation method can be: ; in, for The number of CMs (angle of arrival) of the signal that is reported at all times. For the first The signal arrival angle measured by CM For the first The signal arrival angle measured by CM This represents the total number of CMs.
[0076] Coverage Quality Rewards CH is encouraged to maintain CM within its effective communication and awareness range. Specifically, this includes coverage quality rewards. It can be represented as: ; in, For the first The position of CM, The centroid of CH For reliable communication radius, This is the scale parameter.
[0077] Explore Diversity Rewards CHs are encouraged to actively explore and to diversify their movement. Specifically, there are incentives for exploring diversity. It can be represented as: ; in, For the number of CH, For the first A drone CH in Location at any given moment This is a normalized velocity scale.
[0078] As an example, the reward function of the first policy network can be expressed as: .
[0079] Considering the limitations of a single CM platform's load capacity and operating time, and in order to maximize its monitoring performance, this invention further proposes the DDMPPO algorithm to achieve lightweight and miniaturized CMs by controlling the information processing and networking transmission capabilities of a single platform. By adopting a large-scale deployment approach, while ensuring an effective link with the CH, it uses a distributed method to monitor and locate ultra-shortwave radiation sources, thereby achieving dynamic and sustainable monitoring of the target area and improving the overall monitoring and positioning range and accuracy.
[0080] As an example, the observation space of the second policy network It can be represented as: ; represents the observation space of the first policy network. Includes multiple observed variables: the coordinate position of CM Distance from its parent CH Received signal strength Change in received signal strength Angle of arrival of the signal measured by CM .
[0081] CM's actions are mainly movement actions. It depends on the angle of arrival of the signal detected by the sensor. and received signal strength The measured value. When multiple CMs detect the radiation source simultaneously, the preferred method is... That is, at least three CMs simultaneously obtain effective [conditions / conditions]. Only then can the radiation source be effectively located.
[0082] A scalar signal is generated based on the state and action of the CM and used as a reward feedback. Preferably, the reward function of the second policy network is a signal strength gain reward. Maintaining a certain distance from CH is rewarded. Effective reward for the signal angle of arrival measured by CM The weighted value.
[0083] It should be noted that, ; in, , representing the minimum absolute difference between the signal angles of arrival measured by CM, where, when The smaller the value, meaning multiple CMs converge in the same direction, the lower the reward; when The closer At that time, the reward approaches its maximum.
[0084] As an example, the reward function of the second policy network can be expressed as: .
[0085] On the one hand, based on the near-end policy optimization algorithm, a reward function is designed to guide the agent to learn a cooperative strategy. A distributed training approach is adopted, allowing CHs (Chains Control Centers) to share information, while CMs (Machine Processors) only need to communicate with their respective CHs, reducing communication and computational overhead. On the other hand, through a hierarchical fusion reward mechanism, the reward function can simultaneously optimize the cluster management goals of CHs and the precise detection goals of CMs, ensuring their coordinated alignment towards the final accurate localization objective. Furthermore, the hierarchical structure decomposes the high-dimensional global joint decision-making problem into macroscopic scheduling of CHs and local search of CMs, significantly reducing algorithm complexity and enabling large-scale UAV swarm collaboration.
[0086] In the HMUDRL algorithm, HMUPPO for CH and DDMPPO for CM are both implemented using variants of the PPO algorithm. The PPO algorithm restricts the change of update strategy at each step by introducing a trust interval-based method and is implemented through a shear loss function.
[0087] As an example, the shear loss function can be expressed as: ; in, , representing the probability ratio of the new strategy to the old strategy. The parameters representing the strategy (neural network weight vector). To evaluate the advantage function of actions under the current strategy, The parameters for controlling the update magnitude of the control strategy are obtained through... ,Will Value constraints Within the range.
[0088] The total loss function of the PPO algorithm is the weighted maximization of policy loss, value loss, and entropy gain; therefore, the total loss function can be expressed as: ; in, , is the value loss function, representing the value loss between the estimated return and the actual return; , for target value; This is an entropy gain, which encourages strategy exploration and prevents premature convergence to a local optimum. and These are the weighting coefficients.
[0089] The HMUPPO algorithm is primarily used to optimize the clustering selection and movement of communication links (CHs) to maintain high-quality communication links. Its aim is to approximate a target distribution generated by the observation space. The advantage function can be estimated using the Generalized Advantage Estimation (GAE) method, and it can be expressed as: ; in, This is a discount factor used to measure the value of future rewards; These are hyperparameters of GAE used to weigh the tradeoff between bias and variance; This is the timing difference error (TD Error) calculated for CH.
[0090] Specifically, ; in, For CH at time step The reward received at that time For each observation dimension They all , and Scoring for the observation dimension, The scaling factor for the observation score controls the impact of different observations on the probability distribution; Current state The Critic network evaluation value; For the next state The estimated value of the Critic network.
[0091] The DDMPPO algorithm is mainly used to optimize CM sensing, signal-to-noise ratio, and energy consumption. Its advantage function can be expressed as: ; in, This is the timing difference error (TD Error) calculated for CM.
[0092] Specifically, .
[0093] This invention employs a heterogeneous HMUDRL with CH and CM to learn complex decision-making tasks. To effectively realize a distributed heterogeneous reinforcement learning framework, this invention designs a dual-branch policy network to reduce the complexity of decision-making. The policy network is improved through branch training, enabling the network to simultaneously output control and monitoring perception actions. The branch structure achieves separate optimization of control and monitoring.
[0094] As an example, the first branch is the first policy network, which outputs the clustering selection and movement control actions of the CH (Clustering Center). The second branch is the second policy network, which outputs the monitoring and energy consumption actions of the CM (Movement Control Center). Since the gradient direction of each branch is different, separate network parameters are required for control.
[0095] Based on the shearing loss function, the policy loss function of HMUDRL was modified according to the branching structure, that is: ; .
[0096] The strategy loss function is achieved by pruning parameters. The mechanism controls the magnitude of policy updates to prevent excessive changes between old and new policies. This mechanism is applied independently to the two branch structures of the first policy network and the second policy network to ensure stable updates of the two action spaces.
[0097] As an example, the total policy loss of the first policy network and the second policy network is: .
[0098] In HMUDRL, to ensure effective exploration of the actions of CH and CM, entropy gains are calculated separately for the two types of nodes. The policy entropy losses for CH and CM actions are as follows: ; .
[0099] As an example, the total entropy loss of the first policy network and the second policy network is: .
[0100] At the same time, the total entropy loss is summed with the policy entropy loss of CH and CM to unify the exploration pressure and prevent one branch from converging too early while the other branch still needs extensive exploration.
[0101] Since CH and CM are inherently heterogeneous in terms of state space, action space, reward function, and task objective, it is necessary to design independent value networks and calculate value loss functions separately.
[0102] Specifically, the value loss functions are as follows: ; .
[0103] The total value loss function of HMUDRL is the weighted sum of the value loss functions of the two parallel branches. As an example, the total value loss of the first policy network and the second policy network is: ; in, and These are the weighting coefficients.
[0104] Based on the total loss function of the PPO algorithm, the total loss function of HMUDRL includes policy loss, value loss, and entropy loss. HMUDRL consists of a first policy network and a second policy network.
[0105] As an example, the total loss function of the first policy network and the second policy network is: ; in, , , Let represent the total policy loss, total value loss, and total entropy loss of the first policy network and the second policy network, respectively. The weighting coefficient represents the total entropy loss. By setting appropriate weighting coefficients for balancing, effective optimization of monitoring heterogeneous drone swarms can be achieved.
[0106] It should be noted that, based on the node position and velocity state, node velocity constraints, node acceleration constraints, node position constraints, and positioning triggering conditions, the objective function is to maximize the duration of the positioning radiation source.
[0107] As an example, the movement and information transmission of the control node and monitoring node are adjusted through an objective function. Specifically, the objective function includes: ; ; ; .
[0108] Under various constraints, each CM is within a time range Find the optimal position and speed At any moment When three or more CMs are effectively measured, Only when the count is 1 can three-point positioning be achieved. This is achieved by maximizing... The sum of these can be used to obtain the radiation source. Continuous positioning time.
[0109] Simultaneously, under the constraint of clustering, we seek suitable clustering strategies and UAV CH locations to maximize the overall cluster throughput and minimize intra-cluster communication latency, thereby reducing the impact on the radiation source. The positioning time. Therefore, the signal-to-noise ratio is also used. Total communication latency Further optimize the objective function.
[0110] S300: It uses multiple monitoring nodes to synchronously collect signal characteristic parameters of the radiation source, aggregates the signal characteristic parameters from multiple spatially distributed monitoring nodes, and calculates the location result of the radiation source.
[0111] As one example, multiple monitoring nodes acquire data synchronously. Specifically, the multiple monitoring nodes form a non-collinear geometric distribution in space and simultaneously obtain effective signal characteristic parameters. These signal characteristic parameters include the change in received signal strength and the signal angle of arrival.
[0112] It should be noted that when the preset location triggering condition is met, the signal characteristic parameters from multiple spatially distributed monitoring nodes are aggregated; the preset location triggering condition is: the number of monitoring nodes that effectively detect radiation source signals is greater than or equal to three.
[0113] In other words, when at least three CMs simultaneously detect a valid radiation source from different spatial locations, the CMs upload their measured signal characteristic parameters to their respective CHs. The CHs then use the minimum difference positioning algorithm to jointly estimate the precise location coordinates and path loss index of the radiation source. The changes in received signal strength and signal angles of arrival from different CMs are fused to construct a set of nonlinear equations. By solving the set of nonlinear equations, the two-dimensional plane coordinates of the radiation source and the path loss index of the environment are obtained, thereby achieving high-precision positioning.
[0114] As an example, the minimum gap localization algorithm is used to calculate the localization result of the radiation source, specifically including: Using the change in received signal strength and the angle of arrival as joint observations, a set of nonlinear equations is constructed regarding the coordinates of the radiation source location and the path loss exponent. The nonlinear equations are solved using the weighted least squares method to obtain the location coordinates of the radiation source.
[0115] As an example, the method further includes: dynamically adjusting the deployment of heterogeneous UAV swarms based on the positioning results to achieve continuous tracking and positioning of radiation sources.
[0116] Furthermore, the deployment of heterogeneous UAV swarms is dynamically adjusted based on the positioning results. Specifically, the control nodes adjust their own positions according to the positioning results to optimize the communication link quality between the control nodes and their respective monitoring nodes, and guide the monitoring nodes to move towards the radiation source to improve the positioning signal quality.
[0117] In other words, the location results are also fed back into HMUDRL as feedback information to optimize HMUDRL. Based on the radiation source location results and node energy status, the deployment location and communication topology of the UAV swarm are dynamically adjusted; when the number of monitoring nodes is insufficient, monitoring tasks are also reallocated or the monitoring range is adjusted through control nodes.
[0118] For example, HMUDRL calculates corresponding reward values for CH and CM based on the confidence level of the positioning results (such as the quality of the geometric configuration of the AOA intersection) and the positioning error. Nodes learn and optimize their movement and perception strategies based on the reward values. CH learns how to adjust its position so that the CMs it manages can maintain the best geometric enclosure of the radiation source, thereby continuously obtaining high-precision positioning.
[0119] On the other hand, the present invention also provides a distributed positioning system for ultra-shortwave radiation sources, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of any of the methods described above.
[0120] The consistency of the system's technical features and methods will not be elaborated upon here.
[0121] To comprehensively analyze the performance of the HMUDRL of this invention, simulation experiments were also conducted, and it was compared with existing algorithms.
[0122] This invention simulates the aforementioned monitoring and sensing environment to conduct simulation experiments, evaluating the performance of the HMUDRL of this invention from multiple aspects such as positioning accuracy, reward value, number of steps to first target discovery, and average number of interactions, and comparing and analyzing it with baseline algorithms.
[0123] Consider configuring it in 1000×1000m 2The mission area features a heterogeneous UAV swarm. The CH (Chain Leader) is deployed at an altitude of 200–400 m, and the CM (Mobile Controller) at an altitude of 100–300 m. The CH's speed ranges from 0 to 22 m / s, and the CM's speed ranges from 0 to 14 m / s. The CH's transmit power is 20 dBm, its carrier frequency is 3.0 GHz, its bandwidth is 10 MHz, and the configured Hello message interval is 1 second. The maximum control distance for the CH is set to 500 m. Each CM is equipped with an ultra-shortwave radiation detector, consisting of eight direction-finding units with a 45° direction-finding angle, enabling omnidirectional direction finding to ensure reliable detection. Considering the weight and power consumption limitations of the payload carried by the nodes, a lightweight training scheme using an Intel Core i5-1130G7 processor is employed. Each round lasts 10 seconds, discretized into 100 time steps (0.1 seconds per step). For ease of simulation calculation, the speeds of the CH and CM are set to 10 m / s (1 m / step). For different rewards in a heterogeneous UAV swarm monitoring system, their weights are determined through the training process, and a balance between task control and monitoring is achieved by adjusting specific task scenarios. The simulation parameters are shown in Table 1.
[0124] Table 1 Simulation Parameter Settings
[0125] This invention constructs a heterogeneous UAV swarm, deploying 3 control UAVs (CH) and 14 monitoring UAVs (CM). In each experiment, the location of the ultra-shortwave radiation source is randomly set, and the CH controls the CMs to monitor and locate within the experimental area, with the system recording key data.
[0126] like Figures 5-8 As shown, the red dots and lines represent the trajectory of CH during the positioning process, the blue dots and lines represent the trajectory of CM during the positioning process, the yellow five-pointed star represents the location of the ultra-shortwave radiation source, the green "×" sign represents the location of the radiation source determined by the UAV, and the color change of the concentric circles around the radiation source in the area represents the change of the radiation field intensity gradient.
[0127] Figure 5 The localization process trajectory is shown under the conditions of RSS noise of 2dB and AOA noise of 5°. The complete localization process trajectories of HMUDRL, CH and CM are used. In the figure, CM moves around the radiation source to search for it. Finally, the location of the radiation source is relatively close to the actual radiation source.
[0128] Figure 6 For the positioning process trajectory of CH and CM after random deployment, without dynamic adjustment or collaborative work, the detection effect of UAV swarm on target radiation source is not as good as that of HMUDRL due to the lack of dynamic adjustment mechanism.
[0129] Figure 7To disable the clustering optimization function of CH and keep CH in a distributed, fixed position, such as uniform distribution, and only let CM use the DDMPPO algorithm to move along the localization process trajectory, although CM can autonomously adjust its position to improve its perception capabilities, the overall localization efficiency may be affected due to the lack of effective guidance and support from CH.
[0130] Figure 8 The image shows the trajectory of the localization process under high noise conditions of 6dB RSS noise and 20° AOA noise. The image also shows the performance of HMUDRL under high RSS and AOA noise conditions. Under these more challenging conditions, the localization performance of the system may decrease, but through optimization of HMUDRL, relatively accurate target radiation source localization can still be achieved, demonstrating the robustness and adaptability of HMUDRL in complex environments.
[0131] like Figure 9 The figure shows a comparison of the moving average trend of the total reward between HMUDRL of this invention and four baseline algorithms (COMA, MAA2C, MADDPG, and MASAC) over 1000 training rounds. The horizontal axis represents the number of training rounds, and the vertical axis represents the smoothed value of the cumulative reward, used to reflect the stability and convergence performance of each algorithm in policy optimization during long-term learning.
[0132] Overall, HMUDRL (blue curve) showed a significant learning advantage in the early stages of training. Its reward value rose rapidly and then stabilized after about 150 episodes, eventually settling at around 180, indicating that the algorithm has fast convergence ability and strong policy exploration efficiency. In contrast, while MASAC (purple curve) showed rapid growth in the early stages, its rewards fluctuated significantly and exhibited obvious oscillations in the later stages, indicating policy instability. MAA2C (green curve) maintained a relatively stable growth trend throughout the training process, with rewards remaining between 70 and 90, demonstrating good robustness. However, its maximum reward level was significantly lower than HMUDRL, reflecting room for improvement in its collaborative decision-making ability. COMA (orange curve) experienced drastic fluctuations in reward values, especially significant drops in the middle of training, indicating that its counterfactual baseline mechanism struggles to effectively model individual contributions in complex heterogeneous environments, leading to policy learning instability. MADDPG (red curve) performed the worst, with its reward value hovering around 25, showing almost no significant change with the training process. This suggests that when dealing with large-scale heterogeneous drone swarms, the centralized Critic network faces a severe curse of dimensionality, resulting in large value estimation biases and difficulties in policy updates.
[0133] It should also be noted that HMUDRL, by introducing a hierarchical architecture fusion mechanism, successfully achieved differentiated policy learning for CH and CM agents, avoiding the performance bottleneck caused by the unified policy space in traditional MARL methods. In addition, the hierarchical reward design further enhances the policy's responsiveness to environmental observations, enabling continuous optimization of localization behavior in dynamic electromagnetic environments. HMUDRL not only far surpasses existing mainstream multi-agent reinforcement learning algorithms in terms of final reward level, but also outperforms them in training stability and convergence speed, fully verifying its effectiveness and superiority in heterogeneous UAV cooperative localization tasks.
[0134] Table 2 compares and analyzes the performance of HMUDRL with four baseline algorithms in 1000 episodes, including RMSE value, localization success rate, and number of steps to first target discovery. The results show that HMUDRL significantly outperforms the baseline algorithms in the VHF radiation source localization task.
[0135] Table 2 Performance comparison of HMUDRL with four baseline algorithms
[0136] As shown in Table 2, in terms of the mean RMSE, HMUDRL's 39.34 meters is significantly lower than MADDPG's 620.43 meters, MASAC's 395.88 meters, and COMA's 135.02 meters, and only slightly higher than MAA2C's 82.70 meters. Its median RMSE of 22.83 meters is the lowest among all algorithms, and its maximum error of 678.11 meters is superior to MASAC (1774.58 meters) and MADDPG (1125.30 meters). Furthermore, HMUDRL's first target discovery steps (6 steps) are better than all algorithms except MADDPG (5 steps), and its localization success rate remains high while maintaining speed, reflecting the efficient exploration capability of HMUDRL in dynamic environments. This advantage stems from CH's optimization of cluster selection and movement control through HMUPPO, and CM's implementation of a two-layer architecture design that balances perception and energy consumption through DDMPPO. Combined with a hierarchical fusion reward mechanism, this reduces the risk of dimensionality disaster in policy updates and avoids premature branch convergence through differentiated entropy gain calculation. In contrast, MADDPG's centralized Critic network suffers from dimensionality explosion, leading to biased value estimation; MASAC's increased parameter complexity easily triggers policy oscillations; and MAA2C's shared Critic network struggles to adapt to heterogeneous target differences, resulting in limited positioning accuracy.
[0137] like Figure 10As shown, the data distribution characteristics are further revealed. HMUDRL's RMSE distribution is more concentrated with a smaller interquartile range, while the box plots of MADDPG and MASAC are significantly elongated and accompanied by numerous outliers (e.g., MADDPG's maximum value reaches 1125 meters), indicating drastic fluctuations in positioning error and poor stability. The box plot distributions of COMA and MAA2C are wider than HMUDRL, indicating a decrease in accuracy. In particular, COMA's mean RMSE (135.02 meters) is 3.4 times that of HMUDRL. Although its positioning success rate (96.3%) is close to that of HMUDRL (96.1%), the high error value indicates poor performance in complex environments. It can be said that HMUDRL has achieved a breakthrough in overall performance in terms of accuracy, stability, and efficiency through structural innovation and mechanism optimization.
[0138] Furthermore, this invention analyzes the number of information interactions within the UAV swarm, revealing significant differences in communication overhead across different network structures. For heterogeneous UAV swarms, the CM needs to upload information such as AOA estimation and RSS to the CH. The CH performs source localization estimation based on this information and issues control or clustering commands to the CM. Simultaneously, different CHs share location or estimation results to achieve cross-cluster collaboration.
[0139] In the experimental setup, this invention simplifies by assuming that within each time step, each CM sends state information (including AOA and RSS) to the CH once, and each pair of CHs interacts once if communication is established. In contrast, in a k-NN homogeneous structure, all drones are considered to be of the same type, and due to communication distance limitations, each drone only exchanges information with its five nearest neighbors. In a fully connected homogeneous structure, all drones communicate directly with each other, forming a fully connected network.
[0140] Table 3. Average number of information exchanges and communication savings within the drone swarm
[0141] Table 3 shows the average number of information interactions for UAV swarms under different architectures and their corresponding communication savings per training round. In the minimum configuration (1CH+8CM), the heterogeneous architecture requires only about 900 interactions per training round, while the homogeneous k-NN and fully connected architectures require 4500 and 7200 interactions respectively. This means that the heterogeneous architecture saves 80.0% and 87.5% of communication compared to k-NN and fully connected structures, respectively. As the system scales up to the maximum configuration (7CH+30CM), the number of interactions for the heterogeneous architecture increases to 3100, while the k-NN and fully connected architectures surge to 18500 and 133200, respectively. At this point, the communication savings of the heterogeneous architecture further increase to 83.2% (vs. k-NN) and 97.7% (vs. fully connected).
[0142] This trend demonstrates the significant advantages of heterogeneous architecture in reducing the number of information interactions within a drone swarm, particularly in large-scale drone deployments, where it can substantially reduce communication overhead and improve overall system efficiency. Furthermore, this also verifies the crucial role of the HMUDRL proposed in this invention in optimizing cooperative localization within a drone swarm, especially when performing VHF / UHF radiation source localization tasks in complex electromagnetic environments. Through reasonable role division and sparse communication mechanisms, it can significantly improve overall system energy efficiency and scalability while ensuring positioning accuracy.
[0143] like Figure 11 As shown, with the increase in the number of drones, the number of information exchanges required by heterogeneous architectures is far less than that of homogeneous architectures (whether k-NN or fully connected). This difference is mainly attributed to the introduction of more efficient communication strategies in heterogeneous architectures, such as information aggregation and broadcasting via CH, which reduces redundant information exchange.
[0144] In electromagnetic spectrum monitoring and positioning tasks, the model complexity of multi-agent reinforcement learning algorithms directly affects their training efficiency, communication overhead, and deployment feasibility. The HMUDRL proposed in this invention integrates a hierarchical architecture, RSS variation, and AOA measurement sensing mechanism in its model structure design, resulting in significantly higher complexity compared to classic multi-agent algorithms such as MADDPG, MASAC, COMA, and MAA2C.
[0145] In terms of parameter scale, HMUDRL adopts a heterogeneous dual-network structure: CH (cluster head) uses a standard Actor-Critic network with an input dimension of 13 and a hidden layer size of 128; CM (cluster members) introduces a gated fusion module, whose base network is a 128-dimensional MLP; additionally, it includes an action prior module based on AOA angle soft mapping, with a set action space size of [missing information]. The CH state dimension is 13, and the number of CH network parameters is approximately The CM has 11 state dimensions, its network includes gating logic, and has a slightly higher number of parameters, approximately [number missing]. The total number of parameters for the entire system is ;in, , The total is approximately 750,000.
[0146] In contrast, MADDPG and MASAC employ a centralized training-distributed execution (CTDE) architecture, where each agent needs to maintain an independent policy network and a centralized Q-network, resulting in a total number of parameters. In this case, the input dimension of the centralized Q-network is the sum of the products of the number of agents and the state and action dimensions, respectively. This will cause the number of parameters to increase quadratically with the number of agents, resulting in a computational complexity of up to [value missing]. ,in, For the agent's state or action dimension. When , At that time, the MADDPG Critic network can have up to 306 dimensions of parameters, and each fully connected layer will have up to 10 parameters. 5 The total number of parameters can reach 10. 6 The magnitude is significant. While COMA reduces variance through counterfactual baselines, its critic requires modeling each action combination, resulting in computational complexity reaching [a certain level]. MAA2C is a completely decentralized algorithm. Each agent trains its own Actor network and Critic network independently. Multiple agents cannot cooperate with each other and lack collaborative modeling capabilities.
[0147] From a computational complexity perspective, HMUDRL decouples global collaboration into the macroscopic deployment of the CH (Chain Action) and the local perception and movement of the CM (Commander Action) through a hierarchical structure, avoiding high-dimensional joint action space modeling and maintaining a linear increase in model complexity. At the same time, it can effectively balance collaboration capabilities and computational efficiency, compared to MADDPG. COMA These technologies demonstrate superior scalability and practicality in large-scale electromagnetic spectrum monitoring and positioning tasks.
[0148] In summary, this invention proposes a distributed localization method for ultra-shortwave radiation sources based on the Heterogeneous Multi-Agent Deep Reinforcement Learning (HMUDRL) algorithm. HMUDRL introduces a hierarchical architecture, dividing the UAV into two classes of agents: cluster heads (CH) and cluster members (CM), each with different roles, observation spaces, and reward mechanisms. This enables efficient collaboration, reduces communication overhead, and supports scalable decision-making processes. Combining a hierarchical fusion reward mechanism and an improved HMUPPO / DDMPPO training strategy, HMUDRL achieves stable and efficient policy learning in addressing the inherent heterogeneity and partial observability of agents in tasks. It solves the adaptability and efficiency problems of existing algorithms, achieving high-precision localization and large-scale monitoring of ultra-shortwave radiation sources, significantly improving the localization success rate and reducing localization errors.
[0149] Experimental results demonstrate that HMUDRL exhibits significant advantages in monitoring and positioning accuracy, average number of interactions per UAV unit, and target detection speed. HMUDRL achieved a positioning success rate of 96.1% in 1000 test rounds, an average improvement of 1.8% compared to baseline algorithms. Its mean positioning error (RMSE) was only 39.34 meters, a reduction of approximately 87.3%, significantly outperforming the baseline algorithm and demonstrating clear advantages in both positioning accuracy and robustness. In particular, compared to baseline methods, the cluster-based information aggregation mechanism reduces communication interactions by more than 80%, exhibiting robust performance in complex electromagnetic environments. These results fully validate that HMUDRL can effectively address the core challenges of UAV electromagnetic monitoring, including payload constraints, limited endurance, and complex dynamic electromagnetic environments, thereby effectively solving the data transmission control and monitoring / sensing positioning problems of UAVs in ultra-shortwave monitoring.
[0150] It should be noted that, for the sake of simplicity, the foregoing embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0151] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.
[0152] In the several embodiments provided in this application, it should be understood that the disclosed methods or systems can be implemented in other ways. For example, the embodiments described above are merely illustrative. For instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed.
[0153] The foregoing description is merely an exemplary embodiment of this disclosure and should not be construed as limiting the scope of this disclosure. Any equivalent changes and modifications made in accordance with the teachings of this disclosure shall still fall within the scope of this disclosure. Those skilled in the art will readily conceive of embodiments of this disclosure upon considering the specification and practicing the disclosure herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not described herein. The specification and embodiments are to be considered exemplary only, and the scope and spirit of this disclosure are defined by the claims.
[0154] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0155] Those skilled in the art will readily understand that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A distributed positioning method for ultra-shortwave radiation sources, characterized in that, The method includes: Construct a heterogeneous drone swarm that includes control nodes and monitoring nodes; Based on a multi-agent reinforcement learning algorithm, differentiated decision-making strategies are configured for the control node and the monitoring node respectively, so as to guide the heterogeneous UAV swarm to conduct cooperative search within the mission area; By synchronously collecting signal characteristic parameters of the radiation source using multiple monitoring nodes, aggregating the signal characteristic parameters from multiple spatially distributed monitoring nodes, and calculating the location result of the radiation source; Specifically, differentiating decision-making strategies are configured for the control node and the monitoring node, including: configuring a first strategy network for the control node, whose observation space includes at least one of its own coordinate position, position offset from the cluster centroid, distance to the estimated radiation source, spatial dispersion of its monitoring node, and average position offset between control nodes; and configuring a second strategy network for the monitoring node, whose observation space includes at least one of its own coordinate position, distance from its control node, received signal strength, change in received signal strength, and angle of arrival of the signal.
2. The distributed positioning method for ultra-shortwave radiation sources according to claim 1, characterized in that, Constructing a heterogeneous drone swarm that includes control nodes and monitoring nodes, specifically including: The drones with data processing and communication capabilities are set up as control nodes, serving as cluster heads; Configure drones with monitoring capabilities as monitoring nodes, and treat them as cluster members; Establish a star communication topology centered on the control node, where each monitoring node is associated with only one control node.
3. The distributed positioning method for ultra-shortwave radiation sources according to claim 2, characterized in that, The control node and the monitoring node establish and maintain intra-cluster association links through periodic broadcasting of Hello messages; wherein, the Hello message contains the sending node identifier, association status, transmission carrier frequency and signal power information.
4. The distributed positioning method for ultra-shortwave radiation sources according to claim 1, characterized in that, The first policy network is trained using the HMUPPO algorithm, and its reward function is a weighted sum of location confidence reward, coverage quality reward, and exploration diversity reward. The second policy network is trained using the DDMPPO algorithm, and its reward function is a weighted sum of signal strength gain reward, distance maintenance reward, and effective signal angle of arrival reward.
5. The distributed positioning method for ultra-shortwave radiation sources according to claim 1, characterized in that, The total loss function for the first policy network and the second policy network is: ; in, , , Let represent the total policy loss, total value loss, and total entropy loss of the first policy network and the second policy network, respectively. The weighting coefficient represents the total entropy loss.
6. The distributed positioning method for ultra-shortwave radiation sources according to claim 1, characterized in that, The simultaneous acquisition by multiple monitoring nodes specifically includes: multiple monitoring nodes forming a non-collinear geometric distribution in space, and simultaneously obtaining effective signal characteristic parameters.
7. The distributed positioning method for ultra-shortwave radiation sources according to claim 6, characterized in that, The location result of the radiation source is obtained by using the minimum gap localization algorithm, specifically including: Using the change in received signal strength and the angle of arrival as joint observations, a set of nonlinear equations is constructed regarding the coordinates of the radiation source location and the path loss exponent. The nonlinear equations are solved using the weighted least squares method to obtain the location coordinates of the radiation source.
8. The distributed positioning method for ultra-shortwave radiation sources according to claim 1, characterized in that, The deployment of the heterogeneous UAV swarm is dynamically adjusted based on the positioning results. Specifically, the control node adjusts its own position according to the positioning results to optimize the communication link quality between the control node and its monitoring node, and guides the monitoring node to move towards the radiation source to improve the positioning signal quality.
9. A distributed positioning system for ultra-shortwave radiation sources, comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 8.