Multi-agent unmanned aerial vehicle cluster collaborative search method based on semantic communication technology
By adopting semantic communication technology-based methods in the collaborative search mission of drone clusters, the problems of low communication efficiency and poor adaptability in complex environments are solved, and efficient and precise task execution is achieved.
Patent Information
- Application Number
- CN202510288918.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-06-13
AI Technical Summary
The collaborative search mission of drone clusters faces the problems of low communication efficiency, poor adaptability and insufficient task execution accuracy in complex environments.
The collaborative search method of multi-agent drone cluster based on semantic communication technology is adopted. By defining state space and action space, semantic communication protocols are initialized, and semantic information is summarized and abstract semantic information is carried out, task allocation and path planning is carried out, and dynamic adjustment and real-time task optimization is achieved in combination with semantic information.
It significantly improves the communication efficiency and adaptability of the drone cluster, enhances the accuracy and efficiency of task execution, and reduces the dependence on manual operations.
Smart Images

Figure CN120143879A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical fields of artificial intelligence and multi-agent cooperative control systems, and particularly relates to a multi-agent UAV cluster cooperative search method based on semantic communication technology. Background Art
[0002] With the rapid development of UAV technology, the research on multi-agent UAV clusters has gradually become a hot field at the forefront of technology. Compared with single UAV systems, multi-agent clusters have stronger flexibility, scalability, and task adaptability, and have broad application prospects in fields such as post-disaster rescue, environmental monitoring, urban management, and military reconnaissance. Among the numerous applications of UAV clusters, the cooperative search task has attracted much attention due to its complexity and high-efficiency requirements. Cooperative search generally refers to multiple UAVs cooperating to quickly and efficiently search for targets or cover the entire area within the target region. Its goal is to maximize the search efficiency, reduce the possibility of missing targets, and at the same time reduce the task execution time.
[0003] However, the realization of UAV cluster cooperative search faces various challenges. First is the environmental uncertainty. The search area is often complex and changeable, and may include environmental factors such as dynamic obstacles, irregular terrain, and weather changes. These factors pose extremely high requirements for the path planning and real-time adjustment of UAVs. Second, the cost of communication and cooperation cannot be ignored. The agents in the UAV cluster need to exchange information frequently to achieve cooperation. However, due to the limited communication link bandwidth and latency, traditional data communication modes are difficult to meet the requirements of real-time performance and robustness for complex tasks. Finally, resource limitations and time constraints are also important factors that must be considered when executing the search task. The computing power of UAVs is limited, requiring them to efficiently complete autonomous decision-making and path planning while also being able to handle complex problems such as environmental perception and task allocation. The energy of UAVs is limited, and they need to efficiently complete the search task while minimizing energy consumption and meeting the time requirements of the task.
[0004] To address the above challenges, researchers have proposed various cooperative search methods, which can be roughly divided into the following two categories: methods based on traditional algorithms and methods based on reinforcement learning. Traditional methods usually rely on mathematical optimization or bio-inspired algorithms, such as Voronoi partitioning, ant colony algorithms, and genetic algorithms. These methods achieve task allocation and path planning of UAVs by partitioning the search area or simulating natural behaviors. Although traditional methods have good performance in regular environments, their adaptability to complex and dynamic environments is weak, and their demand for computing resources is high.
[0005] In recent years, reinforcement learning technology has provided new ideas for the collaborative search of UAV swarms. Reinforcement learning conducts policy learning through the interaction between UAVs and the environment, achieving an end-to-end optimization process. Methods such as Deep Q-Network (DQN) and Multi-Agent Deep Deterministic Policy Gradient (MADDPG) have been applied to UAV path planning and collaborative decision-making. Although reinforcement learning methods perform well in certain specific scenarios, they still face a series of challenges that need to be addressed. These include: in complex and changing scenarios, it is particularly difficult to ensure the rapid convergence and excellent performance of the algorithm; due to the low sample utilization rate and the high sampling cost in the actual environment, these algorithms are difficult to be widely promoted in practical applications; as the number of agents increases, the state-action space grows exponentially, making the training process more difficult; the design of the reward function plays a crucial role in real application scenarios, and the parameters need to be continuously adjusted to avoid potential learning strategy errors and prevent the results from deviating significantly from the expectations; the generalization performance of the algorithm is also a major shortcoming and still needs to be significantly improved. Summary of the Invention
[0006] The purpose of the present invention is to provide an innovative solution that is more efficient, reliable, and applicable to the collaborative search tasks of multi-agent UAV swarms in complex environments, aiming at the problems existing in the above-mentioned prior art.
[0007] The technical solution to achieve the purpose of the present invention is: a multi-agent UAV swarm collaborative search method based on semantic communication technology, and the method includes the following steps:
[0008] Step 1, define the state space and action space of the multi-agent system of the UAV swarm according to the task requirements;
[0009] Step 2, initialize the semantic communication protocol for the interaction between the UAV swarm and the environment;
[0010] Step 3, summarize and abstract the semantic information of the UAV swarm;
[0011] Step 4, use the semantic information for the collaborative search task allocation and path planning of the UAV swarm;
[0012] Step 5, combine the semantic information to achieve the dynamic adjustment and real-time task optimization of the multi-agent system;
[0013] Step 6, evaluate the task completion effect and feedback to optimize the system strategy.
[0014] Furthermore, step 1 of defining the state space and action space of the multi-agent system of the UAV swarm according to the task requirements specifically includes:
[0015] Step 1.1. Represent the search task as a multi-agent partially observable Markov decision process, which is represented as a six-tuple:
[0016] MA-POMDP = (S, A, T, R, O, γ)
[0017] where MA-POMDP represents the partially observable Markov decision process, S represents the state space, i.e., the set of all possible world states, representing the true state of the environment, including the target location, obstacle distribution, and UAV state; A represents the action space, i.e., the set of all possible actions that each agent can execute, and the action sets of each agent constitute the action space of the entire cluster; T is represented as T(s'|s, a), which is the state transition probability function, describing the probability distribution of transitioning to the next state given the current state s ∈ S and a set of actions a 1 ,..., a n ∈ A, reflecting the response of the environment to the actions of the agents; R is represented as R(s, a, s') and is the reward function, reflecting the immediate reward or punishment obtained by the agent after taking an action, used to evaluate the quality of the action; O represents the observation space, i.e., the set of all possible observations that the agent can obtain from the environment; γ represents the discount factor, used to balance the importance of current rewards and future rewards;
[0018] Step 1.2. Further refine the state space, represented as the following vector:
[0019] S local = {p i , v i , e i , o i}
[0020] S global = {P progress , T targets , D env}
[0021] where S local represents the state of a single UAV; p i represents the current position of UAV i, denoted as p i = (x i , y i , z i ), where x i , y i , z i represent the horizontal and vertical coordinates and altitude respectively; v i represents the current speed of UAV i, denoted as: v i = (v xi , v yi, v zi ); e i represents the remaining energy of UAV i at present, indicating its current energy state, represented by a scalar; o i represents the target information currently sensed by UAV i, including target position, speed, and obstacles; S global represents the overall state information of the cluster; P progress represents the progress ratio of task completion, including the ratio of the searched area to the total searched area; T targets represents the distribution of global targets, including un-searched targets and covered targets; D env represents dynamic environment information;
[0022] Step 1.3, divide the action space into single-UAV actions and cluster cooperation actions, and define the action A of a single UAV single ={a path , a sensor , a data}; where a path represents path adjustment, including dynamically adjusting the flight trajectory according to the target; a sensor represents sensor control, including adjusting the sensor acquisition range, direction, or working mode; a data represents data upload, including transmitting the acquired target data to the cluster or the ground station; define the cluster cooperation action A cluster ={a share , a cooperate , a avoid}; where a share represents information sharing, including sharing target data, path planning, or task status information among UAVs; a cooperate represents collaborative decision-making, including the cluster dynamically adjusting task allocation according to the current state; a avoid represents obstacle avoidance, including the cluster avoiding obstacles or threats in the environment through a collaborative strategy;
[0023] Step 1.4, design the reward function R, and the reward function R includes the local reward R local and the global reward R global , where the local reward reflects the goal achievement degree of a single UAV, and the global reward evaluates the overall performance of the cluster through cluster efficiency and global target coverage rate. The reward function is expressed as:
[0024] R = αR local + βR global
[0025] where, R local is the local reward, used to calculate the performance of each UAV in task execution, including target discovery, path planning, and energy management. The formula is expressed as: R local = λ1 C target -λ 2 C energy -λ 3 C collision ,C target represents the number of targets successfully sensed or covered by the UAVs, and rewards the UAVs that have successfully detected the targets; C energy represents the energy consumption cost, which is proportional to the flight distance or flight time; C collision represents the penalty cost for collisions or deviations from the planned path; λ 1 ,λ 2 ,λ 3 are the corresponding weighting coefficients; R global is the total reward, used to evaluate the overall task completion and cooperation efficiency of the cluster, and is expressed by the formula: C coverage represents the number of targets effectively covered; C total represents the total number of targets defined in the task; T complete represents the time required to complete the task; C redundancy represents the redundant behaviors generated during the task execution; μ 1 ,μ 2 ,μ 3 are the corresponding weighting coefficients; α and β are weight factors used to balance the influence of local and global rewards.
[0026] Step 1.5, in the task execution process, a multi-agent reinforcement learning optimization strategy π is adopted, and the goal is to maximize the long-term cumulative reward G t , and the expression is:
[0027]
[0028] In the formula, γ represents the discount factor, used to measure the influence degree of future rewards; k represents the time step index, indicating k steps after the current time step t; R t+k represents the immediate reward obtained at time step t + k, and this reward is jointly determined by the previously defined local reward R local and the global reward R global .
[0029] Through the policy optimization algorithm, learn the optimal policy of each UAV while optimizing the cooperation policy Π of the cluster cluster (A cluster |S global ).
[0030] Furthermore, the semantic communication protocol for initializing the interaction between the UAV cluster and the environment described in Step 2 specifically includes:
[0031] Step 2.1: Define the semantic communication framework, initialize the semantic communication module, determine the communication topology structure in the UAV cluster, and initialize the main functions of the semantic communication module, including information extraction, compression, and transmission; configure communication parameters, including bandwidth allocation, maximum transmission delay, and data transmission quality threshold, to ensure the adaptability of semantic communication in different environments;
[0032] Step 2.2: The UAVs collect multi-source perception data, perform data preprocessing operations, remove noise and redundant information, perform timestamp calibration to unify the data time base, and encapsulate the processed data into the following raw data tuple:
[0033] I raw ={I state ,I image ,I lidar ,I timestamp}
[0034] where, I state represents the state data of the current UAV, including the attitude, speed, and position data of the UAV; I image represents the image data collected by the UAV camera; I lidar represents the point cloud data collected by the UAV lidar; I timestamp represents the specific time when the current UAV encapsulates the data, which is used to achieve synchronization;
[0035] Step 2.3: Convert the multi-source perception data of the UAVs into high-level semantic information, and use a deep learning model to extract the semantic features of the perception data;
[0036] Step 2.4: Extract effective semantic information from the high-dimensional perception data, further screen out the effective semantic information related to the task, and use feature fusion technology to integrate the semantic information; perform cross-modal alignment on the semantic features of multi-modal data to ensure the spatial and temporal consistency of feature expressions; and use principal component analysis to reduce the dimension of the high-dimensional features of the semantic information, removing redundant dimensions while retaining the main information; re-encode the features based on the information entropy theory to optimize the semantic information with high confidence; then re-encapsulate the data compressed by the above process into an optimized semantic data tuple;
[0037] Step 2.5: Set the transmission priority of the semantic information according to the task requirements, and give priority to transmitting the semantic information that has a key impact on the collaborative search task. Focus on transmitting environmental scanning and target detection information at the initial stage of the task, pay attention to target distribution and environmental dynamic updates in the middle stage, and give priority to transmitting task progress and summary data in the later stage; adopt a queue management and dynamic bandwidth allocation mechanism to ensure the timeliness and reliability of high-priority information; at the same time, reduce transmission losses through a feedback confirmation mechanism and optimized coding, and finally encapsulate it in the form of a packet sorted by priority.
[0038] Furthermore, in step 2.3, a deep learning model is used to extract the semantic features of the perception data, specifically including:
[0039] For image data, the recognition target category, location, and status are extracted based on a convolutional neural network;
[0040] For point cloud data, based on the point cloud network, three-dimensional environmental structure information is generated;
[0041] After the relevant information is processed, it is encapsulated into a semantic data tuple:
[0042] I semantic ={I state ,I target ,I environment ,I timestamp}
[0043] Among them, I target represents the target information extracted from the image data, and the expression is I target ={i class ,i location ,i motion ,i confidence}, where i class describes the type of the target; i location represents the location information of the target; i motion describes the speed and movement direction of the target; i confidence represents the confidence level of the recognition result; I environment represents the environmental information extracted from the point cloud data, specifically including the semantic features of obstacles and surrounding structures, and the expression is: I environment ={e type ,e boundary ,e density ,e confidence}, where e type represents different types of objects or structure types in the environment; e boundary describes the boundary information of obstacles or environmental objects, expressed as a point cloud bounding box e density represents the point density in a local area of the point cloud data, used to evaluate the complexity of the environmental structure, and is defined as: Among them, N points represents the number of points in the point cloud area, V region represents the volume of the point cloud area; e confidence represents the confidence level of the environmental information extraction result, used to measure the accuracy and reliability of the semantic information.
[0044] Furthermore, the optimized semantic data tuple repackaged in step 2.4 is:
[0045] Isemantic ={I id ,I info ,I timestamp}
[0046] Wherein, I id represents the UAV number, I info represents the semantic information data after fusion and compression, I timestamp represents the timestamp.
[0047] Furthermore, step 3 specifically includes:
[0048] Step 3.1, summarize the UAV semantic information. Each UAV transmits the generated semantic data to the distributed summarization node through the communication link; after receiving the semantic data, the distributed summarization node conducts preliminary sorting on it:
[0049]
[0050] In the formula, I aggregated represents the aggregated semantic information data set, which is the preliminary aggregation result after the distributed summarization node receives the semantic data sent by n UAVs; I semantic,i represents the optimized semantic data tuple of the i-th UAV;
[0051] Step 3.2, extract and classify the received semantic data I aggregated into target information I target , environmental information I environment and status information I state ;
[0052] Step 3.3, summarize the target information, environmental information, and status information respectively;
[0053] Step 3.4, extract and screen the key information, analyze the requirements of the current task stage, and extract high-priority information; calculate the information entropy to screen out duplicate or inefficient information in the data, conduct secondary verification or elimination on data with low confidence to ensure the reliability of the summarized information; use feature selection technology to screen out the key information to reduce the feature dimension and optimize the data structure; at the same time, construct the cluster knowledge base K cluster , and store all semantic information;
[0054] Step 3.5, compress and encode the screened key information, and use the time series alignment algorithm and global coordinate mapping technology to standardize the data; the optimized semantic information is repackaged in a new tuple form, and the expression is: I aggregated ={I id ,I key_info ,I timestamp ,Iquality}, where I id represents the UAV identification number, used to distinguish the data source, I key_info represents the optimized key information, I timestamp represents the timestamp, I quality represents the confidence level and quality index of the information;
[0055] Step 3.6, based on the relevance R i between the UAV position p task and the mission area, distribute the semantic information to the UAVs related to the mission area, and give priority to sending the dynamic environment update information to the UAVs U j in the adjacent areas, and distribute the global target distribution information to the UAVs U target responsible for covering the unsearched areas. The calculation formula is:
[0056]
[0057] where, U k represents the k-th UAV in the cluster; R task,k represents the mission relevance, indicating the relevance k between the UAV U uncovered and the mission area, used to determine whether the UAV needs to receive this information; T
[0058] Step 3.7, use the distributed hash table to locate the data recipient, and the receiving node sends a feedback confirmation message to the sending node after receiving the data. When packet loss or damage occurs during information transmission, enable the automatic retransmission mechanism f retry to retransmit the data packet.
[0059] Furthermore, in Step 3.3, the target information, environmental information, and status information are summarized respectively, specifically including:
[0060] (1) Target information summary: Merge the target information reported by each UAV I target , filter out duplicate targets, and select the most credible target data according to the confidence level i confidence Use the algorithm based on multi-view geometry to match and fuse and adjust the target position for duplicate targets i location , expressed as:
[0061]
[0062] where, i location,fused represents the fused target position information; f MVG represents the multi-view geometry algorithm; i location,i represents the target position information of UAV i;
[0063] (2) Environmental information aggregation: Align and fuse the environmental semantic features I of each UAV environment in space to construct a global environmental map M env :
[0064]
[0065] where f SLAM represents the collaborative mapping function, and I environment,i represents the environmental information of the i-th UAV;
[0066] (3) Status information aggregation: Aggregate the status data I of each UAV state to form a cluster global status table T state,global :
[0067]
[0068] where S local,i represents the local status information of the i-th UAV.
[0069] Furthermore, the use of semantic information for UAV cluster collaborative search task allocation and path planning described in step 4 specifically includes:
[0070] Step 4.1, decompose and initially allocate tasks according to the environmental information obtained through semantic communication;
[0071] Based on the environmental information I obtained through semantic communication environment , divide the task area R task into several small areas, and the division principle is optimized according to the perception ability, flight speed, and task time constraint of the UAV;
[0072] Based on the local status S of the UAV local initially allocate the tasks to obtain an initial task allocation matrix M initial :
[0073]
[0074] Step 4.2, use the multi-agent task allocation optimization algorithm f opt , based on the initial task allocation matrix M initial , perform global optimization adjustment to generate a globally optimized task allocation matrix M opt ; specifically:
[0075] Combine the global status S global and the local status S of each UAV local,i , adjust the initial task allocation plan to solve the task overlap and conflict problems, and the expression of the globally optimized task allocation matrix M opt is:
[0076]
[0077] Step 4.3, after each UAV receives the assigned task area, it generates a flight path P i (t) through a cooperative path planning algorithm; then, it updates the flight trajectory P adjust (t) in real time through an online path correction strategy f i to cope with dynamic environmental changes and generate a corrected flight trajectory P' i (t), and the expression is: P' i (t) = f adjust (P i (t), I environment );
[0078] Step 4.4, through the semantic communication module, low-latency exchange of task information is achieved among UAVs, and key information I key_info is mainly transmitted; according to the data obtained in real time, the task allocation scheme and path planning results are dynamically optimized and adjusted, and the optimized task allocation scheme is encapsulated as a task allocation list TaskList, and the expression is: where, R j is the task area, U i is the UAV number assigned to area R j , and k is the total number of task areas;
[0079] Step 4.5, during the task execution process, each UAV continuously uploads the UAV status information S state and the task completion progress P task through the semantic communication module, and the background dynamically adjusts the task allocation scheme for the uncovered area P uncovered to ensure the integrity of task coverage, and the expression of P uncovered is:
[0080]
[0081] And the adjusted task allocation scheme is recorded as: k' is the total number of adjusted task areas;
[0082] Record the data related to task execution and store it in the system knowledge base K cluster , for subsequent performance optimization and data verification, which is expressed as: K cluster = K cluster ∪ {I execution}, where, I execution represents the task execution result data.
[0083] Furthermore, the dynamic adjustment and real-time task optimization of the multi-agent system by combining semantic information described in Step 5 specifically include:
[0084] Step 5.1, during the task execution, the status S of each UAV in the cluster is monitored in real time through Step 3 state,i ;
[0085] Step 5.2, the UAV cluster periodically updates the local status information S local and the global status information S global , combines the collected real-time semantic data I semantic , and implements the task dynamic optimization of the multi-agent cluster:
[0086] S local,i = f updatelocal (p i , v i , e i , o i , I semantic )
[0087]
[0088] wherein, S local,i represents the local status information of the UAV U i ; f updateLocal represents the local status update function, which calculates the latest local status by integrating various information of the UAV; p i represents the position information of the UAV U i ; v i represents the speed information of the UAV U i ; e i represents the remaining energy of the UAV U i , which is used to evaluate the endurance ability; o i represents the sensor observation data of the UAV U i , which is used to detect the surrounding environment; I semantic represents the semantic information collected by the UAV; f updateGlobal represents the global status update function, which summarizes the local status of all UAVs and generates the overall status of the cluster;
[0089] Step 5.3, the UAV clusters share the latest environmental perception data I environment through semantic communication, and for the path conflicts and path deviations ΔP i during the task execution, they are corrected in real time through semantic information feedback, and the flight path is updated. The correction formula is:
[0090] P’ i (t) = f coordinate (P i (t), P j (t), I environment )
[0091] Among them, P i (t) and P j (t) respectively represent the paths of the drones U i and U j , f coordinate represents the cooperative path planning algorithm, and P’ i (t) represents the path of the updated drone U i ;
[0092] Step 5.4, when an emergency or system anomaly occurs during task execution, obtain the current status information S state and the emergency information E event to quickly evaluate the current status, and readjust the task assignment and flight path based on the evaluation results:
[0093] M adjusted = f reassign (M opt , E event , S global )
[0094] P adjusted,i (t) = f replan (P i (t), E event )
[0095] In the formula, M adjusted represents the task assignment readjusted based on the evaluation results; f reassign represents the task reassignment function, which combines the current optimal task M opt , the emergency information E event and the global status S global to dynamically adjust the task division of the drones; P adjusted,i (t) represents the flight path of the updated drone U i ; f replan represents the path replanning function, which calculates a new flight path based on the original path P i (t) and the emergency E event ; P i (t) represents the original flight path of the drone U i , that is, the path planning before adjustment;
[0096] Step 5.5, according to the real-time task execution data I execution , use the deep learning algorithm f eval to evaluate the task execution effect, and the expression is:
[0097] Q task = f eval (I execution)
[0098] Among them, Q task represents the task completion quality;
[0099] And it is combined with historical data for feedback optimization. The optimized model f optimized is stored in the cluster knowledge base K cluster :
[0100] K cluster = K cluster ∪{f optimize}.
[0101] Furthermore, the evaluation of the task completion effect and the feedback optimization of the system strategy described in step 6 specifically include:
[0102] Step 6.1, after the task is completed, the system collects the drone execution status information S local,i and the cluster global status information S global , and based on the task completion rate P progress and the target coverage rate C coverage conducts a preliminary evaluation of the task effect; at the same time, records all the action sequences A t ={a 1 ,a 2 ,…,a T} and the corresponding state transition sequence S t ={s 1 ,s 2 ,…,s T}, and outputs the task completion score R final ;
[0103] Step 6.2, automatically detects abnormal situations E exception in the task through semantic information, analyzes the causes of abnormal points using the state transition sequence T(s'|s,a), and the abnormal data D anomaly is recorded and transmitted to the central control node C central in the system for subsequent analysis and optimization;
[0104] Step 6.3, based on the state, action, and reward data recorded during the task execution, uses a multi-agent reinforcement learning framework to update the strategy of the drone cluster;
[0105] First, use a multi-agent deep reinforcement learning algorithm to optimize the strategy π i (a|s) of each drone, so that the strategy can complete the task more efficiently under the local state S local,i and the global state S global ;
[0106] Secondly, update the value function Q for each dronei (s, a), the calculation formula is:
[0107]
[0108] where Q i is used to evaluate the long-term reward of the selected action in the current state, and the value function is updated through temporal difference learning; represents the expected value under the policy π i ; a t represents the t-th action in the action sequence A t , and s t represents the state corresponding to the t-th action in the action sequence state transition sequence;
[0109] Combined with the cluster global state S global and the cluster joint action set A cluster to optimize the cooperation strategy Π cluster (A cluster |S global ), to solve the problems of task assignment conflict and uneven resource allocation;
[0110] Step 6.4, through the evaluation of the task execution effect, combined with the historical task data H history and the latest trained reinforcement learning strategy to iteratively optimize the task assignment, path planning and cooperation strategy of the UAV;
[0111] Step 6.5, after completing the task evaluation, by analyzing the historical task execution data H history and the real-time feedback data I feedback , conduct system-level optimization design, specifically:
[0112] Based on the real-time feedback of the multi-agent cluster task execution effect, optimize the resource allocation strategy, path planning algorithm and cooperation strategy for the next round of tasks;
[0113] Construct a closed-loop feedback system for dynamic task optimization based on reinforcement learning, combine semantic information with a dynamic adjustment mechanism, and optimize the multi-agent cooperation strategy through closed-loop feedback. The calculation formula is:
[0114] M optimized = f optimize (H history , I feedback )
[0115] Store the optimization result in the cluster knowledge base K cluster :
[0116] K cluster = K cluster ∪ {f optimize}
[0117] Compared with the prior art, the present invention has the following remarkable advantages:
[0118] (1) Innovatively introduce semantic communication technology into the multi-agent collaborative decision-making process, solve the problem that traditional methods have too high requirements for information transmission and processing efficiency in complex environments, significantly improve the communication efficiency of the UAV cluster, reduce redundant data transmission, and solve the traditional communication bottleneck problem. Through semantic information sharing and collaborative decision-making, improve the adaptability and flexibility of the cluster in complex environments.
[0119] (2) Through the efficient information compression and transmission of semantic communication technology, significantly improve the collaborative ability between clusters, and optimize task allocation and path planning through the reinforcement learning method of the multi-agent system, thereby improving the execution efficiency and adaptability of the UAV cluster in complex tasks.
[0120] (3) Design a task planning mechanism based on the partially observable Markov decision process (MA-POMDP) to solve the problem of difficult sharing of state and decision information in multi-agent collaborative search tasks.
[0121] (4) The present invention dynamically optimizes task planning and resource allocation through continuous evaluation and historical data feedback.
[0122] (5) The present invention reasonably allocates resources and energy, while ensuring efficient execution, meeting the requirements of task time and energy consumption.
[0123] The following further describes the present invention in detail with reference to the accompanying drawings. Description of the Drawings
[0124] Figure 1 It is a flowchart of a multi-agent UAV cluster collaborative search method based on semantic communication technology in an embodiment.
[0125] Figure 2 It is a flowchart of a single UAV for multi-agent collaborative search in an embodiment.
[0126] Figure 3 It is a flowchart of UAV cluster communication using semantic communication technology in an embodiment. Detailed Embodiments
[0127] In order to make the objectives, technical solutions and advantages of the present application clearer, the following further describes the present application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0128] It should be noted that if there are descriptions involving "first", "second", etc. in the embodiments of the present invention, such descriptions of "first", "second", etc. are only for descriptive purposes and should not be construed as indicating or implying their relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include at least one such feature. Additionally, the technical solutions between various embodiments may be combined with each other, but it must be based on what can be achieved by those of ordinary skill in the art. When the combination of technical solutions results in contradictions or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection required by the present invention.
[0129] In view of the problems of low communication efficiency, poor adaptability, and insufficient task execution accuracy in the existing UAV swarm cooperative search method, the present invention proposes a multi-agent UAV swarm cooperative search method based on semantic communication technology.
[0130] In one embodiment, a multi-agent UAV swarm cooperative search method based on semantic communication technology is proposed, and the method includes the following steps:
[0131] Step 1, define the state space and action space of the multi-agent system of the UAV swarm according to the task requirements;
[0132] Step 2, initialize the semantic communication protocol for the interaction between the UAV swarm and the environment;
[0133] Step 3, summarize and abstract the semantic information of the UAV swarm;
[0134] Step 4, use the semantic information for task allocation and path planning of the UAV swarm cooperative search;
[0135] Step 5, combine the semantic information to achieve dynamic adjustment and real-time task optimization of the multi-agent system;
[0136] Step 6, evaluate the task completion effect and feedback to optimize the system strategy.
[0137] Further, in one of the embodiments, step 1 of defining the state space and action space of the multi-agent system of the UAV swarm according to the task requirements specifically includes:
[0138] Step 1.1, represent the search task as a multi-agent partially observable Markov decision process, and represent it as a six-tuple:
[0139] MA-POMDP = (S, A, T, R, O, γ)
[0140] where MA-POMDP represents the partially observable Markov decision process, Denote the state space, which is the set of all possible world states, representing the true state of the environment, including the target location, obstacle distribution, and UAV state, etc.; A denotes the action space, which is the set of all possible actions that each agent can execute, and the action sets of each agent constitute the action space of the entire cluster; T is denoted as T(s'|s,a), which is the state transition probability function, describing the probability distribution of transitioning to the next state given the current state s∈S and a set of actions a 1 ,...,a n ∈A, reflecting the response of the environment to the agent's actions; R is denoted as P(s,a,s') which is the reward function, reflecting the immediate reward or punishment obtained by the agent after taking an action, used to evaluate the quality of the action; O denotes the observation space, which is the set of all possible observations that the agent can obtain from the environment. Due to partial observability, the agent cannot directly obtain the true state of the environment and can only obtain partial information through sensors or communication; γ denotes the discount factor, used to balance the importance of current rewards and future rewards;
[0141] Step 1.2, further refine the state space, represented as the following vector:
[0142] S local ={p i ,v i ,e i ,o i}
[0143] S global ={P progress ,T targets ,D env}
[0144] Among them, S local represents the state of a single UAV; p i represents the current position of UAV i, denoted as p i =(x i ,y i ,z i ), where x i ,y i ,z i represent the horizontal and vertical coordinates and altitude respectively; v i represents the current speed of UAV i, denoted as: v i =(v xi ,v yi ,v zi );e i represents the remaining energy of UAV i, indicating its current energy state, represented by a scalar; o i represents the target information currently perceived by UAV i, including target location, speed, obstacles, etc.; Sglobal Indicates the overall status information of the cluster; P progress Indicates the progress ratio of task completion, including the ratio of the searched area to the total searched area; T targets Indicates the distribution of global targets, including un-searched targets and covered targets; D env Indicates dynamic environment information, including influencing factors such as dynamic changes of obstacles and weather conditions;
[0145] Step 1.3, divide the action space into single drone actions and cluster cooperation actions, and define the action A of a single drone single ={a path ,a sensor ,a data}; where a path Indicates path adjustment, including dynamically adjusting the flight trajectory according to the target, such as turning, accelerating or decelerating; a sensor Indicates sensor control, including adjusting the sensor acquisition range, direction or working mode to optimize the sensing effect; a data Indicates data upload, including transmitting the acquired target data to the cluster or the ground station; define the cluster cooperation action A cluster ={a share ,a cooperate ,a avoid}; where a share Indicates information sharing, including sharing target data, path planning or task status information among drones; a cooperate Indicates collaborative decision-making, including the cluster dynamically adjusting task allocation according to the current state, such as allocating search areas; a avoid Indicates obstacle avoidance, including the cluster avoiding obstacles or threats in the environment through collaborative strategies;
[0146] Step 1.4, design the reward function R, the reward function R includes the local reward R local and the global reward R global , where the local reward reflects the goal completion degree of a single drone, and the global reward evaluates the overall performance of the cluster through cluster efficiency and global target coverage rate. The reward function is expressed as:
[0147] R = αR local + βR global
[0148] where, R local is the local reward, used to calculate the performance of each drone during task execution, including factors such as target discovery, path planning and energy management. The formula is expressed as: R local = λ 1 C target - λ 2 C energy - λ3 C collision , C target represents the number of targets successfully sensed or covered by the UAVs, and rewards the UAVs with successful target detection; C energy represents the energy consumption cost, which is proportional to the flight distance or flight time; C collision represents the penalty cost for collisions or deviations from the planned path; λ 1 , λ 2 , λ 3 is the corresponding weighting coefficient; R global is the total reward, used to evaluate the overall task completion and cooperation efficiency of the cluster, and is expressed by the formula: C coverage represents the number of targets effectively covered; C total represents the total number of targets defined by the task; T complete represents the time required to complete the task; C redundancy represents the redundant behaviors generated during task execution (such as multiple UAVs repeatedly covering the same area); μ 1 , μ 2 , μ 3 is the corresponding weighting coefficient; α and β are weight factors, used to balance the influence of local and global rewards.
[0149] Step 1.5, the multi-agent reinforcement learning (MARL) optimization strategy π is adopted during the task execution process, and the goal is to maximize the long-term cumulative reward G t , and the expression is:
[0150]
[0151] In the formula, γ represents the discount factor, used to measure the influence degree of future rewards; k represents the time step index, indicating k steps after the current time step t; R t+k represents the immediate reward obtained at time step t + k, and this reward is jointly determined by the previously defined local reward R local and the global reward R global .
[0152] Through the policy optimization algorithm, learn the optimal policy of each UAV while optimizing the cooperation policy Π of the cluster cluster (A cluster |S global ).
[0153] Furthermore, in one of the embodiments, the semantic communication protocol for initializing the interaction between the UAV cluster and the environment in step 2 specifically includes:
[0154] Step 2.1, Define the semantic communication framework, initialize the semantic communication module, determine the communication topology structure in the UAV cluster, and initialize the main functions of the semantic communication module, including information extraction, compression, and transmission; configure communication parameters, including bandwidth allocation, maximum transmission delay, and data transmission quality threshold, to ensure the adaptability of semantic communication in different environments;
[0155] Step 2.2, The UAVs collect multi-source perception data, perform data preprocessing operations, remove noise and redundant information, perform timestamp calibration, unify the data time reference, and the processed data is encapsulated into the following raw data tuple:
[0156] I raw ={I state ,I image ,I lidar ,I timestamp}
[0157] Among them, I state represents the state data of the current UAV, including the attitude, speed, and position data of the UAV; I image represents the image data collected by the UAV camera; I lidar represents the point cloud data collected by the UAV lidar; I timestamp represents the specific time when the current UAV encapsulates the data, which is used to achieve synchronization;
[0158] Step 2.3, Convert the multi-source perception data of the UAV into high-level semantic information, and use a deep learning model to extract the semantic features of the perception data;
[0159] Here, for the image data, based on the convolutional neural network, identify the target category, location, and state;
[0160] For the point cloud data, based on the point cloud network, generate three-dimensional environmental structure information;
[0161] After the relevant information is processed, it is encapsulated into a semantic data tuple:
[0162] I semantic ={I state ,I target ,I environment ,I timestamp}
[0163] Among them, I target represents the target information extracted from the image data, and the expression is I target ={i class ,i location ,i motion ,i confidence}, where, i class describes the type of the target; i locationRepresents the position information of the target; i motion Describes the speed and movement direction of the target; i confidence Represents the confidence level of the recognition result; I environment Represents the environmental information extracted from the point cloud data, specifically including the semantic features of obstacles and surrounding structures, and the expression is: I environment ={e type ,e boundary ,e density ,e confidence}, where, e type Represents different types of objects or structure types in the environment, such as "buildings", "trees", "roads", etc.; e boundary Describes the boundary information of obstacles or environmental objects, expressed as a point cloud bounding box e density Represents the point density of the local area in the point cloud data, used to evaluate the complexity of the environmental structure, and is defined as: Among them, T represents the number of points in the point cloud area, V region Represents the volume of the point cloud area; e confidence Represents the confidence level of the environmental information extraction result, used to measure the accuracy and reliability of the semantic information.
[0164] Step 2.4, Extract effective semantic information from the high-dimensional perception data, further screen out the effective semantic information related to the task, and use feature fusion technology to integrate the semantic information; perform cross-modal alignment on the semantic features of multi-modal data to ensure the spatial and temporal consistency of the feature expressions; and use principal component analysis to reduce the dimension of the high-dimensional features of the semantic information, removing redundant dimensions while retaining the main information; re-encode the features based on the information entropy theory to optimize the semantic information with high confidence; then re-package the data compressed by the above process into an optimized semantic data tuple:
[0165] I semantic ={I id ,I info ,I timestamp}
[0166] Among them, I id Represents the UAV number, I info Represents the semantic information data after fusion and compression, I timestamp Represents the timestamp.
[0167] Step 2.5, set the transmission priority of semantic information according to the task requirements, and give priority to transmitting semantic information that has a key impact on the collaborative search task. Focus on transmitting environmental scanning and target detection information in the initial stage of the task, pay attention to target distribution and environmental dynamic updates in the middle stage, and give priority to transmitting task progress and summary data in the later stage. Adopt a queue management and dynamic bandwidth allocation mechanism to ensure the timeliness and reliability of high-priority information. At the same time, reduce transmission losses through a feedback confirmation mechanism and optimized coding, and finally encapsulate it in the form of data packets sorted by priority.
[0168] Further, in one of the embodiments, in combination with Figure 3 , Step 3 specifically includes:
[0169] Step 3.1, summarize the semantic information of the drones. Each drone transmits the semantic data it generates to the distributed summary node through the communication link. After receiving the semantic data, the distributed summary node conducts preliminary collation on it:
[0170]
[0171] In the formula, I aggregated represents the aggregated semantic information data set, which represents the preliminary aggregation result after the distributed summary node receives the semantic data sent by n drones; I semantic,i represents the optimized semantic data tuple of the i-th drone;
[0172] Step 3.2, extract and classify the received semantic data I aggregated , and classify the data into target information I target , environmental information I environment and status information I state ;
[0173] Specifically include:
[0174] (1) Target information summary: Merge the target information I target reported by each drone, filter out duplicate targets, select the most credible target data according to the confidence level i confidence , use an algorithm based on multi-view geometry to match and fuse the duplicate targets and adjust the target position i location , which is expressed as:
[0175]
[0176] Among them, i location,fused represents the fused target position information; f MVG represents the multi-view geometry algorithm; i location,i represents the target position information of the i-th drone;
[0177] (2) Environmental information aggregation: Align and fuse the environmental semantic features I of each UAV environment in space to construct a global environmental map M env :
[0178]
[0179] Among them, f SLAM represents the collaborative mapping function, and I environment,i represents the environmental information of the i-th UAV;
[0180] (3) Status information aggregation: Aggregate the status data I of each UAV state to form a cluster global status table T state,global :
[0181]
[0182] Among them, S local,i represents the local status information of the i-th UAV;
[0183] Step 3.3: Aggregate the target information, environmental information, and status information respectively;
[0184] Step 3.4: Extract and screen the key information, analyze the requirements of the current task stage, and extract high-priority information; Calculate the information entropy to screen out duplicate or inefficient information in the data, perform secondary verification or elimination on data with low confidence to ensure the reliability of the aggregated information; Use feature selection technology to screen out key information to reduce the feature dimension and optimize the data structure; At the same time, construct a cluster knowledge base K cluster to store all semantic information for subsequent task calls or analysis;
[0185] Step 3.5: Compress and encode the screened key information, and use the time series alignment algorithm and global coordinate mapping technology to standardize the data; The optimized semantic information is re-encapsulated in a new tuple form, and the expression is: I aggregated ={I id , I key_info , I timestamp , I quality}, where I id represents the UAV identification number used to distinguish the data source, I key_info represents the optimized key information, I timestamp represents the timestamp, and I quality represents the confidence and quality index of the information;
[0186] Step 3.6: Based on the correlation R between the UAV position p i and the task area task, distribute the semantic information to the drones related to the task area, and preferentially send the dynamic environment update information to the drones U in the adjacent area j , distribute the global target distribution information to the drones U responsible for covering the unsearched area target , the calculation formula is:
[0187]
[0188] wherein, U k represents the k-th drone in the cluster; R task,k represents the task relevance, indicating the relevance between the drone U k and the task area, and is used to judge whether the drone needs to receive this information; T uncovered represents the set of uncovered targets;
[0189] Step 3.7, use the distributed hash table (DHT) to locate the data receiver, and the receiving node sends a feedback confirmation message to the sending node after receiving the data. When packet loss or damage occurs during information transmission, enable the automatic retransmission mechanism f retry , retransmit the data packet.
[0190] Furthermore, in one embodiment, in combination with Figure 2 , the drone swarm collaborative search task allocation and path planning using semantic information described in step 4 specifically includes:
[0191] Step 4.1, decompose and initially allocate the task according to the environmental information obtained by semantic communication;
[0192] Based on the environmental information I obtained by semantic communication environment , divide the task area R task into several small areas, and the division principle is optimized according to the sensing ability, flight speed and task time constraint of the drone;
[0193] Based on the local state S of the drone local initially allocate the task to obtain the initial task allocation matrix M initial :
[0194]
[0195] Step 4.2, use the multi-agent task allocation optimization algorithm f opt , based on the initial task allocation matrix M initial , perform global optimization adjustment to generate the global optimization task allocation matrix M opt ; specifically:
[0196] Combine the global state S global and the local state S of each drone local,i, adjust the initial task allocation plan to solve the problems of task overlap and conflict, and globally optimize the task allocation matrix M opt The expression of
[0197]
[0198] Step 4.3, after each UAV receives the assigned task area, generate a flight path P i (t); then, through the online path correction strategy f adjust , update the flight trajectory P i (t) in real time to cope with dynamic environmental changes and generate a corrected flight trajectory P' i (t), and the expression is: P' i (t) = f adjust (P i (t), I environment );
[0199] Step 4.4, through the semantic communication module, low-latency exchange of task information is realized among UAVs, and key information I key_info is mainly transmitted; according to the real-time acquired data, dynamically optimize and adjust the task allocation plan and path planning results, and encapsulate the optimized task allocation plan as a task allocation list TaskList, and the expression is: Among them, R j is the task area, U i is the UAV number assigned to the area R j , and k is the total number of task areas;
[0200] Step 4.5, during the task execution process, each UAV continuously uploads the UAV status information S state and the task completion progress P task , and the background dynamically adjusts the task allocation plan for the uncovered area P uncovered to ensure the integrity of task coverage, and the expression of P uncovered is:
[0201]
[0202] And record the adjusted task allocation plan as: k' is the adjusted total number of task areas;
[0203] Record the data related to task execution and store it in the system knowledge base K cluster , for subsequent performance optimization and data verification, expressed as: K cluster = K cluster ∪ {I execution}, where I execution represents the task execution result data.
[0204] Furthermore, in one of the embodiments, in combination with Figure 2 , the combination of semantic information in step 5 is used to achieve dynamic adjustment and real-time task optimization of the multi-agent system, specifically including:
[0205] Step 5.1, during the task execution process, the status S of each drone in the cluster is monitored in real time through step 3 state,i ;
[0206] Step 5.2, the drone cluster regularly updates the local status information S local and the global status information S global , and combines the collected real-time semantic data I semantic to implement dynamic task optimization of the multi-agent cluster:
[0207] S local,i = f updatelocal (p i , v i , e i , o i , I semantic )
[0208]
[0209] In the formula, S local,i represents the local status information of the drone U i ; f updateLocal represents the local status update function, which calculates the latest local status by integrating various information of the drone; p i represents the position information of the drone U i ; v i represents the speed information of the drone U i ; e i represents the remaining energy of the drone U i , which is used to evaluate the endurance ability; o i represents the sensor observation data of the drone U i , which is used to detect the surrounding environment; I semantic represents the semantic information collected by the drone; f updateGlobal represents the global status update function, which summarizes the local status of all drones to generate the overall status of the cluster;
[0210] Step 5.3, the drone clusters share the latest environmental perception data I environment through semantic communication, and for the path conflicts and path deviations ΔP i during the task execution, real-time correction is performed through semantic information feedback, and the flight path is updated. The correction formula is:
[0211] P′i $(t) = f$ coordinate (P i (t), P j (t), I environment )
[0212] Among them, P i (t) and P j (t) respectively represent the paths of drones U i and U j , f coordinate represents the cooperative path planning algorithm, P′ i (t) represents the path of the updated drone U i ;
[0213] Step 5.4, when an emergency or system anomaly occurs during the task execution, obtain the current status information S state and the emergency information E event to quickly evaluate the current status, and readjust the task assignment and flight path based on the evaluation result:
[0214] M adjusted = f reassign (M opt , E event , S global )
[0215] P adjusted,i (t) = f replan (P i (t), E event )
[0216] In the formula, M adjusted represents the task assignment readjusted based on the evaluation result; f reassign represents the task reallocation function, which combines the current optimal task M opt , the emergency information E event and the global status S global to dynamically adjust the task division of the drones; P adjusted,i (t) represents the flight path of the drone U i readjusted based on the evaluation result; f replan represents the path replanning function, which calculates a new flight path based on the original path P i (t) and the emergency E event ; P i (t) represents the original flight path of the drone U i , that is, the path planning before adjustment;
[0217] Step 5.5, according to the real-time task execution data I execution , use the deep learning algorithm feval Evaluate the task execution effect, with the expression:
[0218] Q task = f eval (I execution )
[0219] where Q task represents the task completion quality;
[0220] And combine with historical data for feedback optimization. The optimized model f optimized is stored in the cluster knowledge base K cluster :
[0221] K cluster = K cluster ∪ {f optimize}.
[0222] Furthermore, in one of the embodiments, the step of evaluating the task completion effect and feedback optimizing the system strategy described in step 6 specifically includes:
[0223] Step 6.1, after the task is completed, the system collects the drone execution status information S local,i and the cluster global status information S global , and based on the task completion rate P progress and the target coverage rate C coverage conduct a preliminary evaluation of the task effect; at the same time, record all the action sequences A t = {a 1 , a 2 , …, a T} and the corresponding state transition sequence S t = {s 1 , s 2 , …, s T}, and output the task completion score R final ;
[0224] Step 6.2, automatically detect the abnormal situation E exception in the task through semantic information, analyze the cause of the abnormal point using the state transition sequence T(s'|s,a), and the abnormal data D anomaly is recorded and transmitted to the central control node C central in the system for subsequent analysis and optimization;
[0225] Step 6.3, based on the state, action, and reward data recorded during the task execution, use the multi-agent reinforcement learning framework to update the strategy of the drone cluster;
[0226] First, use the multi-agent deep reinforcement learning algorithm to optimize the strategy π i(a|s) such that the policy can complete tasks more efficiently in the local state S local,i and the global state S global ;
[0227] Secondly, update the value function Q i (s,a), and the calculation formula is:
[0228]
[0229] where Q i is used to evaluate the long-term reward of the selected action in the current state, and the value function update is achieved through temporal difference learning; represents the expected value under the policy π i ; a t represents the t-th action in the action sequence A t , and s t represents the state corresponding to the t-th action in the state transition sequence of the action sequence;
[0230] Combine the cluster global state S global and the cluster joint action set A cluster to optimize the cooperation policy Π cluster (A cluster |S global ), and solve the problems of task assignment conflict and uneven resource allocation;
[0231] Step 6.4, through the evaluation of the task execution effect, combine the historical task data H history and the newly trained reinforcement learning policy to iteratively optimize the task assignment, path planning and cooperation policy of the UAVs;
[0232] Step 6.5, after completing the task evaluation, through the analysis of the historical task execution data H history and the real-time feedback data I feedback , perform system-level optimization design, specifically:
[0233] Based on the real-time feedback of the multi-agent cluster task execution effect, optimize the resource allocation strategy, path planning algorithm and cooperation policy for the next round of tasks;
[0234] Construct a closed-loop feedback system for dynamic task optimization based on reinforcement learning, combine semantic information with a dynamic adjustment mechanism, and optimize the multi-agent cooperation policy through closed-loop feedback. The calculation formula is:
[0235] M optimized = f optimize (H history , I feedback )
[0236] The optimized results are stored in the cluster knowledge base K cluster :
[0237] K cluster = K cluster ∪ {f optimize}
[0238] In one embodiment, a multi-agent UAV cluster collaborative search system based on semantic communication technology is provided. The system includes:
[0239] The first module is used to define the state space and action space of the multi-agent system of the UAV cluster according to the mission requirements;
[0240] The second module is used to initialize the semantic communication protocol for the interaction between the UAV cluster and the environment;
[0241] The third module is used to summarize and abstract the semantic information of the UAV cluster;
[0242] The fourth module is used to realize the dynamic adjustment and real-time task optimization of the multi-agent system by combining semantic information;
[0243] The fifth module is used to write a dedicated encirclement prompt word template for the fine-tuned encirclement decision-making large model;
[0244] The sixth module is used to evaluate the task completion effect and feedback to optimize the system strategy.
[0245] For the specific limitations of the multi-agent UAV cluster collaborative search system based on semantic communication technology, reference can be made to the limitations of the multi-agent UAV cluster collaborative search method based on semantic communication technology in the above text, which will not be elaborated here. Each module in the above multi-agent UAV cluster collaborative search system based on semantic communication technology can be implemented in whole or in part by software, hardware and their combination. The above modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.
[0246] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the following steps:
[0247] Step 1, define the state space and action space of the multi-agent system of the UAV cluster according to the mission requirements;
[0248] Step 2, initialize the semantic communication protocol for the interaction between the UAV cluster and the environment;
[0249] Step 3, summarize and abstract the semantic information of the UAV cluster;
[0250] Step 4, using semantic information for UAV swarm collaborative search task allocation and path planning;
[0251] Step 5, realizing dynamic adjustment and real-time task optimization of multi-agent systems by combining semantic information;
[0252] Step 6, evaluating the task completion effect and feeding back to optimize the system strategy.
[0253] For the specific limitations of each step, reference can be made to the limitations of the multi-agent UAV swarm collaborative search method based on semantic communication technology in the above text, which will not be elaborated here.
[0254] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, it realizes:
[0255] Step 1, defining the state space and action space of the multi-agent system of the UAV swarm according to the task requirements;
[0256] Step 2, initializing the semantic communication protocol for the interaction between the UAV swarm and the environment;
[0257] Step 3, summarizing and abstracting the semantic information of the UAV swarm;
[0258] Step 4, using semantic information for UAV swarm collaborative search task allocation and path planning;
[0259] Step 5, realizing dynamic adjustment and real-time task optimization of multi-agent systems by combining semantic information;
[0260] Step 6, evaluating the task completion effect and feeding back to optimize the system strategy.
[0261] For the specific limitations of each step, reference can be made to the limitations of the multi-agent UAV swarm collaborative search method based on semantic communication technology in the above text, which will not be elaborated here.
[0262] The present invention can effectively solve problems such as low communication efficiency, poor adaptability, and insufficient task execution accuracy faced in the actual UAV swarm collaborative search tasks, significantly improve the adaptability and flexibility of the system, and at the same time reduce the dependence on manual operations, achieving more efficient and intelligent automated collaboration.
[0263] The foregoing has shown and described the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments, and what is described in the above embodiments and the specification is only to illustrate the principle of the present invention. Without departing from the spirit and scope of the present invention, any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A multi-agent UAV cluster collaborative search method based on semantic communication technology, characterized in that: The method comprises the following steps: Step 1: Define the state space and action space of the drone swarm multi-agent system according to the task requirements; Step 2: Initialize the semantic communication protocol for the drone cluster to interact with the environment; Step 3: Summarize and abstract the semantic information of the drone cluster; Step 4: Use semantic information to perform collaborative search task allocation and path planning for drone clusters; Step 5: Combine semantic information to achieve dynamic adjustment and real-time task optimization of multi-agent systems; Step 6: Evaluate the task completion results and provide feedback to optimize the system strategy.
2. The multi-agent UAV cluster collaborative search method based on semantic communication technology according to claim 1 is characterized in that: Step 1 defines the state space and action space of the drone swarm multi-agent system according to the task requirements, including: Step 1.1, represent the search task as a multi-agent partially observable Markov decision process, which is represented as a six-tuple: MA-POMDP=(S,A,T,R,O,γ) Among them, MA-POMDP represents a partially observable Markov decision process, represents the state space, that is, the set of all possible world states, representing the real state of the environment, including the target position, obstacle distribution and drone state; A represents the action space, that is, the set of all possible actions that each agent can perform, and the action set of each agent constitutes the action space of the entire cluster; T is denoted as T(s'|s,a), which is the state transition probability function, describing the set of actions a1,...,a taken by all agents given the current state s∈S. n ∈A, the probability distribution of transitioning to the next state reflects the response of the environment to the agent's actions; R, denoted as R(s,a,s'), is the reward function, which reflects the immediate reward or punishment obtained by the agent after taking an action, and is used to evaluate the quality of the action; O represents the observation space, that is, the set of all possible observations that the agent can obtain from the environment; γ represents the discount factor, which is used to balance the importance of current rewards and future rewards; Step 1.2, further refine the state space and express it as the following vector: S local ={p i ,v i ,e i ,o i } S global ={P progress ,T targets ,D env } Among them, S local Indicates the status of a single drone; p i Indicates the current position of drone i, denoted as p i =(x i ,y i ,z i ), where x i ,y i ,z i Respectively represent the horizontal, vertical coordinates and height; v i Indicates the current speed of drone i, expressed as: v i =(v xi ,v yi ,v zi );e i Indicates the current remaining energy of drone i, indicating its current energy state, expressed as a scalar; o i Indicates the target information currently perceived by UAV i, including target position, speed, and obstacles; S global Indicates the overall status information of the cluster; P progress Indicates the progress ratio of task completion, including the ratio of the searched area to the total searched area; T targets represents the distribution of global targets, including unsearched targets and covered targets; D env Represents dynamic environment information; Step 1.3: Divide the action space into single drone actions and cluster collaborative actions, and define the action A of a single drone single ={a path ,a sensor ,a data }; where a path Indicates path adjustment, including dynamic adjustment of flight trajectory according to the target; a sensor Indicates sensor control, including adjusting the sensor acquisition range, direction or working mode; a data Indicates data upload, including transmitting the collected target data to the cluster or ground station; defines cluster collaboration action A cluster ={a share ,a cooperate ,a avoid }; where a share Indicates information sharing, including sharing of target data, path planning or mission status information between drones; a cooperate represents collaborative decision-making, including the cluster dynamically adjusting task allocation according to the current state; a avoid Represents obstacle avoidance, including the cluster avoiding obstacles or threats in the environment through cooperative strategies; Step 1.4, design the reward function R, which includes the local reward R local With the global reward R global , where the local reward reflects the target completion of a single drone, and the global reward evaluates the overall performance of the cluster through cluster efficiency and global target coverage. The reward function is expressed as: R=αR local +βR gloval Among them, R local is a local reward, which is used to calculate the performance of each UAV in task execution, including target discovery, path planning and energy management. The formula is: local =λ1C target -λ2C energy -λ3C collision , C target Indicates the number of targets that the drone successfully perceives or covers, and rewards the drone that successfully detects the target; C energy Represents the energy consumption cost, which is proportional to the flight distance or flight time; C collision represents the penalty cost for collision or deviation from the planned path; λ1, λ2, λ3 are the corresponding weighting coefficients; R global is the total reward, which is used to evaluate the overall task completion and collaboration efficiency of the cluster. The formula is: C coverage Indicates the number of targets effectively covered; C total Indicates the total number of targets defined in the task; T complete Indicates the time required to complete the task; C redundancy represents the redundant behavior generated during task execution; μ1, μ2, μ3 are the corresponding weighting coefficients; α and β are weight factors used to balance the impact of local and global rewards. Step 1.5: The task execution process uses a multi-agent reinforcement learning optimization strategy π, the goal is to maximize the long-term cumulative reward G t , the expression is: In the formula, γ represents the discount factor, which is used to measure the impact of future rewards; k represents the time step index, which represents the k steps after the current time step t; R t+k represents the immediate reward obtained at time step t+k, which is composed of the local reward R defined previously local and the global reward R global Joint decision making; Learn the optimal strategy for each drone through the policy optimization algorithm At the same time, optimize the cluster's collaboration strategy cluster (A cluster |S global ).
3. The multi-agent UAV cluster collaborative search method based on semantic communication technology according to claim 1 or 2, characterized in that: Step 2 initializes the semantic communication protocol for the interaction between the drone cluster and the environment, specifically including: Step 2.1, define the semantic communication framework, initialize the semantic communication module, determine the communication topology in the drone cluster, initialize the main functions of the semantic communication module, including information extraction, compression and transmission; configure communication parameters, including bandwidth allocation, maximum transmission delay and data transmission quality threshold, to ensure the adaptability of semantic communication in different environments; Step 2.2: The drone collects multi-source perception data, performs data preprocessing operations, removes noise and redundant information, performs timestamp calibration, unifies the data time base, and encapsulates the processed data into the following raw data tuples: I raw ={I state ,I image ,I lidar ,I timestamp } Among them, I state Indicates the current state data of the drone, including the drone's attitude, speed, and position data; I image Represents the image data collected by the drone camera; I lidar Represents the point cloud data collected by the UAV’s laser radar; I timestamp Indicates the specific time when the current drone encapsulates data, which is used to achieve synchronization; Step 2.3, convert the multi-source perception data of the drone into high-level semantic information, and use the deep learning model to extract the semantic features of the perception data; Step 2.4, extract effective semantic information from high-dimensional perceptual data, further screen out effective semantic information related to the task, and use feature fusion technology to integrate semantic information; perform cross-modal alignment of semantic features of multimodal data to ensure spatial and temporal consistency of feature expression; and use principal component analysis to reduce the dimensionality of high-dimensional features of semantic information, retaining the main information while eliminating redundant dimensions; re-encode the features based on information entropy theory to optimize high-confidence semantic information; then re-encapsulate the data compressed by the above process into optimized semantic data tuples; Step 2.5, set the transmission priority of semantic information according to task requirements, give priority to the transmission of semantic information that has a key impact on the collaborative search task, focus on the transmission of environmental scanning and target detection information in the early stage of the task, focus on target distribution and dynamic environmental updates in the middle stage, and give priority to the transmission of task progress and summary data in the later stage; use queue management and dynamic bandwidth allocation mechanisms to ensure the timeliness and reliability of high-priority information; at the same time, reduce transmission losses through feedback confirmation mechanisms and optimized coding, and finally encapsulate them in the form of priority-sorted data packets.
4. The multi-agent UAV cluster collaborative search method based on semantic communication technology according to claim 3 is characterized in that: In step 2.3, a deep learning model is used to extract the semantic features of the perception data, including: For image data, the target category, location and state are extracted and recognized based on convolutional neural networks; For point cloud data, three-dimensional environment structure information is generated based on the point cloud network; After the relevant information is processed, it is encapsulated into a semantic data tuple: I semantic ={I state ,I target ,I environment ,I timestamp } Among them, I target Represents the target information extracted from the image data, expressed as I target ={i class ,i location ,i motion ,i confidence }, where i class Describe the type of target; i location Indicates the location information of the target; i motion Describe the speed and direction of movement of the target; i confidence Indicates the confidence of the recognition result; I environment Represents the environmental information extracted from the point cloud data, including the semantic features of obstacles and surrounding structures, and is expressed as: environment ={e type ,e boundary ,e density ,e confidence }, where e type Represents different categories of objects or structure types in the environment; boundary Describes the boundary information of obstacles or environmental objects, represented as a point cloud bounding box e density Represents the point density of a local area in the point cloud data, which is used to evaluate the complexity of the environment structure and is defined as: Among them, N poins Represents the number of points in the point cloud area, V region Represents the volume of the point cloud area; e confidence It represents the confidence of the environmental information extraction result, which is used to measure the accuracy and reliability of semantic information.
5. The multi-agent UAV cluster collaborative search method based on semantic communication technology according to claim 4 is characterized in that: The optimized semantic data tuple repackaged in step 2.4 is: I semantic ={I id ,I info ,I timestamp } Among them, I id Indicates the drone number, I info Represents the semantic information data after fusion and compression, I timestamp Indicates a timestamp.
6. The multi-agent UAV cluster collaborative search method based on semantic communication technology according to claim 3 is characterized in that: Step 3 specifically includes: Step 3.1: Summarize the semantic information of drones. Each drone transmits the semantic data it generates to the distributed aggregation node through the communication link. After receiving the semantic data, the distributed aggregation node performs preliminary sorting: In the formula, I aggregated It represents the aggregated semantic information dataset, which means the preliminary aggregation result after the distributed aggregation node receives the semantic data sent by n drones; I semantic,i represents the optimized semantic data tuple of the i-th drone; Step 3.2: Receive the semantic data I aggregated Perform information extraction and classification to classify data into target information I target 、Environmental Information I environment and status information I state ; Step 3.3, summarize the target information, environment information and status information respectively; Step 3.4, extract and screen key information, analyze the needs of the current task stage, and extract high-priority information; screen out duplicate or inefficient information in the data through information entropy calculation, and perform secondary verification or elimination on low-confidence data to ensure the reliability of the summarized information; use feature selection technology to screen out key information to reduce feature dimensions and optimize data structure; and build a cluster knowledge base K cluster , stores all semantic information; Step 3.5, compress and encode the filtered key information, and use the time alignment algorithm and global coordinate mapping technology to standardize the data; the optimized semantic information is repackaged in the form of a new tuple, expressed as: I aggregated = {I id ,I key_info ,I timestamp ,I quality }, where I id Indicates the drone identification number, which is used to distinguish the data source. key_info Indicates optimized key information, I timestamp Indicates the timestamp, I quality Indicators representing confidence and quality of information; Step 3.6, based on the drone position p i Correlation with the mission area R task , distributes semantic information to UAVs related to the mission area, and dynamic environment update information is sent to UAVs in neighboring areas first. j , the global target distribution information is distributed to the UAV U responsible for covering the unsearched area target , the calculation formula is: Among them, U k represents the kth drone in the cluster; R task,k Indicates task relevance, indicating that the drone U k The correlation with the mission area is used to determine whether the UAV needs to receive the information; T uncovered represents the set of uncovered targets; Step 3.7, use the distributed hash table to locate the data receiver. After receiving the data, the receiving node sends a feedback confirmation message to the sending node. When packet loss or damage occurs during information transmission, the automatic retransmission mechanism is enabled. retry , retransmit the data packet.
7. The multi-agent UAV cluster collaborative search method based on semantic communication technology according to claim 6 is characterized in that: In step 3.3, the target information, environment information and status information are summarized respectively, including: (1) Target information aggregation: Merge the target information reported by each drone. target , filter duplicate targets, according to the confidence i confidence Select the most reliable target data, use the algorithm based on multi-view geometry to match the repeated targets and fuse and adjust the target position i location , expressed as: Among them, i location,fused represents the fused target position information; f MVG Represents the multi-view geometry algorithm; i location,i Indicates the target position information of UAV i; (2) Environmental information aggregation: The environmental semantic features of each drone are aggregated. environment Perform spatial alignment and fusion to build a global environment map M env : Among them, f SLAM represents the collaborative mapping function, I environment,i Represents the environmental information of the i-th UAV; (3) Status information summary: Summarize the status data of each drone. state , forming the cluster global state table T state,global : Among them, S local,i Represents the local state information of the i-th UAV.
8. The multi-agent UAV cluster collaborative search method based on semantic communication technology according to claim 7 is characterized in that: Step 4 uses semantic information to perform UAV cluster collaborative search task allocation and path planning, specifically including: Step 4.1, decompose and preliminarily allocate tasks based on the environmental information obtained by semantic communication; Environmental Information Acquisition Based on Semantic Communication I environment , the task area R task Divide into several small areas, and the division principle is optimized according to the drone's perception ability, flight speed and mission time constraints; Based on the local state S of the UAV local Perform a preliminary assignment of tasks to obtain the initial task assignment matrix M initial : Step 4.2: Use the multi-agent task allocation optimization algorithm f opt , based on the initial task allocation matrix M initial , perform global optimization adjustment and generate the global optimization task allocation matrix M opt ; Specifically: Combined with the global state S global and each UAV's local state S local,i , adjust the initial task allocation plan, solve the problem of task overlap and conflict, and globally optimize the task allocation matrix M opt The expression is: Step 4.3: After receiving the assigned mission area, each UAV generates a flight path P through a collaborative path planning algorithm. i (t); then the online path correction strategy f adjust , real-time update of flight trajectory P i (t) Generate a corrected flight trajectory P' to cope with dynamic environmental changes i (t), the expression is: P' i (t) = f adjust (P i (t),I environment ); Step 4.4: Through the semantic communication module, the UAVs can exchange mission information with low latency, focusing on transmitting key information. key_info ; According to the real-time data, dynamically optimize and adjust the task allocation plan and path planning results, and encapsulate the optimized task allocation plan into a task allocation list TaskList, the expression is: Among them, R j For the mission area, U i To be assigned to region R j The number of the UAV, k is the total number of mission areas; Step 4.5: During the mission execution, each drone continuously uploads drone status information S through the semantic communication module. state and task completion progress P task , the background dynamically adjusts the uncovered area P according to real-time data uncovered The task allocation scheme ensures the completeness of task coverage. uncovered The expression is: And record the adjusted task allocation plan as: k' is adjusted to the total number of task areas; Record task execution related data and store it in the system knowledge base K cluster , for subsequent performance optimization and data verification, expressed as: K cluster =K cluster ∪{I execution }, where I execution Represents task execution result data.
9. The multi-agent UAV cluster collaborative search method based on semantic communication technology according to claim 8 is characterized in that: Step 5 combines semantic information to achieve dynamic adjustment and real-time task optimization of the multi-agent system, specifically including: Step 5.1: During the task execution, the state S of each drone in the cluster is calculated through step 3. state,i Conduct real-time monitoring; Step 5.2: The drone cluster periodically updates the local state information S local and global state information S global , combined with the collected real-time semantic data I semantic , implement dynamic optimization of tasks in multi-agent clusters: In the formula, S local,i Indicates drone U i Local state information of updateLocal represents the local state update function, which integrates various information of the drone to calculate the latest local state; p i Indicates drone U i Location information; i Indicates drone U i Speed information; i Indicates drone U i The remaining energy is used to evaluate the endurance; i Indicates drone U i Sensor observation data is used to detect the surrounding environment; i semantic Represents the semantic information collected by the drone; f updateGlobal represents the global state update function, which aggregates the local states of all drones to generate the overall state of the cluster; Step 5.3: UAV clusters share the latest environmental perception data through semantic communication. environment , path conflicts and path deviations ΔP in task execution i , real-time correction is performed through semantic information feedback to update the flight path. The correction formula is: P′ i (t)=f coordinate (P i (t),P j (t),I environment ) Among them, P i (t) and P j (t) respectively represent the drone U i and U j The path, f coordinate represents the collaborative path planning algorithm, P' i (t) represents the updated UAV U i Path; Step 5.4: When an emergency or system anomaly occurs during task execution, the current status information S is obtained through the semantic communication module. state and emergency information event Perform a quick assessment of the current state and realign task assignments and flight paths based on the assessment results: M adjusted =f reassign (M opt ,E event ,S global ) P adjusted,i (t)=f replan (P i (t),E event ) Where M adjusted represents the task score after readjustment based on the evaluation results; f reassign Represents the task redistribution function, combined with the current optimal task M opt 、Emergency information event and the global state S global , dynamically adjust the task division of UAVs; P adjusted,i (t) represents the UAV after readjustment based on the evaluation results. i The flight path of replan Represents the path replanning function, based on the original path P i (t) and emergency E event Calculate new flight path; P i (t) represents the drone U i The original flight path, i.e. the path planned before adjustment; Step 5.5: Execute task data according to real-time execution , using the deep learning algorithm f eval The task execution effect is evaluated, and the expression is: Q task =f eval (I execution ) Among them, Q task Indicates the quality of task completion; Combined with historical data for feedback optimization, the optimized model f optimized Stored in the cluster knowledge base K cluster : K cluster =K cluster ∪{f optimize }。 10. The multi-agent UAV cluster collaborative search method based on semantic communication technology according to claim 9 is characterized in that: Step 6 evaluates the task completion effect and provides feedback to optimize the system strategy, including: Step 6.1: After the task is completed, the system collects the drone execution status information S local,i and cluster global status information S global , based on the task completion rate P progress and target coverage C coverage Conduct a preliminary evaluation of the task effect; at the same time, record all action sequences generated during the task execution. t ={a1,a2,…,a T } and the corresponding state transition sequence S t ={s1,s2,…,s T }, output the task completion score R final ; Step 6.2: Automatically detect abnormal situations in the task through semantic information E exception , using the state transition sequence T(s'|s,a) to analyze the causes of abnormal points, and the abnormal data D anomaly is recorded and transmitted to the central control node C in the system central , used for subsequent analysis and optimization; Step 6.3, based on the state, action and reward data recorded during the task execution, the strategy of the drone cluster is updated using the multi-agent reinforcement learning framework; First, we use a multi-agent deep reinforcement learning algorithm to optimize the strategy π of each drone. i (a|s), so that the strategy can be in the local state S local,i and the global state S global Complete tasks more efficiently; Secondly, update the value function Q for each drone i (s,a), the calculation formula is: Among them, Q i Used to evaluate the long-term benefits of selecting actions in the current state, the value function update is achieved through time difference learning; In the strategy π i The expected value under t Represents action sequence A t The t-th action in s t Represents the state corresponding to the tth action in the action sequence state transition sequence; Combined with the cluster global state S global and cluster joint action set A cluster Optimize collaboration strategy cluster (A cluster |S global ), solving the problems of task allocation conflicts and uneven resource allocation; Step 6.4, by evaluating the task execution effect, combined with historical task data H history and the latest trained reinforcement learning policy Iteratively optimize the UAV’s task allocation, path planning, and collaboration strategies; Step 6.5: After completing the task evaluation, analyze the historical task execution data H history With real-time feedback data feedback , and conduct system-level optimization design, specifically: Based on the real-time feedback of the multi-agent cluster task execution effect, the resource allocation strategy, path planning algorithm and collaboration strategy for the next round of tasks are optimized; A dynamic task optimization closed-loop feedback system based on reinforcement learning is constructed. Combining semantic information with dynamic adjustment mechanism, the multi-agent collaboration strategy is optimized through closed-loop feedback. The calculation formula is: M optimized =f optimize (H history ,I feedback ) The optimization results are stored in the cluster knowledge base K cluster : K cluster =K cluster ∪{f optimize }。
Citation Information
Cited By
Autonomous intelligent Internet of Things driven unmanned aerial vehicle cluster self-organization method
CN120523232A
An autonomous intelligent internet of things driven unmanned aerial vehicle cluster self-organizing method
CN120523232B
Multi-robot cooperative measurement path planning method based on reinforcement learning
CN120651248A
Air-ground collaborative unified semantic interoperation method and system
CN120725028A
Mobile robot dynamic obstacle avoidance control method and system based on reinforcement learning
CN120848530A