Multi-unmanned aerial vehicle distributed cluster positioning method based on dynamic segmentation and indoor labels

By adopting a multimodal observation method combining dynamic segmentation and indoor labels in multi-UAV systems, the problem of inadequate positioning of indoor multi-UAVs in the prior art is solved, and a higher precision and stable cluster positioning is achieved.

CN120194682APending Publication Date: 2025-06-24ROBOTICS RESEARCH CENTER OF YUYAO CITY +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510059908.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The existing multi-UAV collaboration technology lacks robustness considerations and accuracy optimization for dynamic objects in indoor environments, resulting in insufficient accuracy and stability in cluster positioning.

Method used

The distributed cluster positioning method of multi-UAV based on dynamic segmentation and indoor labels is adopted. Data processing and multi-modal observation are carried out through sensors such as binocular cameras, IMUs, and UWBs carried by the drone. Combined with visual and loopback detection of indoor labels, optimization problems are constructed and cluster positioning is solved using graph optimization.

Benefits of technology

The robustness of cluster positioning for dynamic objects in the environment is improved, the positioning accuracy and stability in the indoor environment is improved, and the robust indoor cluster positioning with distributed perception and collaborative assistance is realized.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120194682A_ABST
    Figure CN120194682A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of unmanned aerial vehicles, and discloses a multi-unmanned aerial vehicle distributed cluster positioning method based on dynamic segmentation and indoor labels. The method comprises the steps of front-end sensor data processing and binocular VIO-based autolocalization observation, front-end dynamic target tracking based on vision, rear-end loopback detection based on vision and indoor labels, and rear-end cluster localization based on graph optimization and multi-modal observation. S1, front-end sensor data processing and binocular VIO-based autologous positioning observation are carried out; s2, front-end dynamic target tracking based on vision; s3, loopback detection based on vision and indoor labels is carried out at the rear end; and S4, carrying out cluster positioning based on graph optimization and multi-modal observation at the rear end. According to the indoor multi-unmanned aerial vehicle distributed cluster positioning method based on dynamic robustness and prior information, cluster sensing is more subdivided, positioning constraints are richer, and a positioning result is more reliable.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of unmanned aerial vehicles, and particularly relates to a multi-UAV distributed cluster positioning method based on dynamic segmentation and indoor tags. Background Art

[0002] Currently, multi-UAV cooperation technology has a wide range of application spaces in various fields such as viewing, industry, and security. For example, using multi-UAV formations to create aerial landscapes, multi-UAV cooperation to complete the task of carrying objects, and multi-UAV cooperation to capture moving targets such as illegally intruding UAVs and suspicious vehicles. During the implementation of collaborative tasks, UAVs need to plan the optimal UAV trajectories according to the target positions and their own position distributions in the environment, which involves environmental perception, target perception, cluster positioning, target positioning, etc. of cluster UAVs. Generally speaking, cluster positioning plays a decisive role in multi-UAVs performing cluster collaborative tasks, which can help cluster UAVs obtain accurate and real-time self-positioning and other UAV positioning, and improve the accuracy and robustness of subsequent task planning and control. However, the current technical solutions have problems such as lack of robust consideration for dynamic objects within the field of view and lack of accuracy optimization for indoor fixed scenarios. Summary of the Invention

[0003] The purpose of the present invention is to provide a multi-UAV distributed cluster positioning method based on dynamic segmentation and indoor tags to solve the above technical problems.

[0004] To solve the above technical problems, the specific technical solution of a multi-UAV distributed cluster positioning method based on dynamic segmentation and indoor tags of the present invention is as follows:

[0005] A multi-UAV distributed cluster positioning method based on dynamic segmentation and indoor tags, including multiple UAVs, each of which is equipped with a binocular camera, an IMU, a UWB, an on-board computer, and the flight control and power set required for takeoff. The method uses the binocular camera, IMU, and UWB carried by the UAVs to transmit data to the on-board computer and performs the following steps:

[0006] S1: Front-end sensor data processing and self-positioning observation based on binocular VIO;

[0007] S2: Front-end vision-based dynamic target tracking;

[0008] S3: Back-end loop detection based on vision and indoor tags;

[0009] S4: Back-end cluster positioning based on graph optimization and multi-modal observations.

[0010] Furthermore, S1 includes the VIO processing of binocular vision and IMU and the reception of UWB information. Each drone in the cluster is equipped with the same set of sensor devices, including an RGBD camera, an IMU, and a UWB. The sensors transmit data back to the on-board computer at frequencies of 30 - 60 Hz for binocular RGB images, 30 - 60 Hz for depth images, 170 - 200 Hz for the IMU, and 170 - 200 Hz for the UWB. The on-board computer runs the VINS-Fusion algorithm distributively to perform VIO calculation and processing on the binocular images and IMU data, and transmits the local positioning result of the VIO calculation to the backend module, and calculates the self-motion observables of the drone:

[0011]

[0012] The UWB data also needs to be processed into ranging observables:

[0013]

[0014] where i is the local drone number; is the observable notation; P represents [R, t], that is, the 4-degree-of-freedom pose of the drone, t represents the {x, y, z} position information, which is the displacement part of P;

[0015] Both calculated observables are sent to the backend for the cluster positioning calculation of the cluster positioning module based on graph optimization and multi-modal observations in step S4.

[0016] Furthermore, S2 includes vision-based drone object target detection, target tracking, PnP-based target positioning, and dynamic target segmentation based on depth and motion information; based on the binocular image data transmitted back in step S1, sequentially perform 2D bounding box detection of the target based on YOLOv8, target tracking based on MOSSE and the Hungarian algorithm, and 3D coordinate estimation processing based on the combination of PnP and prior information for the drone objects within the field of view. According to the visual observations of the local drone on the tracked and positioned drone objects, calculate the visual relative positioning observables:

[0017]

[0018] This observable is also sent to the backend for the cluster positioning calculation of the cluster positioning module based on graph optimization and multi-modal observations in step S4;

[0019] In addition, according to the verification of depth information and dynamic information, the tracked and positioned drone objects are dynamically segmented, and finally the static map points filtered by the dynamic mask are sent to the cluster positioning module based on graph optimization and multi-modal observations in the backend for constructing multi-drone map point positioning constraints.

[0020] Furthermore, the dynamic segmentation includes the following steps:

[0021] First, use the depth image transmitted back by the RGBD camera to obtain the average depth within the target 2D detection box, compare it with the background depth, and if the difference reaches the threshold, verify it through depth information; secondly, use the local positioning result after VIO processing in step S1 Calculate the local motion of the RGBD camera system between two frames And use it to construct a reprojection error for the motion of the feature points of the target object between two frames:

[0022]

[0023] where F obj j is all the feature point pairs k detected for the target object obj j and having a matching relationship between two frames; is the number of matching point pairs in this feature point set; is the matching pixel of the feature point k of the target object obj j in the target 2D detection at time t - 1; is the imaging plane depth and internal parameter matrix of the local camera system c at time t - 1 t-1 ; is the 3D coordinate of the feature point k of the target object obj j in the local camera system c at time t obtained in the 3D coordinate estimation process combining PnP and prior information t ;

[0024] After calculating the residual obj j compare it with the empirical threshold. If it is greater than the threshold, determine that the object obj j is a dynamic object. Therefore, all the feature point sets F covered under the mask of this object obj j are excluded from the feature point set used for subsequent positioning optimization.

[0025] Furthermore, the S3 includes visual descriptor extraction, indoor label detection, database establishment, and loop closure detection; based on the traditional loop closure detection based on feature points, indoor label information pre - arranged in the environment is added, and the loop closure detection robustness is increased through label number matching; first, establish a remote and local loop closure detection database based on Faiss and WiFi communication, and store the key frame descriptors and original images collected by the remote and local drones respectively. In addition, the database also stores the label ID and corner pixel data returned after indoor label detection; whenever a new key frame enters locally, use the K - nearest neighbor algorithm to compare the feature descriptor information and indoor label ID information of the local new key frame with those in the two databases; if a match between key frames occurs, the set of matching corner 2D pixels of one frame and the 3D coordinates of the corresponding mapped points triangulated in the other frame The outliers are removed and the 3D coordinates are restored through PnP with RANSAC for the set, obtaining Then, the gravity consistency test is performed on the point pairs, and finally, the loop closure observations are constructed using the 3D-3D point pairs that pass the test:

[0026]

[0027] Among them and are the 3D coordinates of the same map point observed by two key frames matched in the loop closure detection in the coordinate systems of the two key frames respectively; the constructed observation is based on the loop closure constraint of the same batch of map points jointly observed by drone i at time t0 and drone j at time t1. Subsequently, this observation will be added to the cluster localization optimization problem in step S4.

[0028] Furthermore, S4 includes constructing an optimization problem, solving the optimization problem, and deriving high-frequency pose output; first, an optimization problem is constructed to solve the 4-degree-of-freedom poses of the key frames of all drones in the sliding window in the cluster where i is the number of the pose reference frame, m is the number of key frames participating in the optimization in the optimization problem, and n is the number of drones participating in the optimization. The optimization problem is as follows:

[0029]

[0030] Among them, O, U, V, and L are the sets of drone numbers and timestamps involved in the VIO motion / UWB ranging / visual observation / loop closure detection constraints; ρ(·) is the Huber loss function used to reduce the influence of outlier residual terms; z is the observation calculated from the actual observations in each constraint, and the observation model constructed based on the optimization variables P and t constitutes the residual term together with it; n~(0,σ 2 ) is the Gaussian noise in their respective observation models; the first summation term constrains the self-motion of each drone over time, and the actual observation data comes from the VIO calculation in step S1; the second summation term constrains the ranging between drones based on UWB, and the data comes from the UWB data of local and remote drones; the third summation term constrains the poses of drones based on visual relative positioning, and the data comes from the visual positioning in step S2; the fourth summation term constrains the loop closure constraint based on feature matching between key frames, and the data comes from the loop closure detection in step S3;

[0031] Secondly, after constructing the optimization problem, the Ceres Solver is used to solve this optimization problem and obtain the 4-degree-of-freedom pose estimation of each drone i at time t in the optimization set

[0032] Finally, using this pose and combining with local drones, high-frequency VIO poses promoted forward by the IMU are utilized to predict high-frequency swarm poses through forward promotion:

[0033]

[0034] This pose is the result of the swarm positioning output, and its frequency is approximately equal to the frequency of the IMU. After completing the high-frequency output of the swarm positioning, this high-frequency swarm pose is output to the subsequent planners and controllers of the drones for subsequent planning and execution.

[0035] A multi-UAV distributed swarm positioning method based on dynamic segmentation and indoor tags of the present invention has the following advantages:

[0036] (1) Through target perception and dynamic segmentation, the present invention improves the robustness of swarm positioning to the characteristics of dynamic objects in the environment.

[0037] (2) Through indoor tag detection and loop detection constraints based on indoor tag IDs and features, the present invention improves the accuracy of swarm positioning in indoor environments.

[0038] (3) Through front-end perception and swarm positioning based on four constraints, the present invention realizes robust indoor swarm positioning with distributed perception and collaborative assistance. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 is a schematic flow diagram of the present invention.

[0040] Figure 2 is a schematic diagram of the graph optimization model for swarm positioning based on graph optimization and multi-modal observations at the backend of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0041] To better understand the purpose, structure, and function of the present invention, the following further describes in detail a multi-UAV distributed swarm positioning method based on dynamic segmentation and indoor tags of the present invention in conjunction with the accompanying drawings.

[0042] Example 1: Prepare several drones, each equipped with a binocular camera, IMU, UWB, on-board computer, and flight control, power set, etc. required for takeoff; start the drones, use the above sensors to transmit data to the on-board computer, and implement the multi-UAV distributed swarm positioning method based on dynamic segmentation and indoor tags, as Figure 1 shown, and perform the following steps:

[0043] S1: Front-end sensor data processing and self-localization observation based on binocular VIO, including VIO processing of binocular vision and IMU and reception of UWB information. Each drone in the cluster is equipped with the same set of sensor devices, including an RGBD camera, an IMU, and a UWB. The sensors transmit data back to the on-board computer at frequencies of 30 - 60Hz for binocular RGB images, 30 - 60Hz for depth images, 170 - 200Hz for IMU, and 170 - 200Hz for UWB. The on-board computer runs the VINS-Fusion algorithm distributively to perform VIO calculation and processing on the binocular images and IMU data, and sends the local positioning result (assuming the local drone number is i) to the back-end module, and calculates the self-motion observation quantity of the drone:

[0044]

[0045] The UWB data also needs to be processed into ranging observation quantities:

[0046]

[0047] Among them, is the observation quantity notation; P represents [R, t], that is, the 4-degree-of-freedom pose of the drone, and t represents the {x, y, z} position information, which is the displacement part of P.

[0048] Both calculated observation quantities are sent to the back-end for the cluster positioning calculation of the cluster positioning module based on graph optimization and multi-modal observation in step S4.

[0049] S2: Vision-based dynamic target tracking at the front-end, including vision-based drone object detection, target tracking, PnP-based target positioning, and dynamic target segmentation based on depth and motion information. Based on the binocular image data transmitted back in step S1, perform processing such as 2D bounding box detection of the drone object in the field of view based on YOLOv8, target tracking based on MOSSE and the Hungarian algorithm, and 3D coordinate estimation based on the combination of PnP and prior information in sequence. According to the visual observation of the local drone on the tracked and positioned drone object, calculate the visual relative positioning observation quantity:

[0050]

[0051] This observation quantity is also sent to the back-end for the cluster positioning calculation of the cluster positioning module based on graph optimization and multi-modal observation in step S4.

[0052] In addition, based on the verification of depth information and dynamic information, dynamic segmentation is performed on the tracked and located UAV objects. Finally, the static map points after dynamic mask filtering are sent to the back-end cluster localization module based on graph optimization and multi-modal observations for constructing multi-UAV map point localization constraints.

[0053] Among them, the dynamic segmentation steps are as follows: First, use the depth image transmitted back by the RGBD camera to obtain the average depth within the target 2D detection box, and compare it with the background depth. If the difference reaches the threshold, it passes the depth information verification; Second, use the local positioning result after VIO processing in step S1 Calculate the local motion of the RGBD camera system between two frames And use it to construct the reprojection error for the motion of the feature points of the target object between two frames:

[0054]

[0055] Among them, F obj j Is all the feature point pairs k detected for the target object obj j and having a matching relationship between two frames; Is the number of matching point pairs in this set of feature points; Is the matching pixel of the feature point k of the target object obj j in the target 2D detection at time t-1; Is the imaging plane depth and internal parameter matrix of the local camera system c t-1 At time t-1; Is the 3D coordinate of the feature point k of the target object obj j under the local camera system c t Obtained in the 3D coordinate estimation process combined with PnP and prior information at time t.

[0056] After calculating the residual obj j Let it be compared with the empirical threshold. If it is greater than the threshold, it is determined that the object obj j is a dynamic object. Therefore, all the feature point sets F obj j Covered by the mask of this object are excluded from the set of feature points used for subsequent localization optimization.

[0057] S3: The backend performs loop detection based on vision and indoor tags, including visual descriptor extraction, indoor tag detection, database establishment, and loop detection. Based on traditional feature-point-based loop detection, information about pre-deployed indoor tags in the environment is added, and the robustness of loop detection is increased through tag number matching. First, a remote and a local loop detection database are established based on Faiss and WiFi communication, storing the key frame descriptors and original images collected by the remote drone and the local drone respectively. In addition, the tag ID and corner pixel data returned after indoor tag detection are also stored in the database. Whenever a new key frame arrives locally, the K-nearest neighbor algorithm is used to compare the new local key frame with the feature descriptor information and indoor tag ID information in the two databases; if a match between key frames occurs, the 2D pixel set of the matching corner points of one frame and the 3D coordinates set of the corresponding mapped points that have been triangulated in the other frame are used for outlier rejection and 3D coordinate recovery through PnP with RANSAC (to obtain ), then a gravity consistency check is performed on the point pairs, and finally, the 3D-3D point pairs that pass the check are used to construct the loop observation quantity:

[0058]

[0059] where and are the 3D coordinates of the same mapped point observed in two key frames that are matched in loop detection in the coordinate systems of the two key frames respectively; the constructed observation quantity is based on the loop constraint of the same batch of mapped points jointly observed between drone i at time t0 and drone j at time t1. Subsequently, this observation quantity is added to the cluster localization optimization problem in step S4.

[0060] S4: The backend performs cluster localization based on graph optimization and multi-modal observations, including constructing an optimization problem, solving the optimization problem, and deriving high-frequency pose outputs. First, an optimization problem is constructed to solve the 4-degree-of-freedom poses of the key frames of all drones in the sliding window in the cluster, where i is the number of the pose reference frame, m is the number of key frames participating in the optimization in the optimization problem, and n is the number of drones participating in the optimization. The optimization problem is as follows:

[0061]

[0062] Among them, O, U, V, and L are sets of UAV numbers and timestamps corresponding to the constraints involved in VIO motion / UWB ranging / visual observation / loop detection. ρ(·) is the Huber loss function used to reduce the influence of outlier residual terms. z is the observed quantity calculated from the actual observations in each constraint, and the observed quantity model constructed based on the optimization variables P and t constitutes the residual term together with it. n~(0,σ 2 ) is the Gaussian noise in their respective observation models. The first summation term constrains the self-motion of each UAV over time, and the actual observed data comes from the VIO calculation in step S1; the second summation term constrains the ranging between UAVs based on UWB, and the data comes from the UWB data of local and remote UAVs; the third summation term constrains the poses of UAVs based on visual relative positioning between UAVs, and the data comes from the visual positioning in step S2; the fourth summation term constrains the loop constraint based on feature matching between key frames, and the data comes from the loop detection in step S3.

[0063] Secondly, after constructing the optimization problem, use the Ceres Solver to solve this optimization problem and obtain the 4-degree-of-freedom pose estimation of each UAV i at time t in the optimization set.

[0064] Finally, using this pose combined with the high-frequency VIO pose forward-propagated by the local UAV through the IMU The high-frequency cluster pose can be predicted by forward propagation:

[0065]

[0066] This pose is the result output by the cluster positioning, and the frequency is approximately equal to the frequency of the IMU (170 - 200Hz).

[0067] After completing the high-frequency output of the cluster positioning, this high-frequency cluster pose is output to the subsequent planners and controllers of the UAVs for subsequent planning and execution.

[0068] The present invention is not limited to the above embodiments. Within the knowledge of those skilled in the art, various changes can be made without departing from the purpose of the present invention. For example: applying it to the indoor robot cluster positioning based on vision, IMU, and UWB for any purpose.

[0069] It will be understood that the present invention is described by way of some embodiments, and those skilled in the art will appreciate that various changes or equivalent substitutions can be made to these features and embodiments without departing from the spirit and scope of the present invention. Additionally, under the teachings of the present invention, these features and embodiments can be modified to adapt to specific circumstances and materials without departing from the spirit and scope of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of the present application belong to the scope protected by the present invention.

Claims

1. A multi-UAV distributed cluster positioning method based on dynamic segmentation and indoor tags, comprising multiple UAVs, each of which is equipped with a binocular camera, an IMU, a UWB, an onboard computer, and a flight control and power set required for takeoff, characterized in that: The method uses the binocular camera, IMU and UWB carried by the drone to transmit data to the onboard computer and performs the following steps: S1: Front-end sensor data processing and self-positioning observation based on binocular VIO; S2: Front-end vision-based dynamic target tracking; S3: Backend loop detection based on vision and indoor tags; S4: Backend cluster positioning based on graph optimization and multimodal observation.

2. The multi-UAV distributed cluster positioning method based on dynamic segmentation and indoor tags according to claim 1 is characterized in that: S1 includes binocular vision and IMU VIO processing and UWB information reception. Each drone in the cluster is equipped with the same set of sensor equipment, including RGBD camera, IMU and UWB. The sensor transmits data back to the onboard computer at a frequency of 30-60Hz for binocular RGB images, 30-60Hz for depth images, 170-200Hz for IMU and 170-200Hz for UWB. The onboard computer runs the VINS-Fusion algorithm in a distributed manner to perform VIO calculations on the binocular images and IMU data, and sends the local positioning results calculated by VIO to the onboard computer. Send it to the backend module and calculate the drone's self-motion observation: UWB data also needs to be processed into ranging observations: Among them, i is the machine number; is the symbol of the observation quantity; P represents [R, t], i.e. the 4-DOF posture of the UAV, and t represents the {x, y, z} position information, which is the displacement part of P; The two calculated observation quantities are sent to the back end for cluster positioning calculation of the cluster positioning module based on graph optimization and multimodal observation in the back end in step S4.

3. The multi-UAV distributed cluster positioning method based on dynamic segmentation and indoor tags according to claim 1 is characterized in that: The step S2 includes vision-based drone object target detection, target tracking, PnP-based target positioning, and dynamic target segmentation based on depth and motion information; based on the binocular image data returned by step S1, the drone objects within the field of view are sequentially subjected to YOLOv8-based target 2D bounding box detection, MOSSE-based and Hungarian algorithm-based target tracking, and 3D coordinate estimation based on PnP combined with prior information; and the visual mutual positioning observation is calculated based on the visual observation of the local drone on the tracked and positioned drone object: This observation is also sent to the backend for cluster positioning calculation by the cluster positioning module based on graph optimization and multimodal observation in step S4; In addition, the tracked and located drone objects are dynamically segmented based on the verification of depth information and dynamic information, and finally the static map points after dynamic mask filtering are sent to the back-end cluster positioning module based on graph optimization and multimodal observation to construct multi-UAV map point positioning constraints.

4. The multi-UAV distributed cluster positioning method based on dynamic segmentation and indoor tags according to claim 3 is characterized in that: The dynamic segmentation comprises the following steps: First, the depth image sent back by the RGBD camera is used to obtain the average depth of the target 2D detection frame, and compared with the background depth. If the difference reaches the threshold, the depth information is verified. Secondly, the local positioning result after VIO processing in step S1 is used. Calculate the local motion of the RGBD camera between 2 frames And use it to construct the reprojection error for the motion of the target object feature points between 2 frames: Among them, F objj are all feature point pairs k detected for the target object objj and that have formed a matching relationship between the two frames; Then is the number of matching point pairs of the feature point set; is the matching pixel of the feature point k of the target object objj at time t-1 in the target 2D detection; is the camera system of the machine at time t-1 t-1 The imaging plane depth and the intrinsic parameter matrix; is the camera coordinates of the local camera at time t obtained in the 3D coordinate estimation process based on PnP and prior information. t The 3D coordinates of the feature point k of the target object objj; Calculate the residual objj Then, compare it with the empirical threshold. If it is greater than the threshold, the object objj is judged to be a dynamic object. Therefore, the set of all feature points covered by the mask of the object is F objj Eliminate the feature point set used for subsequent positioning optimization.

5. The multi-UAV distributed cluster positioning method based on dynamic segmentation and indoor tags according to claim 1 is characterized in that: The S3 includes visual descriptor extraction, indoor label detection, database establishment and loop detection; based on the traditional feature point-based loop detection, indoor label information pre-arranged in the environment is added, and the robustness of loop detection is increased by label number matching; first, a remote and local loop detection database is established based on Faiss and WiFi communication, and the key frame descriptors and original images collected by the remote drone and the local drone are stored respectively. In addition, the database also stores the label ID and corner pixel data returned after indoor label detection; whenever a new key frame enters the local area, the K nearest neighbor algorithm is used to compare the local new key frame with the feature descriptor information and indoor label ID information in the two databases; if a match is generated between the key frames, the matching corner 2D pixels of one of the frames are compared. The set of 3D coordinates of the corresponding map points triangulated from another frame The set is subjected to PnP with RANSAC to remove outliers and recover 3D coordinates, and we get Then perform a gravity consistency check on the point pairs, and finally use the 3D-3D point pairs that pass the check to construct the loop closure observation: in and They are the 3D coordinates of the same map point observed in the two key frames matched in the loop detection in the two key frame coordinate systems; the constructed observations Based on the loop constraint that drone i at time t0 and drone j at time t1 observe the same batch of map points, this observation is subsequently added to the cluster positioning optimization problem in step S4.

6. The multi-UAV distributed cluster positioning method based on dynamic segmentation and indoor tags according to claim 1 is characterized in that: S4 includes constructing an optimization problem, solving the optimization problem, and deriving a high-frequency pose output; first, constructing an optimization problem to solve the 4-DOF pose of the key frames of all drones in the cluster within the sliding window. Where i is the number of the pose reference system, m is the number of key frames involved in the optimization problem, and n is the number of drones involved in the optimization. The optimization problem is as follows: Among them, O, U, V, L are the drone numbers and timestamps involved in the corresponding VIO motion / UWB ranging / visual observation / loop detection constraints; ρ(·) is the Huber loss function used to reduce the influence of outlier residual terms; z is the observation calculated from the actual observation in each constraint, and the observation model constructed based on the optimization variables P and t constitutes the residual term together with it; n~(0,σ 2 ) is the Gaussian noise in each observation model; the first summation term constrains the self-motion of each drone over time, and the actual observation data comes from the VIO calculation in step S1; the second summation term constrains the distance measurement between drones based on UWB, and the data comes from the UWB data of local and remote drones; the third summation term constrains the posture of drones based on visual mutual positioning, and the data comes from the visual positioning in step S2; the fourth summation term constrains the loop constraint between key frames based on feature matching, and the data comes from the loop detection in step S3; Secondly, after constructing the optimization problem, the Ceres Solver is used to solve the optimization problem and obtain the 4-DOF pose estimation of each drone i in the optimization set at time t. Finally, this pose is combined with the high-frequency VIO pose promoted by the local drone through the IMU Forward promotion predicts high-frequency cluster poses: This pose is the result of cluster positioning output, and its frequency is approximately equal to the frequency of IMU. After completing the high-frequency output of cluster positioning, the high-frequency cluster pose is output to the subsequent planner and controller of the drone for subsequent planning and execution.