Target intelligent collaborative identification method based on unmanned aerial vehicle cluster
Through dynamic heterogeneous cluster networking and meta-learning optimization of spectrum perception technology, combined with distributed unsupervised contrastive learning and cascaded feature distillation fusion, the problems of recognition accuracy and communication delay of drone clusters in complex environments are solved, efficient multimodal data fusion and model online evolution are achieved, and the real-time recognition capabilities of military reconnaissance and disaster relief are improved.
Patent Information
- Application Number
- CN202510505940.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-09-19
AI Technical Summary
Traditional drone swarms have low recognition accuracy, high communication latency, and low multimodal data fusion efficiency in complex environments. In addition, their collaborative recognition performance degrades in dynamic mission environments and they are difficult to adapt to lighting changes and background interference.
It adopts dynamic heterogeneous cluster networking, meta-learning optimized spectrum sensing technology, distributed unsupervised contrastive learning, cascaded feature distillation fusion strategy and unsupervised federated incremental learning, combined with spatiotemporal attention mechanism and anti-interference communication to achieve efficient fusion of multimodal data and online evolution of models.
Improve target recognition accuracy by 35% in complex environments, reduce latency to 200ms, enhance system robustness, and adapt to dynamic task priorities and resource constraints.
Smart Images

Figure CN120673277A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of unmanned aerial vehicle (UAV) cluster control and computer vision, and in particular to a method for intelligent collaborative target recognition based on UAV clusters. Background Art
[0002] Drone swarm collaborative identification technology, due to its wide coverage and fast response speed, has important strategic value in military reconnaissance, disaster relief, and traffic monitoring. In the military field, swarm systems can achieve rapid positioning and continuous tracking of high-value targets in wide-area battlefield environments, such as dynamic monitoring of time-sensitive targets such as camouflaged mobile vehicles and temporary missile launchers. Such targets often exhibit multi-scale and strong occlusion characteristics in complex environments. Their characteristics are easily affected by changes in lighting and background interference, resulting in high false alarm rates and poor continuous tracking stability in traditional single-drone identification systems. In addition, during the execution of dynamic tasks, drone swarms are constrained by bottlenecks in onboard computing power and limitations in wireless communication bandwidth, making it difficult to achieve efficient fusion of multi-source heterogeneous data, which seriously restricts the real-time perception and decision-making efficiency of complex battlefield situations.
[0003] To address these challenges, collaborative drone recognition technology based on swarm intelligence has gradually become a research hotspot. This technology, through distributed computing and multi-machine collaboration, can improve recognition accuracy while enhancing system robustness. However, existing solutions generally suffer from three drawbacks: First, traditional task allocation algorithms (such as contract network protocols and auction algorithms) use fixed optimization objectives, making it difficult to adapt to dynamically changing task priorities and resource constraints; second, mainstream recognition models are mostly designed based on single-modal data, making the fusion processing of heterogeneous data such as multispectral and lidar inefficient, with cross-modal feature correlation errors as high as 15%-20%; third, existing collaborative learning frameworks use a periodic global model update strategy, which is prone to model drift in communication-restricted environments, resulting in a swarm's overall recognition performance drop of over 30%.
[0004] Therefore, building a drone cluster intelligent recognition system that supports dynamic resource scheduling, efficient cross-modal fusion and online autonomous evolution can not only break through the limitations of single-machine computing power and perception dimensions, but also effectively solve the problem of collaborative stability in complex confrontation environments. It has important military application value for improving the actual combat effectiveness of unmanned systems in strong interference and high-dynamic scenarios. Summary of the Invention
[0005] The purpose of the present invention is to overcome the defects existing in the above-mentioned background technology. The present invention proposes a target intelligent collaborative recognition method based on drone clusters. Compared with traditional methods, the method improves the recognition accuracy when the target is occluded, and reduces the collaborative recognition delay and communication bandwidth occupancy.
[0006] The present invention adopts the following technical solutions to solve the above technical problems:
[0007] A method for intelligent collaborative target recognition based on drone clusters includes the following steps:
[0008] Step S1: Dynamic heterogeneous cluster networking and federation model initialization, building a heterogeneous drone cluster including perception, computing, and relay nodes, and using meta-learning optimized spectrum sensing technology to establish an anti-interference communication link;
[0009] Step S2: Design a multimodal object recognition backbone network and complete global model initialization through distributed unsupervised contrastive learning and differential privacy federated aggregation;
[0010] Step S3: Dynamic task allocation and multimodal feature extraction, generating a task allocation matrix based on the spatiotemporal attention mechanism, and using deep reinforcement learning to optimize resource scheduling strategies;
[0011] Step S4: Use the deformable hybrid network Transformer-CNN to extract multi-view target features, and trigger cross-node feature matching when the target occlusion rate is greater than 30%;
[0012] Step S5: Cascade feature fusion. A cascade feature distillation fusion strategy CFDF is proposed to perform intra-modal compression and cross-modal gated fusion on multi-source data such as multispectral and lidar.
[0013] Step S6: Anti-interference processing, designing a spatiotemporal countermeasure detection module based on gradient feature analysis, and combining dynamic spectrum sensing technology to achieve anti-interference resilient communication;
[0014] Step S7: federated model evolution and recognition result output, establishing an unsupervised federated incremental learning system, and achieving online model evolution through momentum weighted aggregation and differentiated parameter transmission;
[0015] Step S8: Perform multi-level result verification: geometric verification and spectral verification, and output the final recognition spectrum and confidence assessment.
[0016] Furthermore, the step S1 is specifically as follows:
[0017] Hardware deployment and network construction: configure the perception nodes (visible light cameras and integrated lidars), the computing nodes (edge computing units), and the relay nodes (integrated software-defined radios); establish communication links, initialize the meta-learning spectrum perception strategy network, and perform online meta-training.
[0018] Furthermore, step S2 is specifically as follows:
[0019] Step S21: Design a multimodal target recognition network and construct a backbone network, including a visible light branch, a multispectral branch, and a fusion layer;
[0020] Step S22: Load ImageNet pre-training weights and perform unsupervised contrastive learning. ImageNet is a large visualization database.
[0021] Step S23: Initialize the federated model, load pre-trained weights on each node, and perform domain adaptation fine-tuning;
[0022] Step S24: Use differential privacy federation aggregation and initialize federated learning parameters.
[0023] Furthermore, the step S3 is specifically as follows:
[0024] Step S31: Prioritize the spatiotemporal tasks, build a 3D environment map through LiDAR SLAM, and mark the task area priorities;
[0025] Step S32: construct task urgency code, calculate resource status code, generate spatiotemporal attention weight matrix, and perform spatiotemporal attention weight calculation;
[0026] Step S34: Use the PPO algorithm to train the task allocation strategy for deep reinforcement learning optimization.
[0027] Furthermore, the step S4 is specifically as follows:
[0028] Step S41, deformable feature extraction, inputting a visible light image, performing a deformable convolution operation, and outputting a 1024-dimensional feature vector;
[0029] Step S42: input multispectral data, perform 3D convolution spatiotemporal feature extraction, and output a 512-dimensional feature vector;
[0030] Step S43: Perform multi-view association. When the target occlusion rate is greater than 30%, cross-node feature matching is triggered, and the associated multi-view feature set is output.
[0031] Furthermore, the step S5 is specifically as follows: proposing a cascaded feature distillation fusion strategy CFDF, compressing visible light data and multispectral data, building a gating mechanism, and performing feature fusion.
[0032] Furthermore, the step S6 is specifically as follows:
[0033] Step S61: Design a spatiotemporal adversarial detection module based on gradient feature analysis to monitor the gradient distribution of input data in real time. When an anomaly is detected, record the interference signal characteristics and punish the defense mechanism.
[0034] Step S62: Perform dynamic spectrum sensing and switching, scan the 2.4 GHz / 5.8 GHz frequency bands, and execute frequency band switching decisions.
[0035] Furthermore, step S7 specifically includes: establishing an unsupervised federated incremental learning system, loading global model parameters, performing unsupervised contrastive learning, dynamically adjusting the learning rate, collecting local model updates, and performing momentum weighted aggregation.
[0036] Furthermore, the step S8 is specifically as follows:
[0037] Step S81, multi-level result verification, including geometric consistency verification, spectral feature matching and time series motion trajectory analysis;
[0038] Step S82: Recognition result output, including target recognition map generation and system status monitoring.
[0039] Compared with the prior art, the present invention adopts the above technical solution and has the following beneficial effects:
[0040] (1) By using meta-learning optimized spectrum sensing technology to establish an anti-interference communication link, and combining differential privacy federated aggregation to achieve global model initialization, the problems of high communication latency and insufficient privacy protection in traditional drone clusters are solved.
[0041] (2) Based on the spatiotemporal attention mechanism, the task allocation strategy is optimized and the deformable network Transformer-CNN is used to realize multi-view target feature association in occlusion scenarios, which solves the problems of rigid task allocation and poor occlusion adaptability in traditional methods.
[0042] (3) A cascaded feature distillation fusion strategy CFDF is proposed to perform intra-modal compression and cross-modal gated fusion on multi-source data such as multispectral and lidar; a meta-learning-based anti-interference elastic communication mechanism is designed, combined with spatiotemporal adversarial sample detection and dynamic spectrum sensing technology, which significantly improves the system robustness in complex environments.
[0043] (4) Establish an unsupervised federated incremental learning system to achieve online evolution of model parameters through momentum-weighted aggregation; design a multi-level result verification mechanism including geometric verification, spectral verification, and time series trajectory analysis to ensure the reliability of recognition results and provide high-precision real-time recognition support for military reconnaissance, disaster relief and other scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 It is a schematic diagram of the overall process of the present invention;
[0045] Figure 2 Schematic diagram of multimodal target recognition network;
[0046] Figure 3 Schematic diagram of the deformable hybrid network Transformer-CNN;
[0047] Figure 4 Schematic diagram of building a 3D environment map for LiDAR SLAM;
[0048] Figure 5 Schematic diagram of the gating mechanism;
[0049] Figure 6 Schematic diagram of target recognition results. DETAILED DESCRIPTION
[0050] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0051] This application proposes a method for intelligent collaborative target recognition based on drone clusters. It implements cluster networking and model initialization through a dynamic heterogeneous federated learning architecture. It uses a spatiotemporal attention mechanism to optimize task allocation and a deformable network, Transformer-CNN, to extract multi-view target features. It proposes a cascaded feature distillation and fusion strategy, CFDF, to perform modal compression and cross-modal gated fusion on multi-source data, such as multispectral and lidar data. It designs an anti-interference resilient communication mechanism based on meta-learning, combining spatiotemporal adversarial detection with dynamic spectrum sensing technology to enhance system robustness. It also establishes an unsupervised federated incremental learning system, enabling online model evolution through momentum-weighted aggregation. This method overcomes bottlenecks in traditional technologies, such as rigid task allocation, low multimodal fusion efficiency, and insufficient anti-interference capabilities. In complex environments, target recognition accuracy is increased by 35%, and latency is reduced to 200ms, providing high-precision real-time recognition support for scenarios such as military reconnaissance and disaster relief.
[0052] Specifically, the present invention provides a method for intelligent collaborative target recognition based on a drone cluster, comprising the following steps:
[0053] Step S1: Dynamic heterogeneous cluster networking and federated model initialization. A heterogeneous drone cluster consisting of perception, computing, and relay nodes is constructed. Meta-learning-optimized spectrum sensing technology is used to establish an anti-interference communication link. A multimodal target recognition backbone network is designed. Global model initialization is completed through distributed unsupervised contrastive learning and differential privacy federated aggregation.
[0054] Step S11: Hardware deployment and network construction. Configure sensing nodes, computing nodes, and relay nodes. The sensing nodes include visible light cameras (20 megapixels, 30 fps), the computing nodes include edge computing units (NVIDIA Jetson AGX Orin, 32 TOPS computing power, storage capacity ≥ 1 TB, supporting real-time data processing), and the relay nodes include integrated software-defined radios (SDR, 2.4-5.8 GHz adjustable, maximum transmission rate ≥ 100 Mbps). Establish communication links, initialize the meta-learning spectrum sensing strategy network, and perform online meta-training.
[0055] Step S12: Design a multimodal target recognition network, such as Figure 2 As shown in the figure, a backbone network is constructed, including a visible light branch, a multispectral branch, and a fusion layer. The visible light branch includes a deformable ResNet-18, which dynamically adjusts the convolution kernel offset as follows:
[0056] Δp n =RELU(F deform (I RGB ))
[0057] Where Δp n is the offset of the nth convolution kernel (unit: pixel), RELU is the linear rectification function, F deform For the lightweight CNN offset prediction network, I RGB is the input RGB image.
[0058] The multispectral branch uses 3D ConvNeXt-Tiny and inputs 8-band spectral data (400-1000nm); the fusion layer contains a cross-modal attention gating module, loads ImageNet pre-trained weights, and performs unsupervised contrastive learning.
[0059] Step S13: Initialize the federated model, load pre-trained weights on each node, and perform domain adaptation fine-tuning (learning rate 0.001, batch size 32). Use differential privacy federated aggregation and initialize federated learning parameters (60-300 seconds of adaptation). The expression of differential privacy federated aggregation is as follows:
[0060]
[0061] in, represents the local model parameters of the i-th node, N(0,σ 2 I) represents Gaussian noise, and N represents the number of nodes.
[0062] Step S2: Dynamic task allocation and multimodal feature extraction, generating a task allocation matrix based on the spatiotemporal attention mechanism, such as Figure 3As shown in the figure, deep reinforcement learning is used to optimize resource scheduling strategies; a deformable hybrid network Transformer-CNN is used to extract multi-view target features, and cross-node feature matching is triggered when the target occlusion rate is greater than 30%.
[0063] Step S21: Time and space task priority division, build a three-dimensional environment map (resolution 0.1m) through LiDAR SLAM, such as Figure 4 As shown, the task area priority is marked (high priority: dynamic target appearance area, medium priority: static high-value target area, low priority: background area); construct the task urgency code, calculate the resource status code, generate the spatiotemporal attention weight matrix, and perform spatiotemporal attention weight calculation. The expression of the task urgency code is as follows:
[0064] E task =MLP(Concat[N,V,T])
[0065] In the above formula, MLP represents a multi-layer perceptron, Concat represents an operation that combines multiple tensors, vectors, or feature maps into a larger tensor along a specific dimension (such as channel, spatial, or temporal dimension), N represents the target type, V represents the motion speed, and T represents the target appearance time.
[0066] The expression for calculating resource status code is as follows:
[0067] E resource =MLP(Concat[E,C,M])
[0068] In the above formula, MLP represents a multi-layer perceptron, Concat represents an operation that combines multiple tensors, vectors, or feature maps into a larger tensor along a specific dimension (such as the channel, spatial, or temporal dimension), E represents the remaining battery power, C represents the available computing power, and M represents the communication quality.
[0069] The expression of the spatiotemporal attention weight matrix is as follows:
[0070] A t,s =Sigmoid(MLP(Concat[E task ,E resouce ]))
[0071] In the above formula, Sigmoid represents the Sigmoid activation function, MLP represents the multi-layer perceptron, Concat represents the operation of merging multiple tensors, vectors, or feature maps into a larger tensor along a specific dimension (such as channel, spatial, or temporal dimension), and E task Indicates the task urgency code, E resource Indicates the computing resource status code.
[0072] Use the PPO algorithm to train the task allocation strategy for deep reinforcement learning optimization;
[0073] Step S22, deformable feature extraction, input the visible light image (20 million pixel RGB image (30fps)) collected by the drone cluster, perform deformable convolution operation, and output a 1024-dimensional feature vector; input the 8-band spectral data (400-1000nm) collected by the drone cluster, perform 3D convolution spatiotemporal feature extraction, and output a 512-dimensional feature vector; perform multi-view association, and when the target occlusion rate is greater than 30%, trigger cross-node feature matching, and output the associated multi-view feature set.
[0074] Step S3: Cascade feature fusion and anti-interference processing. A cascaded feature distillation fusion strategy (CFDF) is proposed to perform intra-modal compression and cross-modal gated fusion on multi-source data such as multispectral and lidar. A spatiotemporal adversarial detection module based on gradient feature analysis is designed, combined with dynamic spectrum sensing technology to achieve anti-interference resilient communication.
[0075] Step S31: Propose a cascaded feature distillation fusion strategy CFDF to compress visible light data and multispectral data. The visible light data with a 1024-dimensional feature vector is input and the channel attention mechanism is executed. The expression is as follows:
[0076]
[0077] In the above formula, α c is the attention weight of the cth channel, indicating the importance of the channel in feature fusion, H is the information entropy calculation function, F c is the input feature vector of the cth channel, with a dimension of 1024 (visible light data), C is the total number of channels of the input feature, and exp is an exponential function used to convert the value of information entropy into a positive number to facilitate the calculation of weights.
[0078] Input multispectral data with 512-dimensional feature vectors, perform 3D convolution dimensionality reduction (compression rate 60%), and output 256-dimensional compressed features;
[0079] Step S32: construct a gating mechanism, perform feature fusion, and output 768-dimensional fusion features, such as Figure 5 As shown;
[0080] Step S33: Design a spatiotemporal adversarial detection module based on gradient feature analysis to monitor the gradient distribution of input data in real time. When an anomaly is detected, record the interference signal characteristics and punish the defense mechanism; perform dynamic spectrum sensing and switching, scan the 2.4GHz / 5.8GHz frequency bands, and execute frequency band switching decisions.
[0081] Step S4: federated model evolution and recognition result output, establish an unsupervised federated incremental learning system, realize online model evolution through momentum weighted aggregation and differentiated parameter transmission; perform multi-level result verification (geometric verification, spectral verification), and output the final recognition map and confidence assessment.
[0082] Step S41: Establish an unsupervised federated incremental learning system, load global model parameters, perform unsupervised contrastive learning, and dynamically adjust the learning rate; collect local model updates and perform momentum weighted aggregation;
[0083] Step S42: multi-level result verification, including geometric consistency verification, spectral feature matching, and temporal motion trajectory analysis;
[0084] Step S43: Output the target recognition result, such as Figure 6 shown.
[0085] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A target intelligent collaborative identification method based on drone clusters, characterized by: The method comprises the following steps: Step S1: Dynamic heterogeneous cluster networking and federation model initialization, building a heterogeneous drone cluster including perception, computing, and relay nodes, and using meta-learning optimized spectrum sensing technology to establish an anti-interference communication link; Step S2: Design a multimodal object recognition backbone network and complete global model initialization through distributed unsupervised contrastive learning and differential privacy federated aggregation; Step S3: Dynamic task allocation and multimodal feature extraction, generating a task allocation matrix based on the spatiotemporal attention mechanism, and using deep reinforcement learning to optimize resource scheduling strategies; Step S4: Use the deformable hybrid network Transformer-CNN to extract multi-view target features, and trigger cross-node feature matching when the target occlusion rate is greater than 30%; Step S5: Cascade feature fusion. A cascade feature distillation fusion strategy CFDF is proposed to perform intra-modal compression and cross-modal gated fusion on multi-source data such as multispectral and lidar. Step S6: Anti-interference processing, designing a spatiotemporal countermeasure detection module based on gradient feature analysis, and combining dynamic spectrum sensing technology to achieve anti-interference resilient communication; Step S7: federated model evolution and recognition result output, establishing an unsupervised federated incremental learning system, and achieving online model evolution through momentum weighted aggregation and differentiated parameter transmission; Step S8: Perform multi-level result verification: geometric verification and spectral verification, and output the final recognition spectrum and confidence assessment.
2. The method for intelligent collaborative target recognition based on drone clusters according to claim 1 is characterized in that: The step S1 is specifically as follows: Hardware deployment and network construction: configure the perception nodes (visible light cameras and integrated lidars), the computing nodes (edge computing units), and the relay nodes (integrated software-defined radios); establish communication links, initialize the meta-learning spectrum perception strategy network, and perform online meta-training.
3. The method for intelligent collaborative target recognition based on drone clusters according to claim 1 is characterized in that: Step S2 is specifically as follows: Step S21: Design a multimodal target recognition network and construct a backbone network, including a visible light branch, a multispectral branch, and a fusion layer; Step S22: Load ImageNet pre-training weights and perform unsupervised contrastive learning. ImageNet is a large visualization database. Step S23: Initialize the federated model, load pre-trained weights on each node, and perform domain adaptation fine-tuning; Step S24: Use differential privacy federation aggregation and initialize federated learning parameters.
4. The method for intelligent collaborative target recognition based on drone clusters according to claim 1 is characterized in that: The step S3 is specifically as follows: Step S31: Prioritize the spatiotemporal tasks, build a 3D environment map through LiDAR SLAM, and mark the task area priorities; Step S32: construct task urgency code, calculate resource status code, generate spatiotemporal attention weight matrix, and perform spatiotemporal attention weight calculation; Step S34: Use the PPO algorithm to train the task allocation strategy for deep reinforcement learning optimization.
5. The method for intelligent collaborative target recognition based on drone clusters according to claim 1 is characterized in that: The step S4 is specifically as follows: Step S41, deformable feature extraction, inputting a visible light image, performing a deformable convolution operation, and outputting a 1024-dimensional feature vector; Step S42: input multispectral data, perform 3D convolution spatiotemporal feature extraction, and output a 512-dimensional feature vector; Step S43: Perform multi-view association. When the target occlusion rate is greater than 30%, cross-node feature matching is triggered, and the associated multi-view feature set is output.
6. The method for intelligent collaborative target recognition based on drone clusters according to claim 1 is characterized in that: The step S5 specifically includes: proposing a cascaded feature distillation fusion strategy CFDF, compressing visible light data and multispectral data, building a gating mechanism, and performing feature fusion.
7. The method for intelligent collaborative target recognition based on drone clusters according to claim 1 is characterized in that: The step S6 is specifically as follows: Step S61: Design a spatiotemporal adversarial detection module based on gradient feature analysis to monitor the gradient distribution of input data in real time. When an anomaly is detected, record the interference signal characteristics and punish the defense mechanism. Step S62: Perform dynamic spectrum sensing and switching, scan the 2.4 GHz / 5.8 GHz frequency bands, and execute frequency band switching decisions.
8. The method for intelligent collaborative target recognition based on drone clusters according to claim 1 is characterized in that: The step S7 specifically includes: establishing an unsupervised federated incremental learning system, loading global model parameters, performing unsupervised contrastive learning, dynamically adjusting the learning rate, collecting local model updates, and performing momentum weighted aggregation.
9. The method for intelligent collaborative target recognition based on drone clusters according to claim 1 is characterized in that: The step S8 is specifically as follows: Step S81, multi-level result verification, including geometric consistency verification, spectral feature matching and time series motion trajectory analysis; Step S82: Recognition result output, including target recognition map generation and system status monitoring.
Citation Information
Cited By
Detection mechanism guided multi-mode element learning remote sensing reconnaissance target identification method
CN121074378A
Unmanned aerial vehicle closed-loop intelligent cooperative search and rescue method and device
CN121386880A
Heterogeneous unmanned cluster multi-modal multi-layer task modeling method and device
CN121477647A
Cloud side-end collaborative unmanned aerial vehicle cluster intelligent sensing and decision-making system
CN121477980A
Federal learning optimization method and system for multi-source heterogeneous cardiovascular disease risk prediction
CN121502320A