Target positioning and detecting method based on cooperative work of multiple unmanned aerial vehicles
Through the collaborative work of multiple drones and deep learning technology, the acquisition and fused image data by multiple cameras is used to solve the real-time and accuracy of target positioning and detection in complex environments, especially in small-object detection, which significantly improves detection accuracy and robustness.
Patent Information
- Application Number
- CN202510364277.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-06-17
AI Technical Summary
The prior art has problems with real-time and accuracy of data transmission in the target positioning and detection of collaborative work of multiple drones in complex environments, and the YOLOv8 algorithm has limited performance in the detection of small targets of drones.
Through the collaborative work of multiple drones, multi-camera collects image data from different perspectives in real time, and features are extracted and fusion through deep convolutional neural networks and cross-attention mechanisms, multi-scale object detection is performed in combination with feature pyramid networks, and finally consensus algorithm decision fusion is carried out in the decision fusion module.
It significantly improves the accuracy and robustness of target recognition, improves the detection accuracy of small targets, and ensures that high efficiency and high accuracy target recognition capabilities can still be maintained in a changing environment.
Smart Images

Figure CN120164133A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of target recognition and positioning, and particularly to a target positioning and detection method based on multi-UAV collaborative work. Background Art
[0002] The research on multi-UAV collaborative work in the field of target positioning and detection aims to improve the accuracy and efficiency of target detection and positioning through the collaborative operation of multiple UAVs. The UAV system is usually equipped with a variety of sensors, such as cameras, lidar, infrared sensors, etc. These sensors process data through fusion technology to help UAVs achieve accurate target positioning and tracking in complex environments. Multiple UAVs share information and coordinate through a wireless communication network, and can cooperate to conduct real-time monitoring and target recognition in a dynamic environment. For example, multiple UAVs can share the monitoring tasks in different areas, and work together to expand the detection range and improve the positioning accuracy.
[0003] The key challenges in this field include how to ensure the real-time and accuracy of data transmission, especially in the case of large communication delays or limited bandwidth. In addition, the complexity and dynamics of the environment also pose high requirements for multi-UAV collaborative work. How to address these challenges and maintain the stability and efficiency of the system is a hot issue in current research. In this context, deep learning and artificial intelligence technologies are introduced to improve the accuracy of target detection and recognition, especially in complex scenarios.
[0004] The multi-UAV collaborative target positioning method based on distance measurement is a technology that estimates the target position by using the distance between multiple UAVs and the target and the coordinates of the UAV station. This method overcomes the problem that the positioning effect drops sharply when the station error cannot be ignored, and optimizes the formation by the Geometric Dilution of Precision (GDOP) to further improve the positioning effect. In the specific implementation process, at least 3 UAVs are required to collaborate on target positioning. Each UAV is equipped with a GPS and a ranging sensor, and flies around the same target. The target is roughly located below the center point of the formation. The UAV can obtain its own position coordinates by using the GPS, and measure the distance from itself to the target by using the ranging sensor. By assuming time synchronization and target movement, data with similar timestamps are selected to complete the target coordinate calculation. The key advantage of this method is that it can use the collaborative work of multiple UAVs to improve the accuracy of target positioning. In a complex battlefield environment, this technology can effectively reduce the error generated by single-UAV positioning and improve the overall positioning performance through swarm intelligence. In addition, this method also considers various factors in practical applications, such as the station error of the UAV and the mobility of the target, making the positioning result more reliable.
[0005] However, the multi-UAV cooperative target localization method based on distance measurement has some limitations in practical applications. First, when there are random errors in the measurement of the observation point coordinates, its localization performance will drop sharply. In addition, this method does not fully consider the coordinate measurement error of the observation aircraft. When this error is large, the localization performance will also drop sharply. The geometric dilution of precision (GDOP) is closely related to the geometric distribution among observation points. Different distributions will cause changes in the GDOP value, thereby affecting the localization accuracy. These factors limit the application effect of this technology in complex environments.
[0006] The YOLOv8 algorithm is a shining star in the field of object detection, demonstrating excellent performance in multiple object detection tasks with its efficient real-time detection ability. In the application of object detection from the perspective of drones, the YOLOv8 algorithm is used to build a drone object detection system. This system combines PyQt5 to build a user interface and is developed using Python3. It is trained and optimized for the drone object dataset, which contains rich drone object image samples, providing strong guarantees for the accuracy and generalization ability of the model. Through deep learning technology, the model can automatically extract the features of drone objects and classify and identify them. The PyQt5 interface design is simple and intuitive, facilitating user operation and real-time viewing of detection results.
[0007] The application of the YOLOv8 algorithm in drone object detection not only improves the level of drone object recognition but also provides strong support for drone object protection, having important theoretical application value. The core of this technology lies in its ability to quickly and accurately detect objects from the images captured by drones, providing real-time feedback for the operation and control of drones. The application of this technology makes drones more efficient and accurate when performing tasks such as surveillance and reconnaissance.
[0008] However, the YOLOv8 algorithm still faces some challenges in object detection from the perspective of drones. First, there are problems such as small object scales and susceptibility to environmental interference in drone object detection, which limit the object detection performance of the algorithm. Second, the sample quality of the small object dataset is not as good as that of general datasets, which affects the training effect and detection accuracy of the model. In addition, the YOLOv8 algorithm performs poorly in dealing with small objects and needs to be further optimized to improve the detection ability for small objects. These challenges need to be overcome through algorithm improvement and technological innovation. Summary of the Invention
[0009] The technical problem to be solved by the present invention is to provide a target positioning and detection method based on the collaborative work of multiple unmanned aerial vehicles (UAVs) in view of the deficiencies of the above-mentioned prior art. By using the multiple cameras of multiple UAVs, image data from different perspectives are collected in real time, and these data are processed through collaborative computing to improve the efficiency and reliability of target recognition. By fusing multi-perspective information, the accuracy and robustness of target recognition are enhanced. At the same time, the performance limitations of the YOLOv8 algorithm in the detection of small targets by UAVs are also solved. The detection accuracy of small targets is improved by improving the algorithm to ensure high-efficiency and high-accuracy target recognition capabilities in a changing environment.
[0010] To solve the above technical problems, the technical solutions adopted by the present invention are as follows: A target positioning and detection method based on the collaborative work of multiple UAVs, through a system including a UAV cluster, a data transmission module, a data processing module, and a decision fusion module, collects target images captured by multiple cameras of the UAV cluster and transmits them to the data processing module. A deep convolutional neural network is used to extract features from the images collected by each camera to obtain the feature maps of the images; the VGG is selected as the convolutional neural network architecture for the deep convolutional neural network; the convolutional neural network extracts important feature information in the images through convolutional layers, pooling layers, and activation functions; after each image is processed by the convolutional neural network, a corresponding feature map is generated, and the feature map includes the spatial information and semantic information of the image; the feature maps from different perspectives are fused through a designed cross-attention mechanism, and a feature pyramid network is combined for multi-scale target detection and target positioning. Finally, in the decision fusion module, the positioning and detection results from different UAVs are combined, and a consensus algorithm is used for decision fusion. The fused target recognition results are post-processed to output the final recognition results, including the category, location, and confidence of the target.
[0011] Further, the process of collecting the target images is as follows: The UAVs fly in a specified area and use multiple cameras to collect target images from different angles in real time; the cameras of each UAV work synchronously to ensure the time consistency of the data; the collected image data are first stored in the local storage of the UAVs, and then preliminary image preprocessing is performed, including image smoothing and filtering, enhancement, image scaling and normalization, and color space conversion; each UAV is equipped with a GPS and a ranging sensor, which can obtain its own position coordinates and the distance data from the target to ensure the accuracy and real-time nature of the data.
[0012] Further, in the process of transmitting to the data processing module, ZeroMQ is used as the message passing framework; in the data transmission process, a time synchronization protocol is adopted.
[0013] Further, the specific process of fusing feature maps from different perspectives through the designed cross-attention mechanism is as follows: By calculating the correlation between features from different perspectives, information is adaptively selected and fused; feature fusion includes three steps: The first step is feature mapping, which maps the feature maps of each perspective to the same feature space; for the feature maps F1 and F2 of two different perspectives, the linear mapping is expressed as: ; ; where W1 and W2 are the corresponding linear transformation weight matrices, and b1 and b2 are bias terms, and are the feature maps after mapping respectively; The second step is to calculate the correlation, calculate the dot product between the main perspective feature map and the feature maps of other perspectives to generate an attention weight matrix; The dot product calculation formula is as follows: ; where, is the calculated dot product; C is the number of channels; and are the values of the main perspective and the i-th perspective at the k-th channel respectively; Then calculate the attention weight matrix, perform the Softmax operation on the dot product values of all perspectives to obtain the attention weight , as shown in the following formula: ; The third step is weighted fusion, using the attention weights to weight the feature maps to generate the fused feature map : ; where N represents the total number of perspectives.
[0014] Further, in the object detection, the fused feature map is input into the feature pyramid network for multi-scale object detection; for each layer of feature map F i perform 1x1 convolution to obtain a new feature map: ; Then, through the upsampling operation, the high-level feature map and the low-level feature map are fused to obtain multi-scale features: ; Upsample represents upsampling, which processes the high-level features to facilitate addition with the low-level features; Based on the Feature Pyramid Network, multiple detection heads are designed to process feature maps of different scales respectively, and output the bounding boxes and class information of the targets, as follows: Bounding box regression: ; Class prediction: ; Among them, 、 、 and represent the bounding box, class label, offset predictor of the bounding box, and classifier respectively.
[0015] Furthermore, in the target positioning, triangulation or multilateration algorithms are used to calculate the precise position of the target according to the distance between the UAV and the target; According to the distance between the UAV position and the target, calculate the precise position of the target through triangulation; Assume the UAV U is at position , the target T is at position , and the distance from the target is d u ; Use the following triangulation formula to calculate the position of the target: ; ; Among them, θ is the azimuth angle between the UAV and the target, which is further calculated according to other known angle measurements or through multilateration; Then, the geometric dilution of precision is used to evaluate the positioning accuracy, and the position of the UAV is adjusted as needed to optimize the geometric dilution of precision, as follows: ; ; Among them, 、 、 、 represent the abscissa of the target predicted position, the abscissa of the target actual position, the ordinate of the target predicted position, and the ordinate of the target actual position respectively; The complete intersection over union loss for bounding box regression calculates the relative position of the target according to the overlap degree, center point distance, and aspect ratio of the predicted box and the ground truth box, as follows: ; Among them, represents the weighted sum of the center point distance and aspect ratio loss; α 、 β 、 are weight coefficients used to control the influence of each loss; IoU、 , respectively represent the bounding box regression loss, the Euclidean distance between the center point of the predicted box and the center point of the ground truth box, and the aspect ratio loss.
[0016] Further, after the target positioning, error correction is performed to identify and correct the positioning errors caused by sensor errors, environmental factors, or changes in the UAV position; a filtering algorithm is used to smooth the positioning data.
[0017] Further, the specific method for performing decision fusion using the consensus algorithm is as follows: In the global consensus, each agent ai updates its state at time step ak Based on the information of its neighbors, this update is based on the way of weighted average or weighted sum, and the formula is as follows: ; where, is the state of the ai-th agent at time step ak; N ai is the neighbor set of agent ai, representing other agents that have communication connections with agent ai; w aiaj is the communication weight between agent ai and agent aj, representing the information transfer strength from aj to ai, and ; is the information of agent ai itself to ensure that the agent can retain its own state when there are no neighbors; In the local consensus, the agent updates its own state according to the information of its neighbors, which is described by the following local weighted sum formula: ; where, α is the adjustment factor, which controls the update step size; the meaning of this formula is that agent ai updates its own state according to the information of neighbor aj by adjusting α to control the consensus convergence speed; To ensure that all agents reach an agreement, a consistency constraint condition is required: ; where, is the consistency tolerance, indicating that the state difference between agents should be less than a certain threshold; this condition ensures that after multiple iterations, the states of each agent converge to a consistent value; The convergence condition is represented by the number of iterations and the change in the state difference: ; where, is the set tolerance error, indicating that the state difference between two agents must be less than a certain value, indicating that the convergence is completed; Finally, the goal is to make the states of all agents tend to be consistent after a certain number of iterations, so as to achieve global consensus; the global consensus goal is: ; That is, all agents finally tend to be consistent in state at time step .
[0018] The beneficial effects of adopting the above technical solution are as follows: The target positioning and detection method based on multi-UAV collaborative work provided by the present invention equips a UAV cluster with multiple cameras to collect target images in real time from different perspectives, overcoming the limitations of traditional single-perspective methods in complex environments, and significantly improving the accuracy and robustness of target recognition; introducing a cross-attention mechanism to adaptively fuse feature maps from different perspectives can effectively capture the correlation between different perspectives, enhance the expression ability of the main perspective feature map, and thus improve the performance of target detection; combining a feature pyramid network for multi-scale target detection can effectively learn features at different scales, especially in the recognition of small targets, and significantly improve the detection ability; the present invention not only improves the accuracy of target recognition and positioning, but also provides effective technical support for the application of UAVs in fields such as military reconnaissance and outdoor search and rescue, and has broad market prospects and application value. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 is the overall model framework provided by the embodiment of the present invention; Figure 2 is the flowchart of the target positioning and detection method based on multi-UAV collaborative work provided by the embodiment of the present invention; Figure 3 is the schematic diagram of the change of loss value without using the fusion mechanism provided by the embodiment of the present invention; Figure 4 is the car inference result diagram of the addition mechanism provided by the embodiment of the present invention; Figure 5 is the bicycle inference result diagram of the addition mechanism provided by the embodiment of the present invention; Figure 6 is the pedestrian inference result diagram of the addition mechanism provided by the embodiment of the present invention; Figure 7 is the schematic diagram of the change of loss value of the addition mechanism provided by the embodiment of the present invention; Figure 8 is the schematic diagram of the change of loss value of the multiplication mechanism provided by the embodiment of the present invention; Figure 9 is the module performance comparison result diagram provided by the embodiment of the present invention; Figure 10This is the comparison chart of the performance of each model provided by the embodiments of the present invention. Detailed implementation manners
[0020] The following will further describe in detail the specific implementation manners of the present invention in conjunction with the accompanying drawings and embodiments. The following embodiments are used to illustrate the present invention, but are not used to limit the scope of the present invention.
[0021] This embodiment relates to a multi-view collaborative computing target recognition method based on unmanned aerial vehicles (UAVs), aiming to improve the accuracy and robustness of target recognition in complex environments. This method uses multiple cameras of multiple UAVs to collect image data from different perspectives in real time, and processes these data through collaborative computing to improve the efficiency and reliability of target recognition.
[0022] The overall framework of the model in this embodiment is as Figure 1 shown. The system structure of this embodiment mainly includes the following parts: UAV cluster: Multiple UAVs, each equipped with multiple cameras, responsible for collecting image data from different perspectives in real time; Data transmission module: Realize data transmission and information sharing between UAVs through a reliable communication protocol (such as ZeroMQ); Data processing module: Perform preliminary processing on the collected image data, including feature extraction, feature fusion, and target detection; Decision fusion module: Fuse data from different UAVs on the central node to perform final target recognition and decision-making.
[0023] The method of this embodiment mainly includes: Multi-view data collection: Use multiple cameras of the UAV cluster to collect target images from different angles in real time, overcoming the limitations of a single perspective; Feature extraction: Use a deep convolutional neural network (CNN) to extract features from the images collected by each camera to obtain the feature maps of the images; Cross-attention mechanism: Design a cross-attention mechanism to fuse the feature maps from different perspectives to enhance the expression ability of the main perspective feature maps; Multi-scale target detection: Combine a feature pyramid network (FPN) for multi-scale target detection to improve the recognition and positioning ability of small targets.
[0024] As Figure 2 shown, the specific implementation process is as follows.
[0025] Step 1: Data collection.
[0026] The drone flies within the specified area and uses multiple cameras to collect target images in real time from different angles. The cameras of each drone work synchronously to ensure the temporal consistency of the data. The collected image data is first stored in the local storage of the drone and then undergoes preliminary image preprocessing, including operations such as image denoising and enhancement, to improve the effect of subsequent feature extraction. And each drone is equipped with a GPS and a ranging sensor, which can obtain its own position coordinates and the distance data to the target, ensuring the accuracy and real-time nature of the data.
[0027] Step 2: Data transmission.
[0028] ZeroMQ is a high-performance asynchronous messaging library designed for use in distributed or parallel applications. It provides a message queue, but different from message-oriented middleware, the ZeroMQ system can run without a dedicated message broker. Based on this feature of ZeroMQ, the present invention uses ZeroMQ as the messaging framework to ensure fast and reliable transmission of image data and intermediate processing results between drones. During the data transmission process, a time synchronization protocol is adopted to ensure that the data transmission between drones is synchronous, avoiding recognition errors caused by time differences.
[0029] Step 3: Feature extraction.
[0030] In the data processing module, a deep convolutional neural network is used to extract features from the images collected by each camera. VGG is selected as the convolutional neural network architecture, and important feature information in the images is extracted through convolutional layers, pooling layers, and activation functions. After each image is processed by the convolutional neural network, a corresponding feature map is generated, and the feature map contains the spatial information and semantic information of the image. Assuming the input image is I, the operations of the convolutional layer, pooling layer, and activation function of the network can be expressed as: ; where I is the input image, with dimensions H×W×C in (height, width, and number of channels). represents the convolution operation. The convolutional layer convolves the input image through the filter W k to generate a feature map.
[0031] ; where * is the convolution operation, W k is the convolution kernel, b k is the bias, and the generated output is a feature map with dimensions . represents the activation function operation, and usually the ReLU activation function is applied to the elements in each feature map. Represents a pooling operation, usually max pooling or average pooling. The role of pooling is to reduce the spatial dimension (such as reducing the height and width of the feature map).
[0032] ; Among them, the pooling operation calculates the maximum or average value for each region to obtain a feature map with a smaller size. The finally generated feature map F contains the spatial information and semantic information of the image, and the dimension is .
[0033] Step 4: Feature fusion.
[0034] Design and implement a cross-attention mechanism to fuse the main perspective feature map with the feature maps of other perspectives. This mechanism adaptively selects and fuses information by calculating the correlation between features of different perspectives. Feature fusion includes three steps: The first step is feature mapping, which maps the feature maps of each perspective to the same feature space. For two feature maps F1 and F2 of different perspectives, the linear mapping is expressed as: ; ; Among them, W1 and W2 are the corresponding linear transformation weight matrices, b1 and b2 are the bias terms, and are the feature maps after mapping respectively.
[0035] The second step is to calculate the correlation, calculate the dot product between the main perspective feature map and the feature maps of other perspectives, and generate an attention weight matrix.
[0036] The dot product calculation formula is as follows: ; Among them, is the calculated dot product; C is the number of channels; and are the values of the main perspective and the i-th perspective in the k-th channel respectively.
[0037] Then calculate the attention weight matrix, perform the Softmax operation on the dot product values of all perspectives to obtain the attention weight , as shown in the following formula: .
[0038] The third step is weighted fusion, using the attention weight to weight the feature map to generate the fused feature map : ; Among them, N represents the total number of perspectives.
[0039] Step 5: Object detection.
[0040] The object detection inputs the fused feature map into the Feature Pyramid Network for multi-scale object detection. The Feature Pyramid Network can learn features at different levels and improve the detection ability for small objects.
[0041] Perform a 1x1 convolution on each layer of the feature map F i to obtain a new feature map: ; Then, through the upsampling operation, fuse the high-level feature map with the low-level feature map to obtain multi-scale features: ; where Upsample represents upsampling, which processes the high-level features for addition with the low-level features.
[0042] Based on the Feature Pyramid Network, design multiple detection heads to process feature maps at different scales respectively, and output the bounding box and class information of the object as follows: Bounding box regression: ; Class prediction: ; where , , and represent the bounding box, class label, bounding box offset predictor, and classifier respectively.
[0043] Step 6: Object localization.
[0044] During localization, use the triangulation or multilateration algorithm to calculate the exact position of the object based on the position of the UAV and the distance to the object. Calculate the exact position of the object through triangulation according to the UAV position and the distance to the object; assume the UAV U is at position , the object T is at position , and the distance to the object is d u ; Use the following triangulation formula to calculate the position of the object: ; ; where θ is the azimuth angle between the UAV and the object, which is further calculated based on other known angle measurements or through multilateration.
[0045] Then, the geometric dilution of precision (GDOP) is used to evaluate the positioning accuracy, and the position of the UAV is adjusted as needed to optimize the GDOP, as shown in the following formula: ; ; Wherein, 、 、 、 respectively represent the abscissa of the predicted target position, the abscissa of the actual target position, the ordinate of the predicted target position, and the ordinate of the actual target position.
[0046] The complete intersection over union loss for bounding box regression calculates the relative position of the target based on the overlap degree, center point distance, and aspect ratio of the predicted box and the ground truth box, as shown in the following formula: ; Wherein, represents the weighted sum of the center point distance and the aspect ratio loss; α 、 β 、 are weight coefficients used to control the influence of each loss. IoU, 、 respectively represent the bounding box regression loss, the Euclidean distance between the center point of the predicted box and the center point of the ground truth box, and the aspect ratio loss.
[0047] Finally, error correction is performed to identify and correct the positioning errors caused by sensor errors, environmental factors, or UAV position changes. A filtering algorithm is used to smooth the positioning data to improve the stability and accuracy of positioning.
[0048] Assume that the true position of the target is and the estimated position is , the goal of error correction is to reduce the error of . Generally, the sources of errors include sensor noise, environmental factors, etc. To perform error correction, the following error correction formula is used: ; Wherein, is the target position after correction. K is a gain coefficient, usually calculated by methods such as Kalman filtering.
[0049] To further eliminate the positioning error, a smoothing algorithm is used to process the positioning data. The smoothing method used in this embodiment is low-pass filtering, specifically a simple moving average, as shown in the following formula: ; Wherein, N is the size of the sliding window, is the smoothed target position, representing the positions of each point within the sliding window.
[0050] Step 7: Decision fusion.
[0051] At the central node, combining the positioning and detection results from different drones, a consensus algorithm is used for decision fusion. This algorithm ensures that, with the support of data from multiple perspectives, the accuracy of recognition is improved. The fused target recognition results are post-processed to output the final recognition results, including the category, position, and confidence level of the target.
[0052] In global consensus, each agent ai updates its state at time step ak based on the information of its neighbors. This update is based on weighted average or weighted sum, and the formula is as follows: ; where, is the state of the ai-th agent at time step ak (such as target position, speed, etc.). N ai is the set of neighbors of agent ai, representing other agents that have a communication connection with agent ai. w aiaj is the communication weight between agent ai and agent aj, representing the strength of information transfer from aj to ai, and . is the information of agent ai itself, to ensure that the agent can retain its own state when there are no neighbors.
[0053] In local consensus, the agent updates its state based on the information of its own neighbors, which is described by the following local weighted sum formula: ; where, α is the adjustment factor, controlling the update step size. The meaning of this formula is that agent ai updates its state according to the information of neighbor aj, by adjusting α to control the consensus convergence speed.
[0054] To ensure that all agents reach an agreement, a consistency constraint condition is usually required: ; where, is the consistency tolerance, indicating that the state difference between agents should be less than a certain threshold. This condition ensures that after multiple iterations, the states of each agent converge to a consistent value.
[0055] Consensus algorithms usually need to ensure that the states of all agents ultimately converge to a common value. The convergence condition is represented by the number of iterations and the change in state difference: ; Among them, is the set tolerance error, indicating that the state difference between two agents must be less than a certain value, indicating that convergence is completed.
[0056] Ultimately, the goal is to make the states of all agents tend to be consistent after a certain number of iterations, so as to achieve global consensus; the global consensus goal is: ; That is, all agents ultimately tend to be consistent in state at time step .
[0057] To compare with the algorithm designed in this embodiment, a method that does not use any fusion mechanism is given, and it is also tested on the MDMT dataset. The learning rate set in the experiment is 0.01, and the training batch size is set to 64. In this way, during training, the model can make full use of all available video memory for training. As Figure 3 shown, the loss value of the training tends to be stable after about 32 rounds. It can be seen that the model has reached its best state that it can express at this time. At this time, the model is tested using the evaluation index of mean Average Precision (mAP), and the results are shown in Table 1.
[0058] Table 1 Model performance without using the fusion mechanism Category Precision Recall mAP50 Car 0.703 0.641 0.667 Bicycle 0.203 0.052 0.109 Person 0.505 0.106 0.299 To better evaluate the model of this embodiment, it is also tested if only addition is used as the method of multi-view image fusion. The experimental parameters set are the same as those of the algorithm without using the fusion mechanism. The same index mAP is used for testing, and the test results are shown in Table 2.
[0059] Table 2 Model performance using the addition fusion mechanism Category Precision Recall mAP50 Car 0.082 0.011 0.001 Bicycle 0.003 0.000 0.000 Person 0.011 0.001 0.000 As shown in Table 2, the algorithm using the addition mechanism has a very poor effect on the multi-view object detection MDMT dataset. Figure 4 、 Figure 5 and Figure 6 selected a set of pictures for testing, and the detection effect is as Figure 7As shown. By combining the input images from these two perspectives, it can be reasonably speculated that directly fusing the data from these two perspectives through addition will cause chaos in the data from the two perspectives. The model will not be able to distinguish which image the features in the added data come from, and the targets of the two images will be superimposed, ultimately resulting in the deviation of the target. Therefore, it can be speculated that the perspective fusion method using this approach is unreliable.
[0060] Then, the cross-attention based on multiplication was tested. First, in the experiment, the learning rate was set to 0.01. However, according to the change of the loss value obtained during training, the following trend of change occurred as shown in the figure. In the first few rounds of training, the loss value rapidly decreased from a very large value to around 10. Then, during the subsequent training process, the loss value has been hovering around 10. However, when testing the performance of the model at this time, it was found that all indicators were close to the worst situation. Thus, it can be seen that the model is in an underfitting state at this time. Although Adam has the characteristic of adaptive learning rate, too high an initial learning rate may still lead to excessive parameter updates, making the loss unable to decrease stably. Therefore, in subsequent experiments, an attempt was made to test by reducing the learning rate.
[0061] As can be Figure 8 seen, the accuracy rate can decrease correctly at this time. The loss value for epochs less than 2 is hidden in the image to avoid the image being stretched too long. The evaluation results of the experiment are shown in Table 3.
[0062] As can be seen from Table 3, the model with cross-attention based on multiplication is slightly higher than the original model without any fusion mechanism in the ability to recognize cars. However, when recognizing bicycles and pedestrians, the performance decreases in all indicators.
[0063] Table 3 Performance of the model with multiplication-based fusion mechanism Category Precision Recall mAP50 Car 0.745 0.677 0.703 Bicycle 0.183 0.047 0.098 Person 0.455 0.095 0.269
[0064] Next, the cross-attention mechanism based on concatenation was tested. Similarly, when the learning rate was set to 0.01 in this experiment, the loss value stopped decreasing after dropping to around 8. When the learning rate was reset to 0.0001, the model loss could decrease normally. The batch size of this experiment was set to 64, and the training time consumption was similar to that without using the fusion mechanism. The performance of this method in each evaluation index is shown in Table 4.
[0065] As can be seen from Table 4, the fusion mechanism based on concatenation is superior to the model without using the fusion mechanism and the multiplication-based fusion method in all aspects, and the targets of each category have been improved compared with several other methods.
[0066] Table 4 Performance of the model with concatenation-based fusion mechanism Category Precision Recall mAP50 Car 0.824 0.738 0.818 Bicycle 0.214 0.055 0.115 Person 0.532 0.112 0.315
[0067] The following is an analysis of performance comparison.
[0068] During the experiment, it was found that when setting the batch size for the multiplication-based fusion method, if it is set to be greater than 4, an error of CUDA Out of Memory will occur. At the same time, each round of training takes as long as 20 minutes, and such a value is unexpected in a graphics card with 24GB of video memory. In contrast, the cross-attention mechanism based on concatenation has a training time similar to that without adding any fusion mechanism, and the maximum batch size for training can be set to 128. Therefore, in this embodiment, the performance of these two mechanisms is further compared.
[0069] The experiment is set as follows. On the same machine, feature maps of two perspectives with batch sizes of 1, 2, 4, 8, 16, 32, and 64 are randomly generated, and they are respectively moved to CUDA for calculation. Each feature map is calculated 10 times through two mechanisms to test their calculation time. The results of the experiment are as Figure 9 shown.
[0070] Figure 10 shows the performance comparison under the mAP50 metric using various methods. As can be seen from the figure, the recognition capabilities of several methods for vehicles are much greater than those for bicycles and pedestrians. Among them, the method using addition as the fusion method has very poor performance and cannot be used. The effects of not using fusion and using the multiplication-based fusion mechanism are quite similar, while the perspective fusion method using concatenation followed by the attention mechanism achieves the best performance in the detection of each category, and has good efficiency in training and inference performance. Thus, it can be seen that introducing multi-perspective data for object recognition can improve the accuracy of object recognition.
[0071] Comparing Table 1 and Table 4, compared with the traditional single-perspective object recognition method, the algorithm of the present invention has improved the object recognition accuracy and positioning accuracy by about 12% in complex environments, reaching more than 82%, significantly improving the robustness of recognition. Specifically, after adopting the cross-attention mechanism, the feature fusion effect is optimized, the expression ability of the main perspective feature map is enhanced, and thus the detection rate of small targets is improved. In the experiment, the recognition rate of small targets has increased by 1-2%. In addition, as Figure 10 shown, combining the multi-scale detection of the feature pyramid network has increased the efficiency of the system in processing large-scale targets by 12%, significantly reducing the consumption of computing resources. In practical applications, the system architecture of the present invention supports flexible deployment, can quickly adapt to dynamic environments, reduces the need for manual intervention, reduces the labor intensity, and improves the convenience of operation. These advantages make the present invention have higher application value and market potential in fields such as military reconnaissance and outdoor search and rescue.
[0072] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than limiting them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope defined by the claims of the present invention.
Claims
1. A target positioning and detection method based on the collaborative work of multiple UAVs, characterized by: The method uses a system including a drone cluster, a data transmission module, a data processing module and a decision fusion module to collect target images taken by multiple cameras of the drone cluster and transmit them to the data processing module, and uses a deep convolutional neural network to extract features of the images collected by each camera to obtain a feature map of the image; the deep convolutional neural network selects VGG as the convolutional neural network architecture; the convolutional neural network extracts important feature information in the image through a convolution layer, a pooling layer and an activation function; after each image is processed by the convolutional neural network, a corresponding feature map is generated, and the feature map includes spatial information and semantic information of the image; the feature maps from different perspectives are fused through a designed cross-attention mechanism, and multi-scale target detection and target positioning are performed in combination with a feature pyramid network; finally, the positioning and detection results from different drones are combined in the decision fusion module, and a consensus algorithm is used for decision fusion, and the fused target recognition result is post-processed to output a final recognition result, including the category, position and confidence of the target.
2. The target positioning and detection method based on the collaborative work of multiple UAVs according to claim 1 is characterized by: The acquisition process of the target image is as follows: The drone flies in a designated area and uses multiple cameras to collect target images from different angles in real time. The cameras of each drone work synchronously to ensure the temporal consistency of the data. The collected image data is first stored in the drone's local storage, followed by preliminary image preprocessing, including image smoothing and filtering, enhancement, image scaling and normalization, and color space conversion. Each drone is equipped with GPS and ranging sensors, which can obtain its own position coordinates and distance data from the target.
3. The target positioning and detection method based on the collaborative work of multiple UAVs according to claim 1 is characterized by: In the process of transmitting to the data processing module, ZeroMQ is used as the message transmission framework; in the process of data transmission, a time synchronization protocol is adopted.
4. The target positioning and detection method based on the collaborative work of multiple UAVs according to claim 1 is characterized by: The specific process of fusing feature maps from different perspectives through the designed cross attention mechanism is as follows: By calculating the correlation between features from different perspectives, information is adaptively selected and fused; feature fusion includes three steps: The first step is feature mapping, which maps the feature maps of each view into the same feature space; The linear mapping of features F1 and F2 from two different perspectives is expressed as: ; ; Among them, W1 and W2 are the corresponding linear transformation weight matrices, b1 and b2 are bias terms, and They are the feature maps after mapping; The second step is to calculate the correlation, calculate the dot product between the main view feature map and other view feature maps, and generate the attention weight matrix; The dot product calculation formula is as follows: ; in, is the calculated dot product; C is the number of channels; and are the values of the main perspective and the i-th perspective in the k-th channel respectively; Then calculate the attention weight matrix, perform Softmax operation on the dot product values of all perspectives, and get the attention weight , as follows: ; The third step is weighted fusion, which uses attention weights to weight the feature map and generate a fused feature map. : ; Wherein, N represents the total number of viewing angles.
5. The target positioning and detection method based on the collaborative work of multiple UAVs according to claim 4 is characterized in that: In the target detection, the fused feature map is input into the feature pyramid network to perform multi-scale target detection; for each layer of feature map F i Perform 1x1 convolution to obtain a new feature map: ; Then, the high-level feature map is fused with the low-level feature map through upsampling operation to obtain multi-scale features: ; Among them, Upsample means upsampling, which processes high-level features to facilitate addition to low-level features; Based on the feature pyramid network, multiple detection heads are designed to process feature maps of different scales and output the bounding box and category information of the target, as follows: Bounding Box Regression: ; Category prediction: ; in, , , and denote the bounding box, class label, bounding box offset predictor, and classifier, respectively.
6. The target positioning and detection method based on the collaborative work of multiple UAVs according to claim 5 is characterized by: In the target positioning, a triangulation or multilateral positioning algorithm is used to calculate the precise position of the target based on the distance between the position of the UAV and the target; According to the distance between the UAV position and the target, the precise position of the target is calculated by triangulation; assuming that the UAV U is at position , the target T is at position , the distance to the target is d u ; Use the following triangulation formula to calculate the target's position: ; ; in, θ is the azimuth between the UAV and the target, measured from other known angles or further calculated through multilateration; Then the geometric positioning precision factor is used to evaluate the positioning accuracy, and the position of the drone is adjusted as needed to optimize the geometric positioning precision factor, as shown in the following formula: ; ; in, , , , Respectively represent the target predicted position abscissa, the target actual position abscissa, the target predicted position ordinate and the target actual position ordinate; The full intersection-over-union loss for bounding box regression calculates the relative position of the target based on the overlap between the predicted box and the true box, the center point distance, and its aspect ratio, as follows: ; in, Represents the weighted sum of center point distance and aspect ratio loss; α , β , is the weight coefficient used to control the impact of each loss; IoU, , They represent the bounding box regression loss, the Euclidean distance between the center point of the predicted box and the center point of the true box, and the aspect ratio loss respectively.
7. The target positioning and detection method based on the collaborative work of multiple UAVs according to claim 6 is characterized by: After the target is located, error correction is performed to identify and correct positioning errors caused by sensor errors, environmental factors or changes in the position of the drone; and a filtering algorithm is used to smooth the positioning data.
8. The target positioning and detection method based on the collaborative work of multiple UAVs according to claim 7 is characterized by: The specific method of using the consensus algorithm for decision fusion is: In the global consensus, each agent ai updates its state at time step ak Based on the information of its neighbors, this update is based on weighted average or weighted sum, the formula is as follows: ; in, is the state of the ai-th agent at time step ak; N ai is the neighbor set of agent ai, which represents other agents that have communication connections with agent ai; w aiaj is the communication weight between agent ai and agent aj, indicating the strength of information transmission from aj to ai, and ; It is the information of the agent ai itself, to ensure that the agent can retain its own state when there are no neighbors; In local consensus, agents update their states based on information from their neighbors, as described by the following local weighted sum formula: ; in, α is the adjustment factor, which controls the update step size. The meaning of this formula is that the agent ai updates its own state according to the information of its neighbor aj, by adjusting α Control the speed of consensus convergence; In order to ensure that all agents reach consensus, a consistency constraint is required: ; in, is the consistency tolerance, which means that the state difference between agents should be less than a certain threshold; this condition ensures that after multiple iterations, the states between agents converge to a consistent value; The convergence condition is expressed by the change in the number of iterations and the state difference: ; in, is the set tolerance error, indicating that the state difference between the two agents must be less than a certain value, indicating that convergence is complete; Ultimately, the goal is to make the states of all agents After a certain number of iterations, they tend to be consistent, thus reaching a global consensus; the global consensus goals are: ; That is, all agents eventually reach the time step , the status tends to be consistent.
Citation Information
Cited By
Shielding target detection method and system based on multi-view fusion
CN120823376A
An occluded object detection method and system based on multi-view fusion
CN120823376B