Unmanned aerial vehicle target positioning method and device based on cooperation of multiple observation points, and medium

By constructing a target recognition model with a three-dimensional spatial positioning architecture for dynamic and static targets and a bidirectional feature refinement module, and combining it with an end-to-end deep neural network, the problems of small target detection, cross-view matching and sensor bias in UAV target positioning are solved, achieving high-precision and real-time UAV target positioning.

CN121999040APending Publication Date: 2026-05-08CRSC INST OF SMART CITY RES &DESIGN
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CRSC INST OF SMART CITY RES &DESIGN
Filing Date
2025-12-16
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing UAV target positioning technologies suffer from problems such as insufficient accuracy in small target detection, poor reliability of multi-view target association, insufficient data coordination among multiple observation points, and ineffective compensation for optical sensor deviations, resulting in low positioning accuracy and poor reliability.

Method used

By constructing a physical architecture for three-dimensional spatial positioning of dynamic and static targets, and utilizing a target recognition model with a bidirectional feature refinement module and an end-to-end deep neural network, a mapping relationship between multiple observation points and targets is established, thereby improving the detection accuracy of small targets, enhancing the robustness of cross-view matching, and compensating for sensor bias.

Benefits of technology

It significantly improves the real-time positioning accuracy and system reliability of UAV targets in the world coordinate system, and is suitable for scenarios such as airspace management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121999040A_ABST
    Figure CN121999040A_ABST
Patent Text Reader

Abstract

The invention provides an unmanned aerial vehicle target positioning method and device based on cooperation of multiple observation points and a medium, and the method comprises the steps: building a dynamic and static target three-dimensional space positioning physical architecture, and building an unmanned aerial vehicle target recognition model based on bidirectional cross-dimension feature enhancement; an image-based unmanned aerial vehicle multi-target association matching algorithm is proposed to eliminate background interference and correct view angle influence, and an end-to-end deep neural network is adopted to model a mapping relationship between multiple observation points and unmanned aerial vehicle targets, so that multi-source feature fusion and positioning precision improvement are realized; according to the technology, the problems of poor recognition robustness, low positioning precision and the like caused by small target information loss, view angle transformation interference and optical sensor measurement deviation in the actual unmanned aerial vehicle target positioning process are solved, and efficient and reliable unmanned aerial vehicle target positioning technical support is provided for airspace monitoring, unmanned aerial vehicle compliance management, security patrol and other scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This document relates to the field of UAV positioning technology, and in particular to a UAV target positioning method, device and medium based on multi-observation point collaboration. Background Technology

[0002] With the rapid development of the low-altitude economy, drones are increasingly used in various fields, making the demand for accurate and real-time target positioning increasingly urgent. However, existing drone target positioning technologies still have the following limitations: Insufficient accuracy in small target detection: Drones appear as small targets in medium- and long-range imaging, and existing feature extraction networks struggle to effectively integrate spatial details with deep semantic information, resulting in high false negative rates and inaccurate bounding boxes, especially in scenarios where multiple scale targets coexist.

[0003] Poor reliability of multi-view target association: Multi-observation point collaboration depends on cross-view target matching, but existing methods rely too much on appearance features (such as color and texture) that are easily affected by the viewpoint, resulting in a significant decrease in matching performance when the viewpoint changes.

[0004] Insufficient data coordination among multiple observation points: low time synchronization accuracy of multiple cameras, resulting in time lag in observation data; and lack of standardized processing of data such as camera attitude and image coordinates, leading to the accumulation of errors during data fusion and reduced positioning accuracy.

[0005] Optical sensor bias is not effectively compensated: Existing methods fail to adequately model the nonlinear mapping relationship between the observation point and the target, and cannot correct sensor measurement bias in imaging in real time; at the same time, they lack an optimization mechanism that takes into account both position error and distance constraints, resulting in inaccurate positioning results in the world coordinate system.

[0006] The aforementioned problems limit the effectiveness of UAV target localization in complex real-world scenarios. To address this, this invention proposes a multi-observation-point collaborative UAV target localization method, aiming to improve small target identification, cross-view correlation, data collaboration, and deviation compensation capabilities, achieving accurate and real-time localization and providing reliable support for scenarios such as airspace management. Summary of the Invention

[0007] This invention provides a method for constructing a three-dimensional spatial positioning physical architecture for dynamic and static targets, establishing a recognition model based on bidirectional cross-dimensional feature enhancement, and modeling the mapping relationship between multiple observation points and targets through an image multi-target association matching algorithm and an end-to-end deep neural network. This method solves the problems of poor recognition robustness and low positioning accuracy in actual UAV positioning caused by information loss of small targets, viewpoint interference, data asynchrony, and sensor deviation. It achieves accurate and real-time positioning of UAV targets, which is helpful for deployment in scenarios such as airspace monitoring and UAV compliance management.

[0008] According to an embodiment of the present invention, a UAV target localization method based on multi-observation-point collaboration is provided, characterized by comprising: S1. Deploy multiple cameras to form a three-dimensional monitoring network, and perform time synchronization and spatial coordinate calibration on each camera to obtain its three-dimensional position and attitude data. S2. Using a target recognition model that integrates a bidirectional feature refinement module, identify drone targets in images captured by each camera and obtain their two-dimensional coordinates in the images; S3. Based on the extracted target foreground appearance attribute features, perform association matching on the targets identified in different camera images to determine multiple observation data belonging to the same UAV target; S4. Based on the association matching results, the observation data of each UAV target belonging to the same UAV target are fed into an end-to-end deep neural network. The mapping relationship between multi-source observation data and three-dimensional world coordinates is established through the deep neural network, and the three-dimensional position estimate of the UAV target in the world coordinate system is output. The observation data belonging to the same UAV target include: the three-dimensional position and attitude data of the corresponding camera and the two-dimensional coordinate data of the target in the image.

[0009] According to an embodiment of the present invention, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor executes the computer program to implement the steps of the above-described UAV target localization method based on multi-observation point collaboration.

[0010] According to an embodiment of the present invention, a storage medium is provided on which a computer program is stored, characterized in that the computer program, when executed by a processor, implements the steps of the above-described UAV target localization method based on multi-observation point collaboration.

[0011] The UAV target localization method proposed in this application improves the detection accuracy of small targets by integrating a target recognition model with a bidirectional feature refinement module, enhances the robustness of cross-view matching by using a multi-target association algorithm combined with pose correction, and achieves high-precision fusion of multi-source observation data and effective compensation for sensor bias by using an end-to-end deep neural network. This significantly improves the real-time positioning accuracy of UAV targets in the world coordinate system and the reliability of the system. Attached Figure Description

[0012] To more clearly illustrate the technical solutions in one or more embodiments of this specification or in the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0013] Figure 1 This is a flowchart of a UAV target localization method based on multi-observation-point collaboration according to an embodiment of the present invention; Figure 2 This is a physical architecture diagram of the three-dimensional spatial positioning of dynamic and static targets according to an embodiment of the present invention; Figure 3 This is a schematic diagram illustrating the deployment of a multi-observation-point camera array and flight path planning according to an embodiment of the present invention; Figure 4 This is a structural diagram of the bidirectional feature refinement module according to an embodiment of the present invention; Figure 5 This is a schematic diagram of a multi-observation-point collaborative target fusion localization model according to an embodiment of the present invention. Detailed Implementation

[0014] To enable those skilled in the art to better understand the technical solutions in one or more embodiments of this specification, the technical solutions in one or more embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this specification, and not all of the embodiments. Based on one or more embodiments of this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the protection scope of this document.

[0015] Method Implementation Examples According to embodiments of the present invention, a UAV target localization method based on multi-observation-point collaboration is provided. Figure 1 This is a flowchart of a UAV target localization method based on multi-observation-point collaboration according to an embodiment of the present invention. Figure 1 As shown, the UAV target localization method based on multi-observation-point collaboration in this embodiment of the invention specifically includes: S1. Deploy multiple cameras to form a three-dimensional monitoring network, and perform time synchronization and spatial coordinate calibration on each camera to obtain its three-dimensional position and attitude data. Number of cameras selected: N (N>3) cameras are selected and the camera nodes are arranged in a spatially symmetric topology. The N vertices can form a three-dimensional monitoring network. Deploy cameras: Select a location with no obvious obstructions that can cover the basic area of ​​the target monitoring area, install the first camera and adjust its orientation so that the central axis of its lens points in the core direction of the target monitoring area; determine the spatial coordinates of the other vertices based on the spatial range of the target monitoring area to ensure that the three-dimensional monitoring network composed of N vertices fully covers the target area; Synchronize time signals from multiple cameras: Connect all camera nodes to the same time synchronization server, which should be able to provide a stable standard time signal; Set a time synchronization period, and each camera node should perform time calibration with the time synchronization server according to the period to ensure that the internal clock error of each camera is controlled within an acceptable range, so as to ensure that the data of the same event collected by different cameras are consistent in the time dimension. Design the drone flight path: A three-dimensional monitoring network is formed by N vertices. Three vertices are taken from each network to form a plane. To plan the route, the centroid of each plane needs to be determined. Connecting the centroids of every two planes forms the flight path. The drone must pass through the three-dimensional area formed by the monitoring network while flying along these routes. Figure 3 This is a schematic diagram illustrating the deployment of a multi-observation-point camera array and flight path planning according to an embodiment of the present invention.

[0016] Data collection of UAV flight data: The data collected by multiple camera observation points originates from the process of the UAV target traversing the coverage area along the aforementioned specific trajectory; these data constitute the training set and validation set, which are used to train and validate the constructed dynamic and static target three-dimensional spatial positioning system; like Figure 2 This is a physical architecture diagram for the three-dimensional spatial positioning of dynamic and static targets according to an embodiment of the present invention. Step S1, for fixed observation points and dynamic UAV targets, firstly, uses a high-precision observation point positioning module to provide an absolute geographic coordinate reference for all edge computing modules; based on this, by jointly calibrating the observation points of multiple cameras, the three-dimensional spatial positioning of the observation node's own pose is achieved; simultaneously, multiple observation point cameras acquire image data of the UAV target, and each camera obtains the positioning target observation value through the deployed UAV target recognition algorithm, and performs correlation matching on the targets observed by multiple cameras; finally, the edge device sends the processed data to the cloud, and by modeling the mapping relationship between multiple observation points and the UAV target, the relative position of the dynamic target with respect to each observation point is calculated in real time, and the three-dimensional coordinates of the positioning target are obtained; S2. Using a target recognition model, identify drone targets in the images captured by each camera and obtain their two-dimensional coordinates in the images; the target recognition model integrates a bidirectional feature refinement module in the backbone network, which processes feature maps through channel splitting and cross-attention mechanism to fuse spatial details and semantic information; The target recognition model includes a backbone network, a neck network, and a detection head; The backbone network is constructed based on a cross-stage local network architecture and integrates the bidirectional feature refinement module in at least two feature extraction stages at different scales. The neck network adopts a bidirectional feature pyramid structure, which fuses multi-level feature maps from the backbone network through dual-path connections from top to bottom and bottom to top. The detection head includes a classification branch and a regression branch, wherein the regression branch is optimized using the CIoU loss function and the classification branch is optimized using the cross-entropy loss function. Meanwhile, the detection head introduces a zoom loss function as an auxiliary optimization objective during training.

[0017] Figure 4 This is a structural diagram of the bidirectional feature refinement module according to an embodiment of the present invention. The bidirectional feature refinement module processes the input feature map according to the following workflow: The input feature map is split into a first feature subset and a second feature subset along the channel dimension, where the splitting ratio is controlled by the hyperparameter α. Perform convolution operations and channel attention calculations sequentially on the first feature subset to generate channel attention weights; Perform convolution operations and spatial attention calculations on the second feature subset to generate spatial attention weights; The channel attention weights are multiplied element-wise with the second feature subset to obtain the channel enhancement features; The spatial attention weights are multiplied element-wise with the first feature subset to obtain the spatially enhanced features; The channel enhancement features and spatial enhancement features are fused through feature aggregation operations to obtain an output feature map.

[0018] More specifically, a concrete implementation of the bidirectional feature refinement module is as follows: (1) Input feature map (in C The number of channels in the feature map. H and W The feature map (representing its height and width respectively) is divided into two disjoint subsets along the channel dimension: One channel ( )and One channel ( The formula for this stage is as follows: (1); in, Represents a hyperparameter It plays a crucial controlling role in the representation capabilities of the network; (2) After obtaining two branches, an attention mechanism is applied to these two branches to compensate for the missing feature maps in each branch, thereby achieving efficient feature matching; on the one hand, By two standards Convolution and one The branches formed by convolutions extract richer semantic feature information on each channel, resulting in... ;on the other hand, By a standard The branches formed by convolutions extract relatively weak information while retaining a large amount of shallow spatial location information, represented as... The calculation process is as follows: (2); (3); Then, obtain information with richer channel details. Channel attention weights ; Unique weights can be assigned to the key details of each channel, and then mapped to features with low-level spatial location information. To obtain higher-level features The calculation method is as follows: (4); (5); in, This represents the Sigmoid activation function. Indicates batch normalization. Indicates average pooling. This represents element-wise multiplication; Similarly, acquiring richer spatial information Spatial attention weights Then it is mapped to features with rich channel information. To obtain higher-level features The calculation method is as follows: (6); (7); in, This represents global max pooling along the channel dimension. This represents global average pooling along the channel dimension; (3) Finally, connect these two branches to obtain features that include spatial and semantic relationships. The formula for this process is as follows: (8); in, This indicates element-wise summation; Based on this, a UAV target recognition method based on bidirectional cross-dimensional feature enhancement, applicable to the optical camera perspective, is proposed. This method ensures accurate and real-time UAV target recognition while requiring less computational resources. The proposed UAV target recognition model mainly considers improving the detection performance of small targets in multi-viewpoint cameras through network structure design. The recognition model mainly consists of three parts: a backbone network, a neck network, and a head network. First, the backbone network extracts features of different scales layer by layer from the input image along five feature layers. During the feature extraction process, the backbone network adopts the C3k2 module, which is an improved version of the cross-stage local network. This module divides the feature map into two parts, which are then processed through a bottleneck layer and a merging operation. The network is processed to optimize the information flow. Simultaneously, the backbone network integrates a bidirectional feature refinement module, which enhances feature matching and localization accuracy by fusing spatial location and semantic information, thus mitigating the problem of information loss in small targets. Then, the neck network, employing a bidirectional feature pyramid structure, optimizes cross-scale connections and weighted feature fusion to aggregate feature maps from different resolutions, better integrating multi-scale features, and outputs the processed feature maps to the detection head. Finally, the detection head outputs the detection results corresponding to the feature maps after processing through classification and regression branches, including bounding box position parameters, confidence scores, and class probabilities. The losses in the classification and regression stages are the main contributors to the loss error, primarily employing CIoU loss function, zoom loss function, and cross-entropy loss function. The CIoU loss function optimizes the bounding box regression accuracy of the model for occluded targets and targets with unconventional proportions. Its mathematical expression is: (9); in This represents the intersection-union ratio (IoU) between the predicted bounding box and the actual bounding box. It is the center point of the prediction box. It is the center point of the true bounding box. This represents the square of the Euclidean distance between two points. As a penalty for shape similarity, This is the contribution weighting coefficient, and its mathematical expression is as follows: (10); in, This is a shape similarity penalty term used to quantify the difference in aspect ratio between the predicted bounding box and the ground truth bounding box. Its mathematical expression is as follows: (11); in, and These are the width and height of the actual bounding box, respectively. and These are the width and height of the prediction box, respectively; The zoom loss function is a loss function that measures the difference between the prediction results of a recognition model and the ground truth annotations. It guides the model to better learn the location and category information of the target by calculating the similarity between the predicted bounding box and the ground truth bounding box, thereby improving the accuracy and robustness of target detection. Its mathematical expression is as follows: (12); in, The intersection-union ratio (IUU) of the predicted bounding box and the ground truth bounding box. The predicted probability output by the model is the probability when two predicted boxes intersect with the ground truth boxes. That is, if >0, it is a positive sample, and the two boxes do not intersect. =0 indicates a negative sample; Cross-entropy loss is a commonly used loss function in classification problems. It quantifies the difference between the probability distribution predicted by the model and the actual labels, and is used to measure the accuracy of the model's predictions. Cross-entropy loss measures the model's performance by calculating the cross-entropy between the probability distribution predicted by the model and the probability distribution of the true labels. Its mathematical expression is as follows: (13); in, A It is the sample size. M Indicates the number of categories that need to be identified. Indicates sample i Is the predicted category the true category? c , Indicates sample i Belongs to the real category c The predicted probability; S3. Based on the extracted target foreground appearance attribute features, perform association matching on targets identified in images from different cameras to determine multiple observation data belonging to the same UAV target; wherein, each observation data belonging to the same UAV target includes: the three-dimensional position and attitude data of the corresponding camera and the two-dimensional coordinate data of the target in the image. S3 specifically includes: S31. For each target prediction box identified in each camera image, calculate the gradient of the pixels within the box, determine an adaptive threshold based on the sorted gradient values, and mark pixels with gradients greater than or equal to the threshold as foreground. S32. Extract color, texture, or shape features as appearance attribute features of the target based on the foreground pixels; S33. The similarity of the feature vectors of the appearance attributes of the targets between different observation points is calculated by using cosine distance, and the similarity is normalized by using the softmax function to obtain the association probability between the targets at each observation point. The matching is completed based on the association probability.

[0019] After obtaining the location of the target group using the aforementioned UAV target recognition method, the target region sub-image is sent to the appearance attribute feature extraction module to obtain the appearance features of the target and calculate the feature similarity under different viewpoints, so as to realize the association and matching of targets under multiple viewpoints. The appearance attribute features of the target include the target's color, texture, shape, etc., but they are greatly affected by the viewpoint change, and the extraction effect of appearance attribute features and their role in multi-target association are greatly reduced. By introducing the camera position information and observation values ​​of multiple observation points, the influence of viewpoint on the observation results can be corrected to a certain extent, making the introduction of appearance attribute features feasible and effective. The steps for extracting the appearance attribute features such as the color, texture, and shape of the target are as follows: (1) Background interference elimination: In order to reduce the interference of the background on target recognition, it is necessary to distinguish between the target and the background. This can be achieved by calculating the gradient of each pixel in the prediction box, because the gradient reflects whether the pixel belongs to the edge region, and the gradient of the target edge pixels is usually larger. The gradient can be calculated by the Euclidean norm of the horizontal-vertical gradient, and the calculation method is as follows: (14); in, The horizontal gradient of a pixel (grayscale variation along the column direction) is approximated by a first-order difference. The vertical gradient of a pixel (grayscale change along the row direction) is approximated by a first-order difference. Then, the gradient values ​​of all pixels within the prediction box are... Sort in ascending order to obtain a one-dimensional array. This is used for subsequent adaptive threshold calculation: (15); in, This indicates an operation to sort the targets in ascending order; To distinguish between background pixels (small gradient, gradual grayscale change) and target pixels (large gradient, drastic grayscale change), an adaptive threshold is calculated based on the sorted gradient array. T Pixels with gradients less than this threshold are marked as background. Pixels with gradients greater than or equal to the threshold are marked as target regions. Adaptive threshold T The mathematical expression is: (16); in, These are learnable quantile parameters. For the floor function, ensure that the corresponding array S The index is an integer; (2) Feature statistics and correlation probability matching: After background removal, only foreground pixels are selected. The statistical analysis of appearance attribute features (such as mean, variance, etc.) is mathematically represented as the set of effective pixels: (17); Calculate the similarity of appearance attribute features among multiple target regions and assign association probabilities. Cosine distance is used to calculate the similarity of appearance attribute feature vectors. The similarity value ranges from 0 to 1, with higher values ​​indicating higher similarity. The calculation formula is as follows: (18); in, I and J Indicates two different observation points. and These are appearance attribute feature vectors obtained from two different observation points. and These are the L2 norms of the two vectors, respectively. To ensure that the sum of the association probabilities of all measurements from a given observation point to all other observation points equals 1, the similarity is normalized using the softmax function to obtain the association probability: (19); in, Indicates the observation point I Similarity of appearance attribute feature vectors between observation points and other observation points; S4. Based on the association matching results, the observation data of each UAV target belonging to the same UAV target are fed into an end-to-end deep neural network. The mapping relationship between multi-source observation data and three-dimensional world coordinates is established through the deep neural network, and the three-dimensional position estimate of the UAV target in the world coordinate system is output. In step S4, before inputting the observation data into the end-to-end deep neural network, a data preprocessing step is also included: The data collected by each camera are aligned based on timestamps, and the three-dimensional position and attitude data and two-dimensional coordinate data are normalized. In step S4, the end-to-end deep neural network includes an input layer, at least one hidden layer, and an output layer. The hidden layer employs a Sigmoid activation function for nonlinear transformation to learn the nonlinear mapping relationship between the observed data and three-dimensional world coordinates. The end-to-end deep neural network is trained using a composite loss function combining the mean square error loss function and the Euclidean distance loss function to optimize network parameters and compensate for measurement biases of the optical sensor.

[0020] Figure 5 This is a schematic diagram of the multi-observation-point collaborative target fusion localization model of the present invention. In the process of multi-observation-point collaborative target fusion localization, the deep neural network learns from each camera. i The observed target image location estimate and the corresponding camera location information are fused together to accurately estimate the target's position in the world coordinate system. The deep neural network structure includes an input layer, an output layer, and hidden layers. Each neuron is connected to all neurons in the next layer, but there are no connections between neurons within the same layer. The specific structure of this deep neural network is as follows: The input layer is mainly divided into static data and dynamic data. Static data includes the pose and position data of each camera, while dynamic data includes the coordinates of the drone target center point detected by each camera. Preprocessing of the input data is required to ensure that: the data collected by all cameras are time-aligned and synchronized via timestamps; secondly, data normalization is performed to normalize the 3D position and angle data of the cameras, improving the training efficiency and stability of the model; the input data for the network can be obtained by normalizing the data collected by all cameras using the following formula: (20); in, , This indicates that the first point obtained using a high-precision observation point positioning module... i ( The absolute geographic coordinates provided by the edge camera at each observation point. Indicates the first i The camera uses a drone target recognition model to obtain the two-dimensional coordinates and bounding box size of the target in the acquired image. This indicates taking the minimum value of the data. This indicates taking the maximum value of the data; The hidden layer constructs a deep neural network to extract and fuse features from multi-observation attitude and position information, as well as UAV features from multiple observation points, building a non-linear mapping relationship between the input static and dynamic data and the target output. Specifically, neurons in the hidden layer perform non-linear transformations on the input data through activation functions (such as the sigmoid activation function), thereby extracting higher-level features. For the input data, after a total of...n The first hidden layer i The hidden layer transformation can be represented as : (twenty one); (twenty two); (twenty three); in, It is the Sigmoid activation function. It is the first i The weight matrix of the hidden layer, It is the first i The bias term of the hidden layer; The output layer outputs the geographical locations of each observation point and the positioning target. The mapping relationship is expressed mathematically as follows: (twenty four); in, It is the first i The geographic coordinates of each camera; this mapping relationship is solved based on features extracted from the hidden layer, and the final location is estimated through the network's output layer; for n The output of the hidden layer is The prediction result of the output layer The mathematical expression is: (25); in, It is the weight matrix of the output layer. It is the bias term of the output layer; To evaluate model performance and optimize network parameters, it is necessary to choose an appropriate loss function. Loss functions include the mean squared error loss function and the Euclidean distance loss function, as detailed below: Mean Squared Error Loss Function: The mean squared error loss function measures the performance of a model by calculating the squared difference between the predicted and actual values. Its formula is as follows: (26); in, It is the first i The model prediction value for each sample. It is the first i The model's true value for each sample. A It is the sample size; Euclidean distance loss function: The Euclidean distance loss function calculates the Euclidean distance between the predicted location and the actual location. The formula is as follows: (27); in, It is the first i The predicted location coordinates of each sample. It is the first i The true location coordinates of each sample; An end-to-end multi-observation point collaborative target fusion positioning estimation method based on deep neural network is used to model the mapping relationship between multiple observation points and UAV targets. The method uses deep neural network to output dynamic target three-dimensional position estimation (24). According to the mean square error loss function (26) and Euclidean distance loss function (27), the measurement deviation caused by optical sensor is compensated in real time, so as to realize the accurate position estimation of UAV target in world coordinate system.

[0021] By employing the embodiments of the present invention, the following beneficial effects are achieved: The target recognition model, which integrates a bidirectional feature refinement module, improves the detection accuracy of small targets. The multi-target association algorithm combined with pose correction enhances the robustness of cross-view matching. Furthermore, the end-to-end deep neural network enables high-precision fusion of multi-source observation data and effective compensation for sensor biases, thereby significantly improving the real-time positioning accuracy and system reliability of UAV targets in the world coordinate system.

[0022] Device Example 1 According to an embodiment of the present invention, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor executes the computer program to implement the steps of the above-described UAV target localization method based on multi-observation point collaboration.

[0023] Device Example 2 According to an embodiment of the present invention, a storage medium is provided on which a computer program is stored, characterized in that the computer program, when executed by a processor, implements the steps of the above-described UAV target localization method based on multi-observation point collaboration.

[0024] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A UAV target localization method based on multi-observation-point collaboration, characterized in that... include: S1. Deploy multiple cameras to form a three-dimensional monitoring network, and perform time synchronization and spatial coordinate calibration on each camera to obtain its three-dimensional position and attitude data. S2. Using a target recognition model, identify drone targets in the images captured by each camera and obtain their two-dimensional coordinates in the images; the target recognition model integrates a bidirectional feature refinement module in the backbone network, which processes feature maps through channel splitting and cross-attention mechanism to fuse spatial details and semantic information; S3. Based on the extracted target foreground appearance attribute features, perform association matching on the targets identified in different camera images to determine multiple observation data belonging to the same UAV target; S4. Based on the association matching results, the observation data of each UAV target belonging to the same UAV target are fed into an end-to-end deep neural network. The mapping relationship between multi-source observation data and three-dimensional world coordinates is established through the deep neural network, and the three-dimensional position estimate of the UAV target in the world coordinate system is output. The observation data belonging to the same UAV target include: the three-dimensional position and attitude data of the corresponding camera and the two-dimensional coordinate data of the target in the image.

2. The method according to claim 1, characterized in that, S1 specifically includes: S11. Select N cameras, where N>3, and deploy them using a spatially symmetrical topology to form a three-dimensional monitoring network covering the target monitoring area; S12. Connect all camera nodes to the same time synchronization server and perform time calibration according to the set synchronization period to achieve time signal synchronization of each camera. S13. Acquire and record the spatial coordinates and attitude parameters of each camera node to complete the spatial coordinate calibration. S14. Plan the flight path of the UAV. The path is formed by calculating and connecting the centroids of each plane formed by the vertices of the cameras in the three-dimensional monitoring network to ensure that the flight path of the UAV runs through the area covered by the monitoring network. S15. Collect image data simultaneously captured by multiple cameras during the flight of the drone along the flight path to form a dataset for training and verification.

3. The method according to claim 1, characterized in that, The target recognition model includes a backbone network, a neck network, and a detection head; The backbone network is constructed based on a cross-stage local network architecture and integrates the bidirectional feature refinement module in at least two feature extraction stages at different scales. The neck network adopts a bidirectional feature pyramid structure, which fuses multi-level feature maps from the backbone network through dual-path connections from top to bottom and bottom to top. The detection head includes a classification branch and a regression branch, wherein the regression branch is optimized using the CIoU loss function and the classification branch is optimized using the cross-entropy loss function. Meanwhile, the detection head introduces a zoom loss function as an auxiliary optimization objective during training.

4. The method according to claim 3, characterized in that, The bidirectional feature refinement module processes the input feature map according to the following workflow: The input feature map is split into a first feature subset and a second feature subset along the channel dimension, where the splitting ratio is controlled by the hyperparameter α. Perform convolution operations and channel attention calculations sequentially on the first feature subset to generate channel attention weights; Perform convolution operations and spatial attention calculations on the second feature subset to generate spatial attention weights; The channel attention weights are multiplied element-wise with the second feature subset to obtain the channel enhancement features; The spatial attention weights are multiplied element-wise with the first feature subset to obtain the spatially enhanced features; The channel enhancement features and spatial enhancement features are fused through feature aggregation operations to obtain an output feature map.

5. The method according to claim 1, characterized in that, S3 specifically includes: S31. For each target prediction box identified in each camera image, calculate the gradient of the pixels within the box, determine an adaptive threshold based on the sorted gradient values, and mark pixels with gradients greater than or equal to the threshold as foreground. S32. Extract color, texture, or shape features as appearance attribute features of the target based on the foreground pixels; S33. The similarity of the feature vectors of the appearance attributes of the targets between different observation points is calculated by using cosine distance, and the similarity is normalized by using the softmax function to obtain the association probability between the targets at each observation point. The matching is completed based on the association probability.

6. The method according to claim 1, characterized in that, In step S4, before inputting the observation data into the end-to-end deep neural network, a data preprocessing step is also included: The data collected by each camera are aligned based on timestamps, and the 3D position and attitude data and 2D coordinate data are normalized using the following formula: ; in, , This indicates that the first point obtained using a high-precision observation point positioning module... i ( The absolute geographic coordinates provided by the edge camera at each observation point. Indicates the first i The camera uses a drone target recognition model to obtain the two-dimensional coordinates and bounding box size of the target in the acquired image. This indicates taking the minimum value of the data. This indicates taking the maximum value of the data.

7. The method according to claim 1, characterized in that, In step S4, the end-to-end deep neural network includes an input layer, at least one hidden layer, and an output layer; the hidden layer uses the Sigmoid activation function for nonlinear transformation to learn the nonlinear mapping relationship between the observed data and the three-dimensional world coordinates.

8. The method according to claim 7, characterized in that, The end-to-end deep neural network is trained using a composite loss function that combines the mean square error loss function and the Euclidean distance loss function to optimize network parameters and compensate for measurement biases of the optical sensor.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the UAV target localization method based on multi-observation point collaboration as described in any one of claims 1 to 8.

10. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the UAV target localization method based on multi-observation point collaboration as described in any one of claims 1 to 8.