A target detection method, device, equipment and medium based on unmanned aerial vehicles (UAVs)

By constructing a bird's-eye view feature map on the UAV and refining the grid scale and enhancing the multi-view features, the problem of insufficient accuracy of multi-scale target detection by UAV in complex environments is solved, and high-precision target detection is achieved.

CN121095822BActive Publication Date: 2026-01-30FUTENG TECH BRANCH OF QUZHOU GUANGMING POWER INVESTMENT GRP CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511641227.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-11
Publication Date
2026-01-30
Estimated Expiration
2045-11-11

AI Technical Summary

Technical Problem

In existing technologies, the accuracy of UAVs in detecting multi-scale targets in complex environments still needs to be improved, making it difficult to achieve high-precision three-dimensional environmental perception.

Method used

By constructing a bird's-eye view feature map and refining the grid scale and enhancing features from multiple perspectives, the grid related to the target object is processed with fine detail using multi-view image data, thereby improving detection accuracy.

Benefits of technology

It achieves accurate capture of key details of the target object, improves detection accuracy, reduces computational resource costs, and enhances the efficiency of target detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121095822B_ABST
    Figure CN121095822B_ABST
Patent Text Reader

Abstract

This application relates to the field of unmanned aerial vehicle (UAV) technology, and discloses a UAV-based target detection method, apparatus, device, and medium. The method involves classifying and predicting targets based on a bird's-eye view feature map of the target scene, and then refining the grids related to the target object in the bird's-eye view feature map based on the prediction results to obtain a refined feature map. Next, multi-view feature enhancement is performed on the grids related to the target object in the refined feature map based on multi-view image data to obtain an enhanced feature map. Finally, target detection is performed based on the enhanced feature map to obtain the target detection result. The beneficial effect is that after refining the scale of the grids related to the target object in the bird's-eye view feature map, multi-view feature enhancement is performed on the corresponding grids in the bird's-eye view feature map based on multi-view data of the target scene, thereby actively adapting to the scale of the target object and effectively improving the detection accuracy of the target object.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of unmanned aerial vehicle (UAV) technology, and in particular to a target detection method, apparatus, equipment and medium based on UAVs. Background Technology

[0002] With the development of related technologies, Unmanned Aerial Vehicle (UAV) technology has been widely used in monitoring and reconnaissance missions. Aerial collaborative systems composed of multiple UAVs, with their advantages of comprehensive environmental monitoring and autonomous decision-making, have shown great application potential in urban surveillance, disaster response, traffic management, and agricultural and forestry surveys. In practical scenarios, aerial collaborative systems need to possess high-precision three-dimensional environmental perception capabilities, requiring them not only to identify the type of target but also to accurately estimate its position, size, and orientation in the environmental space. However, the detection accuracy of related technologies based on UAVs for detecting multi-scale targets in complex environments still needs improvement. Summary of the Invention

[0003] This application provides a target detection method, apparatus, device, and medium based on unmanned aerial vehicles (UAVs). After refining the scale of the grid related to the target object in the bird's-eye view feature map, the corresponding grid in the bird's-eye view feature map is enhanced with multi-view features based on multi-view data of the target scene, thereby actively adapting to the scale of the target object and effectively improving the detection accuracy of the target object.

[0004] To achieve the above objectives, the main technical solutions adopted in this application include:

[0005] In a first aspect, embodiments of this application provide a target detection method based on an unmanned aerial vehicle (UAV), the method comprising:

[0006] Target classification prediction is performed based on the bird's-eye view feature map of the target scene, and the grid scale of the grid related to the target object in the bird's-eye view feature map is refined based on the prediction results to obtain a refined feature map; wherein, the bird's-eye view feature map is constructed based on multi-view image data of the target scene, and the multi-view image data is obtained by the UAV in image acquisition of the target scene;

[0007] Based on the multi-view image data, the grid related to the target object in the thinned feature map is enhanced using multi-view features to obtain an enhanced feature map;

[0008] Target detection is performed based on the enhanced feature map to obtain target detection results for the target object.

[0009] Optionally, each grid in the refined feature map related to the target object corresponds to a sub-grid; the step of performing multi-view feature enhancement on the grids related to the target object in the refined feature map based on the multi-view image data to obtain an enhanced feature map includes:

[0010] For any subgrid, a spatial reference point is selected in the target scene based on the bird's-eye view coordinates of the subgrid, and the spatial reference point is projected onto any view feature image in the multi-view image data to obtain a view reference point;

[0011] Based on the viewpoint reference point, predictive dynamic sampling and multi-viewpoint feature aggregation are performed on all viewpoint feature images to obtain the aggregated image features of any sub-grid;

[0012] The refined feature map is obtained by fusing target features with the aggregated image features of all sub-grids.

[0013] Optionally, the step of performing predictive dynamic sampling and multi-view feature aggregation on all view feature images based on the view reference point to obtain the aggregated image features of any sub-grid includes:

[0014] For any viewpoint feature image, based on the bird's-eye view feature corresponding to any sub-grid in the bird's-eye view feature image, feature offset prediction is performed on the viewpoint reference point in the any viewpoint feature image to obtain the offset reference point;

[0015] Based on the offset reference point, feature sampling is performed on the feature image of any viewpoint to obtain the offset feature of the feature image of any viewpoint;

[0016] Based on the offset features of each of the feature images from all viewpoints, feature aggregation is performed on any one of the sub-grids to obtain the aggregated image features of any one of the sub-grids.

[0017] Optionally, the step of fusing target features into the thinned feature map based on the aggregated image features of all sub-grids to obtain the enhanced feature map includes:

[0018] For any target-related grid in the refined feature map that is related to the target object, the aggregated image features of all sub-grids corresponding to any target-related grid are aggregated to obtain the enhanced features of the target-related grid;

[0019] The enhanced feature map is obtained based on the enhanced features and the bird's-eye view features corresponding to any grid unrelated to the target object in the refined feature map.

[0020] Optionally, the prediction result of the target classification prediction includes a target detection map and a scale policy map for the target object. The target detection map represents the distribution of the target object in the bird's-eye view feature map, and the scale policy map represents the scale refinement strategy of each grid in the bird's-eye view feature map. The step of refining the grids related to the target object in the bird's-eye view feature map according to the prediction result to obtain a refined feature map includes:

[0021] Based on the target detection map, the target objects are distributed and located to obtain candidate grids related to the target objects;

[0022] Based on the scale refinement strategy corresponding to the candidate grid in the scale strategy map, the candidate grid is scaled to obtain the refined feature map; wherein, the refined feature map contains the refined candidate grid.

[0023] Optionally, the target detection map is obtained in the following way:

[0024] The bird's-eye view feature map is compressed and extracted to obtain binary image features indicating whether the target object is included in each grid.

[0025] Based on the binary image features, each grid in the bird's-eye view feature map is probabilistically activated to obtain the target detection map.

[0026] Optionally, the scale strategy map can be obtained in the following way:

[0027] The bird's-eye view feature map is subjected to feature dimensionality reduction extraction to obtain strategy type features that represent the scale refinement strategies corresponding to each grid in the bird's-eye view feature map;

[0028] Based on the policy type features, normalize the policy activation for each grid in the bird's-eye view feature map to obtain the scale policy map.

[0029] Secondly, embodiments of this application provide a target detection device based on an unmanned aerial vehicle (UAV), the device comprising:

[0030] The grid prediction refinement module is used to perform target classification prediction based on the bird's-eye view feature map of the target scene, and refine the grid scale of the grid related to the target object in the bird's-eye view feature map according to the prediction result to obtain a refined feature map; wherein, the bird's-eye view feature map is constructed based on the multi-view image data of the target scene, and the multi-view image data is obtained by the UAV to acquire images of the target scene;

[0031] The target feature fusion module is used to perform multi-view feature enhancement on the grid related to the target object in the refined feature map based on the multi-view image data to obtain an enhanced feature map.

[0032] The target object detection module is used to perform target detection based on the enhanced feature map to obtain target detection results for the target object.

[0033] Thirdly, embodiments of this application provide a computer device, including: a memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the computer instructions to perform the method described in any one of the above technical solutions.

[0034] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer instructions, which are used to cause a computer to perform the method described in any one of the above technical solutions.

[0035] Fifthly, embodiments of this application provide a computer program product, including computer instructions, which are used to cause a computer to execute the method described in any of the above technical solutions.

[0036] The target detection method based on UAVs proposed in this application, targeting a target scene containing a target object, uses a UAV to acquire multi-view images to construct a bird's-eye view feature map. Target classification prediction is then performed based on the bird's-eye view feature map. The target classification prediction based on the bird's-eye view feature map refines the grids related to the target object in the bird's-eye view feature map, obtaining a refined feature map. Multi-view feature enhancement is then performed on the grids related to the target object in the refined feature map based on multi-view image data, resulting in an enhanced feature map that reflects the key details of the target object. This enhanced feature map is then used for target detection, yielding a target detection result. Compared with related technologies, this application, after refining the grids related to the target object in the bird's-eye view feature map to obtain a refined feature map, further enhances the finely divided grids in the refined feature map based on multi-view image data of the target scene. This improves the feature resolution of the refined feature map regarding the target object, thereby enabling more accurate capture of the key details of the target object and effectively improving the detection accuracy. Attached Figure Description

[0037] To more clearly illustrate the technical solutions in the specific embodiments of this application or the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0038] Figure 1 A flowchart illustrating the steps of a target detection method based on an unmanned aerial vehicle (UAV) provided in an embodiment of this application;

[0039] Figure 2 This is a flowchart illustrating the steps involved in obtaining the enhanced feature map in an embodiment of this application.

[0040] Figure 3 This is a step diagram illustrating the process of obtaining aggregated image features of any sub-grid in an embodiment of this application;

[0041] Figure 4 This is a flowchart illustrating the steps involved in obtaining the enhanced feature map in an embodiment of this application.

[0042] Figure 5 This is a flowchart illustrating the steps involved in obtaining the refined feature map in an embodiment of this application.

[0043] Figure 6a This is a schematic diagram of the target detection map in an embodiment of this application;

[0044] Figure 6b This is a schematic diagram of the scale strategy map in an embodiment of this application;

[0045] Figure 6c This is a schematic diagram of the enhanced feature map in an embodiment of this application;

[0046] Figure 7 This is a flowchart illustrating the steps involved in obtaining the target detection map in an embodiment of this application.

[0047] Figure 8 This is a flowchart illustrating the steps involved in obtaining the scale strategy map in an embodiment of this application.

[0048] Figure 9 A block diagram of a target detection device based on an unmanned aerial vehicle (UAV) provided in an embodiment of this application;

[0049] Figure 10 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation

[0050] This embodiment provides a target detection method based on unmanned aerial vehicles (UAVs), which can be used to detect targets using image data acquired by UAVs, thereby improving the detection accuracy of three-dimensional target objects. (Refer to...) Figure 1 As shown, the method includes:

[0051] S100. Perform target classification prediction based on the bird's-eye view feature map of the target scene, and refine the grid scale of the grid related to the target object in the bird's-eye view feature map according to the prediction results to obtain a refined feature map; wherein, the bird's-eye view feature map is constructed based on multi-view image data of the target scene, and the multi-view image data is obtained by the UAV to acquire images of the target scene.

[0052] S200. Based on the multi-view image data, perform multi-view feature enhancement on the grid related to the target object in the thinned feature map to obtain the enhanced feature map.

[0053] S300. Perform target detection based on the enhanced feature map to obtain the target detection results for the target object.

[0054] The target scene can be the scene where the target object is located. The target object can be a 3D object that needs to be detected. The type of the target object can be determined according to actual requirements. The number of target objects can be one or more. The spatial scales corresponding to different target objects can be the same or different.

[0055] Specifically, an aerial collaborative system is constructed using multiple drones. Each drone acquires images of the target scene from different perspectives relative to the target object, resulting in multi-view image data of the target scene. Understandably, each drone acquires its own camera intrinsic and extrinsic parameter matrices simultaneously when acquiring images. The camera intrinsic parameter matrix may include parameters such as the focal length and principal point of the camera mounted on the drone, while the extrinsic parameter matrix may include the rotation and translation of the drone relative to the global coordinate system of the target scene. Based on the drone's camera intrinsic and extrinsic parameter matrices, the perspective of each drone relative to the target scene can be determined and mapped to the corresponding perspective feature images in the multi-view image data.

[0056] Furthermore, based on the multi-view image data and the viewpoint corresponding to each viewpoint feature image in the multi-view image data, a bird's-eye view feature map of the target scene is constructed. In some embodiments, the process of constructing the bird's-eye view feature map may include: extracting semantic features from each viewpoint feature image in the multi-view image data to obtain a set of multi-scale depth image feature maps based on the image semantic features of each viewpoint feature image; obtaining a set of learnable and regularly arranged initial bird's-eye view query vectors, each initial bird's-eye view query vector corresponding to any bird's-eye view coordinate in the target scene, where the bird's-eye view coordinate can be the center region of any horizontal grid in the target scene; and combining the visual features at the corresponding positions in the set of multi-scale depth image feature maps at the bird's-eye view coordinates of each initial bird's-eye view query vector through a spatial cross-attention mechanism to aggregate the visual features at the corresponding positions onto the initial bird's-eye view query vector, forming a bird's-eye view feature map of the target scene, thereby providing a unified contextual basis for subsequent target detection.

[0057] Furthermore, semantic features are extracted from the bird's-eye view feature map, and target classification prediction is performed based on the feature extraction results to obtain one or more semantic classification results for the target object in the bird's-eye view feature map. It should be noted that the semantic classification results used as target classification prediction results are soft classification results, which can include one or more types such as target detection maps, scale strategy maps, position offset maps, and category heatmaps, representing the probability that different grids in the bird's-eye view feature map belong to a certain semantic category. Based on the prediction results of the target classification prediction, grids related to the target object can be identified from the bird's-eye view feature map, and targeted grid scale refinement can be performed on these target-related grids to obtain a refined feature map. It is understood that when refining the grid scale of grids related to the target object in the bird's-eye view feature map, the scale refinement strategy used can include fine-grained strategies, wide-area feature strategies, and global feature strategies. The scale refinement strategy corresponding to each target-related grid can be preset or determined through target classification prediction. Different target-related grids can use the same scale refinement strategy or different scale refinement strategies.

[0058] Furthermore, each target-related grid in the refined feature map is further refined into multiple sub-grids. The bird's-eye view feature corresponding to each sub-grid is the same as that corresponding to the target-related grid. The bird's-eye view coordinates of each sub-grid can be obtained by updating the bird's-eye view coordinates of the corresponding target-related grid according to the refinement method of the scale refinement strategy. For each sub-grid, multi-view feature enhancement is performed on the bird's-eye view feature corresponding to the sub-grid based on the image features corresponding to the sub-grid in the multi-view image data to enhance the ability of the bird's-eye view feature corresponding to the sub-grid to represent the target object. After feature enhancement is performed on all sub-grids of any target-related grid, feature enhancement is performed on any target-related grid based on the enhanced bird's-eye view features corresponding to all sub-grids. After feature enhancement is completed on all target-related grids, an enhanced feature map is obtained.

[0059] It should be noted that, in this embodiment, after refining the target object-related grid scale of the bird's-eye view features to obtain a refined feature map, multi-view feature enhancement is further performed on the refined sub-grids based on multi-view image data of the target scene. This yields accurate features related to the target object, thereby achieving proactive adaptation to the spatial scale of the target object, improving the feature accuracy of the target object in the enhanced feature map, and thus improving the detection accuracy of the target object. Furthermore, this embodiment can also refine the scale and enhance the features of grids related to the target object, while reducing the processing of grids unrelated to the target object. This achieves adaptive allocation of computational resources for the target object, reducing the computational resource cost required for image processing and improving the efficiency of target detection.

[0060] Furthermore, feature decoding is performed on the enhanced feature map to perform target detection on the target object, obtaining the target detection result. It is understood that each target-related grid in the enhanced feature map has undergone grid-scale refinement and multi-view feature enhancement based on the spatial scale of the target object. This allows for more accurate capture of the key details of the target object during target detection based on the enhanced feature map, effectively improving the detection accuracy.

[0061] The UAV-based target detection method provided in this embodiment, for a target scene containing a target object, uses a UAV to acquire multi-view images to construct a bird's-eye view feature map, and performs target classification prediction based on the bird's-eye view feature map. The target classification prediction based on the bird's-eye view feature map refines the grid related to the target object in the bird's-eye view feature map to obtain a refined feature map. Based on the multi-view image data, the grid related to the target object in the refined feature map is enhanced with multi-view features to obtain an enhanced feature map that can reflect the key details of the target object, which is then used for target detection to obtain the target detection result for the target object.

[0062] Compared with related technologies, this application refines the grid related to the target object in the bird's-eye view feature map to obtain a refined feature map. Based on the multi-view image data of the target scene, it enhances the finely divided grid in the refined feature map with multi-view features to improve the feature resolution of the target object in the refined feature map. This enables more accurate capture of the key details of the target object and effectively improves the detection accuracy of the target object.

[0063] Reference Figure 2 As shown, in one embodiment of this application, each grid related to the target object in the thinned feature map corresponds to a sub-grid; multi-view feature enhancement is performed on the grids related to the target object in the thinned feature map based on multi-view image data to obtain an enhanced feature map, including:

[0064] S210. For any subgrid, select a spatial reference point in the target scene based on the bird's-eye view coordinates of any subgrid, and project the spatial reference point onto any view feature image in the multi-view image data to obtain the view reference point.

[0065] S220. Based on the viewpoint reference point, perform predictive dynamic sampling and multi-viewpoint feature aggregation on all viewpoint feature images to obtain the aggregated image features of any sub-grid.

[0066] S230. Based on the aggregated image features of all sub-grids, perform target feature fusion on the thinned feature map to obtain the enhanced feature map.

[0067] The refined feature map, like the bird's-eye view feature map, displays the target scene from a top-down perspective. Each target-related grid in the refined feature map is refined into multiple sub-grids. The bird's-eye view features corresponding to each sub-grid are the same as those corresponding to the target-related grid. The bird's-eye view coordinates of each sub-grid can be obtained by updating the bird's-eye view coordinates of the corresponding target-related grid according to the refinement method of the scale refinement strategy. Based on the bird's-eye view features and coordinates of the sub-grid, a query vector corresponding to that sub-grid is constructed.

[0068] Specifically, for any sub-grid, its location in the target scene is determined based on its bird's-eye view coordinates. A spatial reference point is then selected at this location along a vertical direction for feature sampling in the multi-view image data based on that sub-grid. The number of spatial reference points can be determined according to actual requirements. Each spatial reference point corresponds to spatial coordinates, with the horizontal coordinate being the same as the bird's-eye view coordinates of that sub-grid. Based on the camera intrinsic and extrinsic parameter matrices of the UAV corresponding to the feature images of each viewpoint in the multi-view image data, a coordinate system transformation is performed on the spatial coordinates of the spatial reference points. These spatial reference points are then projected onto the plane corresponding to each feature image of the viewpoint, resulting in multiple viewpoint reference points corresponding to each spatial reference point.

[0069] Furthermore, the bird's-eye view features corresponding to each sub-grid are the same as those corresponding to the target-related grid. However, there is a deviation between the actual sampling position of each sub-grid and the target object, resulting in a difference between the actual bird's-eye view features of the sub-grid and those of the target-related grid. This causes the edges or details of the target object to be inaccurately detected, affecting the detection accuracy of the target object. Based on the above reasons, this embodiment predicts the degree of deviation between any spatial reference point and the target object in the view feature image where that spatial reference point is located, based on any viewpoint reference point corresponding to any spatial reference point. Dynamic sampling is then performed in the view feature image based on the prediction result to enhance the image features of the target object in any sub-grid, improving the feature accuracy in the refined feature map. It can be understood that by predicting the degree of deviation between any spatial reference point and the target object, the target object can be accurately located during sampling in the view feature image, resulting in accurate image features of the target object. This improves the sampling accuracy of any spatial reference point and provides an accurate data foundation for feature aggregation of the sub-grids.

[0070] Furthermore, for any given subgrid, multi-view feature aggregation is performed on the image features corresponding to all its spatial reference points at different viewpoints to obtain the aggregated image features of that subgrid. It can be understood that predictive dynamic sampling and multi-view feature aggregation enable any given subgrid to extract high-precision features of the target object, resulting in more accurate bird's-eye view features. This enhances the bird's-eye view features of the given subgrid for use in subsequent feature map enhancement processes.

[0071] Furthermore, for the target-related grid corresponding to any given sub-grid, feature fusion is performed on the aggregated image features of all sub-grids corresponding to that target-related grid to obtain the actual bird's-eye view feature corresponding to that target-related grid. This feature is then used to enhance the refined feature map, resulting in an enhanced feature map. It is understood that the target feature fusion method can be average pooling or similar techniques. By enhancing the bird's-eye view feature corresponding to the target-related grid based on the aggregated image features of each sub-grid, the ability of the target-related grid to represent key details of the target object is improved, thereby increasing the detail of the target object in the enhanced feature map and improving the detection accuracy of the target object.

[0072] Reference Figure 3 As shown, in one embodiment of this application, predictive dynamic sampling and multi-view feature aggregation are performed on all view feature images based on the view reference point to obtain the aggregated image features of any sub-grid, including:

[0073] S222. For any viewpoint feature image, based on the bird's-eye view feature corresponding to any subgrid in the bird's-eye view feature image, perform feature offset prediction on the viewpoint reference point in any viewpoint feature image to obtain the offset reference point.

[0074] S224. Based on the offset reference point, perform feature sampling in the feature image of any viewpoint to obtain the offset features of the feature image of any viewpoint.

[0075] S226. Based on the offset features of each of the feature images from all viewpoints, perform feature aggregation on any sub-grid to obtain the aggregated image features of any sub-grid.

[0076] Specifically, feature queries are performed on any viewpoint feature image based on the bird's-eye view features of any sub-grid to determine the corresponding position of the target object in that viewpoint feature image, thereby determining the offset reference point. In some embodiments, feature offset prediction can be performed by inputting multi-view image data and the query vector of any sub-grid into a spatial cross-attention network. The spatial cross-attention network may include an image value transformation layer, an offset prediction branch, and a weight prediction branch. The image value transformation layer may be a linear layer acting on the image features of each viewpoint feature image in the multi-view image data, outputting the part of the viewpoint feature image related to the target object as the value vector and key vector of the attention mechanism. The offset prediction branch may be a linear layer acting on the query vector of the sub-grid, outputting the degree of deviation between the viewpoint reference point and the target object. The weight prediction branch may be a linear layer acting on the query vector of the sub-grid, outputting the attention weight corresponding to the viewpoint reference point. Exemplarily, both the offset prediction branch and the weight prediction branch may adopt a multilayer perceptron structure. The structure of the offset prediction branch may be a linear layer, a non-linear activation function ReLU, a linear layer, and an offset prediction layer Logits_offset. The structure of the weight prediction branch may be a linear layer, a non-linear activation function ReLU, a linear layer, and a weight prediction layer Logits_weight.

[0077] In feature offset prediction, multi-view image data and the query vector of any sub-grid are input into a spatial cross-attention network, and then into the offset prediction branch and the weight prediction branch, respectively. The offset prediction branch queries based on the query vector and outputs the predicted offset score of the viewpoint reference point in any viewpoint feature image, which serves as the feature offset of that viewpoint reference point. The weight prediction branch queries based on the query vector and outputs the attention weight score corresponding to this offset prediction score. The attention weight score is normalized using the Softmax function to obtain the attention weight corresponding to the feature offset. The offset reference point is determined based on its position in the feature image of any viewpoint, a process that can be represented by the following formula:

[0078]

[0079] in, Use the offset reference point; As a reference point for perspective; This is the feature offset corresponding to the viewpoint reference point.

[0080] Further, in the feature image of any viewpoint, feature sampling is performed on the offset reference point to obtain the image features of the feature image of any viewpoint at the offset reference point, which are used as offset features. It is understood that the number of offset features corresponding to any sub-grid is related to the number of spatial reference points and the number of views corresponding to the viewpoint feature image. After obtaining all the offset features corresponding to any sub-grid, feature aggregation is performed on the sub-grid based on these offset features to obtain the aggregated image features of the sub-grid. In some embodiments, feature aggregation can be obtained by weighted summation of the corresponding offset features according to the attention weight corresponding to each feature offset, and the process can be expressed by the following formula:

[0081]

[0082] in, Aggregate image features; For offset features; Attention weights; The total number of viewpoints corresponding to the viewpoint feature image; This represents the total number of spatial reference points.

[0083] Reference Figure 4 As shown, in one embodiment of this application, the thinned feature map is fused with target features based on the aggregated image features of all sub-grids to obtain an enhanced feature map, including:

[0084] S232. For any target-related grid in the refined feature map that is related to the target object, aggregate the aggregated image features of all sub-grids corresponding to any target-related grid to obtain the enhanced features of the target-related grid.

[0085] S234. Obtain the enhanced feature map based on the bird's-eye view feature corresponding to any grid unrelated to the target object in the enhanced and refined feature maps.

[0086] Specifically, for any target-related grid, the aggregated image features of all sub-grids corresponding to that target-related grid are aggregated to obtain the enhanced features of that target-related grid. It is understood that the target feature fusion method can be average pooling or similar techniques. Based on the aggregated image features of each sub-grid, the bird's-eye view features corresponding to the target-related grid are enhanced, improving the target-related grid's ability to represent key details of the target object. This, in turn, improves the detail of the target object in the enhanced feature map and enhances the detection accuracy of the target object.

[0087] Furthermore, for target-independent grids in the refined feature map that are unrelated to the target object, the bird's-eye view features of each target-independent grid remain unchanged. An enhanced feature map is obtained based on the enhanced features of the target-related grids and the bird's-eye view features of the target-independent grids, which is then used for target detection.

[0088] Reference Figure 5 As shown in one embodiment of this application, the prediction result of target classification prediction includes a target detection map and a scaling strategy map for the target object. The target detection map represents the distribution of the target object in the bird's-eye view feature map, and the scaling strategy map represents the scaling strategy of each grid in the bird's-eye view feature map. Based on the prediction result, the grid scale of the grid related to the target object in the bird's-eye view feature map is refined to obtain a refined feature map, including:

[0089] S110. Based on the target detection map, the target objects are distributed and located to obtain candidate grids related to the target objects.

[0090] S120. Based on the scale refinement strategy corresponding to the candidate grid in the scale strategy map, the candidate grid is scaled to obtain a refined feature map; wherein the refined feature map contains the refined candidate grid.

[0091] Specifically, the object detection map can be obtained based on the probability of a target object existing within each grid in the bird's-eye view feature map, representing the distribution of target objects in the bird's-eye view feature map. Target objects are located based on the distribution of the object detection map. A preset probability threshold is set, and all grids in the object detection map are traversed. Grids with probabilities exceeding the probability threshold are selected as candidate grids associated with the target object. For example, the probability threshold can be 0.5.

[0092] Reference Figure 6a As shown, Figure 6a The color of each grid cell represents the probability of a target object existing within that cell; a brighter grid cell has a higher probability of containing a target object than a darker grid cell. Figure 6a In the process, the candidate grid can be a grid with positions (2,2), (2,3), (3,2), (3,3), (1,6), (1,7), (2,6), (2,7), (5,5), (5,6), (6,5) and (6,6).

[0093] Furthermore, the scale policy map represents the scale refinement strategy for each grid in the bird's-eye view feature map. The scale refinement strategy can include fine-grained strategies, wide-area feature strategies, and global feature strategies. The scale refinement strategy corresponding to each target-related grid can be preset or determined through target classification prediction. Different target-related grids can use the same scale refinement strategy or different scale refinement strategies. Based on the candidate grids, the scale refinement strategy with the highest probability among all strategies corresponding to the corresponding grid is determined in the scale policy map, and the candidate grids are scale refined according to the scale refinement strategy to obtain the refined feature map.

[0094] For example, the process of determining the scale refinement strategy from the scale strategy map can be represented by the following formula:

[0095]

[0096] in, For the corresponding scale refinement strategy; For scale strategy maps; For target-related grids; An index for the scale refinement strategy. This represents the total number of scaling strategies.

[0097] Reference Figure 6b As shown, Figure 6b The color of each grid cell represents a different scale refinement strategy. Strategy 1 is used for candidate grid cells at positions (2,2), (2,3), (3,2), and (3,3); Strategy 2 is used for candidate grid cells at positions (5,5), (5,6), (6,5), and (6,6); Strategy 3 is used for candidate grid cells at positions (1,6), (1,7), (2,6), and (2,7); and Strategy 0 is used for other non-candidate grid cells. Strategy 0 may not refine the grid, Strategy 1 may refine the grid into 2×2 sub-grids, Strategy 2 may refine the grid into 4×4 sub-grids, and Strategy 3 may refine the grid into 8×8 sub-grids. (See reference...) Figure 6c As shown, according to Figure 6b The scale refinement strategy shown refines each grid into multiple sub-grids, resulting in a refined feature map. Based on the refined feature map, multi-view feature enhancement is performed on the grids related to the target object in the refined feature map using multi-view image data, resulting in an enhanced feature map.

[0098] In some embodiments, the bird's-eye view feature map can be processed by an adaptive thinning model to obtain a thinned feature map. The adaptive thinning model may include a feature extraction backbone and object detection and scaling policy branches connected to the backbone network. The feature extraction backbone may include multiple consecutive convolutional blocks, each containing a 3×3 convolutional layer, a batch normalization layer, and a ReLU nonlinear activation function. The 3×3 convolutional layer has 256 input and 256 output channels. The adaptive thinning model can be trained using a task loss function, which may include object detection loss, detection assistance loss, and policy assistance loss. The object detection loss can be obtained based on the error between the object detection result of the enhanced feature map during training and the ground truth value of the object in the target scene.

[0099] The detection auxiliary loss can be obtained based on the error between the target detection map and the ground truth map of the target scene during training, and can be expressed by the following formula:

[0100]

[0101] in, To detect auxiliary losses; For target detection maps; This is a truth map of the target scene, where the location of the target object is represented as 1, and other locations are represented as 0; The loss function can be represented by other loss functions that can handle class imbalance problems, including but not limited to the Dice loss function.

[0102] The policy-assisted loss can be obtained based on the error between the scaled policy map and the supervised policy map during training, and can be expressed by the following formula:

[0103]

[0104] in, To assist in strategy loss; For monitoring the strategy map, there are preset scale labels assigned based on the area occupied by each 3D truth box in the target truth map; This represents the cross-entropy loss function.

[0105] The task loss function is derived from the target detection loss, detection assistance loss, and policy assistance loss, and its form can be expressed by the following formula:

[0106]

[0107] in, The task loss function; Loss for target detection; and These are the hyperparameters for balancing the weights. Based on the task loss function described above, the total task loss function is minimized using the backpropagation algorithm to complete the training of the adaptive refinement model.

[0108] Reference Figure 7 As shown, in one embodiment of this application, the target detection map is obtained in the following manner:

[0109] S130. Perform feature compression extraction on the bird's-eye view feature map to obtain binary image features indicating whether each grid contains a target object.

[0110] S140. Perform probabilistic logic activation on each grid in the bird's-eye view feature map based on the binary image features to obtain the target detection map.

[0111] Specifically, the bird's-eye view feature map is input into the trained adaptive refinement model. After semantic feature extraction via the feature extraction backbone, the semantic feature map of the bird's-eye view feature map is input into the object detection branch. In some embodiments, the object detection branch can be a lightweight network module, including a 1×1 convolutional layer and a sigmoid activation function. The 1×1 convolutional layer has 256 input channels and 1 output channel. In the object detection branch, the number of feature channels of the semantic feature map of the bird's-eye view feature map is reduced from 256 to 1, resulting in a binary image feature representing whether each grid cell contains a target object, and the corresponding object detection logic map is obtained. The object detection logic map is then passed through the sigmoid activation function to perform probabilistic logical activation on each grid cell, resulting in the object detection map. Exemplarily, the process of generating the object detection map in the object detection branch can be represented by the following formula:

[0112]

[0113] in, A bird's-eye view of the features; For object detection branch, Learnable network parameters for the object detection branch; This is the Sigmoid activation function.

[0114] Reference Figure 8 As shown, in one embodiment of this application, the scale strategy map is obtained in the following manner:

[0115] S150. Perform feature dimensionality reduction extraction on the bird's-eye view feature map to obtain the strategy type feature representing the scale refinement strategy corresponding to each grid in the bird's-eye view feature map.

[0116] S160. Normalize the policy activation of each grid in the bird's-eye view feature map according to the policy type characteristics to obtain the scale policy map.

[0117] Specifically, the bird's-eye view feature map is input into the trained adaptive refinement model. After semantic feature extraction via the feature extraction backbone, the semantic feature map of the bird's-eye view feature map is input into the scaling policy branch. The scaling policy branch can be a lightweight network module, consisting of a 1×1 convolutional layer and a Softmax function along the channel dimension. This 1×1 convolutional layer has 256 input channels and 5 output channels. In the scaling policy branch, the number of feature channels in the semantic feature map of the bird's-eye view feature map is reduced from 256 to 5, resulting in 5 policy type features representing the scaling refinement policy corresponding to each grid in the bird's-eye view feature map, and the corresponding scaling policy logic map is obtained. The scaling policy logic map is then passed along the channel dimension through the Softmax activation function to normalize the activation policy for each grid, resulting in the scaling policy map.

[0118] For example, the process of generating a scale policy map in the scale policy branch can be represented by the following formula:

[0119]

[0120] in, This is a branch of scale strategy. Its network parameters; This refers to the Softmax function.

[0121] Accordingly, please refer to Figure 9 This application provides a target detection device based on an unmanned aerial vehicle (UAV), the device comprising:

[0122] The mesh prediction refinement module 910 is used to perform target classification prediction based on the bird's-eye view feature map of the target scene, and refine the mesh scale of the mesh related to the target object in the bird's-eye view feature map according to the prediction results to obtain a refined feature map; wherein, the bird's-eye view feature map is constructed based on multi-view image data of the target scene, and the multi-view image data is obtained by the UAV to acquire images of the target scene.

[0123] The target feature fusion module 920 is used to enhance the grid related to the target object in the refined feature map based on multi-view image data, so as to obtain an enhanced feature map.

[0124] The target object detection module 930 is used to perform target detection based on the enhanced feature map and obtain the target detection result about the target object.

[0125] In some alternative implementations, the target feature fusion module 920 includes:

[0126] The reference selection projection unit is used to select a spatial reference point in the target scene based on the bird's-eye view coordinates of any sub-grid, and project the spatial reference point onto any view feature image in the multi-view image data to obtain the view reference point.

[0127] The dynamic sampling aggregation unit is used to perform predictive dynamic sampling and multi-view feature aggregation on all view feature images based on the view reference point to obtain the aggregated image features of any sub-grid.

[0128] The target feature fusion unit is used to perform target feature fusion on the thinned feature map based on the aggregated image features of all sub-grids to obtain an enhanced feature map.

[0129] In some alternative implementations, the dynamic sampling aggregation unit includes:

[0130] The feature offset prediction sub-unit is used to predict the feature offset of the view reference point in any view feature image based on the view feature corresponding to any sub-grid in the view feature image, and obtain the offset reference point.

[0131] The offset feature sampling subunit is used to perform feature sampling in any viewpoint feature image based on the offset reference point to obtain the offset features of the feature image in any viewpoint.

[0132] The offset feature aggregation sub-unit is used to perform feature aggregation on any sub-grid based on the offset features of all view feature images, so as to obtain the aggregated image features of any sub-grid.

[0133] In some optional implementations, the target feature fusion unit includes:

[0134] The image feature aggregation subunit is used to aggregate the aggregated image features of all sub-grids corresponding to any target-related grid in the thinned feature map that is related to the target object, so as to obtain the enhanced features of the target-related grid.

[0135] The feature image acquisition subunit is used to obtain the enhanced feature map based on the bird's-eye view features corresponding to any grid unrelated to the target object in the enhanced feature map and the thinned feature map.

[0136] In some alternative implementations, the grid prediction refinement module 910 includes:

[0137] Candidate grid localization units are used to distribute and locate target objects based on the target detection map, and obtain candidate grids related to the target objects.

[0138] The candidate mesh refinement unit is used to refine the candidate mesh based on the scale refinement strategy corresponding to the candidate mesh in the scale strategy map, so as to obtain a refined feature map; wherein the refined feature map contains the refined candidate mesh.

[0139] In some optional implementations, the grid prediction refinement module 910 further includes:

[0140] The binary feature extraction unit is used to compress and extract features from the bird's-eye view feature map to obtain binary image features indicating whether each grid contains a target object.

[0141] The target logic activation unit is used to perform probabilistic logic activation on each grid in the bird's-eye view feature map based on the binary image features to obtain the target detection map.

[0142] In some optional implementations, the grid prediction refinement module 910 further includes:

[0143] The dimension reduction feature extraction unit is used to perform dimension reduction extraction on the bird's-eye view feature map to obtain the strategy type feature representing the scale refinement strategy corresponding to each grid in the bird's-eye view feature map.

[0144] The normalized policy activation unit is used to perform normalized policy activation on each grid in the bird's-eye view feature map according to the policy type characteristics, so as to obtain the scale policy map.

[0145] Further functional descriptions of the above modules and units are the same as those in the corresponding embodiments described above, and will not be repeated here.

[0146] In this embodiment, the target detection device based on UAV is presented in the form of a functional unit. Here, a unit refers to an ASIC (Application Specific Integrated Circuit), a processor and memory that execute one or more software or fixed programs, and / or other devices that can provide the above functions.

[0147] Please see Figure 10 , Figure 10 This is a schematic diagram of a computer device according to an embodiment of this application. As shown in the figure, the computer device includes one or more processors 10, a memory 20, and interfaces for connecting the various components, including high-speed interfaces and low-speed interfaces. The various components communicate with each other using different buses and can be mounted on a common motherboard or otherwise installed as needed. The processors can process instructions executed within the computer device, including instructions stored in or on memory to display graphical information of a GUI on an external input / output device (such as a display device coupled to the interface). In some alternative embodiments, multiple processors and / or multiple buses can be used with multiple memories and multiple memory modules, if desired. Similarly, multiple computer devices can be connected, each providing some of the necessary operations (e.g., as a server array, a group of blade servers, or a multiprocessor system). Figure 10 Take a processor 10 as an example.

[0148] Processor 10 may be a central processing unit, a network processor, or a combination thereof. Processor 10 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The programmable logic device may be a complex programmable logic device (CAMP), a field-programmable gate array (FPGA), a general-purpose array logic (GPA), or any combination thereof.

[0149] The memory 20 stores instructions executable by at least one processor 10 to cause the at least one processor 10 to perform the method shown in the above embodiments.

[0150] The memory 20 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the computer device. Furthermore, the memory 20 may include high-speed random access memory and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some alternative embodiments, the memory 20 may optionally include memory remotely located relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0151] The memory 20 may include volatile memory, such as random access memory; the memory may also include non-volatile memory, such as flash memory, hard disk or solid-state drive; the memory 20 may also include a combination of the above types of memory.

[0152] The computer device also includes a communication interface 30 for communicating with other devices or communication networks.

[0153] This application also provides a computer-readable storage medium. The methods described in this application can be implemented in hardware or firmware, or implemented as recordable on a storage medium, or implemented as computer code downloaded over a network and originally stored on a remote storage medium or a non-transitory machine-readable storage medium and subsequently stored on a local storage medium. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium can also include combinations of the above types of memory. It is understood that computers, processors, microprocessor controllers, or programmable hardware include storage components capable of storing or receiving software or computer code. When the software or computer code is accessed and executed by the computer, processor, or hardware, the methods shown in the above embodiments are implemented.

[0154] This application provides a computer program product including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the method of any embodiment of this application.

Claims

1. A method for target detection based on a UAV, characterized in that, The method comprises: According to the bird's eye view feature map of the target scene, target classification prediction is performed, and grid scale refinement is performed on the grid related to the target object in the bird's eye view feature map according to the prediction result, to obtain a refined feature map; wherein the bird's eye view feature map is constructed based on multi-view image data of the target scene, the multi-view image data is obtained by image acquisition of the target scene by a UAV, any target-related grid in the refined feature map corresponding to the target object has a sub-grid, and any sub-grid in the sub-grid corresponds to an image feature in the multi-view image data; For any sub-grid, a spatial reference point is selected in the target scene based on the bird's eye view coordinates of the any sub-grid, and the spatial reference point is projected onto any view feature image in the multi-view image data to obtain a view reference point; according to the deviation degree between the view reference point and the target object, prediction dynamic sampling is performed on all view feature images, and multi-view feature aggregation is performed on the sampling results to obtain the aggregated image feature of the any sub-grid; According to the aggregated image features of all sub-grids, target feature fusion is performed on the refined feature map to obtain an enhanced feature map; and target detection is performed according to the enhanced feature map to obtain a target detection result about the target object.

2. The method of claim 1, wherein, According to the deviation degree between the view reference point and the target object, prediction dynamic sampling is performed on all view feature images, and multi-view feature aggregation is performed on the sampling results to obtain the aggregated image feature of the any sub-grid, comprising: For any view feature image, according to the bird's eye view feature corresponding to the any sub-grid in the bird's eye view feature map, feature offset prediction is performed on the view reference point in the any view feature image to obtain an offset reference point; Based on the offset reference point, feature sampling is performed in the any view feature image to obtain offset features of the any view feature image; According to the offset features of all view feature images, feature aggregation is performed on the any sub-grid to obtain the aggregated image feature of the any sub-grid.

3. The method of claim 1, wherein, According to the aggregated image features of all sub-grids, target feature fusion is performed on the refined feature map to obtain the enhanced feature map, comprising: For any target-related grid related to the target object in the refined feature map, the aggregated image features of all sub-grids corresponding to the any target-related grid are aggregated to obtain enhanced features of the target-related grid; According to the enhanced features and the bird's eye view features corresponding to any grid unrelated to the target object in the refined feature map, the enhanced feature map is obtained.

4. The method of claim 1, wherein, The prediction result of the target classification prediction includes a target detection map and a scale strategy map about the target object, the target detection map represents the distribution of the target object in the bird's eye view feature map, and the scale strategy map represents the scale refinement strategy of each grid in the bird's eye view feature map; According to the prediction result, grid scale refinement is performed on the grid related to the target object in the bird's eye view feature map to obtain a refined feature map, comprising: According to the target detection map, the target object is distributedly positioned, and a candidate grid related to the target object is obtained; Based on the corresponding scale refinement strategy of the candidate grid in the scale strategy map, the candidate grid is refined in scale, and a refined feature map is obtained; wherein the refined feature map contains a refined candidate grid.

5. The method of claim 4, wherein, The target detection map is obtained in the following manner: The aerial feature map is feature-compressed to obtain a binary image feature representing whether the target object is included in each grid; According to the binary image feature, each grid in the aerial feature map is activated by probability logic, and the target detection map is obtained.

6. The method of claim 4, wherein, The scale strategy map is obtained in the following manner: The aerial feature map is feature-reduced to obtain a strategy type feature representing the scale refinement strategy corresponding to each grid in the aerial feature map; According to the strategy type feature, each grid in the aerial feature map is activated by a normalization strategy to obtain the scale strategy map.

7. An unmanned aerial vehicle based target detection apparatus, comprising: The device comprises: A grid prediction refinement module is configured to perform target classification prediction based on an aerial feature map of a target scene, and to perform grid scale refinement on a grid related to a target object in the aerial feature map based on a prediction result, to obtain a refined feature map; wherein the aerial feature map is constructed based on multi-view image data of the target scene, the multi-view image data is obtained by image acquisition of the target scene by a UAV, any target-related grid related to the target object in the refined feature map corresponds to a sub-grid, and any sub-grid in the sub-grid corresponds to an image feature in the multi-view image data; A target feature fusion module is configured to, for any sub-grid, select a spatial reference point in the target scene based on the aerial coordinates of the any sub-grid, project the spatial reference point onto any view feature image in the multi-view image data to obtain a view reference point, perform prediction dynamic sampling on all view feature images according to the deviation degree between the view reference point and the target object, and aggregate the sampling results to obtain aggregated image features of the any sub-grid; A target object detection module is configured to perform target feature fusion on the refined feature map based on the aggregated image features of all sub-grids to obtain an enhanced feature map, and to perform target detection based on the enhanced feature map to obtain a target detection result about the target object.

8. A computer device, comprising: It comprises: A memory and a processor are communicatively connected, the memory stores computer instructions, and the processor executes the computer instructions to perform the method of any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer instructions for causing a computer to execute the method of any one of claims 1-6.

Citation Information

Patent Citations

  • Target detection method and device, model training method and device, equipment and storage medium

    CN117746133A

  • Image enhancement method, system and equipment for automatic driving vision task

    CN120563789A