Unmanned aerial vehicle target identification method and system

The target recognition model is deployed through the cloud-edge collaborative platform, combining image cropping and position information restoration and screening of detection scenarios and route settings, the problem of insufficient small target recognition accuracy and efficiency of drones in complex environments is solved, and stable and efficient small target recognition is achieved.

CN120259919APending Publication Date: 2025-07-04CISDI INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510375138.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

UAVs have poor accuracy and efficiency in complex environments. The existing methods are insufficiently adaptable and have high computing overhead in different environments, making it difficult to deploy efficiently on resource-constrained edge devices.

Method used

The target recognition model is deployed through the cloud-edge collaborative platform, and the original image is cropped according to the detection scenario and route settings, sub-maps are generated and position information is restored and filtered, and target recognition is optimized with the non-maximum suppression algorithm with cross-size ratio.

Benefits of technology

It improves the accuracy and efficiency of small target recognition, reduces the missed detection rate, ensures the stability of the identification effect in various environments, adapts to different detection scenarios, and reduces calculation overhead.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259919A_ABST
    Figure CN120259919A_ABST
Patent Text Reader

Abstract

The invention provides an unmanned aerial vehicle target identification method and system, and the method comprises the steps: obtaining an original image, cutting the original image according to a target size, obtaining a plurality of sub-images and sub-image information, and setting the target size according to a scene and a route detected by an unmanned aerial vehicle; the original image and the sub-image are input into a target recognition model, first position information of the to-be-detected target in the original image and second position information of the to-be-detected target in the sub-image are obtained, and the target recognition model is deployed according to the detected scene and route; restoring the second position information of the to-be-detected target in the subgraph to obtain third position information; and screening target position information of the to-be-detected target from the first position information and the plurality of pieces of third position information. The original image is cut and reconstructed based on the detected scene and route setting size, and the corresponding target recognition model is deployed according to the specific scene and route, so that the stability of the recognition effect in various environments is ensured, and the precision and efficiency of small target recognition are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of visual detection, and in particular to a method and system for drone target recognition. Background Art

[0002] With the continuous progress of technology, drones have gradually become an important tool in the fields of smart parks, smart transportation, environmental monitoring, public safety, etc. due to their high mobility, flexible perspective, low cost, and convenient deployment. However, when drones perform tasks, they often encounter complex environmental limitations such as building occlusion, thick vegetation, or dense overhead power lines. At the same time, since the captured image data must cover the entire inspection area, the acquisition distance of the video data is relatively far and the focal length cannot be adjusted arbitrarily, resulting in great challenges in recognizing small targets during the inspection task. Therefore, it is an urgent problem to improve the small target recognition ability of drones in complex environments to improve the accuracy and efficiency of the inspection task.

[0003] In related technologies, the improvement of target detection methods for small targets usually starts from multiple aspects such as network structure modification and multi-scale learning. For example, HyperNet (HyperNetwork, super network) fuses information at different levels through a multi-branch network to improve the accuracy of small target detection; the feature pyramid network structure can obtain feature pyramids of different sizes through a top-down feature extraction network and a bottom-up feature fusion network, thereby outputting a feature map with more semantic information; based on the feature pyramid network structure, PANet (Path Aggregation Network) connects a bottom-up enhancement path from the bottom of the feature pyramid to shorten the transmission path of information fusion, so as to input the position information with fine-grained features into the fusion network to improve the detection ability of the feature pyramid structure. Although the above methods have improved the accuracy of small target detection to a certain extent, there are still the following deficiencies: one is the insufficient adaptability to general scenarios, and the detection effect is unstable in different environments; the other is that the model has a large computational cost for high-resolution images and is difficult to be efficiently deployed on resource-constrained edge devices. As a result, the accuracy and efficiency of drone small target recognition are not good. Summary of the Invention

[0004] To provide a basic understanding of some aspects of the disclosed embodiments, a simple summary is given below. This summary is not a comprehensive review, nor is it intended to identify key / important elements or delineate the scope of protection of these embodiments. Instead, it serves as a preamble to the subsequent detailed description.

[0005] In view of the above-mentioned disadvantages of the prior art, the present invention discloses a method and system for unmanned aerial vehicle (UAV) target recognition, which is used to solve the technical problems of poor accuracy and efficiency of UAV small target recognition in the prior art.

[0006] In a first aspect, the present application provides a method for UAV target recognition. The method includes: obtaining an original image of a to-be-detected area, and cropping the original image according to a preset target size to obtain multiple sub-images and sub-image information of each sub-image, where the target size is set according to the detection scenario and detection route of the UAV; inputting the original image and the multiple sub-images into a target recognition model to obtain a recognition result, where the recognition result includes first position information of a to-be-detected target in the original image and second position information of the to-be-detected target in the multiple sub-images, and the target recognition model is deployed by a cloud-edge collaborative platform according to the detection scenario and the detection route; restoring the second position information of the to-be-detected target in each sub-image according to each sub-image information to obtain multiple third position information of the to-be-detected target in the original image; and screening the position information of the to-be-detected target among the first position information and the multiple third position information to obtain the target position information of the to-be-detected target.

[0007] In an embodiment of the present invention, the cropping of the original image according to the preset target size to obtain multiple sub-images and sub-image information of each sub-image includes: setting a sub-image size and an overlapping rate of adjacent sub-images according to the target size; and slidingly cropping the original image according to the sub-image size and the overlapping rate of adjacent sub-images to obtain the multiple sub-images and the sub-image information of each sub-image, where the sub-image information includes first coordinate information of each sub-image in the original image.

[0008] In an embodiment of the present invention, the obtaining manner of the multiple third position information includes: obtaining first coordinate information of each sub-image in the original image and second coordinate information of the to-be-detected target in each sub-image, where the second position information includes the second coordinate information; and calculating the first coordinate information and the second coordinate information to obtain the third position information of the to-be-detected target in the original image, where the third position information includes third coordinate information.

[0009] In an embodiment of the present invention, there is at least one target to be measured, and each target to be measured corresponds to the first position information and the multiple third position information. The first position information is at least one. The method for obtaining the target position information includes: screening at least one piece of the first position information of each target to be measured to obtain first candidate position information; constructing at least one first position information set according to the type of the target to be measured. The position information in each first position information set includes the first candidate position information and the multiple third position information corresponding to all targets to be measured under one type of target to be measured; for each first position information set, sort the multiple position information according to the confidence level. The recognition result further includes the confidence levels of the first position information and the multiple second position information; according to the position information with the highest confidence level in the first position information set, screen the position information of each target to be measured in the first position information set to obtain the target position information.

[0010] In an embodiment of the present invention, the step of screening the position information of each target to be measured in the first position information set according to the position information with the highest confidence level in the first position information set to obtain the target position information includes: constructing a second position information set, and moving the position information with the highest confidence level in the first position information set into the second position information set; calculating the intersection-over-union ratio between the position information with the highest confidence level and each remaining position information in the first position information set; removing the position information in the first position information set with an intersection-over-union ratio greater than a preset threshold, and moving the position information with the highest confidence level in the first position information set after removal into the second position information set; repeating the foregoing steps, continuing to calculate the intersection-over-union ratio and remove the position information. When the first position information set is empty, the obtained second position information set is the target position information set, and the target position information set includes the target position information of all targets to be measured under one type of target to be measured.

[0011] In an embodiment of the present invention, before restoring the second position information of the target to be measured in each sub-graph, it further includes: sorting at least one piece of the second position information corresponding to the target to be measured in each sub-graph according to the confidence level to construct a third position information set; iteratively screening the second position information in the third position information set, and moving the position information with the highest confidence level in the third position information set into a pre-constructed fourth position information set in each iteration; when the third position information set is empty, using the position information in the fourth position information set as the second candidate position information of the target to be measured in each sub-graph.

[0012] In an embodiment of the present invention, the training of the target detection model includes: obtaining a sample image set, where the sample image set includes multiple aerial images captured by a drone of various targets in the detection scene and under the detection flight path, as well as multiple top-down images captured by a fixed camera; preprocessing the sample image set, and the preprocessing includes quality enhancement, data enhancement, target annotation, and construction of a training set and a validation set; setting up a training environment and a training network, constructing an initial model for the detection scene and under the detection flight path, and constructing a training script according to the training set, a preset number of iterations, the number of batch samples, and the learning rate; training the initial model according to the training script, and validating the trained initial model using the validation set, and obtaining the target detection model when both the average precision and recall of the target meet the threshold requirements.

[0013] In an embodiment of the present invention, before inputting the original image and the multiple sub-images into the target recognition model, it further includes: adjusting the sizes of the original image and the multiple sub-images according to a preset image size; normalizing the pixel values of the size-adjusted original image and the multiple sub-images; and reconstructing the normalized original image and the multiple sub-images into the batch input format according to the batch input format of the target recognition model.

[0014] In an embodiment of the present invention, after obtaining the target position information of the target to be detected, it further includes: obtaining event display content, where the event display content includes the detection scene, the detection flight path, the detection time of the target to be detected, the original image, and the target position information of the target to be detected; generating a detection event according to the event display content, and pushing the detection event to the cloud-edge collaboration platform, so that the cloud-edge collaboration platform centrally displays the detection event, searches for the event processing method of the detection event, monitors the event status of the detection event, and centrally displays the event processing method and the event status.

[0015] Second aspect, the present application provides a drone target recognition system, the system comprising: an image processing module, configured to obtain an original image of a to-be-detected area, and crop the original image according to a preset target size to obtain multiple sub-images and sub-image information of each sub-image, the target size being set according to the detection scenario and detection flight path of the drone; a target recognition module, configured to input the original image and the multiple sub-images into a target recognition model to obtain a recognition result, the recognition result including first position information of a to-be-detected target in the original image and second position information of the to-be-detected target in the multiple sub-images, the target recognition model being deployed by a cloud-edge collaborative platform according to the detection scenario and the detection flight path; a position restoration module, configured to restore the second position information of the to-be-detected target in each sub-image according to each sub-image information to obtain multiple third position information of the to-be-detected target in the original image; a position screening module, configured to screen the position information of the to-be-detected target among the first position information and the multiple third position information to obtain the target position information of the to-be-detected target.

[0016] Third aspect, the present application provides an electronic device, the electronic device comprising: one or more processors; a storage device, configured to store one or more programs, when the one or more programs are executed by the one or more processors, enabling the electronic device to implement the drone target recognition method described in the first aspect.

[0017] As described above, a drone target recognition method and system provided by an embodiment of the present invention have the following beneficial effects:

[0018] First, obtain the original image of the area to be detected, crop the original image according to the preset target size to obtain multiple sub-images and the sub-image information of each sub-image. The target size is set according to the detection scenario and detection route of the drone. Then, input the original image and the multiple sub-images into the target recognition model to identify the first position information of the target to be detected in the original image and the second position information in the multiple sub-images. The target recognition model is deployed by the cloud-edge collaborative platform according to the detection scenario and detection route. Then, according to the information of each sub-image, restore the second position information of the target to be detected in each sub-image to obtain multiple third position information of the target to be detected in the original image. Finally, screen the position information of the target to be detected from the first position information and the multiple third position information, and crop and reconstruct the original image based on the specific detection scenario and detection route to set the target size, so that the target to be detected, whether it is a conventional-size target or a small target, is distributed in the sub-images, ensuring that the small targets in the sub-images are accurately identified after being input into the model. Moreover, the target recognition model is deployed according to the specific detection scenario and detection route, adapting to the current detection environment, thereby making the recognition effect of small targets stable in various environments, improving the accuracy and efficiency of small target recognition, being able to effectively reduce the missed detection rate of small targets, and ensuring the inspection accuracy in the intelligent recognition of drone targets.

[0019] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit this application. Brief Description of the Drawings

[0020] The accompanying drawings here are incorporated into the specification and form a part of this specification, showing the embodiments consistent with this application, and are used together with the specification to explain the principles of this application. Obviously, the accompanying drawings in the following description are only some embodiments of this application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts. In the drawings:

[0021] Figure 1 is a schematic diagram of the implementation environment of a drone target recognition system shown in an exemplary embodiment of this application;

[0022] Figure 2 is a flowchart of a drone target recognition method shown in an exemplary embodiment of this application;

[0023] Figure 3 is a schematic diagram of the cropping result of an original image shown in an exemplary embodiment of this application;

[0024] Figure 4 is a schematic diagram of batch integration of pictures shown in an exemplary embodiment of this application;

[0025] Figure 5It is a flowchart of a specific UAV target recognition method shown in an exemplary embodiment of the present application;

[0026] Figure 6 It is a block diagram of a UAV target recognition system shown in an exemplary embodiment of the present application;

[0027] Figure 7 It is a schematic structural diagram of an electronic device shown in an exemplary embodiment of the present application. Detailed implementation manners

[0028] The following will describe the embodiments of the present invention with reference to the accompanying drawings and preferred embodiments. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be understood that the preferred embodiments are only for explaining the present invention, rather than for limiting the protection scope of the present invention.

[0029] It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. Therefore, only the components related to the present invention are shown in the diagrams, rather than being drawn according to the number, shape, and size of the components in actual implementation. The type, quantity, and ratio of each component in actual implementation can be arbitrarily changed, and the component layout type may also be more complex.

[0030] In the following description, a large number of details are discussed to provide a more thorough explanation of the embodiments of the present invention. However, it is obvious to those skilled in the art that the embodiments of the present invention can be implemented without these specific details. In other embodiments, well-known structures and devices are shown in the form of block diagrams rather than in detail to avoid making the embodiments of the present invention difficult to understand.

[0031] Drones are widely used in smart parks, smart transportation, environmental monitoring, public safety and other fields. The improvement of the recognition accuracy and efficiency of small targets of drones in complex environments is often achieved by optimizing the network structure and enhancing multi-scale learning capabilities. For example, HyperNet fuses information at different levels through a multi-branch network; the feature pyramid network structure uses a top-down feature extraction network and a bottom-up feature fusion network; PANet, based on the feature pyramid network structure, connects a bottom-up enhancement path from the bottom of the feature pyramid to input the location information with fine-grained features into the fusion network. However, the inventors of this application have found that the above methods still have shortcomings: first, they are not adaptable enough to general scenarios, and the detection effect is unstable in different environments; second, the model has a large computational overhead for high-resolution images, and it is difficult to deploy efficiently on resource-constrained edge devices. As a result, the accuracy and efficiency of drone small target recognition are poor.

[0032] Therefore, see Figure 1 , Figure 1 FIG. 1 is a schematic diagram of an implementation environment of a drone target recognition system shown in an exemplary embodiment of the present application. Figure 1 As shown, the implementation environment may include a drone 110, an edge computing device 120, a cloud-edge collaboration platform 130, and a drone target recognition system 140, wherein the drone target recognition system 140 is integrated in the edge computing device 120. The drone 110 is used to capture the original image of the area to be detected and transmit it to the drone target recognition system 140 in the edge computing device 120, and the cloud-edge collaboration platform 130 is used to build and train target detection models under various detection scenarios and detection routes, and when the drone target recognition system 140 starts recognition, the corresponding target detection model is issued and deployed to the drone target recognition system 140 based on the detection scenario and detection route. The drone target recognition system 140 sets the target size based on the specific detection scene and detection route, and crops and reconstructs the original image, so that the target to be detected, whether it is a regular-sized target or a small target, is distributed in a sub-image, ensuring that the small target in the sub-image is accurately identified after the sub-image is input into the model, and the target recognition model is adapted to the current detection environment according to the specific detection scene and detection route deployment, thereby making the recognition effect of small targets in various environments stable, improving the accuracy and efficiency of small target recognition, and being able to effectively reduce the missed detection rate of small targets. In the drone target intelligent recognition, the inspection accuracy is guaranteed.

[0033] See also Figure 2 , Figure 2 is a flowchart of a method for identifying a drone target according to an exemplary embodiment of the present application. The method can be applied to Figure 1The illustrated implementation environment is specifically executed by the UAV target recognition system in this implementation environment. It should be understood that this method can also be applied to other exemplary implementation environments and specifically executed by devices in other implementation environments. This embodiment does not limit the implementation environment applicable to this method.

[0034] As Figure 2 shown, in an exemplary embodiment, the UAV target recognition method at least includes steps S210 to S240, which are introduced in detail as follows:

[0035] Step S210: Obtain the original image of the area to be detected, and crop the original image according to a preset target size to obtain multiple sub-images and the sub-image information of each sub-image. The target size is set according to the detection scenario and detection route of the UAV.

[0036] Among them, the original image is captured by the UAV.

[0037] In this embodiment, the target size is set according to the detection scenario and detection route of the UAV, that is, under different detection scenarios and detection routes, corresponding target sizes are set. Adaptively, the target size is the size of the small target presented in the original image under the corresponding detection scenario and detection route, that is, the preset small target size under the corresponding detection scenario and detection route. For example, the pixel area corresponding to the size is 32 2 .

[0038] It should be noted that the detection scenario and the detection route are in a one-to-many correspondence relationship, and the detection route and the target detection model are in a one-to-one correspondence relationship, that is, the detection scenario and the target detection model are also in a one-to-many correspondence relationship.

[0039] Exemplarily, the detection scenario includes but is not limited to environmental detection, safety detection, and rescue detection, etc. Environmental detection such as river garbage detection, and the detection target can be floating garbage in the river. Safety detection such as traffic safety detection, and the detection target can be pedestrians, vehicles, or other moving objects on the road. Rescue detection such as river area rescue detection, and the detection target can be people, vehicles, and boats in the river area.

[0040] In an embodiment, cropping the original image according to a preset target size to obtain multiple sub-images and the sub-image information of each sub-image includes: setting the sub-image size and the overlapping rate of adjacent sub-images according to the target size; performing sliding cropping on the original image according to the sub-image size and the overlapping rate of adjacent sub-images to obtain multiple sub-images and the sub-image information of each sub-image. The sub-image information includes the first coordinate information of each sub-image in the original image.

[0041] Exemplarily, the first coordinate information of each sub - figure in the original image can be the coordinates (x, y) of the upper - left corner of the sub - figure in the original image. Then, according to the size of the sub - figure, the specific position of the sub - figure in the original image can be obtained. It should be noted that the embodiments of the present application do not specifically limit the first coordinate information.

[0042] In this embodiment, setting the sub - figure size according to the target size can ensure that small targets are distributed in the sub - figures, and setting the overlapping rate of adjacent sub - figures according to the target size can ensure that small targets are completely presented in the sub - figures. In this way, when the subsequent target detection model performs target recognition, it can efficiently and accurately identify small targets.

[0043] Please refer to Figure 3 , Figure 3 which is a schematic diagram of the cropping result of an original image shown in an exemplary embodiment of the present application. As Figure 3 shown, it is a sub - figure after the original image is cropped by sliding. Exemplarily, the length and width of the sub - figure size are both 256 pixels, and the overlapping rate of adjacent sub - figures is 0.3.

[0044] Step S220: Input the original image and multiple sub - figures into the target recognition model to obtain the recognition result. The recognition result includes the first position information of the target to be detected in the original image and the second position information of the target to be detected in multiple sub - figures. The target recognition model is deployed by the cloud - edge collaborative platform according to the detection scenario and detection route.

[0045] In this embodiment, there is at least one target to be detected, and the recognition result includes the first position information of each target to be detected in the original image and the second position information of the target to be detected in multiple sub - figures. In addition, since multiple sub - figures are obtained by sliding - cropping and reconstructing the original image through the sub - figure size and the overlapping rate of adjacent sub - figures, the targets to be detected recognized by the model include both regular - size targets and small targets.

[0046] It should also be noted that various target detection models for different detection scenarios and detection routes are trained and stored in the cloud - edge collaborative platform. Each target detection model is trained according to the sample images of a specific detection scenario and detection route. The target recognition model is deployed by the cloud - edge collaborative platform according to the detection scenario and detection route, and this target detection model is trained according to the sample images of this detection scenario and detection route. Therefore, this target recognition model can be applicable to the recognition of the original image and multiple sub - figures in this detection scenario and detection route, and can achieve accurate recognition of small targets.

[0047] In one embodiment, the training of the target detection model includes: obtaining a sample image set, which includes multiple aerial images captured by a drone of various targets under a detection scenario and a detection flight path, as well as multiple overhead images captured by a fixed camera; preprocessing the sample image set, where the preprocessing includes quality enhancement, data augmentation, target annotation, and the construction of a training set and a validation set; setting up a training environment and a training network, constructing an initial model for the detection scenario and the detection flight path, and constructing a training script according to the training set, the preset number of iterations, the number of batch samples, and the learning rate; training the initial model according to the training script, validating the trained initial model using the validation set, and obtaining the target detection model when both the average precision and recall rate of the target meet the threshold requirements.

[0048] In this embodiment, the sample image set includes multiple aerial images captured by a drone of various targets under a detection scenario and a detection flight path, as well as multiple overhead images captured by a fixed camera. It not only includes sample images of various types of detection targets under the detection scenario and the detection flight path, but also uses the overhead images captured by the fixed camera to supplement the aerial images, such as providing overhead images for the entire time period, overhead images under different weather conditions or different lighting conditions, etc. In this way, the diversity of the sample images is ensured, so as to ensure that the target recognition model trained based on this sample image set can initially have the ability to recognize small targets and targets of conventional sizes, thereby ensuring the model accuracy.

[0049] In this embodiment, setting up the training environment includes: an operating system, a CPU (Central Processing Unit), an NPU (Neural Processing Unit), a heterogeneous computing architecture, and a deep learning framework.

[0050] For example, the operating system is EulerOS-2.8-aarch64, the CPU is 192 cores with 755GB, the NPU is a commonly used chip on the market, the heterogeneous computing architecture is CANN6.0.RC1, the deep learning framework is MindSpore2.0.0a0, and the main dependencies are Python3.9.13, Open-cv4.7.0, and Numpy1.26.4. It should be noted that the training environment can be set up according to requirements, and this application does not further limit the specific training environment.

[0051] In this embodiment, the training network is set up, for example, based on MindSpore; the initial model for the detection scenario and the detection flight path can be constructed based on any target detection algorithm. In addition, the number of iterations is 1000, the number of batch samples is 32, and the learning rate is 0.0008. The embodiments of this application do not limit the values.

[0052] In this embodiment, the model validation is performed based on the average precision and recall of the target, which includes regular-sized targets and small targets. Specifically, AP small (small Average Precision, small target average precision), AR small l The small average recall (small average recall) measures the average precision and recall of the model for small targets as the key point of model evaluation. The model parameters are tuned by referring to the average precision and recall of the model for all targets (AP (Average Precision)) and AR (Average Recall).

[0053] Exemplarily, the quality enhancement in preprocessing includes: performing similarity filtering on each image in the sample image set to filter highly similar data to avoid overfitting of the model to specific samples, and filtering, denoising and adjusting the brightness of each image in the sample image set to improve image quality. The similarity calculation method can be the SSIM (StructuralSimilarity Index Measure) algorithm.

[0054] Exemplarily, the data enhancement in preprocessing is: to address the problems of data imbalance and too little data in the sample image set, by expanding the minority class samples through data enhancement means, or by using synthetic minority class sample technology to balance the data, including random rotation, cropping and scaling based on geometric transformation, and adjustment of brightness, contrast and saturation based on color space transformation.

[0055] Exemplarily, the target annotation in the preprocessing is to mark the positions of all regular-sized targets and small targets in each image in the sample image set. In addition, after the small targets and regular-sized targets in each image are annotated, they can be exported in the form of an annotation file to facilitate model learning and training.

[0056] Exemplarily, the construction of the training set and the validation set in the preprocessing is to randomly divide the images in the sample image set into the training set and the validation set, for example, the division ratio of the training set to the validation set is 8:2.

[0057] In a possible embodiment, the YOLOv8l model (You Only Look Once version 8 large, a large model in the YOLOv8 series) is used as the initial model. The YOLOv8l model includes the Backbone, Neck, and Head parts. The Backbone adopts the C2f module and enhances the feature extraction ability through the Bottleneck Block and SPPF (Spatial Pyramid Pooling Fast) module; the Neck part adopts the PAN-FAN structure and performs multi-scale feature fusion through PAN (Path Aggregation Network) and FAN (Feature Aggregation Network); the Head part is responsible for object detection, including predicting the bounding box, category, and confidence.

[0058] As an embodiment, after obtaining the target recognition model, it further includes: offline model export, that is, using a model conversion tool to convert the target recognition model into an offline model in a format supported by various processors. In this way, the target recognition model can be applicable to different hardware platforms and application scenarios, ensuring that the target recognition model can run efficiently on various devices.

[0059] In an embodiment, before inputting the original image and multiple sub-images into the target recognition model, it further includes: adjusting the sizes of the original image and multiple sub-images according to a preset image size; normalizing the pixel values of the resized original image and multiple sub-images; and reconstructing the normalized original image and multiple sub-images into a batch input format according to the batch input format of the target recognition model.

[0060] In this embodiment, all sub-images and the original image are adjusted to a unified size, where the preset image size is represented as (model_hight, model_width), where model_hight is the image height and model_width is the image width. In this way, adjusting all images to the same size in advance enables the model to directly accept and process these images, making it easier to perform batch processing and improving the speed of model training and inference.

[0061] In this embodiment, the pixel values of the resized original image and multiple sub-images are normalized, that is, the pixel values of all resized images are mapped to a fixed range of [-1, 1]. In this way, the speed of model training and inference can be improved. Additionally, the normalization calculation formula is:

[0062]

[0063] Among them, y represents the output pixel value of the target pixel point, and x represents the current pixel value of the target pixel point. represents the pixel average value of all pixel points in the image, and σ represents the pixel standard deviation of all pixel points in the image.

[0064] In this embodiment, according to the batch input format of the target recognition model, the normalized original image and multiple sub-images are reconstructed into the batch input format to achieve batch integration of the image, reduce the model inference time consumed exponentially due to the increase in the number of sub-images, thereby accelerating the model inference process and improving the recognition efficiency. For example, the original image and multiple sub-images are reconstructed into the batch input format of (n, model_hight, model_width), where n is the total number of images, that is, the number of sub-images + 1, model_hight is the image height, and model_width is the image width.

[0065] Please refer to Figure 4 , Figure 4 which is a schematic diagram of batch integration of pictures shown in an exemplary embodiment of the present application. As Figure 4 shown, each original image will obtain a certain number of picture sets after sliding cropping and reconstruction. For example, 54 pictures, plus the original picture is 55 pictures. Each picture is normalized and its size is adjusted. For example, both the length and width are 256 pixels, and then reconstructed into the batch input format, as Figure 4 shown on the right.

[0066] Step S230: According to each sub-image information, restore the second position information of the target to be measured in each sub-image to obtain multiple third position information of the target to be measured in the original image.

[0067] Among them, the sub-image information includes the first coordinate information of the sub-image in the original image.

[0068] In one embodiment, the method for obtaining multiple third position information includes: obtaining the first coordinate information of each sub-image in the original image and the second coordinate information of the target to be measured in each sub-image, and the second position information includes the second coordinate information; calculating the first coordinate information and the second coordinate information to obtain the third position information of the target to be measured in the original image, and the third position information includes the third coordinate information.

[0069] In this embodiment, the coordinate reduction of the target to be measured in the sub-image is completed based on the sub-image information, so as to perform redundant result screening under the original image coordinates later.

[0070] In this embodiment, for example, the first coordinate information of the sub - image in the original image is represented as the coordinates (x, y) of the upper - left corner of the sub - image in the original image; the second coordinate information of the target to be detected in the sub - image is represented as (x1, y1, w, h), where (x1, y1) are the coordinates of the upper - left corner of the target to be detected in the sub - image, w represents the width of the target to be detected, and h represents the height of the target to be detected; adding the upper - left corner coordinates of the target to be detected to the coordinates of the upper - left corner of the sub - image where it is located in the original image, the coordinates of the target to be detected in the original image are obtained as (x + x1, y + y1, w, h), that is, the third position coordinates.

[0071] In a possible embodiment, before inputting the original image and multiple sub - images into the target recognition model, if the sizes of the original image and multiple sub - images are adjusted according to the preset image size, then before restoring the second position information of the target to be detected in each sub - image according to each sub - image information, it is necessary to first restore the size of each sub - image.

[0072] Step S240, screen the position information of the target to be detected among the first position information and multiple third position information to obtain the target position information of the target to be detected.

[0073] In this embodiment, screen the position information of the target to be detected to obtain more accurate target position information of the target to be detected.

[0074] In an embodiment, the target to be detected is at least one, and each target to be detected corresponds to first position information and multiple third position information, and the first position information is at least one; the method for obtaining the target position information includes: screening at least one first position information of each target to be detected to obtain the first candidate position information; constructing at least one set of first position information according to the type of the target to be detected, and the position information in each set of first position information includes all the first candidate position information and multiple third position information corresponding to the targets to be detected under one type of the target to be detected; for each set of first position information, sort the multiple position information according to the confidence level, and the recognition result also includes the confidence levels of the first position information and multiple second position information; according to the position information with the largest confidence level in the set of first position information, screen the position information of each target to be detected in the set of first position information to obtain the target position information.

[0075] Among them, the target to be detected is at least one. When there are multiple targets to be detected, it may correspond to at least one type of target to be detected. For example, people and vehicles are different types of targets to be detected. In addition, there is a one - to - one correspondence between the set of first position information and the type of the target to be detected.

[0076] It should be noted that, considering various reasons such as multiple first position information being generated for the same target to be measured due to overlap or occlusion in the original image, the target recognition model usually recognizes targets at different scales, resulting in multiple first position information under detection frames of different sizes being generated at the same position, and the confidence of the target recognition model for the same target to be measured fluctuates, generating multiple similar but different-confidence first position information, etc. Therefore, it is necessary to screen at least one first position information of each target to be measured to determine the first candidate position information, thereby improving the efficiency and accuracy of subsequent target position information screening.

[0077] In this embodiment, based on the confidence of the first position information corresponding to the target to be measured and multiple third position information, the position information of the target to be measured is screened to obtain more accurate target position information of the target to be measured, ensuring the accuracy of small target recognition.

[0078] Further, according to the position information with the highest confidence in the first position information set, the position information of each target to be measured in the first position information set is screened to obtain the target position information, including: constructing a second position information set, and moving the position information with the highest confidence in the first position information set into the second position information set; calculating the intersection over smallest ratio of the position information with the highest confidence and each remaining position information in the first position information set; removing the position information in the first position information set with an intersection over smallest ratio greater than the preset threshold, and moving the position information with the highest confidence in the first position information set after removal into the second position information set; repeating the above steps, continuing to calculate the intersection over smallest ratio and remove position information. When the first position information set is empty, the obtained second position information set is the target position information set, and the target position information set includes the target position information of all targets to be measured under one type of target to be measured.

[0079] Among them, there is a one-to-one correspondence between the second position information set and the type of target to be measured.

[0080] In this embodiment, a result filtering method more suitable for small targets is selected, that is, the redundant inference results in multiple sub-images and the original image are filtered based on the intersection over smallest ratio non-maximum suppression algorithm. In this way, the intersection over smallest ratio is used to replace the commonly used intersection over union ratio in the non-maximum suppression algorithm, reducing the suppression of small targets. When there are overlapping or nested targets to be measured, small targets are more likely to be retained, thereby improving the accuracy of small target recognition.

[0081] Exemplarily, one first position information corresponding to each target to be measured and multiple third position information are represented by the target detection frame bbox i The confidence of the first position information and multiple third position information is represented by conf iIt means that the corresponding detection boxes are placed in different sets according to the type of target to be measured, that is, different sets of first position information D x In one set of first position information is a class set D x , the class set D x includes all detection boxes of all targets to be measured under one type of target to be measured. In the class set D x sort all the detection boxes in descending order of confidence. All subsequent filtering steps are carried out within the class set D x inside.

[0082] For a class set D x , take the detection box bbox x with the largest confidence in D i , denoted as bbox1, and move it to the result set of this type of target to be measured, that is, the second position information set, set as the result set D. Calculate the intersection over union of bbox1 and all the remaining detection boxes in D x , and the calculation formula is:

[0083]

[0084] Among them, represents the area of bbox i , represents the overlapping area of bbox i and the remaining detection box bbox x in D j . Then remove all the detection boxes in D x whose intersection over union with bbox1 is greater than the preset threshold. Then continue to take the detection box bbox x with the largest confidence from D i , denoted as bbox2, and move it to the result set D. Repeat the above steps until there are no detection boxes in D x . At this time, the result set D is the target detection boxes of all targets to be measured under one type of target to be measured, that is, the target position information is obtained.

[0085] In a possible embodiment, screening at least one first position information of each target to be measured to obtain first candidate position information, including: constructing at least one fifth position information set according to the type of the target to be measured, where the position information in each fifth position information set includes at least one first position information corresponding to all targets to be measured under one type of target to be measured; for each fifth position information set, sorting the position information in the fifth position information set according to the confidence level; constructing a sixth position information set, and moving the position information with the largest confidence level in the fifth position information set into the sixth position information set; calculating the intersection-small ratio between the position information with the largest confidence level and each remaining position information in the fifth position information set; removing the position information in the fifth position information set with an intersection-small ratio greater than a preset threshold, and moving the position information with the largest confidence level in the fifth position information set after removal into the sixth position information set; repeating the foregoing steps, continuing to calculate the intersection-small ratio and remove the position information, and when the fifth position information set is empty, taking the position information in the sixth position information set as the first candidate position information of all targets to be measured under one type of target to be measured.

[0086] Wherein, there is a one-to-one correspondence between the fifth position information set and the type of the target to be measured, and there is a one-to-one correspondence between the sixth position information set and the type of the target to be measured.

[0087] In an embodiment, before restoring the second position information of the target to be measured in each sub-graph, it further includes: sorting at least one second position information corresponding to the target to be measured in each sub-graph according to the confidence level to construct a third position information set; iteratively screening the second position information in the third position information set, and moving the position information with the largest confidence level in the third position information set into a pre-constructed fourth position information set in each iteration; when the third position information set is empty, taking the position information in the fourth position information set as the second candidate position information of the target to be measured in each sub-graph.

[0088] In this embodiment, considering that there may be overlaps or occlusions of the target to be measured in each sub-graph, resulting in multiple second position information being generated for the same target to be measured, the target recognition model usually recognizes the target at different scales, resulting in multiple second position information under detection frames of different sizes being generated at the same position, and the confidence level of the target recognition model for the same target to be measured fluctuates, resulting in multiple similar but different-confidence second position information, etc. Therefore, it is necessary to screen at least one second position information of the target to be measured in each sub-graph to determine the second candidate position information, thereby improving the efficiency and accuracy of subsequent target position information screening.

[0089] In addition, there is at least one target to be measured in each sub - figure. Therefore, according to the type of the target to be measured, there is at least one third - position - information set and at least one fourth - position - information set. The third - position - information set has a one - to - one correspondence with the type of the target to be measured, and the fourth - position - information set has a one - to - one correspondence with the type of the target to be measured.

[0090] It should also be noted that in this embodiment, the process of iteratively screening the second - position information is the same as the process of screening the first - position information and the target - position information. The non - maximum suppression algorithm using the intersection - over - union is used for screening. That is, for each third - position - information set, the position information with the highest confidence is moved into the fourth - position - information set; the intersection - over - union between the position information with the highest confidence and each remaining position information in the third - position - information set is calculated; the position information in the third - position - information set with an intersection - over - union greater than the preset threshold is removed, and the position information with the highest confidence in the third - position - information set after removal is moved into the fourth - position - information set; the above steps are repeated to continue the intersection - over - union calculation and position - information removal. When the third - position - information set is empty, the position information in the fourth - position - information set is used as the second - candidate position information of all targets to be measured under one type of target to be measured in each sub - figure.

[0091] In a possible embodiment, the recognition result further includes the category of the target to be measured, that is, the output result is the first - position information of the target to be measured in the original image, the second - position information of the target to be measured in multiple sub - figures, the confidence levels of the first - position information and multiple second - position information, and the type of the target to be measured.

[0092] In one embodiment, the edge - computing device has the capabilities of establishing a secure and reliable communication mechanism with the cloud - edge collaborative platform, processing image data, sending the processed image into a model for recognition to obtain a recognition result, and organizing the recognition result into an event and pushing it to the cloud - edge collaborative platform.

[0093] In one embodiment, after obtaining the target - position information of the target to be measured, it further includes: obtaining event - display content, where the event - display content includes the detection scenario, detection route, detection time of the target to be measured, the original image, and the target - position information of the target to be measured; generating a detection event according to the event - display content and pushing the detection event to the cloud - edge collaborative platform, so that the cloud - edge collaborative platform centrally displays the detection event, searches for the event - processing method of the detection event, monitors the event status of the detection event, and centrally displays the event - processing method and the event status.

[0094] Among them, the cloud-edge collaborative platform displays at least one detection event. In addition, the event processing method refers to the processing method of the event in the detection scenario, for example, the event processing method in the river garbage detection scenario is to clean up the river garbage, and the event status refers to the processing status of the event in the detection scenario, for example, whether the garbage detected in the river garbage detection scenario is cleaned up.

[0095] In this embodiment, the detection results are organized into events and pushed to the cloud-edge collaborative platform, which centrally displays the events, automatically archives the drone inspection results, and generates inspection reports, which is conducive to improving management efficiency and ensuring long-term, effective, stable, and low-cost operation of the inspection area.

[0096] In one embodiment, the functions of the cloud-edge collaboration platform include but are not limited to detection scene management, target recognition model management, detection target management and detection event management.

[0097] In this embodiment, in the detection scene management, each detection scene corresponds to multiple detection routes. In the target recognition model management, each target recognition model is mapped and managed with the corresponding detection scene and detection route. The detection target management is to manage the corresponding target to be detected in each inspection scene, and the detection event management is to manage the detection events obtained in each detection. In this way, the detection scenes, models, events, etc. are uniformly managed through cloud-edge collaboration, which improves management efficiency and reduces management costs.

[0098] See also Figure 5 , Figure 5 FIG. 1 is a flowchart of a specific method for identifying a drone target according to an exemplary embodiment of the present application. Figure 5 As shown, the specific drone target recognition method includes at least steps S510 to S590, which are described in detail as follows:

[0099] Step S510, the drone captures a plurality of aerial images of a specific detection scene and detection route, and the fixed camera captures a plurality of overhead images of the specific detection scene and detection route;

[0100] Step S520: The edge computing device obtains multiple aerial images and multiple bird's-eye view images, processes them to form a sample image set, and pushes them to the cloud-edge collaborative platform;

[0101] Step S530: the cloud-edge collaborative platform pre-processes the sample image set to obtain a training set and a validation set;

[0102] Step S540: The cloud-edge collaborative platform performs model training based on the training set and the validation set to obtain a target recognition model, and associates and stores the target recognition model with a specific detection scenario and a detection route;

[0103] Step S550: Based on the specific detection scenario and detection route of the drone, the corresponding target recognition model is sent to the edge computing device for deployment;

[0104] Step S560: The edge computing device acquires the original images of the area to be detected under the specific detection scenario and detection route captured by the drone, and crops the original images according to the preset target size to obtain multiple sub-images and the sub-image information of each sub-image;

[0105] Step S570: After processing the original images and multiple sub-images, they are input into the target recognition model to obtain the first position information of the target to be detected in the original image and the second position information in the multiple sub-images;

[0106] Step S580: According to the information of each sub-image, the second position information of the target to be detected in each sub-image is restored to obtain multiple third position information of the target to be detected in the original image;

[0107] Step S590: Based on the non-maximum suppression algorithm of intersection over small union, the target position information of the target to be detected is screened from the first position information and multiple third position information.

[0108] In this specific drone target recognition method, by combining computer vision technology, deep learning algorithms, image processing technology, edge computing technology and network transmission technology, drone target recognition is realized. Adopting a cloud-edge collaborative working mode, real-time and rapid target recognition is carried out at the edge using edge computing devices, reducing the network bandwidth requirement and alleviating the burden on the central platform. The edge devices and event data are aggregated and centrally managed on the cloud-edge collaborative platform, reducing the management cost. The combination of the two forms a complementary advantage. In addition, the target recognition model can be replaced with any lightweight or complex target detection algorithm according to application scenarios with high timeliness requirements or high detection accuracy requirements, achieving the same accurate detection effect for small targets and targets of conventional sizes. It is more flexible in application, convenient for model deployment, simple to operate and easy to get started, providing convenience in management, reducing management costs and improving management efficiency.

[0109] The above-mentioned UAV target recognition method first obtains the original image of the area to be detected, crops the original image according to the preset target size to obtain multiple sub-images and the sub-image information of each sub-image. The target size is set according to the detection scene and detection route of the UAV. Then, the original image and the multiple sub-images are input into the target recognition model to identify the first position information of the target to be detected in the original image and the second position information in the multiple sub-images. The target recognition model is deployed by the cloud-edge collaborative platform according to the detection scene and detection route. Then, according to the sub-image information of each sub-image, the second position information of the target to be detected in each sub-image is restored to obtain multiple third position information of the target to be detected in the original image. Finally, the position information of the target to be detected is screened from the first position information and the multiple third position information, and the original image is cropped and reconstructed based on the specific detection scene and detection route to set the target size, so that the target to be detected, whether it is a conventional-size target or a small target, is distributed in the sub-images, ensuring that the small targets in the sub-images are accurately recognized after being input into the model. Moreover, the target recognition model is deployed according to the specific detection scene and detection route, adapting to the current detection environment, thereby making the recognition effect of small targets in various environments stable, improving the accuracy and efficiency of small target recognition, effectively reducing the missed detection rate of small targets, and ensuring the inspection accuracy in UAV target intelligent recognition.

[0110] Please refer to Figure 6 , Figure 6 which is a block diagram of a UAV target recognition system shown in an exemplary embodiment of the present application. This system can be applied to Figure 1 the implementation environment shown. It should be understood that this system can also be applicable to other exemplary implementation environments, and the implementation environment applicable to this system is not limited in this embodiment.

[0111] As Figure 6 shown, in an exemplary embodiment, the UAV target recognition system 600 at least includes an image processing module 610, a target recognition module 620, a position restoration module 630, and a position screening module 640, which are introduced in detail as follows:

[0112] The image processing module 610 is configured to obtain the original image of the area to be detected and crop the original image according to the preset target size to obtain multiple sub-images and the sub-image information of each sub-image. The target size is set according to the detection scene and detection route of the UAV;

[0113] The target recognition module 620 is configured to input the original image and the multiple sub-images into the target recognition model to obtain the recognition result. The recognition result includes the first position information of the target to be detected in the original image and the second position information of the target to be detected in the multiple sub-images. The target recognition model is deployed by the cloud-edge collaborative platform according to the detection scene and detection route;

[0114] A position restoration module 630, configured to restore the second position information of the target to be measured in each sub-graph according to each sub-graph information, so as to obtain multiple pieces of third position information of the target to be measured in the original image;

[0115] A position screening module 640, configured to screen the position information of the target to be measured from the first position information and the multiple pieces of third position information, so as to obtain the target position information of the target to be measured.

[0116] It should be noted that the drone target recognition system provided in the above embodiment and the drone target recognition method provided in the above embodiment belong to the same concept. The content of the operations performed by each module has been described in detail in the method embodiment, and will not be repeated here.

[0117] Please refer to Figure 7 , Figure 7 which is a schematic structural diagram of an electronic device shown in an exemplary embodiment of the present application. Figure 7 shows a schematic structural diagram of a computer system of an electronic device suitable for implementing the embodiments of the present application. It should be noted that Figure 7 the computer system 700 of the electronic device shown is only an example, and should not impose any limitation on the functions and usage scope of the embodiments of the present application.

[0118] As Figure 7 shown, the computer system 700 includes a central processing unit (CPU) 701, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 702 or the program loaded from the storage section 708 into the random access memory (RAM) 703, such as executing the method in the above embodiment. In the RAM 703, various programs and data required for system operation are also stored. The CPU 701, ROM 702, and RAM 703 are connected to each other through a bus 704. The input / output (I / O) interface 705 is also connected to the bus 704.

[0119] The following components are connected to the I / O interface 705: an input section 706 including a keyboard, a mouse, etc.; an output section 707 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 708 including a hard disk, etc.; and a communication section 709 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 709 performs communication processing via a network such as the Internet. A drive 710 is also connected to the I / O interface 705 as required. A removable medium 711, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 710 as required so that a computer program read therefrom is installed into the storage section 708 as required.

[0120] Specifically, according to an embodiment of the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment of the present application includes a computer program product that includes a computer program carried on a computer-readable medium, and the computer program includes a computer program for performing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication section 709, and / or installed from the removable medium 711. When the computer program is executed by a central processing unit (CPU) 701, various functions defined in the system of the present application are executed.

[0121] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. Among them, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the above-mentioned module, program segment, or part of code includes one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order from that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0122] The units involved in the embodiments of the present invention can be implemented in software or in hardware, and the described units can also be provided in a processor. Among them, the names of these units do not constitute a limitation to the unit itself in some cases.

[0123] The above embodiments are only used to exemplarily illustrate the principles and effects of the present invention, rather than to limit the present invention. Any person familiar with this technology can modify or change the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or changes completed by those with ordinary knowledge in the technical field without departing from the spirit and technical idea disclosed by the present invention should still be covered by the claims of the present invention.

Claims

1. A method for identifying an unmanned aerial vehicle target, characterized in that, The method includes: Obtaining an original image of the area to be detected, and cropping the original image according to a preset target size to obtain multiple sub-images and sub-image information of each sub-image, where the target size is set according to the detection scenario and detection route of the drone; Inputting the original image and the multiple sub-images into a target recognition model to obtain a recognition result, where the recognition result includes first position information of the target to be detected in the original image and second position information of the target to be detected in the multiple sub-images, and the target recognition model is deployed by the cloud-edge collaborative platform according to the detection scenario and the detection route; Restoring the second position information of the target to be detected in each sub-image according to each sub-image information to obtain multiple third position information of the target to be detected in the original image; Performing position information screening on the first position information and the multiple third position information to obtain the target position information of the target to be detected.

2. The drone target recognition method according to claim 1, characterized in that, The cropping the original image according to a preset target size to obtain multiple sub-images and sub-image information of each sub-image includes: Setting the sub-image size and the overlapping rate of adjacent sub-images according to the target size; Performing sliding cropping on the original image according to the sub-image size and the overlapping rate of adjacent sub-images to obtain the multiple sub-images and the sub-image information of each sub-image, where the sub-image information includes the first coordinate information of each sub-image in the original image.

3. The drone target recognition method according to claim 1, characterized in that, The obtaining manner of the multiple third position information includes: Obtaining the first coordinate information of each sub-image in the original image and the second coordinate information of the target to be detected in each sub-image, where the second position information includes the second coordinate information; Calculating the first coordinate information and the second coordinate information to obtain the third position information of the target to be detected in the original image, where the third position information includes third coordinate information.

4. The drone target recognition method according to claim 1, characterized in that, There is at least one target to be detected, and each target to be detected corresponds to the first position information and the multiple third position information, and the first position information is at least one; The obtaining manner of the target position information includes: Screening at least one of the first position information of each target to be detected to obtain first candidate position information; Constructing at least one first position information set according to the type of the target to be detected, where the position information in each first position information set includes the first candidate position information and the multiple third position information corresponding to all the targets to be detected under one type of the target to be detected; For each first position information set, sorting the multiple position information according to the confidence level, and the recognition result further includes the confidence levels of the first position information and the multiple second position information; Screening the position information of each target to be detected in the first position information set according to the position information with the largest confidence level in the first position information set to obtain the target position information.

5. The method for identifying an unmanned aerial vehicle target according to claim 4, wherein The screening the position information of each target to be detected in the first position information set according to the position information with the largest confidence level in the first position information set to obtain the target position information includes: Constructing a second location information set, and moving the location information with the highest confidence in the first location information set into the second location information set; Calculate an intersection ratio of the position information with the highest confidence and each remaining position information in the first position information set; Remove the location information with a cross-reflection ratio greater than a preset threshold in the first location information set, and move the location information with the highest confidence in the first location information set after the removal into the second location information set; Repeat the above steps to continue calculating the cross-reduction ratio and removing the position information. When the first position information set is empty, the second position information set obtained is a target position information set, and the target position information set includes the target position information of all targets to be measured under one target type to be measured.

6. The method for identifying a drone target according to claim 1, wherein, Before restoring the second position information of the target to be detected in each sub-image, the method further includes: Sort at least one piece of the second position information corresponding to the target to be detected in each sub-image according to the confidence level to construct a third position information set; Iteratively screening the second position information in the third position information set, and in each iteration, moving the position information with the highest confidence in the third position information set into a pre-constructed fourth position information set; When the third position information set is empty, the position information in the fourth position information set is used as the second candidate position information of the target to be detected in each sub-image.

7. The drone target recognition method according to claim 1, characterized in that, The training of the target detection model includes: Acquire a sample image set, wherein the sample image set includes a plurality of aerial images of the detection scene and various targets on the detection route captured by a drone and a plurality of overhead images captured by a fixed camera; Preprocessing the sample image set, wherein the preprocessing includes quality enhancement, data enhancement, object annotation, and construction of a training set and a validation set; Building a training environment and a training network, constructing an initial model for the detection scenario and the detection route, and building a training script based on the training set and a preset number of iterations, batch sample quantity, and learning rate; The initial model is trained according to the training script, and the trained initial model is verified using the verification set, and when the average precision and recall of the target both meet the threshold requirements, the target detection model is obtained.

8. The UAV target recognition method according to claim 1, wherein Before inputting the original image and the plurality of sub-images into the target recognition model, the method further includes: According to a preset image size, resizing the original image and the plurality of sub-images; Normalizing the pixel values ​​of the original image and the plurality of sub-images after resizing; According to the batch input format of the target recognition model, the normalized original image and the multiple sub-images are reconstructed into the batch input format.

9. The method for identifying a target of an unmanned aerial vehicle according to any one of claims 1 to 8, wherein, After obtaining the target position information of the target to be measured, the method further includes: Acquire event display content, wherein the event display content includes the detection scene, the detection route, the detection time of the target to be detected, the original image, and the target position information of the target to be detected; Generate a detection event according to the event display content, and push the detection event to the cloud-edge collaboration platform, so that the cloud-edge collaboration platform centrally displays the detection event, searches for the event handling method of the detection event, monitors the event status of the detection event, and centrally displays the event handling method and the event status.

10. A drone target recognition system, characterized in that, The system includes: An image processing module, configured to obtain an original image of the area to be detected, and crop the original image according to a preset target size to obtain multiple sub-images and sub-image information of each sub-image, where the target size is set according to the detection scenario and detection route of the unmanned aerial vehicle; A target recognition module, configured to input the original image and the multiple sub-images into a target recognition model to obtain a recognition result, where the recognition result includes first position information of the target to be detected in the original image and second position information of the target to be detected in the multiple sub-images, and the target recognition model is deployed by the cloud-edge collaboration platform according to the detection scenario and the detection route; A position restoration module, configured to restore the second position information of the target to be detected in each sub-image according to each sub-image information to obtain multiple third position information of the target to be detected in the original image; A position screening module, configured to screen the position information of the target to be detected from the first position information and the multiple third position information to obtain the target position information of the target to be detected.