Target detection method and system and computer readable storage medium
By collaborating between drones and cloud platforms, and using the feature information of local images to adaptively adjust the data enhancement strategy, data enhancement detection is performed on aerial images, which solves the problems of low efficiency and accuracy in aerial image target detection and achieves more efficient and accurate target detection.
Patent Information
- Application Number
- CN202510705720.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-09-16
AI Technical Summary
Existing aerial image target detection methods have shortcomings in data transmission efficiency, model optimization and accuracy of detection results, especially in achieving real-time and efficient target detection on UAV platforms.
By working collaboratively between the drone and the cloud platform, the feature information of the local image is used to adaptively adjust the data enhancement strategy. After data enhancement of the local image, the pre-trained target detection model is used for detection, reducing the amount of data transmission and improving detection accuracy and efficiency.
It improves the accuracy and efficiency of target detection in aerial images, can capture target features more accurately, and solves the problems of low detection efficiency and accuracy.
Smart Images

Figure CN120655892A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of target detection technology, and in particular to a target detection method, system, and computer-readable storage medium. Background Art
[0002] With the rapid development of drone technology, aerial images captured by drones have been widely used in fields such as agricultural monitoring, traffic management, urban planning, and disaster response due to their wide viewing angles and rich information. However, the high resolution of aerial images also leads to large data volumes, large variations in target scales, and complex backgrounds, which poses significant challenges for object detection in aerial images.
[0003] Currently, cloud-based aerial imagery object detection methods still have shortcomings in data transmission efficiency, model optimization, and detection accuracy. For example, uploading large numbers of aerial images to the cloud consumes significant time and bandwidth resources during data transmission. Furthermore, leveraging the distributed computing capabilities of the cloud to improve detection efficiency and accuracy during model training and detection is an urgent issue. Summary of the Invention
[0004] The present application provides a target detection method, system, and computer-readable storage medium to at least solve the problem of low detection efficiency and detection accuracy in target detection on aerial images.
[0005] The present application provides a target detection method, which is applied to a cloud platform, including: obtaining at least one partial image of an original image to be detected; extracting feature information of the at least one partial image, and determining a target data enhancement strategy that matches the feature information; for any target partial image in the at least one partial image, performing data enhancement on the target partial image according to the target data enhancement strategy corresponding to the target partial image, and obtaining a data enhanced image corresponding to the target partial image; using a pre-trained first target detection model, performing target detection on the at least one data enhanced image, and obtaining a target detection result.
[0006] The present application also provides a target detection method, which is applied to a drone, including: obtaining at least one original image to be detected; using a pre-trained second target detection model to perform target detection on the at least one original image to obtain position information of the target to be detected in the at least one original image; according to the position information of the target to be detected, cropping the area where the target to be detected is located in the at least one original image to obtain at least one partial image; sending the at least one partial image to a cloud platform so that the cloud platform performs target detection on the at least one partial image.
[0007] The present application also provides a target detection system, which includes a drone and a cloud platform, wherein the drone is used to obtain at least one original image to be detected; using a pre-trained second target detection model, target detection is performed on the at least one original image to obtain position information of the target to be detected in the at least one original image; according to the position information of the target to be detected, the area where the target to be detected is located is cropped in the at least one original image to obtain at least one local image; the at least one local image is sent to the cloud platform so that the cloud platform performs target detection on the at least one local image; the cloud platform is communicatively connected with the drone to obtain at least one local image of the original image to be detected; feature information of the at least one local image is extracted, and a target data enhancement strategy matching the feature information is determined; for any target local image in the at least one local image, data enhancement is performed on the target local image according to the target data enhancement strategy corresponding to the target local image to obtain a data enhanced image corresponding to the target local image; using the pre-trained first target detection model, target detection is performed on the at least one data enhanced image to obtain a target detection result.
[0008] The present application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned target detection methods are implemented.
[0009] Through the target detection method of the present application, the target data enhancement strategy of the local image is determined according to the feature information of the local image, and the target data enhancement strategy is used to perform data enhancement on the local image to obtain a data enhanced image corresponding to the local image, and then the data enhanced image is input into the pre-trained first target detection model to obtain the target detection result by performing target detection on the data enhanced image. Since the present solution can adaptively adjust the target data enhancement strategy of the local image according to the feature information of the local image, it can enhance the adaptability of the first target detection model to various complex aerial images, and the first target detection model does not need to detect the entire original image, but only needs to perform target detection on the data enhanced image, which can more accurately capture the characteristics of the target, thereby improving the accuracy and efficiency of target detection, and solving the problem of poor detection efficiency and detection accuracy when performing target detection on aerial images. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0011] Figure 1A schematic diagram of a target detection system provided according to an embodiment of the present application;
[0012] Figure 2 A schematic diagram of another target detection system provided according to an embodiment of the present application;
[0013] Figure 3 A flowchart of a target detection method provided according to an embodiment of the present application is provided;
[0014] Figure 4 A schematic diagram of a first target detection model obtained through training according to an embodiment of the present application;
[0015] Figure 5 A schematic diagram of a flow chart of another target detection method provided according to an embodiment of the present application;
[0016] Figure 6 A schematic diagram of the structure of an attention module provided according to an embodiment of the present application;
[0017] Figure 7 A flowchart of another target detection method provided according to an embodiment of the present application;
[0018] Figure 8 A flowchart of another target detection method provided according to an embodiment of the present application;
[0019] Figure 9 A schematic diagram of a process for extracting a feature fusion map from an original image by a drone or an edge device configured on the drone according to an embodiment of the present application;
[0020] Figure 10 A schematic diagram of a process for detecting a target based on cloud-edge collaboration according to an embodiment of the present application;
[0021] Figure 11 A schematic structural diagram of a target detection device provided according to an embodiment of the present application;
[0022] Figure 12 A schematic structural diagram of another target detection device provided according to an embodiment of the present application;
[0023] Figure 13 The figure is a schematic structural diagram of an electronic device provided according to an embodiment of the present application. DETAILED DESCRIPTION
[0024] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0025] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.
[0026] With the rapid development of drone technology, aerial images captured by drones have been widely used in fields such as agricultural monitoring, traffic management, urban planning, and disaster response due to their wide viewing angles and rich information. However, the high resolution of aerial images also leads to large data volumes, large variations in target scales, and complex backgrounds, which poses significant challenges for object detection in aerial images.
[0027] Currently, methods for object detection in aerial imagery fall into two main categories: traditional machine learning-based methods rely on handcrafted features and are less adaptable to complex scenarios. Deep learning-based methods, while achieving significant improvements in detection accuracy, typically require extensive computing and storage resources, making them difficult to run in real time on edge devices such as drones. To address these challenges, leveraging the powerful computing and storage capabilities of cloud platforms for object detection in aerial imagery has become a trend. However, cloud-based methods for aerial imagery object detection still face limitations in data transmission efficiency, model optimization, and detection accuracy. For example, uploading large numbers of aerial images to cloud platforms consumes significant time and bandwidth resources during data transmission. Furthermore, leveraging the distributed computing capabilities of cloud platforms to improve detection efficiency and accuracy during model training and detection is an urgent issue.
[0028] In view of this, the present application proposes a target detection method, system, and computer-readable storage medium. The target detection method includes: obtaining at least one partial image of an original image to be detected; extracting feature information of the at least one partial image and determining a target data enhancement strategy that matches the feature information; performing data enhancement on any target partial image in the at least one partial image according to the target data enhancement strategy corresponding to the target partial image to obtain a data-enhanced image corresponding to the target partial image; and performing target detection on the at least one data-enhanced image using a pre-trained first target detection model to obtain a target detection result.
[0029] The target detection method of the present application determines the target data enhancement strategy of the local image based on the feature information of the local image, and uses the target data enhancement strategy to perform data enhancement on the local image to obtain a data enhanced image corresponding to the local image, and then inputs the data enhanced image into a pre-trained first target detection model to obtain a target detection result by performing target detection on the data enhanced image. Since the present solution can adaptively adjust the target data enhancement strategy of the local image based on the feature information of the local image, it can enhance the adaptability of the first target detection model to various complex aerial images, and the first target detection model does not need to detect the entire original image, but only needs to perform target detection on the data enhanced image, which can more accurately capture the characteristics of the target, thereby improving the accuracy and efficiency of target detection, and solving the problem of poor detection efficiency and detection accuracy when performing target detection on aerial images.
[0030] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0031] In conjunction with the specific application environment architecture or specific hardware architecture on which the execution of the target detection method depends, the specific application environment architecture or specific hardware architecture is described here.
[0032] like Figure 1 As shown in FIG, a schematic diagram of the target detection system provided by this application. The target detection system includes a drone and a cloud platform, and the drone and the cloud platform are in communication connection. Figure 2As shown, the drone is deployed with a pre-trained second target detection model. The drone can perform preliminary target detection on the original image taken during the flight mission to obtain a local image in the original image, and send at least one local image to the cloud platform. The cloud platform is deployed with a pre-trained first target detection model. The cloud platform uses the first target detection model to perform target detection again on at least one local image to achieve accurate detection and identification of the target. The target detection system based on cloud-edge collaboration provided in this application uses edge devices such as drones to perform preliminary screening of targets in the original image, and only transmits the local image to the cloud platform. Compared with transmitting the complete original image, it can reduce the amount of data transmission, reduce the demand for network bandwidth, and save transmission time costs. At the same time, the powerful computing power of the cloud platform is utilized to accurately perform target detection on local images. This solution not only gives play to the real-time advantages of edge devices such as drones, but also improves the overall detection efficiency and accuracy.
[0033] According to an embodiment of the present invention, an embodiment of a target detection method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0034] In this embodiment, a target detection method is provided, which can be used in the above-mentioned cloud platform. Figure 3 is a flow chart of a target detection method according to an embodiment of the present invention. Figure 3 As shown, the process includes the following steps:
[0035] Step S302: Acquire at least one partial image of the original image to be detected.
[0036] The original image can be an aerial image captured by a drone during a flight mission, specifically captured by an image acquisition device on the drone. The partial image can be a portion of the original image, i.e., the original image includes the partial image. In a specific embodiment, the partial image can be sent from the drone to the cloud platform via a network, or an edge device on the drone can send the partial image to the cloud platform.
[0037] Step S304: extract feature information of at least one local image and determine a target data enhancement strategy that matches the feature information.
[0038] The feature information of the local image may include, but is not limited to, Histogram of Oriented Gradient (HOG), Scale-Invariant Feature Transform (SIFT), Oriented FAST and Rotated BRIEF (ORB), etc. Specifically, the feature information of the local image may be extracted using a feature extraction method based on key point detection, a feature extraction method based on a region, or a feature extraction method based on a deep learning method, and this application does not impose any restrictions on this.
[0039] Data enhancement strategies include, but are not limited to, brightness, contrast, and saturation adjustments, random cropping, scaling, color dithering, and noise addition. Before extracting features from a local image, the correspondence between feature parameters and data enhancement strategies can be pre-set. After extracting the feature information of the local image, the target data enhancement strategy for that local image can be determined based on the feature parameters corresponding to the feature information.
[0040] Step S306 : For any target partial image in the at least one partial image, data enhancement is performed on the target partial image according to the target data enhancement strategy corresponding to the target partial image to obtain a data enhanced image corresponding to the target partial image.
[0041] If there are multiple local images, any local image is determined as the target local image, and the target data enhancement strategy corresponding to the target local image is used to perform data enhancement on the target local image to obtain a data enhanced image corresponding to the target local image, that is, one local image corresponds to one data enhanced image.
[0042] This application uses a target data enhancement strategy to perform data enhancement on the target local image, only to reduce the interference information in the target detection process compared with the corresponding target local image in the data enhanced image after data enhancement, so as to detect the target in the target local image more accurately and efficiently. For example, if the contrast between the target in the target local image and the target local image is small, it indicates that the feature information related to the target is weak at this time. In this way, the contrast between the target and the target local image is enhanced, and the subsequent first target detection model can perform target detection more accurately and efficiently. It should be understood that this application does not enhance the interference of the data to improve the generalization ability of the model, but reduces some noise to achieve accurate and efficient target detection.
[0043] Step S308: Use the pre-trained first target detection model to perform target detection on at least one data augmented image to obtain a target detection result.
[0044] The first target detection model can be a high-precision target detection model deployed based on methods such as neural networks, deep learning or machine learning. For example, the first target detection model can be a Faster R-CNN model, a real-time target detection algorithm (You Only Look Once, YOLO), or a single-stage target detection algorithm (Single Shot MultiBox Detector, SSD). The first target detection model is trained based on a model optimization method for dynamic data enhancement, that is, according to the characteristics of the original image and the needs of the actual detection scene, different initial data enhancement strategies are dynamically generated, such as random cropping, rotation, scaling, color jittering, etc., and the parameters and strategies of data enhancement are automatically adjusted according to the model performance feedback during the training process. Figure 4 As shown, the specific process of training to obtain the first target detection model specifically includes steps S401 to S406.
[0045] in,
[0046] Step S401: Obtain a training data set.
[0047] Step S402: perform data enhancement on the training data set using the initial data enhancement strategy to obtain a data-enhanced training data set.
[0048] Step S403: Using the data-enhanced training data set, perform model training on the untrained first object detection model.
[0049] Step S404: Evaluate the model performance on the validation dataset. Model performance includes but is not limited to mean average precision (mAP), intersection over union (IoU), etc., which are not limited in this application.
[0050] In step S405, if the model performance meets the standard, that is, meets the preset requirements, the model parameters at this time are saved to obtain the trained first object detection model. If the model performance does not meet the standard, step S406 is executed.
[0051] Step S406: Adjust the data enhancement strategy.
[0052] The target detection results may include but are not limited to the type of target, the coordinate information of the target in the data-enhanced image, the size of the target, etc. This application does not impose specific restrictions on this and can be flexibly set and adjusted according to the needs of target detection.
[0053] The target detection method provided in this embodiment determines the target data enhancement strategy of the local image based on the feature information of the local image, and uses the target data enhancement strategy to perform data enhancement on the local image to obtain a data enhanced image corresponding to the local image, and then inputs the data enhanced image into a pre-trained first target detection model to obtain a target detection result by performing target detection on the data enhanced image. Since this solution can adaptively adjust the target data enhancement strategy of the local image based on the feature information of the local image, it can enhance the adaptability of the first target detection model to various complex aerial images, and the first target detection model does not need to detect the entire original image, but only needs to perform target detection on the data enhanced image, which can more accurately capture the characteristics of the target, thereby improving the accuracy and efficiency of target detection, and solving the problem of poor detection efficiency and detection accuracy when performing target detection on aerial images.
[0054] In this embodiment, a target detection method is provided, which can be used in the above-mentioned cloud platform. Figure 5 is a flow chart of a target detection method according to an embodiment of the present invention. Figure 5 As shown, the process includes the following steps:
[0055] Step S502: Obtain at least one partial image of the original image to be detected. Figure 3 Step S302 of the illustrated embodiment will not be described in detail here.
[0056] Step S504: extract feature information of at least one local image and determine a target data enhancement strategy that matches the feature information.
[0057] Specifically, the above step S504 includes:
[0058] Step S5042: extract feature information of at least one local image and obtain feature parameters corresponding to the feature information.
[0059] Step S5044: determine the data enhancement strategy corresponding to the feature parameter as the target data enhancement strategy that matches the feature information.
[0060] The feature parameters corresponding to the feature information are used to quantify the specific numerical indicators that describe the image features. Different types of features correspond to different parameters, so appropriate data enhancement strategies can be determined based on the feature parameters corresponding to the feature information. For example, if the feature information is a HoG feature, the corresponding feature parameters can be the number of bins in the gradient direction histogram, the size of the cell (Cell), and the size of the block (Block). If the feature information is a SIFT feature, the corresponding feature parameters can be parameters related to the feature descriptor.
[0061] As a specific embodiment, for each local image, the feature information of the local image can also be statistically analyzed to obtain the size information and contrast of the target in the local image, wherein the contrast is used to characterize the degree of difference between the target and the background of the local image. If the size information of the target is less than the first preset value, it indicates that the size of the target in the local image is small, so it can be determined that the target data enhancement strategy of the local image is cropping, that is, the target is enlarged. If the size information of the target is greater than the second preset value, it indicates that the size of the target in the local image is large, so it can be determined that the data enhancement strategy of the local image is reduction, that is, the target is reduced. If the contrast is less than the third preset value, it indicates that the difference between the target and the background of the local image is small, and the first target detection model cannot extract the target from the background of the local image well, so it can be determined that the data enhancement strategy of the local image is an enhanced contrast strategy, that is, increasing the degree of difference between the target and the background of the local image.
[0062] By using the feature parameters corresponding to the feature information, the target data enhancement strategy corresponding to the local image can be determined more accurately. Subsequently, the target data enhancement strategy corresponding to the local image is used to perform data enhancement for the local image, which can improve the accuracy and efficiency of target detection by the first target detection model.
[0063] Step S506: For any target partial image in the at least one partial image, data enhancement is performed on the target partial image according to the target data enhancement strategy corresponding to the target partial image to obtain a data enhanced image corresponding to the target partial image. Figure 3 Step S306 of the illustrated embodiment will not be described in detail here.
[0064] In an optional embodiment, the first target detection model includes an attention module, wherein the attention module includes a channel attention module and a spatial attention module, the number of channels of the fully connected layer in the channel attention module can be adaptively adjusted, and the spatial attention module includes a multi-scale convolutional layer.
[0065] Specifically, as network depth increases, the level of feature abstraction increases. Therefore, the number of channels in the fully connected layer can be dynamically and adaptively adjusted based on the network depth. For example, in shallow networks, a higher number of channels can be retained (e.g., a 2 / 3 ratio) to preserve more detailed information; in deeper networks, the number of channels can be appropriately reduced (e.g., a 1 / 3 ratio) to focus on more abstract features. The number of channels in the fully connected layer can also be dynamically and adaptively adjusted based on certain statistics of the feature map (e.g., variance and mean in the channel dimension). For example, the variance of each channel feature can be calculated. If the variance is large, it indicates that the channel feature is rich in information, and the proportion retained in the fully connected layer can be appropriately increased; otherwise, the proportion can be reduced. Of course, the number of channels in the fully connected layer can also be adaptively adjusted by referring to information from other modules in the attention module. For example, if the features output by global average pooling and global maximum pooling differ significantly, the number of channels in the fully connected layer can be appropriately increased to incorporate more information; if the difference is small, the number of channels can be reduced.
[0066] As a specific example, since small objects contain relatively subtle feature information, the number of channels in the fully connected layer can be set to a smaller value. That is, the size of the fully connected layer channels in the shared multi-layer perceptron (MLP) can be reduced in the channel attention module. Using a smaller fully connected layer allows the model to focus more on the features of small objects, thereby reducing information loss and better capturing the importance of small objects in different channels. Since small objects may appear at different scales in an image, multi-scale convolution can be introduced in the spatial attention module. That is, convolution kernels of different sizes are set in the multi-scale convolution to capture feature information of different scales and more comprehensively capture small objects.
[0067] As a specific example, an attention module can be embedded between two adjacent convolutional layers of the feature extraction network of the first object detection model to better adapt to the small object detection task in the original image. Figure 6 Figure 2 shows the structure of the attention module. The attention module includes an input feature layer, a channel attention module, a spatial attention module, and an output feature layer. In order of execution, the channel attention module includes global average pooling, global maximum pooling, a fully connected layer, an activation function, a sigmoid function, and an output layer. The spatial attention module includes channel-by-channel maximum pooling, channel-by-channel average pooling, concatenation, multi-scale convolution, a sigmoid function, and an output layer.
[0068] Since the original images taken by drones are rich in information but have large variations in target scale and complex background, the number of channels of the fully connected layer is adaptively adjusted in the channel attention module, and a multi-scale convolutional layer is set in the spatial attention module. This makes the first target detection better adapted to the detection task of small targets, further improving the efficiency and accuracy of model detection.
[0069] Step S508: Use the pre-trained first target detection model to perform target detection on at least one data augmented image to obtain a target detection result. Figure 3 Step S308 of the illustrated embodiment will not be described in detail here.
[0070] Step S510: Obtain coordinate information of an object in at least one data-enhanced image from the object detection result.
[0071] If there are multiple targets in the data-enhanced image, the coordinate information of each target in the data-enhanced image is obtained one by one. This coordinate information can be the coordinate information of the target in the image coordinate system corresponding to the data-enhanced image, or the coordinate information of the target in the image coordinate system corresponding to the original image. For example, when the drone sends a partial image to the cloud platform, it can also send the coordinate information of the partial image in the original image to the cloud platform. After detecting the coordinate information of the target in the data-enhanced image corresponding to the partial image, the cloud platform can combine the coordinate information of the partial image in the original image to further convert the coordinate information of the target in the image coordinate system corresponding to the original image.
[0072] Step S512: Convert the coordinate information of the target from the image coordinate system to the geographic coordinate system to obtain the geographic coordinates of the target.
[0073] Geographic coordinates are the coordinates of an object described by longitude and latitude. There are several ways to convert the coordinates of an object from an image coordinate system to a geographic coordinate system. For example, a perspective transformation matrix can be used to convert the coordinates of an object from an image coordinate system to a geographic coordinate system. Alternatively, the pinhole imaging principle can be used in conjunction with the internal and external parameters of the image acquisition device to convert the coordinates of an object from an image coordinate system to a geographic coordinate system.
[0074] Step S514: Mark the target on the map using the target's geographic coordinates, and send the marked map to the target device.
[0075] Mark the target on the map and send the marked map to the target device. The user can query the target device to more clearly view the target's location. As a specific example, the target can be marked on the map using a bounding box, text, or a corresponding icon.
[0076] Specifically, the target device can be a personal computer, tablet computer, mobile phone, etc., and this application does not impose specific restrictions on this. In addition, the cloud platform can also accept user feedback information sent by the user through the target device. For example, if the cloud platform detection result is incorrect, the user can send a message to inform the cloud platform that the target is not the correct object for this target detection, etc. After receiving the user feedback information, the cloud platform can continue to train and optimize the first target detection model, so that the detection results of the first target detection model are more and more accurate.
[0077] Through coordinate conversion, the geographic coordinates of the target can be obtained more quickly with less calculation. The target can be marked on the map using the geographic coordinates, so that the user can more clearly and intuitively determine the location of the target, giving the user a better user experience.
[0078] The object detection method provided in this embodiment determines a target data enhancement strategy for a local image based on feature parameters corresponding to the feature information of the local image. The target data enhancement strategy is then used to perform data enhancement on the local image, resulting in a data-enhanced image of the local image. The data-enhanced image is then input into a pre-trained first object detection model to perform object detection on the data-enhanced image to obtain an object detection result. Because the first object detection model in this solution is equipped with an attention module, it can more accurately capture the features of small objects, improving the accuracy and efficiency of object detection.
[0079] In this embodiment, a target detection method is provided, which can be used in the above-mentioned drone or an edge device configured on the drone. Figure 7 is a flow chart of a target detection method according to an embodiment of the present invention. Figure 7 As shown, the process includes the following steps:
[0080] Step S702: Acquire at least one original image to be detected.
[0081] During the flight of the drone, an image acquisition device configured on the drone may capture at least one original image. After capturing the at least one original image, the image acquisition device may send the at least one original image to the drone or an edge device configured on the drone.
[0082] Step S704: Utilize the pre-trained second target detection model to perform target detection on at least one original image to obtain position information of the target to be detected in the at least one original image.
[0083] The second object detection model can be a lightweight object detection model built based on a neural network model, such as, but not limited to, MobileNetv3, MobileNet, or ShuffleNet. The lightweight model on the drone or on an edge device deployed on the drone can use lightweight convolution operations such as depthwise separable convolution and grouped convolution to reduce the computational complexity and parameter count of the second object detection model, enabling real-time execution on resource-limited drones or edge devices.
[0084] The position information of the target to be detected in the corresponding original image may be the coordinate information of the target to be detected in the corresponding original image, or may be the coordinate information of a bounding box surrounding the target to be detected.
[0085] Step S706 : According to the position information of the target to be detected, the area where the target to be detected is located is cropped in at least one original image to obtain at least one partial image.
[0086] If the original image corresponds to multiple targets to be detected, then the areas where the targets to be detected are located can be cropped one by one in the original image according to the position information of the targets to be detected to obtain multiple local images.
[0087] Step S708: Send the at least one partial image to the cloud platform so that the cloud platform performs target detection on the at least one partial image.
[0088] The drone or the edge device configured on the drone can send the local image to the cloud platform through the 4G / 5G network. Of course, as a specific embodiment, the coordinate information of the local image in the original image and the feature information corresponding to the local image can also be packaged and sent to the cloud platform. In this way, the first target detection model deployed on the cloud platform can focus on detecting the target without the need for feature extraction, and the coordinates of the detected target in the local image can be converted to the coordinates in the original image later, so that the target can be marked on the map more efficiently and accurately.
[0089] The target detection method provided in this embodiment utilizes a pre-trained second target detection model to perform target detection on one or more original images captured by an image acquisition device. After detecting the target to be detected, the area of the target to be detected in the original image is cropped according to the location information of the target to be detected, thereby obtaining a local image corresponding to the target to be detected. The one or more local images are then sent to a cloud platform for the cloud platform to perform secondary target detection. Compared to transmitting the complete original image to the cloud platform, transmitting only the local image of the target to be detected with the initial screening can greatly reduce the amount of data transmitted, reduce the demand for network bandwidth, save transmission time and cost, and is suitable for scenarios with poor network conditions or limited data traffic.
[0090] In this embodiment, a target detection method is provided, which can be used in the above-mentioned drone or an edge device configured on the drone. Figure 8 is a flow chart of a target detection method according to an embodiment of the present invention. Figure 8 As shown, the process includes the following steps:
[0091] Step S802: Obtain at least one original image to be detected. Figure 7 Step S702 of the illustrated embodiment will not be described in detail here.
[0092] Step S804: Utilize the pre-trained second target detection model to perform target detection on at least one original image to obtain position information of the target to be detected in the at least one original image.
[0093] Specifically, the above step S804 includes:
[0094] Step S8042: Obtain flight parameters of the drone when capturing at least one original image.
[0095] Flight parameters include, but are not limited to, altitude, speed, attitude, and the time and location of the original image captured by the image acquisition device on the drone. Flight attitude includes, but is not limited to, the yaw, roll, and pitch angles of the drone during flight.
[0096] As a specific embodiment, the flight altitude, flight speed, and flight attitude of the drone when capturing the at least one original image can be obtained using a laser radar, a gyroscope, an accelerometer, and a magnetometer. The shooting time and location of the drone when capturing the at least one original image can also be obtained using a global positioning system module (GPS).
[0097] Step S8044: extract visual features of at least one original image, perform feature fusion on the visual features of at least one original image and its flight parameters to obtain at least one feature fusion graph.
[0098] In the field of computer vision, visual features are an abstract representation of image content that describes representative information in an image and helps computers understand and distinguish different images. Visual features include but are not limited to HOG features, SIFT features, and ORB features.
[0099] Specifically, a fully connected layer can be used to perform feature fusion of the visual features of the original image and the flight parameters corresponding to the original image; or a feature fusion model constructed based on a neural network model can be used to perform feature fusion of the visual features of the original image and the flight parameters corresponding to the original image, wherein the feature fusion model can be a fully connected neural network.
[0100] As a specific example, before extracting visual features from the original image, the original image can be preprocessed. For example, the original image can be converted to a color space, i.e., from RGB to YUV. Another example is noise removal. Of course, other preprocessing methods are also possible, and this application does not specifically limit them. The preprocessing method can be flexibly selected based on the characteristics of the target and actual detection requirements.
[0101] Step S8046: Use the second target detection model to perform target detection on at least one feature fusion image to obtain position information of the target to be detected in at least one original image.
[0102] The position information is used to describe the coordinates of the target to be detected in the original image. In addition, the position information can be the coordinate information of the bounding box surrounding the target to be detected, which is not limited in this application.
[0103] When drones capture raw images, the image acquisition device's field of view is easily affected, making the extracted visual features susceptible to occlusion and lighting changes. Flight parameters (such as altitude, speed, attitude, and the location and time of capture) provide a stable spatial reference. These two parameters are collected independently but complement each other through fusion. If one sensor fails, the parameters provided by the other sensor can still provide some scene information, maintaining system robustness and improving the reliability of drone missions.
[0104] In some optional implementations, the above step S8044 includes:
[0105] Step a1: extract block features and key point features of at least one original image.
[0106] Step a2: normalizing the block features, key point features, and flight parameters of at least one original image to obtain target block features, target key point features, and target flight parameters of at least one original image.
[0107] Step a3: performing feature fusion on the target block features, target key point features and target flight parameters of at least one original image to obtain at least one feature fusion graph.
[0108] As a specific example, block features can be HoG features, and key point features can be SIFT features. Block features can capture the texture, color, etc. of an image from the perspective of a local area of the image. For example, texture features such as the grayscale co-occurrence matrix of a local area in an image can reflect the detailed features of that area. Key point features (such as those extracted by algorithms such as SIFT and ORB) focus on stable and representative feature points in the image, such as corner points, which can reflect the structural characteristics of the image. The combination of the two can comprehensively and meticulously describe the image, providing a richer information basis for subsequent processing.
[0109] The numerical ranges and dimensions of block features, key point features, and flight parameters are often different, and the numerical ranges of flight parameters (such as the altitude and speed of the drone) are quite different from those of image features. By normalizing them and mapping them to the same numerical interval (such as [0, 1] or [-1, 1]), the impact of data scale differences on subsequent calculations can be eliminated, making the resulting feature fusion map more accurate and further improving the accuracy and efficiency of model detection.
[0110] In some optional implementations, step S8046 includes:
[0111] Step b1: Using the second target detection model, perform target detection on at least one feature fusion image to obtain position information of at least one optional bounding box, where the optional bounding box is used to enclose the target to be detected.
[0112] Step b2: Scoring at least one optional bounding box using a preset scoring rule to obtain at least one scoring value.
[0113] Step b3: determining the optional bounding box with a score greater than a preset value as the target bounding box, and determining the position information of the target bounding box as the position information of the target to be detected in at least one original image.
[0114] As a specific example, the optional bounding box can be a rectangular box, a circular box, a triangular box, and the like. There are many ways to score multiple optional bounding boxes using preset scoring rules to obtain multiple scoring values. For example, the confidence scores of the optional bounding boxes are sorted, and the optional bounding boxes with the highest rankings are retained. For the retained optional bounding boxes, the optional bounding boxes whose intersection over union (IoU) ratio exceeds a set threshold (such as 0.5) are removed. For another example, the score of the optional bounding box is determined by using the size of the optional bounding box and the ratio of the length to the width of the optional bounding box.
[0115] The larger the score value, the closer the optional target in the optional bounding box is to the target to be detected in this target detection task. Therefore, the optional bounding box with a score greater than the preset value is determined as the target bounding box. When the cloud platform subsequently performs target detection on the local image corresponding to the target bounding box, the probability of false detection can be reduced, thereby further improving the accuracy of target detection.
[0116] like Figure 9 As shown, it is a schematic diagram of a process of extracting a feature fusion map using an original image by a drone or an edge device configured on the drone, specifically including steps S901 to S908.
[0117] Step S901: The image acquisition device of the UAV captures the original image.
[0118] Step S902: pre-process the acquired original image, for example, perform conventional pre-processing operations such as color space conversion and noise removal on the original image.
[0119] Step S903: extract visual features of the pre-processed original image, for example, extract HOG features and SIFT features of the original image.
[0120] Step S904: normalize the extracted visual features.
[0121] Step S905: Obtain the flight parameters of the drone when capturing the original image, such as the drone's flight altitude, flight speed, flight attitude, and the time and location when the original image was captured.
[0122] Step S906: normalize the flight parameters.
[0123] Step S907: performing feature fusion on the visual features after feature normalization and the flight parameters after information normalization, for example, by using a feature fusion model built based on a neural network model to obtain a fused feature vector.
[0124] Step S908: output the feature fusion map.
[0125] Step S806: According to the location information of the target to be detected, the area where the target to be detected is located is cropped in at least one original image to obtain at least one partial image. Figure 7 Step S706 of the illustrated embodiment will not be described in detail here.
[0126] Step S808: Send at least one partial image to the cloud platform so that the cloud platform can perform target detection on the at least one partial image. Figure 7 Step S708 of the illustrated embodiment will not be described in detail here.
[0127] The target detection method of this embodiment fuses the visual features and flight parameters of at least one original image to obtain at least one feature fusion map, and uses a pre-trained second target detection model to perform target detection on the at least one feature fusion map. After detecting the target to be detected, the area where the target to be detected is located in the original image is cropped according to the location information of the target to be detected to obtain a local image corresponding to the target to be detected, and then one or more local images are sent to the cloud platform so that the cloud platform can perform target detection again. Compared with transmitting the complete original image to the cloud platform, transmitting only the local image of the target to be detected with the initial screening can greatly reduce the amount of data transmission, reduce the demand for network bandwidth, save transmission time and cost, and is suitable for scenarios with poor network conditions or limited data traffic.
[0128] As a specific embodiment of the present application, Figure 10 The following is a flow chart of target detection based on cloud-edge collaboration. The specific process is as follows:
[0129] Step S1001: During the flight of the UAV, an image acquisition device configured on the UAV captures at least one original image in real time.
[0130] In step S1002, the drone receives at least one original image sent by the image acquisition device, and the sensors on the drone (such as lidar, GPS sensor) obtain flight parameters such as the flight altitude, flight speed, flight attitude (such as pitch angle, yaw angle, roll angle), shooting time and shooting position of the image acquisition device when shooting the at least one original image.
[0131] In step S1003, an edge device installed on the drone receives at least one original image and flight parameters corresponding to the original image. The edge device preprocesses the at least one original image. First, the at least one original image is converted from RGB to YUV to reduce the redundancy of color information and improve the efficiency of subsequent processing. Then, a median filter is used to denoise the at least one original image after color conversion to remove salt and pepper noise. Then, the HOG features and SIFT features of the at least one denoised image are extracted. The HOG features, SIFT features, and flight parameters are normalized to the range [0,1]. Then, a fully connected neural network is used to fuse the normalized visual features and the normalized flight parameters. The input layer of the fully connected neural network is the concatenation of the feature vectors, the hidden layer uses the ReLU activation function, and the output layer obtains the fused feature fusion map. The edge device can use the NVIDIA Jetson Nano development board.
[0132] In step S1004, the second target detection model deployed on the edge device performs feature detection on at least one feature fusion map to obtain the location information of the target bounding box. The second target detection model can be a lightweight target detection model, and can be improved based on the MobileNetv3 architecture and add a candidate region generation network (RPN) on the basis of MobileNetv3 to generate optional bounding boxes. The operating environment of the second target detection model on the edge device is the TensorRT acceleration engine to improve the inference speed.
[0133] In step S1005, for at least one pre-processed feature fusion map, the second target detection model on the edge device first performs feature extraction to obtain a feature map, then generates optional bounding boxes through the RPN, calculates the coordinates and score of each optional bounding box, and filters out target bounding boxes with a score higher than a threshold (such as 0.5). Based on the position information of the target bounding box, the corresponding original image is cropped to obtain at least one partial image corresponding to the original image. The partial image and its position information in the original image are packaged and transmitted to the cloud platform via the 4G / 5G network.
[0134] In step S1006, the cloud platform deploys the Faster R-CNN high-precision target detection model based on the attention mechanism (i.e., the first target detection model). In the model training phase, public aerial image datasets (such as UCAS-AOD, DOTA, etc.) and the actual collected original image data are used as training datasets. During the training process, the data enhancement strategy is dynamically adjusted based on the current performance feedback of the model on the verification dataset. For example, when it is found that the high-precision target detection model has low detection accuracy for small targets, random cropping and scaling operations are automatically added, and the small target area is enlarged to a certain proportion and then input into the high-precision detection model for training; when the high-precision detection model has poor detection effect under complex backgrounds, enhancement operations such as Gaussian noise and color jitter are added to simulate complex shooting environments. In the detection phase, the cloud platform receives at least one local image transmitted by the edge device, and uses the first target detection model to re-inspect the at least one local image.
[0135] Step S1007, obtain the target detection result according to the output of the first target detection model. From the target detection result, the position of the target is converted from the image coordinate system to the geographic coordinate system to obtain the geographic coordinates of the target. According to the geographic coordinates of the target, the target is marked on the map to display the location, category and other information of the target on the map. It is returned to the target device through the API interface to facilitate user monitoring and decision-making. For example, in agricultural monitoring, users can quickly understand the crop growth conditions, pest and disease areas, etc. in the farmland through the detection results; in traffic management, the number of vehicles on the road, driving status, etc. can be monitored in real time. Among them, the target device can be a mobile phone, personal computer, etc. The target device can be in accordance with the corresponding application, and the user can view the target detection results through the application.
[0136] In step S1008, the user can also provide feedback based on the viewed target detection results. For example, if the detected target is incorrect, information such as the target detection error can be fed back to the cloud platform or the drone so that the cloud platform or the drone receives the user feedback information.
[0137] In step S1009 , the cloud platform optimizes the model based on user feedback information, for example, fine-tuning the parameters of the model to continuously improve detection accuracy.
[0138] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.
[0139] The embodiment of the present application also provides a target detection device, such as Figure 11As shown, the target detection device can be applied in a cloud platform, and the target detection device includes:
[0140] The first acquisition module 1110 is configured to acquire at least one partial image of the original image to be detected.
[0141] The extraction module 1120 is configured to extract feature information of at least one local image and determine a target data enhancement strategy that matches the feature information.
[0142] The data enhancement module 1130 is configured to perform data enhancement on any target local image in the at least one local image according to a target data enhancement strategy corresponding to the target local image, to obtain a data enhanced image corresponding to the target local image.
[0143] The first detection module 1140 is configured to perform target detection on at least one data-augmented image using a pre-trained first target detection model to obtain a target detection result.
[0144] In an optional implementation, the extraction module is further configured to obtain feature parameters corresponding to the feature information; and determine the data enhancement strategy corresponding to the feature parameters as a target data enhancement strategy that matches the feature information.
[0145] In an optional embodiment, the first target detection model includes an attention module, wherein the attention module includes a channel attention module and a spatial attention module, the number of channels of the fully connected layer in the channel attention module can be adaptively adjusted, and the spatial attention module includes a multi-scale convolutional layer.
[0146] In an optional embodiment, the target detection device also includes a second acquisition module, a coordinate conversion module and a marking module, wherein the second acquisition module is used to perform target detection on at least one data-enhanced image using a pre-trained first target detection model, and after obtaining the target detection result, obtain the coordinate information of the target in at least one data-enhanced image from the target detection result; the coordinate conversion module is used to convert the coordinate information of the target from the image coordinate system to the geographic coordinate system to obtain the geographic coordinates of the target; the marking module is used to use the geographic coordinates of the target to mark the target in the map, and send the marked map to the target device.
[0147] The target detection device provided by the present application determines the target data enhancement strategy of the local image based on the feature information of the local image, and uses the target data enhancement strategy to perform data enhancement on the local image to obtain a data enhanced image corresponding to the local image, and then inputs the data enhanced image into a pre-trained first target detection model to obtain a target detection result by performing target detection on the data enhanced image. Since this solution can adaptively adjust the target data enhancement strategy of the local image based on the feature information of the local image, it can enhance the adaptability of the first target detection model to various complex aerial images, and the first target detection model does not need to detect the entire original image, but only needs to perform target detection on the data enhanced image, which can more accurately capture the characteristics of the target, thereby improving the accuracy and efficiency of target detection, and solving the problem of poor detection efficiency and detection accuracy when performing target detection on aerial images.
[0148] The embodiment of the present application also provides a target detection device, such as Figure 12 As shown, the target detection device can be applied to a drone, and the target detection device includes:
[0149] The third acquisition module 1210 is configured to acquire at least one original image to be detected.
[0150] The second detection module 1220 is configured to perform target detection on at least one original image using a pre-trained second target detection model to obtain position information of a target to be detected in the at least one original image.
[0151] The cropping module 1230 is configured to crop the area where the target to be detected is located in at least one original image according to the position information of the target to be detected, to obtain at least one partial image.
[0152] The sending module 1240 is configured to send the at least one partial image to the cloud platform so that the cloud platform performs target detection on the at least one partial image.
[0153] In an optional embodiment, the second detection module is also used to obtain the flight parameters of the drone when shooting at least one original image; extract the visual features of the at least one original image, perform feature fusion on the visual features of the at least one original image and its flight parameters to obtain at least one feature fusion map; use the second target detection model to perform target detection on the at least one feature fusion map to obtain the position information of the target to be detected in the at least one original image.
[0154] In an optional embodiment, the second detection module is also used to extract block features and key point features of at least one original image; normalize the block features, key point features and flight parameters of at least one original image to obtain target block features, target key point features and target flight parameters of at least one original image; and perform feature fusion on the target block features, target key point features and target flight parameters of at least one original image to obtain at least one feature fusion map.
[0155] In an optional embodiment, the second detection module is further used to use a second target detection model to perform target detection on at least one feature fusion image to obtain position information of at least one optional bounding box, where the optional bounding box is used to enclose the target to be detected; use a preset scoring rule to score the at least one optional bounding box to obtain at least one scoring value; determine the optional bounding box with a scoring value greater than the preset value as the target bounding box, and determine the position information of the target bounding box as the position information of the target to be detected in at least one original image.
[0156] The target detection device provided by the present application uses a pre-trained second target detection model to perform target detection on one or more original images taken by an image acquisition device. After detecting the target to be detected, the area where the target to be detected is located in the original image is cropped according to the location information of the target to be detected, and a local image corresponding to the target to be detected is obtained. Then, one or more local images are sent to the cloud platform so that the cloud platform can perform secondary target detection. Compared with transmitting the complete original image to the cloud platform, transmitting only the local image of the target to be detected with the initial screening can greatly reduce the amount of data transmission, reduce the demand for network bandwidth, save transmission time and cost, and can be applied to scenarios with poor network conditions or limited data traffic.
[0157] For the description of the features in the embodiment corresponding to the target detection device, please refer to the relevant description of the embodiment corresponding to the target detection method, and will not be repeated here.
[0158] The embodiment of the present application also provides an electronic device, such as Figure 13 As shown, it is applied to the above-mentioned drone or cloud platform. The electronic device includes a memory 1310 and a processor 1320. The memory 1310 stores a computer program, and the processor 1320 is configured to run the computer program to perform the steps in any of the above-mentioned target detection method embodiments.
[0159] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps of any of the above-mentioned target detection method embodiments when run.
[0160] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.
[0161] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps in any of the above-mentioned target detection method embodiments are implemented.
[0162] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in any of the above-mentioned target detection method embodiments are implemented.
[0163] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0164] The above is a detailed introduction to a target detection method, system and computer-readable storage medium provided by the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method and core ideas of the present application. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.
Claims
1. A target detection method, characterized in that: Applied to a cloud platform, the method includes: Acquire at least one partial image of the original image to be detected; Extracting feature information of the at least one partial image and determining a target data enhancement strategy that matches the feature information; For any one target partial image in the at least one partial image, data enhancement is performed on the target partial image according to the target data enhancement strategy corresponding to the target partial image to obtain a data enhanced image corresponding to the target partial image; Using the pre-trained first target detection model, target detection is performed on at least one of the data augmented images to obtain a target detection result.
2. The method according to claim 1, characterized in that Determine a target data enhancement strategy that matches the feature information, including: Obtaining characteristic parameters corresponding to the characteristic information; The data enhancement strategy corresponding to the feature parameter is determined as the target data enhancement strategy that matches the feature information.
3. The method according to claim 1, characterized in that The first target detection model includes an attention module, wherein the attention module includes a channel attention module and a spatial attention module, the number of channels of the fully connected layer in the channel attention module can be adaptively adjusted, and the spatial attention module includes a multi-scale convolutional layer.
4. The method according to claim 1, wherein After performing target detection on at least one of the data augmented images using the pre-trained first target detection model to obtain a target detection result, the method further includes: Obtaining coordinate information of at least one target in the data-augmented image from the target detection result; Converting the coordinate information of the target from the image coordinate system to the geographic coordinate system to obtain the geographic coordinates of the target; The target is marked on a map using the geographic coordinates of the target, and the marked map is sent to the target device.
5. A target detection method, characterized in that: Applied to a drone, the method comprises: Acquire at least one original image to be detected; Performing target detection on the at least one original image using a pre-trained second target detection model to obtain position information of the target to be detected in the at least one original image; According to the position information of the target to be detected, cropping the area where the target to be detected is located in the at least one original image to obtain at least one partial image; The at least one partial image is sent to a cloud platform so that the cloud platform performs target detection on the at least one partial image.
6. The method according to claim 5, characterized in that Using a pre-trained second object detection model, performing object detection on the at least one original image to obtain position information of the object to be detected in the at least one original image includes: Obtaining flight parameters of the drone when capturing the at least one original image; extracting visual features of the at least one original image, and performing feature fusion on the visual features of the at least one original image and the flight parameters thereof to obtain at least one feature fusion graph; Utilize the second target detection model to perform target detection on the at least one feature fusion image to obtain position information of the target to be detected in the at least one original image.
7. The method according to claim 6, characterized in that Extracting visual features of the at least one original image, and performing feature fusion on the visual features of the at least one original image and the flight parameters thereof to obtain at least one feature fusion graph, including: extracting block features and key point features of the at least one original image; Normalizing the block features, the key point features, and the flight parameters of the at least one original image to obtain target block features, target key point features, and target flight parameters of the at least one original image; Feature fusion is performed on the target block features, the target key point features and the target flight parameters of the at least one original image to obtain the at least one feature fusion graph.
8. The method according to claim 6, characterized in that Performing target detection on the at least one feature fusion image using the second target detection model to obtain position information of the target to be detected in the at least one original image includes: Using the second target detection model, perform target detection on the at least one feature fusion image to obtain position information of at least one optional bounding box, where the optional bounding box is used to enclose the target to be detected; Scoring the at least one optional bounding box using a preset scoring rule to obtain at least one scoring value; The optional bounding box with the score greater than the preset value is determined as the target bounding box, and the position information of the target bounding box is determined as the position information of the target to be detected in the at least one original image.
9. A target detection system, characterized in that: include: A drone is configured to acquire at least one original image to be detected; perform target detection on the at least one original image using a pre-trained second target detection model to obtain position information of a target to be detected in the at least one original image; crop a region of the target to be detected in the at least one original image according to the position information of the target to be detected to obtain at least one partial image; and transmit the at least one partial image to a cloud platform so that the cloud platform performs target detection on the at least one partial image. The cloud platform is communicatively connected to the drone and is used to obtain at least one partial image of the original image to be detected; extract feature information of the at least one partial image and determine a target data enhancement strategy that matches the feature information; for any target partial image in the at least one partial image, perform data enhancement on the target partial image according to the target data enhancement strategy corresponding to the target partial image to obtain a data enhanced image corresponding to the target partial image; and perform target detection on at least one of the data enhanced images using a pre-trained first target detection model to obtain a target detection result.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the steps of the target detection method according to any one of claims 1 to 4 or the steps of the target detection method according to any one of claims 5 to 8 are implemented.
Citation Information
Patent Citations
Multi-scale target detection method based on self-attention mechanism
CN110533084A
Adaptive channel attention three-dimensional reconstruction method based on deep learning
CN114463492A
Target detection data enhancement method for aerial image of unmanned aerial vehicle
CN116416139A
Unmanned aerial vehicle image high-efficiency identification system
CN119148755A
Ship target detection method, device and equipment and storage medium
CN119559374A