Unmanned aerial vehicle-based image recognition method, device, equipment, storage medium and product
Patent Information
- Application Number
- CN202610853233.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-12
- Publication Date
- 2026-09-01
AI Technical Summary
[0005]本申请的主要目的在于提供一种基于无人机的图像识别方法、装置、设备、存储介质及产品,旨在解决,如何在无人机端侧有限的算力条件下实现大规模图像的快速准确识别的技术问题
本申请可先获取无人机采集的巡检图像数据;然后将所述巡检图像数据输入至轻量化识别模型,所述轻量化识别模型通过将标准目标检测框架中的骨干网络替换为轻量级骨干网络并进行量化压缩获得;再通过所述轻量化识别模型对所述巡检图像数据进行前向推理计算,得到所述巡检图像数据中每个目标对象的位置信息;最后根据所述每个目标对象的位置信息生成图像识别结果。相比现有的,本申请通过端侧轻量化模型替代云端识别,实现了无人机端侧大规模图像的实时识别。
Smart Images

Figure CN122676385A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image recognition technology, and in particular to an image recognition method, apparatus, device, storage medium and product based on unmanned aerial vehicles (UAVs). Background Technology
[0002] With the widespread application of drone inspection technology in fields such as power transmission lines, river monitoring, and infrastructure inspection, drones need to process large-scale image data in real time to identify faulty targets or safety hazards. Traditional drone inspection systems typically transmit the collected image data back to ground stations or the cloud for processing, relying on high-performance computing platforms to complete target detection and recognition tasks.
[0003] However, the aforementioned existing technologies have the following drawbacks: First, image data transmission relies on wireless communication links, and long-distance transmission has significant delays, making it difficult to meet real-time recognition requirements; Second, the computing power and memory resources of the UAV edge computing platform are limited, making it impossible to directly run standard deep learning models for large-scale image processing; Third, in the event of weak or interrupted communication signals, the edge recognition function completely fails, leading to the failure of the inspection task; Fourth, inspection scenarios such as river channels and power transmission lines require the simultaneous processing of multimodal images, including visible light and infrared, which multiplies the amount of data and further exacerbates the problem of insufficient edge computing power.
[0004] Therefore, how to achieve rapid and accurate recognition of large-scale images under the limited computing power of drones is an urgent problem to be solved. Summary of the Invention
[0005] The main objective of this application is to provide an image recognition method, apparatus, device, storage medium, and product based on unmanned aerial vehicles (UAVs), aiming to solve the technical problem of how to achieve rapid and accurate recognition of large-scale images under the limited computing power conditions of UAVs.
[0006] To achieve the above objectives, this application proposes an image recognition method based on unmanned aerial vehicles (UAVs), which is applied to an UAV inspection system. The method includes: Acquire inspection image data collected by drones; The inspection image data is input into a lightweight recognition model, which is obtained by replacing the backbone network in the standard target detection framework with a lightweight backbone network and performing quantization compression. The lightweight recognition model is used to perform forward inference calculations on the inspection image data to obtain the location information of each target object in the inspection image data. Image recognition results are generated based on the location information of each target object.
[0007] In one embodiment, the lightweight backbone network includes a MobileNetV3 network, and the lightweight recognition model is obtained by replacing the backbone network in the standard object detection framework with a lightweight backbone network and performing quantization compression, including the following steps: Determine the total number of target object categories that need to be identified in the inspection scenario, and use the total number of categories as the preset number of categories; The original backbone network in the standard target detection framework is replaced with the MobileNetV3 network to obtain the recognition model after backbone network replacement. Based on the preset number of categories, channel pruning is performed on the detection heads in the recognition model after the backbone network is replaced, and the output dimension of the classification branch is compressed to the preset number of categories to obtain the pruned recognition model. The weight parameters of the pruned recognition model are compressed from floating-point numbers to integers using the INT8 quantization method, generating a compressed lightweight recognition model.
[0008] In one embodiment, the step of performing channel pruning on the detection heads in the replaced recognition model of the backbone network according to the preset number of categories further includes: Obtain a pre-trained standard object detection model; The knowledge distillation method is used to train the recognition model after the backbone network is replaced using the pre-trained standard target detection model; The recognition model replaced by the trained backbone network is used as the object to perform the channel pruning process.
[0009] In one embodiment, the step of compressing the weight parameters of the pruned recognition model from floating-point numbers to integers using INT8 quantization includes: Acquire actual images in the inspection scenario and use the actual images as a calibration dataset. The quantization-perception training method is used to train the pruned recognition model with quantization error compensation through the calibration dataset to obtain the recognition model after quantization error compensation. The INT8 quantization method is used to compress the weight parameters of the recognition model after quantization error compensation from floating-point numbers to integers.
[0010] In one embodiment, the step of performing forward inference calculations on the inspection image data using the lightweight recognition model to obtain the location information of each target object in the inspection image data includes: The inspection image data is input into the lightweight recognition model, and the lightweight backbone network is used to perform depthwise separable convolution operations on the inspection image data to extract multi-scale feature maps. The multi-scale feature maps are fused using a feature pyramid network to generate a fused feature map. Regression prediction is performed on the fused feature map to obtain the location information of each target object in the inspection image data.
[0011] In one embodiment, the step of generating image recognition results based on the location information of each target object includes: Based on the location information of each target object, the corresponding target area image is cropped from the inspection image data; Perform classification and recognition on the target region image to determine the category label for each target object; The category label and location information of each target object are associated to generate image recognition results.
[0012] Furthermore, to achieve the above objectives, this application also proposes an image recognition device based on a drone, the device comprising: The image acquisition module is used to acquire inspection image data collected by the drone; The model loading module is used to input the inspection image data into the lightweight recognition model, which is obtained by replacing the backbone network in the standard target detection framework with a lightweight backbone network and performing quantization compression. The inference calculation module is used to perform forward inference calculation on the inspection image data through the lightweight recognition model to obtain the location information of each target object in the inspection image data. The result generation module is used to generate image recognition results based on the location information of each target object.
[0013] In addition, to achieve the above objectives, this application also proposes an image recognition device based on unmanned aerial vehicles (UAVs), the device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the UAV-based image recognition method described above.
[0014] In addition, to achieve the above objectives, this application also proposes a storage medium, which is a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the image recognition method based on UAVs as described above.
[0015] In addition, to achieve the above objectives, this application also provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the image recognition method based on unmanned aerial vehicles as described above.
[0016] One or more technical solutions proposed in this application have at least the following technical effects: This application first acquires inspection image data collected by a drone; then, it inputs the inspection image data into a lightweight recognition model, which is obtained by replacing the backbone network in the standard target detection framework with a lightweight backbone network and performing quantization compression; next, it performs forward inference calculations on the inspection image data through the lightweight recognition model to obtain the position information of each target object in the inspection image data; finally, it generates image recognition results based on the position information of each target object. Compared with existing methods, this application achieves real-time recognition of large-scale images from a drone's edge by replacing cloud-based recognition with an edge-side lightweight model. Attached Figure Description
[0017] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0018] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 This is a flowchart illustrating the first embodiment of the image recognition method based on unmanned aerial vehicles (UAVs) of this application. Figure 2 This is a flowchart illustrating the second embodiment of the image recognition method based on unmanned aerial vehicles (UAVs) in this application; Figure 3 This is a flowchart illustrating the third embodiment of the image recognition method based on unmanned aerial vehicles (UAVs) in this application. Figure 4 This is a schematic diagram of the module structure of the image recognition device based on an unmanned aerial vehicle (UAV) according to an embodiment of this application; Figure 5 This is a schematic diagram of the device structure of the hardware operating environment involved in the image recognition method based on UAV in the embodiments of this application.
[0020] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0021] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.
[0022] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.
[0023] With the widespread application of drone inspection technology in fields such as power transmission lines, river monitoring, and infrastructure inspection, drones need to process large-scale image data in real time to identify faulty targets or safety hazards. Traditional drone inspection systems typically transmit the collected image data back to ground stations or the cloud for processing, relying on high-performance computing platforms to complete target detection and identification tasks.
[0024] However, the aforementioned existing technologies have the following drawbacks: First, image data transmission relies on wireless communication links, and long-distance transmission has significant delays, making it difficult to meet real-time recognition requirements; Second, the computing power and memory resources of the UAV edge computing platform are limited, making it impossible to directly run standard deep learning models for large-scale image processing; Third, in the event of weak or interrupted communication signals, the edge recognition function completely fails, leading to the failure of the inspection task; Fourth, inspection scenarios such as river channels and power transmission lines require the simultaneous processing of multimodal images, including visible light and infrared, which multiplies the amount of data and further exacerbates the problem of insufficient edge computing power.
[0025] Therefore, how to achieve rapid and accurate recognition of large-scale images under the limited computing power of drones is an urgent problem to be solved.
[0026] The main solution of this application embodiment is as follows: First, acquire inspection image data collected by the UAV; then, input the inspection image data into a lightweight recognition model, which is obtained by replacing the backbone network in the standard target detection framework with a lightweight backbone network and performing quantization compression; then, perform forward inference calculation on the inspection image data through the lightweight recognition model to obtain the position information of each target object in the inspection image data; finally, generate image recognition results based on the position information of each target object. Compared with existing methods, this application achieves real-time recognition of large-scale images from the UAV edge by replacing cloud-based recognition with an edge-side lightweight model.
[0027] It should be noted that the executing entity in this embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, or mobile phone, or an electronic device or image recognition device capable of performing the above functions. The following description uses an image recognition device (hereinafter referred to as the device) as an example to illustrate this embodiment and the subsequent embodiments.
[0028] Based on this, embodiments of this application provide a facility fault identification method based on unmanned aerial vehicles (UAVs), referring to... Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the image recognition method based on unmanned aerial vehicles (UAVs) in this application.
[0029] In this embodiment, the method is applied to an unmanned aerial vehicle (UAV) inspection system, and the method includes steps S10 to S40: Step S10: Acquire inspection image data collected by the drone.
[0030] Step S20: Input the inspection image data into the lightweight recognition model, which is obtained by replacing the backbone network in the standard target detection framework with a lightweight backbone network and performing quantization compression.
[0031] It should be noted that the aforementioned inspection image data can be raw image information collected by the UAV from the inspected environment through its onboard visible light camera or infrared thermal imaging camera while the UAV is flying along a preset route.
[0032] The aforementioned inspection image data typically includes at least one of visible light image data and infrared thermal imaging data. For example, visible light image data is acquired under sufficient daylight conditions to obtain high-resolution texture information, while infrared thermal imaging data is acquired at night or under low light conditions to capture the thermal radiation characteristics of the target object.
[0033] The aforementioned lightweight recognition model can be a deep learning target detection model that has been compressed and optimized and can run in real time on a UAV with limited computing power. Unlike traditional solutions that require sending images back to the ground station for processing, the aforementioned lightweight recognition model is directly deployed on the embedded computing platform of the UAV.
[0034] The aforementioned standard object detection framework can be an object detection network architecture based on a single forward propagation, such as the YOLO series framework or the SSD framework. This framework typically consists of three components: a backbone network, a feature fusion network, and a detection head.
[0035] The aforementioned backbone network can be used to extract multi-scale depth features from input images, such as the CSPDarknet network or the ResNet network.
[0036] The aforementioned lightweight backbone network can be a convolutional neural network structure with a small number of parameters and computational cost, such as the MobileNet series, ShuffleNet series, or EfficientNet-Lite series. The aforementioned lightweight backbone network replaces standard convolution with operators such as depthwise separable convolution, which significantly reduces the amount of computation while maintaining the feature extraction capability.
[0037] The quantization compression mentioned above can be a process of converting model weight parameters from floating-point numbers with a higher bit width to integers with a lower bit width, such as compressing them from 32-bit floating-point numbers (FP32) to 8-bit integers (INT8) to reduce model size and accelerate inference computation, while keeping model accuracy within an acceptable range.
[0038] In its implementation, the aforementioned UAV acquires real-time inspection image data via its onboard image acquisition equipment while flying along the inspection path. Visible light images are continuously acquired at a preset frame rate using a visible light camera, which can employ a high-resolution CMOS sensor and automatically adjust exposure parameters and sensitivity under different lighting conditions to ensure image quality. Infrared thermal imaging data can also be acquired using an infrared thermal imaging camera. This camera uses an uncooled vanadium oxide detector capable of capturing thermal radiation signals in the range of -20°C to 150°C, serving as a supplement to the visible light images in nighttime or low-light conditions. The acquired visible light and infrared thermal imaging data are timestamped according to the acquisition time and temporarily stored in the UAV's onboard memory buffer for subsequent processing. Finally, the acquired inspection image data is input into a lightweight recognition model pre-deployed on the UAV's embedded computing platform. The aforementioned lightweight recognition model was pre-trained and compressed on the ground: First, a pre-trained standard object detection model was acquired, such as a YOLOv8 model trained on a large public dataset; then, the original backbone network in the standard object detection framework was replaced with a lightweight backbone network, for example, replacing the CSPDarknet backbone network in the YOLOv8 framework with a MobileNetV3 network; next, the model after backbone network replacement underwent knowledge distillation training and channel pruning to further compress the model size; finally, INT8 quantization technology was used to compress the model weight parameters, converting the weight parameters in the model from 32-bit floating-point numbers to 8-bit integers. After the above replacement, pruning, distillation, and quantization compression processes, the size of the aforementioned lightweight recognition model was reduced by about four times compared to the original standard model, the computational load was reduced to about one-tenth of the original, and the inference speed was increased by two to three times, enabling real-time forward inference on a UAV-based platform with limited computing power. After the inspection image data is input into the lightweight recognition model, the model is triggered to perform subsequent forward inference calculations. The processing latency of each image frame can be controlled within 200 milliseconds, which meets the requirements of real-time inspection.
[0039] For example, suppose a power transmission line inspection scenario involves a drone equipped with a 48-megapixel visible light camera and a 640×512 pixel infrared thermal imaging camera, flying along a pre-set route 5 kilometers long. During flight, the drone simultaneously acquires visible light and infrared thermal images at a rate of 2 frames per second, with each visible light image having a resolution of 1920×1080 pixels. The drone temporarily stores the acquired inspection image data in a memory buffer, and then inputs the image data frame by frame into a pre-deployed lightweight recognition model. This lightweight recognition model is obtained by replacing the CSPDarknet backbone network in the YOLOv8 framework with a MobileNetV3 network, and performing channel pruning and INT8 quantization compression. The model size is compressed from the original 6.5 megabytes to 1.6 megabytes, and the inference time for a single frame is approximately 180 milliseconds. After acquiring a frame of inspection image data, the UAV immediately sends the frame of image data into a lightweight recognition model and waits for the model to output the detection result. The entire process does not require the image to be transmitted back to the ground station, realizing real-time recognition at the edge.
[0040] Step S30: Perform forward inference calculation on the inspection image data using the lightweight recognition model to obtain the location information of each target object in the inspection image data.
[0041] It should be noted that the aforementioned forward inference computation can be a process of passing input data forward layer by layer from the model's input layer, obtaining the output result after computation by each layer of the network. In deep learning models, forward inference computation includes basic operations such as convolution, activation function computation, pooling, and fully connected layer computation.
[0042] The aforementioned target objects can be objects or abnormal situations that need to be identified in drone inspection scenarios, such as damaged insulators, broken strands in cables, and tilted towers in power transmission lines, or illegal fishing, swimming, large animals entering, and garbage accumulation in river channels in river inspection scenarios.
[0043] The aforementioned location information can be parameters describing the position of the target object in the inspection image data. It is usually represented in the form of bounding box coordinates, such as the pixel coordinates of the upper left and lower right corners of the bounding box, or the coordinates of the center point of the bounding box combined with the width and height of the bounding box.
[0044] In its implementation, after the inspection image data is input into the lightweight recognition model, forward inference computation is performed on the inspection image data through this lightweight recognition model. The lightweight recognition model receives the inspection image data as an input tensor. The dimension of the input tensor is typically (number of channels × height × width). For example, for a visible light image, the input dimension is 3 × 640 × 640, where 3 represents the three color channels: red, green, and blue. The lightweight backbone network in the lightweight recognition model first extracts features from the input image. This lightweight backbone network uses depthwise separable convolution instead of standard convolution. Depthwise separable convolution decomposes standard convolution into two independent operations: depthwise convolution and pointwise convolution. Depthwise convolution performs spatial convolution independently on each input channel, using an independent convolution kernel for each channel to extract its spatial features. Pointwise convolution uses a 1×1 convolution kernel to linearly combine the multi-channel feature maps output by the depthwise convolution along the channel dimension for cross-channel feature fusion. The aforementioned lightweight backbone network extracts multi-scale feature maps from shallow to deep layers through multiple stacked depthwise separable convolutional modules. Shallow feature maps have high spatial resolution and less semantic information, making them suitable for detecting small-sized target objects, while deep feature maps have lower spatial resolution and richer semantic information, making them suitable for detecting large-sized target objects.
[0045] Subsequently, the feature fusion network in the aforementioned lightweight recognition model fuses multi-scale feature maps to generate a fused feature map. This feature fusion network employs a feature pyramid structure, upsampling deep feature maps to restore them to the same spatial resolution as shallow feature maps, and then element-wise adding or concatenating the upsampled deep and shallow feature maps. This fusion process allows shallow feature maps to acquire semantic information from deep feature maps, and deep feature maps to acquire spatial location information from shallow feature maps, thus achieving feature representations with both high resolution and strong semantics at different scales. The detection head in the lightweight recognition model performs regression and classification predictions on the fused feature map. The detection head predicts multiple candidate bounding boxes on each grid cell of the fused feature map using convolutional operations. Each candidate bounding box corresponds to a set of prediction parameters, including the offset of the bounding box center coordinates relative to the top-left corner of the grid cell, the scaling of the bounding box width and height relative to the preset anchor box size, the target confidence score, and the conditional probability of each target object category. The target confidence score represents the probability that the current grid cell contains a target object, and the category conditional probability represents the probability of belonging to each category given the presence of a target object. The lightweight recognition model multiplies the target confidence score by the category conditional probability to obtain the comprehensive confidence score for each candidate bounding box. The lightweight recognition model uses a non-maximum suppression algorithm to filter the prediction results. For multiple overlapping bounding boxes generated for the same target object, it retains the bounding box with the highest comprehensive confidence score and removes other redundant bounding boxes. After non-maximum suppression filtering, the lightweight recognition model outputs the location information of each target object, presented in the form of bounding box coordinates.
[0046] For example, suppose a drone captures a frame of inspection image data with a resolution of 640×640 pixels in a power transmission line inspection scenario. The image contains a target of insulator damage. The drone inputs this frame of image into a lightweight recognition model, which uses MobileNetV3 as the backbone network with an input dimension of 3×640×640. The MobileNetV3 network performs depthwise separable convolution operations on the input image, extracting feature maps at three scales with resolutions of 80×80, 40×40, and 20×20, respectively. The feature pyramid network upsamples the 20×20 deep feature map to 40×40, merges it with the original 40×40 feature map, and then upsamples the merged 40×40 feature map to 80×80, merges it with the original 80×80 feature map, generating a fused feature map at three scales. The detection head predicts on a 40×40 scale fused feature map, outputting multiple candidate bounding boxes. The detection result corresponding to the candidate bounding box with the highest confidence is as follows: the target category is insulator damage, the bounding box coordinates are (top left x=300 pixels, y=250 pixels, bottom right x=340 pixels, y=320 pixels), and the overall confidence score is 0.89. After non-maximum suppression filtering, the above lightweight recognition model outputs the location information of the insulator damage target, i.e., the bounding box coordinates (300, 250, 340, 320).
[0047] Step S40: Generate image recognition results based on the location information of each target object.
[0048] It should be noted that the image recognition results described above can be a structured dataset describing the target object information contained in the inspection image data. These image recognition results typically include at least one of the following: the target object's category label, the target object's location information in the image, and the confidence score of the detection result.
[0049] The above category labels can be used to identify the specific type of the target object. For example, in the scenario of power transmission line inspection, the category labels may include insulator damage, cable strand breakage, tower tilting, bird nests, etc.; in the scenario of river inspection, the category labels may include anglers, swimmers, large animals, river garbage, etc.
[0050] The confidence score mentioned above can represent the credibility of the detection result. The value range is usually between 0 and 1. The higher the value, the more confident the model is in the detection result.
[0051] In the specific implementation, after obtaining the location information of each target object, an image recognition result is generated based on this information. The detection result for each target object is obtained from the output of the lightweight recognition model. Each detection result includes at least the target object's location information, category label, and confidence score. Based on the bounding box coordinates of each target object, the corresponding target region image is cropped from the original inspection image data. This cropping operation extracts pixel data within the pixel region defined by the bounding box coordinates, generating an independent sub-image. This target region image retains the visual features of the target object and can be used for subsequent detailed analysis or manual verification. The UAV then performs classification recognition on the target region image to determine the category label for each target object. If the lightweight recognition model has already output the category label, the category label output by the model is directly read as the classification recognition result. To further improve recognition accuracy, a more refined classification model can be called to perform secondary classification verification on the target region image. For example, a lightweight classification network can be used to reclassify the cropped target region image to confirm the specific category of the target object. Finally, the category label and location information of each target object are associated and stored to generate image recognition results. This associated storage can use the category label as the key and the location information and confidence score as the value to construct a structured record entry. The image recognition results can also be associated with the inspection image data itself; for example, the bounding box of the detected target object can be drawn on the original inspection image, and the category label and confidence score can be annotated next to the bounding box to generate an annotated result image. The generated image recognition results are stored in the airborne memory and uploaded to the ground control platform in real time or periodically via a wireless communication link for ground operators to view and further process.
[0052] This embodiment first acquires inspection image data collected by a UAV; then, the inspection image data is input into a lightweight recognition model, which is obtained by replacing the backbone network in the standard target detection framework with a lightweight backbone network and performing quantization compression; next, the lightweight recognition model performs forward inference calculations on the inspection image data to obtain the position information of each target object in the inspection image data; finally, image recognition results are generated based on the position information of each target object. Compared with existing methods, this embodiment achieves real-time recognition of large-scale images from the UAV edge by replacing cloud-based recognition with an edge-side lightweight model.
[0053] Based on the first embodiment of this application, a second embodiment of this application is proposed. In the second embodiment, content that is the same as or similar to that in the first embodiment described above can be referred to the above description, and will not be repeated hereafter. Based on this, please refer to... Figure 2 , Figure 2This is a flowchart illustrating the second embodiment of the image recognition method based on unmanned aerial vehicles (UAVs) in this application.
[0054] In this embodiment, the lightweight backbone network includes the MobileNetV3 network. The lightweight recognition model is obtained by replacing the backbone network in the standard object detection framework with a lightweight backbone network and performing quantization compression. Step S201: Determine the total number of target object categories that need to be identified in the inspection scenario, and use the total number of categories as the preset number of categories.
[0055] Step S202: Replace the original backbone network in the standard target detection framework with the MobileNetV3 network to obtain the recognition model after backbone network replacement.
[0056] It should be noted that the aforementioned inspection scenarios can refer to the specific environmental and task types in which the drone performs inspection missions, such as power line inspection scenarios, river inspection scenarios, dam inspection scenarios, or highway inspection scenarios. Different inspection scenarios require the identification of different target object categories. For example, power line inspection scenarios require the identification of fault types such as insulator damage, broken cable strands, and tilted towers, while river inspection scenarios require the identification of target types such as anglers, swimmers, large animals, and river garbage.
[0057] The total number of target object categories mentioned above can be the number of all possible target object categories that the drone needs to identify in a specific inspection scenario. For example, in the scenario of power transmission line inspection, if only two faults, insulator damage and cable strand breakage, are needed to be identified, the total number of target object categories is 2; if five faults, such as insulator damage, cable strand breakage, tower tilting, bird nests, and vibration damper detachment, are needed to be identified, the total number of target object categories is 5.
[0058] The aforementioned preset number of categories can be the target output dimension used to guide the pruning of the detection head during model compression, i.e., the number of categories that the model's classification branch needs to output. The MobileNetV3 network can be a lightweight convolutional neural network structure built on depthwise separable convolution and channel attention mechanisms, consisting of multiple stacked inverted residual modules. Each inverted residual module first expands the low-dimensional input to a high-dimensional channel space through a 1×1 convolution, then extracts spatial features through depthwise separable convolution, and finally compresses the high-dimensional features back to the low-dimensional channel space through a 1×1 convolution. The recognition model after the backbone network replacement can be a recognition model formed by removing the original backbone network (e.g., CSPDarknet) from the standard object detection framework and connecting the MobileNetV3 network as the new backbone network.
[0059] In the specific implementation, when constructing and compressing the lightweight recognition model, the total number of target object categories to be identified in the inspection scenario is first determined, and this total number of categories is used as the preset category quantity. Specifically, based on the specific inspection task to be performed by the drone, a pre-configured list of recognition categories is obtained. This list of recognition categories can be preset by the operator according to the inspection scenario and stored in a configuration file. For example, for a power transmission line inspection task, the list of recognition categories may include insulator damage, cable strand breakage, tower tilting, bird nests, vibration damper detachment, and equipotential ring damage. After reading the list of recognition categories, the number of category entries in the list is counted, and the counted value is used as the total number of target object categories. This value is also stored in memory as the preset category quantity for subsequent head pruning operations.
[0060] Next, the original backbone network in the standard object detection framework is replaced with a MobileNetV3 network to obtain a recognition model with a replaced backbone network. Specifically, a pre-trained standard object detection framework is first loaded, such as a YOLOv8 model pre-trained on a large public dataset. This model consists of three components: a CSPDarknet backbone network, a feature pyramid network, and a detection head. The CSPDarknet backbone network is removed from the model, while the structural definitions of the feature pyramid network and the detection head are retained. A pre-trained MobileNetV3 network is obtained, which can be pre-trained on large image classification datasets such as ImageNet and has good feature extraction capabilities. The MobileNetV3 network is connected to the original backbone network, i.e., the output of the MobileNetV3 network is connected to the input of the feature pyramid network. The output feature map of the MobileNetV3 network is used as the input of the feature pyramid network, which then performs multi-scale feature fusion operations. After the above replacement, a recognition model with a replaced backbone network is obtained. The backbone network of this model is MobileNetV3, which has a much lower computational cost than the original CSPDarknet backbone network.
[0061] Step S203: Perform channel pruning on the detection head in the recognition model after the backbone network is replaced according to the preset number of categories, compress the output dimension of the classification branch to the preset number of categories, and obtain the pruned recognition model.
[0062] Step S204: Use the INT8 quantization method to compress the weight parameters of the pruned recognition model from floating-point numbers to integers to generate a compressed lightweight recognition model.
[0063] It should be noted that the aforementioned detection head can be a network module in a standard object detection framework used to output prediction results, typically consisting of two parts: a classification branch and a regression branch. The classification branch predicts the probability that the target object belongs to each category, and its output dimension is equal to the total number of target object categories to be identified; the regression branch predicts the bounding box position parameters of the target object, and its output dimension is usually related to the number of bounding box coordinate parameters.
[0064] The aforementioned channel pruning process can be a compression technique that reduces the number of model parameters and computational cost by removing unimportant convolutional kernels or feature channels in a convolutional neural network.
[0065] The output dimension of the aforementioned classification branch can be the length of the category score vector output in the last fully connected layer or convolutional layer of the classification branch of the detection head. For example, an output dimension of 80 indicates that 80 different target objects can be identified.
[0066] The INT8 quantization method described above is a technique that converts the weight parameters and activation values in a neural network model from a high-precision floating-point representation to an 8-bit integer representation. By mapping 32-bit floating-point numbers (FP32) to an integer range of -128 to 127 or 0 to 255, it achieves model volume compression and inference acceleration.
[0067] In the specific implementation, after obtaining the preset number of categories and the recognition model after replacing the backbone network, channel pruning is performed on the detection head in the recognition model after the backbone network replacement according to the preset number of categories. The detection head module in the recognition model after the backbone network replacement is located, and the detection head module typically contains two sub-networks: a classification branch and a regression branch. The last convolutional layer or fully connected layer of the classification branch is selected. The number of output channels of this layer corresponds to the number of categories in the original dataset. For example, when YOLOv8 is pre-trained on the COCO dataset, the output dimension of the classification branch is 80. The number of output channels of the last layer of the classification branch is modified from the original value to the preset number of categories, for example, from 80 to 5. Channel pruning can also be further performed on the intermediate layers in the classification branch by removing redundant convolutional kernels by calculating the importance index of each convolutional kernel. A norm-based pruning strategy is adopted, calculating the L1 norm or L2 norm of each convolutional kernel weight, and marking convolutional kernels with a norm less than a preset threshold as unimportant and removing them. After removing the convolutional kernels, the channel dimensions of subsequent network layers are adjusted accordingly to maintain the connectivity of the network structure. After the above channel pruning process, the output dimension of the classification branch is compressed to the preset number of categories, while the regression branch retains its original structure, resulting in the pruned recognition model.
[0068] Next, the weight parameters of the pruned recognition model are compressed from floating-point numbers to integers using INT8 quantization, generating a compressed lightweight recognition model. Specifically, all learnable parameters in the pruned recognition model are read, including the weight parameters of convolutional layers, fully connected layers, and the scaling factor and bias parameters of batch normalization layers. These parameters are typically stored in 32-bit floating-point (FP32) format. Statistical analysis is performed on the weight parameters of each layer to calculate the maximum and minimum values, determining the quantization scaling factor and zero-point offset. The quantization scaling factor is calculated as the maximum value minus the minimum value divided by 255, and the zero-point offset is calculated as the negative minimum value divided by the scaling factor and then rounded down. Based on the calculated scaling factor and zero-point offset, each floating-point weight parameter is mapped to an 8-bit integer using the formula: the integer value is the floating-point value divided by the scaling factor plus the zero-point offset and then rounded down. The quantized integer weight parameters are stored in the model file in INT8 format, while the scaling factor and zero-point offset of each layer are retained for dequantization operations during inference. After INT8 quantization, a compressed lightweight recognition model is generated. The storage space of the weight parameters of this model is compressed from four bytes per parameter in FP32 format to one byte per parameter in INT8 format, reducing the model size by about four times.
[0069] Furthermore, in this embodiment, before the step of performing channel pruning on the detection heads in the replaced recognition model of the backbone network according to the preset number of categories, the method further includes: Step S2031: Obtain a pre-trained standard object detection model.
[0070] Step S2032: Using the knowledge distillation method, the pre-trained standard target detection model is used to train the recognition model after the backbone network is replaced.
[0071] Step S2033: Use the recognition model after the training backbone network is replaced as the object to perform the channel pruning process.
[0072] It should be noted that the aforementioned pre-trained standard object detection models can be standard object detection models that have been pre-trained on large-scale public datasets (such as the COCO dataset, ImageNet dataset, or OpenImages dataset), and typically have a large model size and high detection accuracy.
[0073] The aforementioned pre-trained standard target detection models employ heavy backbone networks, such as CSPDarknet, ResNet, or VGG networks, which have a large number of parameters and high computational complexity, making them difficult to deploy directly on the UAV edge. However, they have high detection accuracy and can be used as teacher models in knowledge distillation.
[0074] The aforementioned knowledge distillation method can be a model compression technique that transfers knowledge from the teacher model to the student model by having a student model with fewer parameters mimic the output behavior of a teacher model with more parameters.
[0075] The aforementioned recognition model after backbone network replacement can be a model trained as a student model during the knowledge distillation process. This model has already completed backbone network replacement (e.g., replacing the original backbone network with a MobileNetV3 network), but has not yet undergone sufficient training and optimization.
[0076] In the specific implementation, before performing channel pruning, the recognition model after backbone network replacement is trained and optimized using knowledge distillation. A pre-trained standard object detection model is obtained, which serves as the teacher model in knowledge distillation. This teacher model has a large number of parameters and high detection accuracy, enabling accurate detection of target objects in various complex scenarios. Simultaneously, the recognition model after backbone network replacement is obtained, serving as the student model in knowledge distillation. This student model uses a lightweight backbone network with far fewer parameters than the teacher model, but its detection accuracy is not yet optimal and needs to learn from the teacher model through knowledge distillation. Knowledge distillation is used to train the recognition model after backbone network replacement using the pre-trained standard object detection model. A training image dataset is prepared, containing actual images collected in inspection scenarios, with each image labeled with the true category and true location coordinates of the target object. The training images are simultaneously input into both the teacher and student models. The teacher model performs forward inference calculations on the input images and outputs the teacher model's prediction results, including the probability distribution of the target category and the position parameters of the bounding box. The student model performs forward inference on the same input image and outputs its prediction. A distillation loss function is calculated, typically consisting of hard and soft losses. The hard loss is calculated based on the difference between the student model's prediction and the ground truth label, using either cross-entropy loss or mean squared error loss. The soft loss is calculated based on the difference between the student model's prediction and the teacher model's prediction, using either KL divergence or mean squared error loss. The soft loss allows the student model to mimic the teacher model's output behavior. The hard and soft losses are then weighted and summed to obtain the total loss function. The weight coefficients for the soft loss are typically set between 0.7 and 0.9, and the weight coefficients for the hard loss are set between 0.1 and 0.3. The gradient is calculated based on the total loss function, and an optimizer (e.g., a stochastic gradient descent optimizer or an Adam optimizer) is used to update the student model's weight parameters. This process is repeated for multiple rounds of iterative training on the training dataset until the total loss function converges or a predetermined number of training rounds is reached.
[0077] The trained recognition model with the replaced backbone network is then used as the target for channel pruning. After knowledge distillation training, the detection accuracy of the student model (i.e., the recognition model with the replaced backbone network) is significantly improved, approaching or reaching the level of the teacher model, while maintaining the computational efficiency advantage of the lightweight backbone network. The trained recognition model with the replaced backbone network is then passed to the subsequent channel pruning module as the input model for the pruning operation.
[0078] For example, in a power transmission line inspection scenario, the teacher model is a YOLOv8 model pre-trained on the COCO dataset. This model uses the CSPDarknet backbone network, has approximately 25 megabytes of parameters, and achieves a detection accuracy (mAP) of 0.65. The student model is a recognition model with a replaced backbone network, using MobileNetV3, with approximately 4 megabytes of parameters and an initial detection accuracy of 0.52. 5000 power transmission line inspection images are prepared as the training dataset, each labeled with the category and location information of fault targets such as insulator damage and cable strand breakage. The training images are simultaneously input into both the teacher and student models, and a distillation loss function is calculated, with the soft loss weight set to 0.8 and the hard loss weight set to 0.2. The Adam optimizer is used, with an initial learning rate of 0.001, and 100 iterations of training are performed on the training dataset. After training, the student model's detection accuracy improves to 0.61, approaching the level of the teacher model, while maintaining a small model size of 4 megabytes and a relatively fast inference speed. The recognition model after replacing the trained backbone network is used as the object of channel pruning for subsequent pruning and compression operations.
[0079] Furthermore, in this embodiment, the step of compressing the weight parameters of the pruned recognition model from floating-point numbers to integers using the INT8 quantization method includes: Step S2041: Obtain the actual collected images in the inspection scenario, and use the actual collected images as the calibration dataset.
[0080] Step S2042: Using the quantization-perception training method, the pruned recognition model is trained with quantization error compensation using the calibration dataset to obtain the recognition model after quantization error compensation.
[0081] Step S2043: Using the INT8 quantization method, the weight parameters of the recognition model after quantization error compensation are compressed from floating-point numbers to integers.
[0082] It should be noted that the aforementioned actual acquired images can be real-world images taken by the UAV during actual inspection missions, including images from different inspection scenarios such as power transmission lines, rivers, and dams. The difference between the aforementioned actual acquired images and the training images is that the actual acquired images realistically reflect the lighting conditions, weather conditions, background complexity, and target scale distribution encountered by the UAV in the actual flight environment, and can better represent the deployment environment of the model.
[0083] The aforementioned calibration dataset can be a small set of images used for quantization-aware training, typically containing 100 to 500 actual acquired images, used to statistically analyze the distribution range of activation values in each layer of the model and to compensate for quantization errors.
[0084] The aforementioned quantization-aware training method can be a technique that simulates quantization operations during model training. By simulating the effect of INT8 quantization in the forward propagation of training and using pseudo-gradients to update the weights in the backward propagation, the model weights gradually adapt to the accuracy loss caused by quantization.
[0085] The aforementioned quantization error compensation training can be achieved through quantization-aware training, allowing the model to learn how to compensate for the precision loss caused by the conversion from floating-point numbers to integers.
[0086] The recognition model after quantization error compensation can be a model obtained through quantization perception training. Its weight parameters are still stored in floating-point form, but have been adaptively adjusted for quantization error, so that they can better adapt to subsequent INT8 quantization operations.
[0087] In the specific implementation, before performing INT8 quantization, the pruned recognition model is trained using quantization-aware training to compensate for quantization errors. Actual images from inspection scenarios are acquired; for example, 200 representative images are selected from historical UAV inspection data. These images cover different lighting conditions (sunny, cloudy, dusk), different weather conditions (no rain, light rain, fog), and different background complexities (simple sky background, complex vegetation background). These actual images are used as a calibration dataset to simulate various input distributions the model will encounter in the deployment environment. Then, quantization-aware training is used to train the pruned recognition model using the calibration dataset to compensate for quantization errors. Simulated quantization nodes are inserted after each weight layer and activation layer in the pruned recognition model. These simulated quantization nodes perform pseudo-quantization operations during forward propagation, simulating the conversion of floating-point weights and activation values to low-precision integers and then back to floating-point numbers, simulating the precision loss effect of INT8 quantization. The formula for calculating pseudo-quantization is: the quantized value equals the original value divided by the scaling factor, rounded down, and then multiplied by the scaling factor, where the scaling factor is dynamically calculated based on the numerical range of the weights or activation values of that layer. Several rounds of fine-tuning training are performed on the model with inserted simulated quantization nodes on the calibration dataset, typically requiring only a few rounds (e.g., 10 to 20 rounds) and a low learning rate (e.g., one-tenth of the initial learning rate). During this fine-tuning training, the model's forward propagation experiences the accuracy loss due to simulated quantization, while the backpropagation uses a pass-through estimator to calculate the gradient, assuming the gradient of the quantization operation is 1, directly passing the loss gradient to the weights before quantization. After this quantization-aware training, the model learns to adjust the weight parameters to adapt to the accuracy loss caused by quantization; for example, weights that were originally sensitive to quantization will gradually adjust to a numerical range with smaller quantization errors. This yields a quantization-error-compensated recognition model. The weight parameters of this model are still stored in floating-point form, but have undergone adaptive quantization adjustment. Then, the INT8 quantization method is used to compress the weight parameters of the quantization-error-compensated recognition model from floating-point numbers to integers. Read the floating-point values of the weight parameters for each layer in the recognition model after quantization error compensation, and calculate the numerical range of each layer's weights. Calculate the quantization parameters for each layer based on the numerical range, including the scaling factor and zero-point offset. Map each floating-point weight parameter to an 8-bit integer using the formula: the quantized integer value equals the original floating-point value divided by the scaling factor plus the zero-point offset, rounded down, with the result limited to the range of -128 to 127 or 0 to 255. Store the quantized integer weight parameters in INT8 format, replacing the original floating-point weight parameters, while retaining the scaling factor and zero-point offset for each layer's dequantization operation during inference.After the above INT8 quantization, a compressed lightweight recognition model is obtained. The storage space of the weight parameters of this model is compressed from four bytes per parameter in FP32 format to one byte per parameter in INT8 format, reducing the model size by about four times and increasing the inference speed by two to three times.
[0088] For example, suppose in a power transmission line inspection scenario, 200 actual images collected from historical drone inspection data are selected as a calibration dataset. These images include 100 images taken on sunny days, 60 on cloudy days, and 40 on foggy days. The pruned recognition model has approximately 3 megabytes of parameters, stored in FP32 format. The pruned recognition model is trained on the calibration dataset using quantization perception. Simulated quantization nodes are inserted after each layer of the model, with 15 training epochs and a learning rate of 0.0001. After training, the model's detection accuracy on the validation set slightly decreases from 88.5% before quantization to 88.1%, a loss of only 0.4 percentage points, far lower than the more than 2 percentage point accuracy loss that direct quantization might cause. Next, the quantization error-compensated recognition model is subjected to INT8 quantization, converting approximately 800,000 weight parameters from FP32 format to INT8 format. The converted model size is 0.75 megabytes. The quantized lightweight recognition model was deployed on the drone side, reducing the inference time per frame from 300 milliseconds to 120 milliseconds, thus meeting the real-time recognition requirements.
[0089] Based on the first and second embodiments of this application, a third embodiment of this application is proposed. In this third embodiment, content that is the same as or similar to the first and second embodiments described above can be referred to the above description, and will not be repeated hereafter. Based on this, please refer to... Figure 3 , Figure 3 This is a flowchart illustrating the third embodiment of the image recognition method based on unmanned aerial vehicles (UAVs) in this application.
[0090] The step of performing forward inference calculations on the inspection image data using the lightweight recognition model to obtain the location information of each target object in the inspection image data includes: Step S301: Input the inspection image data into the lightweight recognition model, and perform depthwise separable convolution operation on the inspection image data through the lightweight backbone network to extract multi-scale feature maps.
[0091] Step S302: Perform feature pyramid network fusion processing on the multi-scale feature map to generate a fused feature map.
[0092] Step S303: Perform regression prediction on the fused feature map to obtain the location information of each target object in the inspection image data.
[0093] It should be noted that the above-mentioned depthwise separable convolution operation can be a convolution calculation method that decomposes standard convolution into two independent steps: depthwise convolution and pointwise convolution.
[0094] The aforementioned depthwise convolution performs spatial convolution independently on each input channel. Each channel uses an independent two-dimensional convolution kernel to extract the spatial features of that channel. Depthwise convolution does not change the number of channels for the input features.
[0095] The pointwise convolution described above uses a 1×1 kernel to linearly combine the multi-channel feature maps output by the depthwise convolution along the channel dimension, enabling cross-channel feature fusion and adjusting the number of output channels. These multi-scale feature maps can be feature representations output at different depth levels of a lightweight backbone network. Shallow feature maps have higher spatial resolution and less semantic information, while deep feature maps have lower spatial resolution and richer semantic information.
[0096] The aforementioned feature pyramid network fusion process can be considered a top-down feature fusion mechanism. It transfers strong semantic information from deep feature maps to shallow feature maps through upsampling operations, and then horizontally connects and fuses feature maps from different levels element-wise. The regression prediction described above can be the process of predicting the bounding box's positional parameters for each location in the fused feature map, including the bounding box's center coordinates, width, and height offsets relative to preset anchor boxes.
[0097] In its implementation, after the inspection image data is input into the lightweight recognition model, the model performs depthwise separable convolution operations on the image data through a lightweight backbone network to extract multi-scale feature maps. The lightweight backbone network (e.g., MobileNetV3) in this model receives the inspection image data as an input tensor, typically with dimensions of 3×H×W, where 3 represents the red, green, and blue color channels, and H and W represent the image's height and width, respectively. This lightweight backbone network is composed of multiple stacked inverted residual modules, each performing a depthwise separable convolution operation. In this depthwise separable convolution operation, the lightweight backbone network first performs depthwise convolution, applying an independent two-dimensional convolution kernel to each channel of the input feature map. Each kernel is typically 3×3 or 5×5 in size, and the number of channels in the output feature map is the same as the number of channels in the input feature map. Subsequently, the aforementioned lightweight backbone network performs pointwise convolution, that is, it uses a 1×1 convolution kernel to perform convolution operations on the feature maps output by the depthwise convolution. The 1×1 convolution linearly combines the information of each channel, adjusting the number of output channels to the target number. This lightweight backbone network progressively downsamples and extracts features from the input image by stacking multiple inverted residual modules. Features are extracted at the output position of each module with a specific downsampling factor, resulting in multiple feature maps with different spatial resolutions. These feature maps constitute a multi-scale feature map set. For example, feature maps are extracted at downsampling factors of 8, 16, and 32, respectively, with spatial resolutions of 1 / 8, 1 / 16, and 1 / 32 of the input image. The lightweight recognition model then performs feature pyramid network fusion processing on the multi-scale feature maps to generate fused feature maps. The lightweight recognition model first constructs a top-down path, starting from the deepest layer (lowest spatial resolution) feature map, and performs a 2x upsampling through bilinear interpolation or nearest neighbor interpolation to ensure that the spatial resolution of the deep feature map is consistent with that of the adjacent shallow feature maps. The lightweight recognition model described above horizontally connects the upsampled deep feature maps with their corresponding shallow feature maps, fusing them through element-wise addition or concatenation. Before fusion, the model can perform a 1×1 convolution on the shallow feature maps to adjust the number of channels, matching the number of channels in the upsampled deep feature maps. The model repeats the upsampling and fusion operations from deep to shallow, generating fused feature maps at each scale. These fused feature maps typically contain three scales, used to detect target objects of different sizes: a large-scale fused feature map (high resolution) for detecting small targets, a medium-scale fused feature map for detecting medium-sized targets, and a small-scale fused feature map (low resolution) for detecting large targets. The model then performs regression prediction on the fused feature maps to obtain the location information of each target object in the inspected image data.The detection head in the aforementioned lightweight recognition model performs convolution operations on the fused feature map at each scale to generate a prediction tensor. For each grid cell in the fused feature map, the detection head predicts multiple anchor boxes, each prediction including the bounding box's positional parameters. These parameters include the offset of the bounding box's center coordinates relative to the top-left corner of the grid cell, and the scaling of the bounding box's width and height relative to a preset anchor box size. The prediction results output by the detection head are decoded to convert the offset and scaling into actual bounding box coordinates. The lightweight recognition model then aggregates the prediction results at each scale and uses a non-maximum suppression algorithm to select the prediction boxes with the highest confidence, ultimately outputting the positional information of each target object in the inspected image data, represented by bounding box coordinates.
[0098] Furthermore, in this embodiment, the step of generating image recognition results based on the location information of each target object includes: Step S401: Based on the location information of each target object, crop out the corresponding target area image from the inspection image data.
[0099] Step S402: Perform classification recognition on the target region image to determine the category label of each target object.
[0100] Step S403: Associate the category label and location information of each target object to generate image recognition results.
[0101] It should be noted that the above cropping can be an operation of extracting sub-images from the original inspection image data according to a specified coordinate range.
[0102] The target region image mentioned above can be a small sub-image obtained after cropping, containing only the local area where the target object is located, and excluding the background or other irrelevant objects.
[0103] The above classification and recognition can be a process of judging the category of the input image and outputting the category label to which the image belongs. For example, it can identify the fault type in the target area image as a person, broken insulator, broken cable strand, or tilted tower.
[0104] The aforementioned category labels can be identifiers used to identify the specific type of the target object. For example, in a power transmission line inspection scenario, category labels could include "insulator damage," "cable strand breakage," "tower tilt," and "bird's nest." The aforementioned association can be a process of binding and integrating information from different dimensions (such as category labels and location information) according to their corresponding relationships to form a complete inspection record.
[0105] In the specific implementation, after obtaining the location information of each target object, the corresponding target region image is cropped from the inspection image data based on the location information of each target object. The location information of each target object output by the lightweight recognition model is obtained. This location information is represented in bounding box coordinates, such as the coordinates of the top-left and bottom-right corners of the bounding box, or the coordinates of the center point of the bounding box combined with the width and height of the bounding box. Based on these bounding box coordinates, the pixel region where the target object is located is located in the original inspection image data. All pixel values within this pixel region are extracted from the original inspection image data to generate an independent sub-image, i.e., the target region image. The target region image maintains the color channel format and pixel precision of the original image, and its size is equal to the width multiplied by the height of the bounding box. The target region image can be scaled to a fixed size, such as 224×224 pixels or 227×227 pixels, to facilitate input to the subsequent classification and recognition model. Classification and recognition are performed on the target region image to determine the category label of each target object. An independent lightweight classification model can be used to classify the target region image. The lightweight classification model described above can be a classification network pre-trained on an inspection scenario dataset, such as MobileNetV3, EfficientNet-Lite, or ShuffleNet. The target region image is input into this lightweight classification model, which performs forward inference calculations and outputs a probability vector. The dimension of the probability vector equals the total number of preset categories, and the value in each dimension represents the probability that the input image belongs to the corresponding category. The category corresponding to the highest probability value in the probability vector is selected as the category label for the target region image. To improve processing efficiency, the category labels output by the lightweight recognition model can be directly reused without calling the additional classification model. The lightweight recognition model has already output the category label and confidence score for each detection box during the target detection process, and this category label can be directly read as the classification result. The category label and location information of each target object are associated to generate the image recognition result. A detection record is constructed for each detected target object; the data structure of the detection record can be a set of key-value pairs or a structure. The aforementioned detection record includes at least a category label field and a location information field. The category label field stores the category label of the target object, and the location information field stores the bounding box coordinates of the target object in the inspected image data. The detection record may also include a confidence score field, storing the model's confidence score for the detection result. All detection records in the current frame image are summarized to generate a detection result list, which is the image recognition result.Image recognition results can be associated and stored with the original inspection image data. For example, the list of detection results can be appended as metadata to the image file, or the detected bounding boxes can be drawn on the original inspection image and labeled with category labels and confidence scores next to the bounding boxes to generate labeled result images. The generated image recognition results are stored in the airborne memory and uploaded to the ground control platform in real time or periodically via a wireless communication link.
[0106] For example, suppose a frame of an inspection image has a size of 1920×1080 pixels. The lightweight recognition model detects two target objects in this frame: the first target object is labeled "insulator damage," with location information in bounding box coordinates (top left x=800 pixels, y=450 pixels, bottom right x=880 pixels, y=520 pixels), and a confidence score of 0.89; the second target object is labeled "cable strand breakage," with location information in bounding box coordinates (top left x=1200 pixels, y=600 pixels, bottom right x=1280 pixels, y=660 pixels), and a confidence score of 0.76. Based on the bounding box coordinates of the first target object, the drone crops a pixel region with coordinates (800, 450, 880, 520) from the original inspection image, generating an 80×70 pixel target region image. This sub-image clearly shows the detailed features of the insulator skirt damage. The aforementioned drone used a lightweight classification model to classify and identify the target area image. The probability vector output by the classification model showed a probability of 0.92 for the insulator damage category, confirming the category label as "insulator damage". The drone also cropped the corresponding target area image based on the location information of the second target object. The category labels, location information, and confidence scores of the two target objects were correlated to generate the image recognition result, as follows: the detection timestamp was 10:30:15 AM on June 11, 2025, the image frame number was frame 15, and two targets were detected. Target 1 was classified as insulator damage, located at (800, 450, 880, 520), with a confidence score of 0.89; Target 2 was classified as cable strand breakage, located at (1200, 600, 1280, 660), with a confidence score of 0.76. The aforementioned drone stores the image recognition results in its onboard memory and marks the positions of two target objects with red rectangles on the original inspection image, labeling "insulator damage 0.89" and "cable strand breakage 0.76" next to the rectangles, generating an annotated result image. Subsequently, the drone uploads the image recognition results and the annotated result image to the ground control platform.
[0107] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the image recognition method based on UAVs in this application. Any simple modifications based on this technical concept are within the protection scope of this application.
[0108] This application also provides an image recognition device based on a drone; please refer to [reference needed]. Figure 4 The device includes: Image acquisition module 10 is used to acquire inspection image data collected by the UAV; The model loading module 20 is used to input the inspection image data into the lightweight recognition model, which is obtained by replacing the backbone network in the standard target detection framework with a lightweight backbone network and performing quantization compression. The inference calculation module 30 is used to perform forward inference calculation on the inspection image data through the lightweight recognition model to obtain the location information of each target object in the inspection image data. The result generation module 40 is used to generate image recognition results based on the location information of each target object.
[0109] The UAV-based image recognition device provided in this application, employing the UAV-based image recognition method described in the above embodiments, can solve the technical problem of how to achieve rapid and accurate recognition of large-scale images under limited computing power conditions on the UAV side. Compared with the prior art, the beneficial effects of the UAV-based image recognition device provided in this application are the same as those of the UAV-based image recognition method provided in the above embodiments, and other technical features in the UAV-based image recognition device are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.
[0110] This application provides an image recognition device based on a drone, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, which are executed by the at least one processor to enable the at least one processor to perform the image recognition method based on the drone in Embodiment 1 described above.
[0111] The following is for reference. Figure 5 The diagram illustrates a structural schematic of an image recognition device based on a drone suitable for implementing embodiments of this application. The drone-based image recognition device in these embodiments may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital radio receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Description), PMPs (Portable Media Players), and in-vehicle terminals (e.g., in-vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. Figure 5The drone-based image recognition device shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.
[0112] like Figure 5 As shown, the UAV-based image recognition device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM) 1004. The RAM 1004 also stores various programs and data required for the operation of the UAV-based image recognition device. The processing unit 1001, ROM 1002, and RAM 1004 are interconnected via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to the I / O interface 1006: input devices 1007 including, for example, a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 1003 including, for example, magnetic tape, hard disk, etc.; and communication devices 1009. Communication device 1009 allows the UAV-based image recognition device to communicate wirelessly or wiredly with other devices to exchange data. While the figures show UAV-based image recognition devices with various systems, it should be understood that implementation or possession of all the systems shown is not required. More or fewer systems may be implemented alternatively.
[0113] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from ROM 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.
[0114] The UAV-based image recognition device provided in this application, employing the UAV-based image recognition method described in the above embodiments, can solve the technical problem of how to achieve rapid and accurate recognition of large-scale images under limited computing power conditions on the UAV side. Compared with the prior art, the beneficial effects of the UAV-based image recognition device provided in this application are the same as those of the UAV-based image recognition method provided in the above embodiments, and other technical features in this UAV-based image recognition device are the same as those disclosed in the previous embodiment method, and will not be repeated here.
[0115] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.
[0116] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0117] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, the computer-readable program instructions being used to execute the image recognition method based on a drone in the above embodiments.
[0118] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.
[0119] The aforementioned computer-readable storage medium may be included in an image recognition device based on a drone; or it may exist independently and not assembled into an image recognition device based on a drone.
[0120] The aforementioned computer-readable storage medium carries one or more programs that, when executed by an image recognition device based on a drone, cause the image recognition device based on a drone to: acquire inspection image data collected by the drone; input the inspection image data into a lightweight recognition model, which is obtained by replacing the backbone network in a standard target detection framework with a lightweight backbone network and performing quantization compression; perform forward inference calculations on the inspection image data through the lightweight recognition model to obtain the position information of each target object in the inspection image data; and generate an image recognition result based on the position information of each target object.
[0121] Computer program code for performing the operations of this application can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0122] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0123] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.
[0124] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the above-described UAV-based image recognition method. This solves the technical problem of how to achieve rapid and accurate recognition of large-scale images under limited computing power conditions on the UAV side. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the UAV-based image recognition method provided in the above embodiments, and will not be repeated here.
[0125] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the image recognition method based on unmanned aerial vehicles as described above.
[0126] The computer program product provided in this application can solve the technical problem of how to achieve fast and accurate recognition of large-scale images under the limited computing power of the UAV. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as those of the UAV-based image recognition method provided in the above embodiments, and will not be repeated here.
[0127] The above description is only a part of the embodiments of this application and does not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.
Claims
1. An image recognition method based on unmanned aerial vehicles (UAVs), characterized in that, The method is applied to an unmanned aerial vehicle (UAV) inspection system, and the method includes: Acquire inspection image data collected by drones; The inspection image data is input into a lightweight recognition model, which is obtained by replacing the backbone network in the standard target detection framework with a lightweight backbone network and performing quantization compression. The lightweight recognition model is used to perform forward inference calculations on the inspection image data to obtain the location information of each target object in the inspection image data. Image recognition results are generated based on the location information of each target object.
2. The method as described in claim 1, characterized in that, The lightweight backbone network includes the MobileNetV3 network. The lightweight recognition model is obtained by replacing the backbone network in the standard object detection framework with a lightweight backbone network and performing quantization compression. Determine the total number of target object categories that need to be identified in the inspection scenario, and use the total number of categories as the preset number of categories; The original backbone network in the standard target detection framework is replaced with the MobileNetV3 network to obtain the recognition model after backbone network replacement. Based on the preset number of categories, channel pruning is performed on the detection heads in the recognition model after the backbone network is replaced, and the output dimension of the classification branch is compressed to the preset number of categories to obtain the pruned recognition model. The weight parameters of the pruned recognition model are compressed from floating-point numbers to integers using the INT8 quantization method, generating a compressed lightweight recognition model.
3. The method as described in claim 2, characterized in that, Before the step of performing channel pruning on the detection heads in the recognition model after replacing the backbone network according to the preset number of categories, the following steps are also included: Obtain a pre-trained standard object detection model; The knowledge distillation method is used to train the recognition model after the backbone network is replaced using the pre-trained standard target detection model; The recognition model replaced by the trained backbone network is used as the object to perform the channel pruning process.
4. The method as described in claim 2, characterized in that, The step of compressing the weight parameters of the pruned recognition model from floating-point numbers to integers using the INT8 quantization method includes: Acquire actual images in the inspection scenario and use the actual images as a calibration dataset. The quantization-perception training method is used to train the pruned recognition model with quantization error compensation through the calibration dataset to obtain the recognition model after quantization error compensation. The INT8 quantization method is used to compress the weight parameters of the recognition model after quantization error compensation from floating-point numbers to integers.
5. The method as described in claim 1, characterized in that, The step of performing forward inference calculations on the inspection image data using the lightweight recognition model to obtain the location information of each target object in the inspection image data includes: The inspection image data is input into the lightweight recognition model, and the lightweight backbone network is used to perform depthwise separable convolution operations on the inspection image data to extract multi-scale feature maps. The multi-scale feature maps are fused using a feature pyramid network to generate a fused feature map. Regression prediction is performed on the fused feature map to obtain the location information of each target object in the inspection image data.
6. The method as described in claim 1, characterized in that, The step of generating image recognition results based on the location information of each target object includes: Based on the location information of each target object, the corresponding target area image is cropped from the inspection image data; Perform classification and recognition on the target region image to determine the category label for each target object; The category label and location information of each target object are associated to generate image recognition results.
7. An image recognition device based on a drone, characterized in that, The device includes: The image acquisition module is used to acquire inspection image data collected by the drone; The model loading module is used to input the inspection image data into the lightweight recognition model, which is obtained by replacing the backbone network in the standard target detection framework with a lightweight backbone network and performing quantization compression. The inference calculation module is used to perform forward inference calculation on the inspection image data through the lightweight recognition model to obtain the location information of each target object in the inspection image data. The result generation module is used to generate image recognition results based on the location information of each target object.
8. An image recognition device based on a drone, characterized in that, The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the image recognition method based on a drone as described in any one of claims 1 to 6.
9. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the image recognition method based on unmanned aerial vehicles as described in any one of claims 1 to 6.
10. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the steps of the image recognition method based on unmanned aerial vehicles as described in any one of claims 1 to 6.