Defect image incremental learning method and system under unmanned aerial vehicle cloud edge cooperation
Through the defect image incremental learning method of drone cloud-edge collaboration, combined with edge computing and cloud computing, the problem of low defect detection efficiency and accuracy in drone inspection is solved, and efficient and accurate defect identification and model update are achieved.
Patent Information
- Application Number
- CN202510686716.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-05-27
AI Technical Summary
In the existing drone inspection, the artificial intelligence platform model cannot self-learning and iteratively improve defect recognition capabilities, and cloud computing cannot meet real-time and security requirements, resulting in low defect detection efficiency and accuracy.
The defect image incremental learning method of drone cloud-edge collaboration is adopted to realize real-time defect detection through edge computing nodes, and the computing power in the cloud is used to perform incremental updates of models to improve the adaptability and accuracy of models.
While ensuring real-time detection, it significantly improves the accuracy of defect recognition and system adaptability, simplifies operations, and is suitable for applications in more scenarios.
Smart Images

Figure CN120219923A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of defect image incremental learning and annotation, and particularly to a defect image incremental learning method and system under the cooperation of an unmanned aerial vehicle (UAV) between the cloud and the edge. Background Art
[0002] With the development of UAV technology, it has become increasingly common to use UAVs to replace humans for power transmission line inspection. Currently, during the inspection of power transmission UAVs, the mainstream method is to send the pictures and video data captured by the equipment to an artificial intelligence platform and call the artificial intelligence platform model for power transmission fault detection and defect identification. However, this method has the following problems: 1. The artificial intelligence platform model is a static model and does not currently support the self-learning iteration improvement of the defect identification ability. 2. The cloud computing and channel network pressure are too high. Currently, the cloud computing in the artificial intelligence platform cannot meet the application services with too high requirements for real-time performance and security, such as fire detection in defect detection, which requires a relatively high real-time performance, and currently the cloud artificial intelligence platform cannot meet the requirements.
[0003] In the task of power transmission line defect identification, accurate and efficient object detection annotation is a key step. The traditional manual annotation method is not only time-consuming and laborious, but also prone to human errors. At the same time, with the expansion of the scale of power transmission lines and the increase in data volume, it has become an urgent need to build a distributed cloud-edge collaborative identification platform.
[0004] Therefore, designing an efficient data annotation and model update mechanism for timely and accurate detection of power transmission lines and defect troubleshooting is an important means and an urgent task to ensure the safe and stable operation of the power grid. Summary of the Invention
[0005] In view of this, it is necessary to provide a defect image incremental learning method under the cooperation of an unmanned aerial vehicle between the cloud and the edge. Real-time defect detection is achieved through an edge computing node, and the powerful computing power of the cloud is used for model incremental update. While ensuring the real-time performance of detection, it can effectively adapt to newly emerging defect types, significantly improve the accuracy of defect identification and the adaptability of the system, is easy to promote and apply, is relatively simple to operate, and is suitable for popularization and application in more scenarios.
[0006] The present application provides a method and system for incremental learning of defective images under the cooperation of drones between the cloud and the edge. A distributed cloud-edge collaborative power transmission line defect recognition platform is constructed. An inference and semi-automatic annotation system is arranged at the edge side, and a multi-model incremental training system is constructed on the cloud through digital twin technology. The central cloud and the edge nodes cooperate with each other, and the annotation efficiency and model accuracy can be improved through incremental learning. At the same time, the samples are assisted in annotation by humans or other systems, and then the model is retrained incrementally. If the model after incremental training has a significant improvement in accuracy compared with the original model, it can be deployed to the edge nodes to update the original model. Through the crowdsourcing method, a mechanism for pipeline tasks is realized, and the semi-automatic annotation task is divided into multiple small pipeline tasks. The pipeline operation not only makes the annotation work more efficiently executed, but also combines the characteristics of different models to make the obtained annotation results more accurate.
[0007] In the first aspect, an embodiment of the present application provides a method for incremental learning of defective images under the cooperation of drones between the cloud and the edge. The method S1: Construct a distributed cloud-edge collaborative power transmission line defect recognition platform. Arrange an inference and semi-automatic annotation system at the edge side, and construct a multi-model incremental training system on the cloud through digital twin technology; S2: Use the low-power wide-area Internet of Things technology through drones to realize the real-time acquisition of line temperature, sag, and vibration state quantities, and upload the acquired data to the edge side; S3: Load multiple pre-trained models required by the inference and semi-automatic annotation system at the edge side, distribute the uploaded acquired data in the first crowdsourcing method, perform pre-annotation tasks, and summarize the samples with low confidence in the automatic recognition inference results, and package and send them to the central cloud; S4: Assist in annotating samples by humans or Label Studio software in the central cloud, and then retrain multiple incremental training models constructed through digital twin technology; S5: Calculate the accuracy of the model after incremental training respectively, and compare it with the accuracy of the original model; S6: If the comparison result is higher than the threshold, deploy the trained model to each edge-side node to update the original model.
[0008] Optionally, in an implementation manner of the first aspect of the present invention, in step S1, constructing a distributed cloud-edge collaborative power transmission line defect recognition platform, arranging an inference and semi-automatic annotation system at the edge side, and constructing a multi-model incremental training system on the cloud through digital twin technology includes: The distributed cloud-edge collaborative power transmission line defect identification platform includes: a central cloud, an edge device, and a drone. Among them, the edge device includes on-site devices, an edge gateway, and an edge cloud. A multi-level distributed management architecture is adopted to realize the interconnection between edge nodes. Based on converting the private communication protocols adopted by various on-site devices into a standardized OPC UA protocol, the computing tasks are distributed to multiple nodes for parallel processing, realizing the access, data collection, transmission, and processing of heterogeneous devices. The inference and semi-automatic annotation system includes a data input module, a model annotation module, an RTMDet inference service module, and an annotation output module; the data input module is used to receive image data to be annotated; the model annotation module uses different pre-trained models to perform preliminary annotation on the images; the RTMDet inference service module optimizes and adjusts the preliminary annotation results; the annotation output module saves and outputs the final annotation results. The multi-model incremental training system includes a digital twin layer, a model management layer, and an incremental training engine; the digital twin layer is used to mirror the model physical structure according to the models of the model annotation module to obtain multiple mirrored models; the model management layer is used to perform version control, performance monitoring, and model dependency management on the multiple mirrored models; the incremental training engine is used to adaptively adjust the learning rate through an online learning algorithm library and a distributed training framework.
[0009] Optionally, in an implementation manner of the first aspect of the present invention, S3: Install the inference and semi-automatic annotation system on the edge side to provide multiple required pre-trained models, and use the first crowdsourcing method to allocate the uploaded acquisition data, perform pre-annotation tasks, and summarize the samples with low confidence in the automatic recognition inference results, and package and send them to the central cloud, including: The multiple pre-trained models at least include: a CNN model, an improved double Unet model, and a Logits distillation model; the first crowdsourcing method is a pipeline task implementation mechanism. Among them, the semi-automatic annotation task is divided into multiple pipeline small tasks; according to the number of the divided small tasks, the corresponding number of pre-trained models are selected from the multiple pre-trained models for permutation and combination to obtain multiple candidate model combinations, and the pipeline tasks are respectively executed. Select the optimal model combination according to the confidence of the automatic recognition inference results, and summarize the samples with low confidence in the recognition results corresponding to the optimal model combination, and package and send them to the central cloud.
[0010] Optionally, in an implementation of the first aspect of the present invention, the method of selecting the corresponding number of pre-trained models from the multiple pre-trained models according to the number of segmented small tasks, performing permutations and combinations to obtain multiple candidate model combinations, and respectively executing the pipeline tasks; selecting the optimal model combination according to the confidence of the automatically recognized inference result, and summarizing the samples with low confidence in the recognition result corresponding to the optimal model combination and packing and sending them to the central cloud includes: The number of the segmented small tasks is three; Respectively select the first model in the combination from the multiple candidate model combinations to perform the task of annotating the type of the target object, the second model to perform the task of annotating the number and position of the target object, and the third model to perform the task of annotating the contour boundary of the target object; Select the optimal candidate model combination according to the weighted sum of the accuracies of the three execution results as the confidence; Take the optimal candidate model combination as the target model combination, establish the dependency relationship between the models in the target model combination, and use the first model in the target model combination again to perform the task of annotating the type of the target object, the second model to perform the task of annotating the number and position of the target object, and the third model to perform the task of annotating the contour boundary of the target object according to the execution result of annotating the type of the target object combined with the execution result of annotating the number and position of the target object, and summarize the samples with low confidence in the recognition result corresponding to the optimal model combination and pack and send them to the central cloud.
[0011] Optionally, in an implementation of the first aspect of the present invention, the step S4: assisting in annotating samples manually or by LabelStudio software in the central cloud, and then retraining multiple incremental training models constructed by digital twin technology includes: Collect an image data set containing transmission line defects and divide it into a training set and a test set; Preprocess the images, resize them according to different pre-trained models, and perform normalization operations to meet the input requirements of different models; at the same time, perform preliminary screening and compression on the collected images at the edge end to reduce the data transmission volume; Select different pre-trained models and use the interface provided by MMDETECTION to perform preliminary annotation on the images, including: Load the pre-trained model: use the configuration file and pre-trained weights of MMDETECTION to load the model; input the preprocessed images into the model for inference to obtain preliminary annotation boxes; manually view the inference results, drag the boxes manually, correct the positions of the boxes to obtain the corrected annotation results; save the annotation results; Retrain multiple incremental training models constructed by digital twin technology using the annotated data.
[0012] Optionally, in an implementation of the first aspect of the present invention, the improved dual Unet model structure is an asymmetric Unet model, including UNet1 and UNet2. Among them, low-level and high-level semantic features are respectively extracted through the encoders of UNet1 and UNet2, and low-level and high-level semantic features are extracted in the decoding layer for interaction and fusion, and a prediction result is output. The loss function is a cross-entropy loss function and the Dice loss function to form a combined loss function. The formula is: ; wherein, represents the true distribution of the sample , represents the predicted distribution of the sample , represents the total number of samples; ;
[0013] wherein, respectively represent the probability value of the class prediction and the true label value, is the number of classes, is the total number of classes, represents the class weight; ; The Logits distillation model structure is composed of three parallel network models, namely a standard teacher network model, a student network model, and an adversarial teacher network model. Among them, the standard teacher network model uses a UNet network structure based on the attention mechanism as the backbone network, and embeds the attention mechanism into the decoding part of the UNet network. The student network model uses a lightweight UNet network structure. The adversarial teacher network model uses a U-CliqueNet model structure, replaces the convolutional layer of the UNet with a Clique Block, and the loss function is composed of a hard loss and a soft loss. Among them, the hard loss function is: ; wherein, represents the true distribution of the sample , represents the predicted distribution of the sample , represents the total number of samples; The soft loss function is : ; wherein, represents the temperature parameter, respectively represent the probability distributions after high-temperature softening of the standard teacher network model, the student network model, and the adversarial teacher network model, is the total number of categories, is the divergence; The total loss function is : ; Among them, represents the hyperparameter for adjusting the weights of the hard loss and the soft loss, .
[0014] Optionally, in an implementation manner of the first aspect of the present invention, the method of calculating the accuracy of the model after incremental training respectively and comparing it with the accuracy of the original model includes: Evaluating the accuracy of each model through the mean intersection over union, and the formula is: ; Among them, represents the total number of categories, represents that the true category is and the predicted category is also the number of samples, represents that the true category is and the predicted category is the number of samples.
[0015] In a second aspect, an embodiment of the present application provides a defect image incremental learning system under drone cloud-edge collaboration, which is applied to the defect image incremental learning method under drone cloud-edge collaboration as described in the first aspect, and is characterized by including: Cloud-edge collaboration platform construction module: Construct a distributed cloud-edge collaboration power transmission line defect recognition platform, arrange an inference and semi-automatic annotation system at the edge side, and construct a multi-model incremental training system through digital twin technology on the cloud; Data acquisition module: Use the low-power wide-area Internet of Things technology through drones to realize the real-time acquisition of line temperature, sag, and vibration status quantities, and upload the acquired data to the edge side; First annotation module: Carry multiple pre-trained models required by the inference and semi-automatic annotation system on the edge side, allocate the uploaded acquired data in the first crowdsourcing manner, execute the pre-annotation task, and summarize the samples with low confidence in the automatic recognition inference results and send them to the central cloud in a package; Second annotation module: Manually or assisted by Label Studio software to annotate samples in the central cloud, and then re-train multiple incremental training models constructed through digital twin technology; Accuracy calculation module: Calculate the accuracy of the model after incremental training respectively, and compare it with the accuracy of the original model; Model update module: If the comparison result is higher than the threshold, deploy the trained model to each edge node to update the original model.
[0016] In a third aspect, an embodiment of the present application provides an electronic device, which is characterized by including: A processor; A memory for storing instructions executable by the processor; Wherein, when the processor is configured to execute the instructions, it implements the defect image incremental learning method under the drone cloud-edge collaboration as described in the first aspect.
[0017] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which is characterized in that the computer-readable storage medium stores a program, and the program instructs the device to execute the defect image incremental learning method under the drone cloud-edge collaboration as described in the first aspect. The present application provides a defect image incremental learning method and system under drone cloud-edge collaboration, constructs a distributed cloud-edge collaborative transmission line defect identification platform, arranges an inference and semi-automatic annotation system on the edge side, and constructs a multi-model incremental training system through digital twin technology on the cloud side; uses drones to adopt low-power wide-area Internet of Things technology to realize real-time collection of line temperature, sag, and vibration state quantities, and upload the collected data to the edge side; on the edge side, an inference and semi-automatic annotation system is carried to provide multiple pre-trained models required, and the uploaded collected data is distributed in the first crowdsourcing manner, pre-annotation tasks are executed, and samples with low confidence in the automatic recognition inference results are summarized and packaged and sent to the central cloud; on the central cloud, samples are assisted in annotation by humans or Label Studio software, and then multiple incremental training models constructed through digital twin technology are re-trained; calculate the accuracy of the models after incremental training respectively, and compare it with the accuracy of the original model; if the comparison result is higher than the threshold, deploy the trained models to each edge node to update the original model.
[0018] Beneficial effects:
[0019] (1) The central cloud and the edge nodes cooperate with each other, and the accuracy of the model can be improved through incremental learning. After the model is deployed on the edge node for inference operation, samples with low confidence in the automatic recognition inference results are sent to the central cloud, and the samples are assisted in annotation by humans or other systems, and then the model is re-incrementally trained. If the model after incremental training has a significant improvement in accuracy compared with the original model, it can be deployed to the edge node to update the original model.
[0020] (2) Through the edge-side model inference and semi-automatic annotation technology and equipment, the annotation work is made more efficient, and the samples are assisted in annotation by humans or other systems, making the annotation work more accurate.
[0021] (3) Implement a mechanism for pipeline tasks through crowdsourcing, split the semi-automated annotation tasks into multiple small pipeline tasks, and execute the pipeline tasks separately through a combination of candidate models. The pipeline operation not only enables the annotation work to be executed more efficiently, but also combines the characteristics of different models to make the obtained annotation results more accurate.
[0022] (4) It is easy to promote and apply, and the operation is relatively simple, making it suitable for promotion and application in more application scenarios. Brief Description of the Drawings
[0023] Figure 1 It is a schematic flow chart of the method for incremental learning of defective images under the cooperation of drones between cloud and edge provided by an embodiment of the present application.
[0024] Figure 2 It is a framework diagram of a transmission line defect recognition platform for distributed cloud-edge cooperation provided by an embodiment of the present application.
[0025] Figure 3 It is an architecture diagram of the multi-level distributed management of each edge node provided by an embodiment of the present application.
[0026] Figure 4 It is an architecture diagram of the double Unet model structure provided by an embodiment of the present application.
[0027] Figure 5 It is an architecture diagram of the Logits distillation model structure provided by an embodiment of the present application.
[0028] Figure 6 It is a schematic diagram of the system module for incremental learning of defective images under the cooperation of drones between cloud and edge provided by an embodiment of the present application.
[0029] Figure 7 It is a schematic diagram of an electronic device provided by an embodiment of the present application. Detailed Embodiments
[0030] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application.
[0031] It should be noted that "at least one" in the embodiments of the present application refers to one or more, and multiple refers to two or more. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used in the description of this application are only for the purpose of describing specific embodiments and are not intended to limit this application.
[0032] It should be noted that, in the embodiments of the present application, words such as "first" and "second" are only used for the purpose of distinguishing descriptions, and cannot be understood as indicating or implying relative importance, nor can they be understood as indicating or implying order. Features defined as "first" and "second" may explicitly or implicitly include one or more of the features. In the description of the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design described as "exemplary" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a concrete way.
[0033] Based on the implementations in this application, all other implementations obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0034] Embodiment 1.
[0035] The present application provides a method and system for incremental learning of defective images under the collaboration of drones and cloud-edge, constructs a distributed cloud-edge collaborative power transmission line defect recognition platform, deploys reasoning and semi-automatic annotation systems on the edge, and constructs a multi-model incremental training system on the cloud through digital twin technology; the central cloud and edge nodes cooperate with each other, and the annotation efficiency and model accuracy can be improved through incremental learning. At the same time, samples are annotated manually or with the assistance of other systems, and then the model is incrementally trained again. If the incrementally trained model has a significant improvement in accuracy over the original model, it can be deployed to the edge node to update the original model. The mechanism for realizing assembly line tasks is implemented through crowdsourcing, and the semi-automatic annotation tasks are divided into multiple small assembly line tasks. The assembly line operation not only makes the annotation work more efficient, but also integrates the characteristics of different models to make the obtained annotation results more accurate.
[0036] Figure 1 A flowchart of a method for incremental learning of defect images under drone-cloud-edge collaboration is provided for one embodiment of the present application.
[0037] like Figure 1 As shown, a defect image incremental learning method under UAV cloud-edge collaboration includes: S1: Build a distributed cloud-edge collaborative transmission line defect identification platform, deploy reasoning and semi-automatic annotation systems on the edge, and build a multi-model incremental training system on the cloud through digital twin technology.
[0038] It can be understood that in this embodiment, in step S1, a distributed cloud-edge collaborative transmission line defect identification platform is constructed, an inference and semi-automatic annotation system is deployed on the edge, and a multi-model incremental training system is constructed on the cloud through digital twin technology, including: Figure 2 This is a framework diagram of a distributed cloud-edge collaborative power transmission line defect identification platform provided by an embodiment of the present application. As Figure 2 shown, the distributed cloud-edge collaborative power transmission line defect identification platform includes: a central cloud, an edge terminal, and an unmanned aerial vehicle (UAV). Among them, the edge terminal includes on-site devices, an edge gateway, and an edge cloud. Figure 3 This is a multi-level distributed management architecture diagram of each edge node provided by an embodiment of the present application. As Figure 3 shown, a multi-level distributed management architecture is adopted to realize the interconnection between each edge node. Based on converting the private communication protocols adopted by various on-site devices into a standardized OPC UA protocol, the computing tasks are dispersed to multiple nodes for parallel processing, realizing the access, data collection, transmission, and processing of heterogeneous devices; The inference and semi-automatic annotation system includes a data input module, a model annotation module, an RTMDet inference service module, and an annotation output module; the data input module is used to receive image data to be annotated; the model annotation module preliminarily annotates the image using different pre-trained models; the RTMDet inference service module optimizes and adjusts the preliminary annotation results; the annotation output module saves and outputs the final annotation results; The multi-model incremental training system includes a digital twin layer, a model management layer, and an incremental training engine; the digital twin layer is used to mirror the model physical structure according to the model of the model annotation module to obtain multiple mirror models; the model management layer is used to perform version control, performance monitoring, and model dependency management on the multiple mirror models; the incremental training engine is used to adaptively adjust the learning rate through an online learning algorithm library and a distributed training framework.
[0039] Specifically, building a distributed cloud-edge collaborative global management platform needs to be based on resource management, centered on intelligent scheduling, and extended by an open ecosystem. Through cloud-native technology, agile deployment of applications is realized, and collaborative strategies are optimized in combination with industry scenarios, ultimately achieving the goal of "real-time edge response and global cloud optimization".
[0040] S2: The UAV uses low-power wide-area Internet of Things technology to realize the real-time collection of line temperature, sag, and vibration status quantities, and uploads the collected data to the edge side.
[0041] It is understandable that in this embodiment, the drone can adopt LPWAN technology based on the LoRa protocol (such as LoRaWAN or LoPo-IoT), which has long-distance transmission (up to 15 kilometers in suburbs) and low power consumption, and is suitable for large-scale power transmission line inspection. LoRa modules (such as RN2483) perform well in complex terrain and can adapt to the dynamic movement of drones. At the same time, the drone is equipped with temperature, sag, and vibration sensors (such as the temperature detection unit in) to collect data in real time through LPWAN. For example, the LoRa module can support multi-node concurrent transmission to ensure high-frequency data upload. At the same time, as a mobile gateway, the drone can optimize coverage blind areas and improve collection efficiency.
[0042] The data collected by the drone is uploaded to the edge computing node (such as deployed in a substation or mobile edge server) through LPWAN. The edge side is equipped with a lightweight AI model, which can complete fault feature extraction, sag abnormality judgment (such as through temperature-sag relationship curve analysis) and vibration status evaluation in real time.
[0043] Cloud functions: deploy multi-model incremental training systems to support dynamic model updates, algorithm library management, and large-scale data processing. Edge functions: lightweight inference systems (such as TensorFlow Lite) to achieve real-time defect detection, and semi-automatic annotation systems to assist manual correction of low-confidence samples through interactive interfaces. Collaboration mechanism: MQTT / HTTP protocols are used to achieve cloud-edge communication, edge nodes are responsible for data filtering, and the cloud focuses on model optimization.
[0044] Specifically, the distributed cloud-edge collaborative transmission line defect identification platform includes: a central cloud, an edge terminal, and a drone, wherein the edge terminal includes field equipment, an edge gateway, and an edge cloud, which is used to realize the access and data collection of heterogeneous devices based on the conversion of the private communication protocols adopted by various field equipment and devices into standardized OPC UA protocols; through a lightweight computing framework and related algorithms, the data is cleaned and format converted for pre-processing operations on the edge side, and redundant data is effectively eliminated from the massive data generated on the edge side, thereby reducing the amount of data uploaded to the cloud and reducing the cloud platform load and transmission bandwidth pressure.
[0045] S3: The edge side is equipped with an inference and semi-automatic annotation system to provide the required multiple pre-trained models, distribute the uploaded collected data in a crowdsourcing manner, perform pre-annotation tasks, and summarize samples with low confidence in automatic identification reasoning results, package them and send them to the central cloud.
[0046] It can be understood that, in this embodiment, S3: Load multiple pre-trained models required by the inference and semi-automatic annotation system on the edge side, distribute the uploaded collected data in the first crowdsourcing manner, perform pre-annotation tasks, and summarize and package the samples with low confidence in the automatic recognition inference results and send them to the central cloud, including: The multiple pre-trained models at least include: three models, namely, a CNN model, an improved double Unet model, and a Logits distillation model; the first crowdsourcing manner is a pipeline task implementation mechanism; Among them, the semi-automatic annotation task is divided into multiple pipeline sub-tasks; according to the number of the divided sub-tasks, the corresponding number of pre-trained models are selected from the multiple pre-trained models for permutation and combination to obtain multiple candidate model combinations, and the pipeline tasks are respectively executed; An optimal model combination is selected according to the confidence of the automatic recognition inference result, and the samples with low confidence are screened from the recognition results corresponding to the optimal model combination, summarized and packaged, and sent to the central cloud.
[0047] Specifically, the corresponding number of pre-trained models are selected from the multiple pre-trained models according to the number of the divided sub-tasks for permutation and combination to obtain multiple candidate model combinations, and the pipeline tasks are respectively executed; an optimal model combination is selected according to the confidence of the automatic recognition inference result, and the samples with low confidence are screened from the recognition results corresponding to the optimal model combination, summarized and packaged, and sent to the central cloud, including: The number of the divided sub-tasks is three; From the multiple candidate model combinations, the first model in the combination is respectively selected to perform the task of annotating the type of the target object, the second model performs the task of annotating the quantity and position of the target object, and the third model performs the task of annotating the contour boundary of the target object; An optimal candidate model combination is selected according to the weighted sum of the accuracies of the three execution results as the confidence; The optimal candidate model combination is used as the target model combination, the dependency relationship between the models in the target model combination is established, the first model in the target model combination is used again to perform the task of annotating the type of the target object, the second model performs the task of annotating the quantity and position of the target object, and the third model performs the task of annotating the contour boundary of the target object according to the execution result of annotating the type of the target object combined with the execution result of annotating the quantity and position of the target object, and the samples with low confidence are screened from the recognition results corresponding to the optimal model combination, summarized and packaged, and sent to the central cloud.
[0048] Specifically, Figure 4 This is the structural architecture diagram of the double Unet model provided by an embodiment of the present application. As Figure 4As shown in the figure, the improved double Unet model structure is an asymmetric Unet model, including UNet1 and UNet2. Among them, low-level and high-level semantic features are respectively extracted through the encoders of UNet1 and UNet2, and low-level and high-level semantic features are extracted in the decoding layer for interaction and fusion, and the prediction result is output. The loss function is the cross-entropy loss function and the Dice loss function to form a combined loss function, and the formula is: ; where represents the true distribution of the sample , represents the predicted distribution of the sample , represents the total number of samples; ;
[0049] where respectively represent the probability value and the true label value of the class prediction, is the number of classes, is the total number of classes, represents the class weight; ; Figure 5 is the architecture diagram of the Logits distillation model structure provided by an embodiment of the present application. As Figure 5 shown, the Logits distillation model structure is composed of three parallel network models, namely a standard teacher network model, a student network model, and an adversarial teacher network model. Among them, the standard teacher network model uses a UNet network structure based on the attention mechanism as the backbone network, and embeds the attention mechanism into the decoding part of the UNet network. The student network model uses a lightweight UNet network structure. The adversarial teacher network model uses a U-CliqueNet model structure, replaces the convolutional layer of the UNet with a Clique Block, and the loss function is composed of a hard loss and a soft loss. Among them, the hard loss function is: ; where represents the true distribution of the sample , represents the predicted distribution of the sample , represents the total number of samples; The soft loss function is : ; Among them, represents the temperature parameter, respectively represent the probability distributions after high-temperature softening of the standard teacher network model, the student network model, and the adversarial teacher network model, is the total number of categories, is the divergence; The total loss function is : ; Among them, represents the hyperparameter for adjusting the weights of the hard loss and the soft loss, .
[0050] Specifically, after using the CNN model to extract features from the input image, the class Logits are output through the fully connected layer, and the class probability distribution is generated by combining with Softmax. The results with high confidence can be directly used to label the target type; the dual Unet model extracts multi-scale features through the encoder, and the decoder generates the position heat map and quantity prediction of the target; according to the labeled target type and the generated position heat map and quantity prediction results of the target by the distillation model, the output of the Logits segmentation task of the teacher model is used as the supervision signal, and the sorting information of the boundary pixels is optimized by using the loss to refine the contour annotation.
[0051] S4: Manually or assisted by Label Studio software to label samples in the central cloud, and then retrain multiple incremental training models constructed by digital twin technology.
[0052] Specifically, MMDETECTION is an open-source object detection toolbox based on PyTorch, which provides rich pre-trained models and efficient training and inference interfaces; LABEL-STUDIO is a powerful annotation tool that supports multiple annotation tasks. Combining these two tools can achieve semi-automated object detection annotation, and then integrating it into the distributed cloud-edge collaborative architecture can effectively improve the annotation efficiency and accuracy of transmission line defect recognition.
[0053] It can be understood that in this embodiment, the S4: Manually or assisted by Label Studio software to label samples in the central cloud, and then retrain multiple incremental training models constructed by digital twin technology, includes: Collect an image dataset containing transmission line defects and divide it into a training set and a test set; Preprocess the images, resize them according to different pre-trained models, and perform normalization operations to meet the input requirements of different models; at the same time, preliminarily screen and compress the collected images at the edge side to reduce the data transmission volume; Select different pre-trained models and use the interfaces provided by MMDETECTION to perform preliminary annotation on images, including: Load the pre-trained model: Use the configuration file and pre-trained weights of MMDETECTION to load the model; input the pre-processed image into the model for inference to obtain the preliminary annotation boxes; manually view the inference results, drag the boxes, and correct the positions of the boxes to obtain the corrected annotation results; save the annotation results; Use the annotated data to retrain multiple incremental training models constructed through digital twin technology.
[0054] S5: Calculate the accuracy of the incrementally trained models respectively and compare it with the accuracy of the original model.
[0055] It can be understood that in this embodiment, calculating the accuracy of the incrementally trained models respectively and comparing it with the accuracy of the original model includes: Evaluate the accuracy of each model through the mean intersection over union. The formula is: ; where, represents the total number of categories, represents the number of samples where the true category is and the predicted category is also ; represents the number of samples where the true category is and the predicted category is ;
[0056] S6: If the comparison result is higher than the threshold, deploy the trained model to each edge node to update the original model.
[0057] Specifically, in the scenario of incremental learning, when the accuracy of the incrementally trained model is significantly improved, it will be deployed to the edge node to update the original model. A specific threshold can be set, and the accuracy is compared with the threshold. The comparison of the accuracy with the threshold determines whether to participate in the update of the global model.
[0058] It can be understood that in this embodiment, the samples for AI model training mainly come from data collection on the edge side. The central cloud and the edge nodes cooperate with each other, and the accuracy of the model can be improved through incremental learning. After deploying the model to the edge node for inference, samples with low confidence in the inference results are automatically identified and sent to the central cloud. The samples are assisted in annotation by humans or other systems, and then the model is retrained incrementally. If the incrementally trained model has a significant improvement in accuracy compared to the original model, it can be deployed to the edge node to update the original model.
[0059] Embodiment 2.
[0060] As shown Figure 6 The present application provides a defect image incremental learning system under unmanned aerial vehicle (UAV) cloud-edge collaboration, which is applied to the defect image incremental learning method under UAV cloud-edge collaboration described in Embodiment 1, and includes: a cloud-edge collaboration platform construction module 11, a data acquisition module 12, a first annotation module 13, a second annotation module 14, an accuracy calculation module 15, and a model update module 16.
[0061] It can be understood that, in this embodiment, the cloud-edge collaboration platform construction module 11 is used to construct a distributed cloud-edge collaboration power transmission line defect identification platform, arrange an inference and semi-automatic annotation system at the edge side, and construct a multi-model incremental training system through digital twin technology on the cloud side.
[0062] It can be understood that, in this embodiment, the data acquisition module 12 is used to use the low-power wide-area Internet of Things technology through a UAV to realize the real-time acquisition of line temperature, sag, and vibration state quantities, and upload the acquired data to the edge side.
[0063] It can be understood that, in this embodiment, the first annotation module 13 is used to carry multiple pre-trained models required by the inference and semi-automatic annotation system on the edge side, allocate the uploaded acquired data in the first crowdsourcing manner, execute the pre-annotation task, and summarize the samples with low confidence in the automatic recognition inference results, and package and send them to the central cloud.
[0064] It can be understood that, in this embodiment, the second annotation module 14 is used to manually or assist in annotating samples by LabelStudio software in the central cloud, and then re-train multiple incremental training models constructed through digital twin technology.
[0065] It can be understood that, in this embodiment, the accuracy calculation module 15 is used to calculate the accuracy of the model after incremental training respectively, and compare it with the accuracy of the original model.
[0066] It can be understood that, in this embodiment, the model update module 16 is used to deploy the trained model to each edge side node to update the original model if the comparison result is higher than the threshold.
[0067] This application provides a method and system for incremental learning of defective images under the cloud-edge collaboration of unmanned aerial vehicles (UAVs). A distributed cloud-edge collaborative power transmission line defect recognition platform is constructed. An inference and semi-automatic annotation system is arranged at the edge side, and a multi-model incremental training system is constructed through digital twin technology on the cloud side. The UAV uses low-power wide-area Internet of Things technology to realize real-time collection of line temperature, sag, and vibration state quantities, and upload the collected data to the edge side. An inference and semi-automatic annotation system is deployed on the edge side to provide multiple pre-trained models required. The uploaded collected data is distributed in the first crowdsourcing method to perform pre-annotation tasks, and samples with low confidence in the automatic recognition inference results are summarized and packaged and sent to the central cloud. On the central cloud, samples are assisted in annotation by humans or Label Studio software, and then multiple incremental training models constructed through digital twin technology are re-trained. The accuracy of the models after incremental training is calculated respectively and compared with the accuracy of the original models. If the comparison result is higher than the threshold, the trained models are deployed to each edge-side node to update the original models.
[0068] The problems in the UAV image recognition process are effectively improved mainly by completing the following two parts of work. One is to realize the incremental training of the power transmission line defect recognition model under cloud-edge collaboration; the other is to study the edge-side model inference and semi-automatic annotation technology and equipment.
[0069] Through the cooperation between the central cloud and the edge nodes, the model accuracy can be improved through incremental learning. After the model is deployed and runs inference at the edge node, samples with low confidence in the automatic recognition inference results are sent to the central cloud, and the samples are assisted in annotation by humans or other systems, and then the model is re-incrementally trained. If the model after incremental training has a significant improvement in accuracy compared with the original model, it can be deployed to the edge node to update the original model.
[0070] Through the edge-side model inference and semi-automatic annotation technology and equipment, the annotation work is made more efficient, and the samples are assisted in annotation by humans or other systems, making the annotation work more accurate.
[0071] Through the crowdsourcing method as the implementation mechanism of the pipeline task, the semi-automatic annotation task is divided into multiple pipeline small tasks, and the pipeline tasks are respectively executed through the combination of candidate models. The pipeline operation not only makes the annotation work more efficiently executed, but also combines the characteristics of different models to make the obtained annotation results more accurate.
[0072] It is easy to promote and apply, and the operation is relatively simple, suitable for promotion and application in more application scenarios.
[0073] Aiming at the challenges faced by defect image recognition in UAV inspection, an incremental learning method for defect images based on a cloud-edge collaborative architecture is proposed. This method realizes real-time defect detection through edge computing nodes, uses the powerful computing power of the cloud for model incremental update, and designs an efficient data transmission and model update mechanism. Experimental results show that while ensuring the real-time detection, this method can effectively adapt to newly emerging defect types, significantly improving the accuracy of defect recognition and the adaptability of the system.
[0074] Figure 7 This is an electronic device provided by an embodiment of the present application. As Figure 7 shown, the electronic device at least includes the following parts: a processor 101, a memory 100, a communication interface 103, and a bus 102.
[0075] In the embodiment of the present application, the memory 100 is used to store executable instructions of the processor 101, and the processor 101 is configured to implement a device module for incremental learning of defect images under UAV cloud-edge collaboration as Figure 6 shown.
[0076] In the embodiment of the present application, a computer-readable storage medium includes instructions that direct the device to execute the method of the first aspect. For example, the instructions direct the device to execute the method shown in the process steps of Figure 1 .
[0077] The program operating in the electronic device involved in an embodiment of the present application can be a program that controls a central processing unit (CPU) etc. to implement the functions of the above-mentioned embodiments involved in a solution of the present invention (a program that makes a computer function). Then, the information processed by these devices is temporarily stored in a random access memory (RAM) during its processing, and then stored in various ROMs such as a read-only memory (FlashROM), a hard disk drive (HDD), etc. It is read out, corrected, and written by the CPU as needed.
[0078] It should be noted that a part of the electronic device of the above-mentioned embodiment can also be implemented by a computer. In this case, the program for implementing the control function can be recorded on a computer-readable recording medium, and the program recorded on the recording medium is read into the computer and executed to implement it.
[0079] Note that the "computer" mentioned here refers to a computer built into an electronic device, which is a computer using hardware including an OS, peripheral devices, etc. In addition, the "computer-readable recording medium" refers to removable media such as floppy disks, magneto-optical disks, ROMs, CD-ROMs, and storage devices such as hard disks built into a computer.
[0080] Moreover, the "computer-readable recording medium" may include: a medium that stores a program dynamically for a short period of time, such as a communication line in the case of transmitting a program via a network such as the Internet or a communication line such as a telephone line; and a medium that stores a program for a fixed period of time, such as a volatile memory inside a computer serving as a server or a client in this case. In addition, the above program may be a program for implementing a part of the above functions, and may also be a program that can implement the above functions by combining with a program already recorded in a computer.
[0081] In addition, the electronic device in the above embodiment can also be implemented as an aggregate (device group) composed of multiple devices. Each device constituting the device group may have some or all of the functions or function blocks of the electronic device in the above embodiment. As the device group, it is sufficient to have all the functions or function blocks of the electronic device.
[0082] Those of ordinary skill in the art of this technology should recognize that the above embodiments are only used to illustrate the present application, rather than to limit the present application. As long as it is within the scope of the essential spirit of the present application, appropriate changes and variations made to the above embodiments fall within the scope of protection required by the present application.
Claims
1. A method for incremental learning of defect images under the cloud-edge collaboration of unmanned aerial vehicles, characterized in that, The method includes: S1: Construct a distributed cloud-edge collaborative power transmission line defect identification platform. Arrange an inference and semi-automatic annotation system at the edge side, and construct a multi-model incremental training system through digital twin technology in the cloud; S2: Use a drone to adopt low-power wide-area Internet of Things technology to realize real-time collection of line temperature, sag, and vibration status quantities, and upload the collected data to the edge side; S3: Install multiple pre-trained models required by the inference and semi-automatic annotation system on the edge side. Distribute the uploaded collected data in the first crowdsourcing method, perform pre-annotation tasks, and summarize the samples with low confidence in the automatic recognition inference results, and package and send them to the central cloud; S4: Manually or assisted by Label Studio software to annotate samples in the central cloud, and then retrain multiple incremental training models constructed through digital twin technology; S5: Calculate the accuracy of the models after incremental training respectively, and compare it with the accuracy of the original models; S6: If the comparison result is higher than the threshold, deploy the trained models to each edge side node to update the original models.
2. The defect image incremental learning method under the cloud-edge collaboration of an unmanned aerial vehicle according to claim 1, wherein In the above S1, constructing a distributed cloud-edge collaborative power transmission line defect identification platform, arranging an inference and semi-automatic annotation system at the edge side, and constructing a multi-model incremental training system through digital twin technology in the cloud includes: The distributed cloud-edge collaborative power transmission line defect identification platform includes: a central cloud, an edge side, and a drone. Among them, the edge side includes on-site devices, an edge gateway, and an edge cloud. A multi-level distributed management architecture is adopted to realize the interconnection between edge nodes. Based on converting the private communication protocols adopted by various on-site devices into a standardized OPC UA protocol, the computing tasks are dispersed to multiple nodes for parallel processing to realize the access, data collection, transmission, and processing of heterogeneous devices; The inference and semi-automatic annotation system includes a data input module, a model annotation module, an RTMDet inference service module, and an annotation output module; the data input module is used to receive image data to be annotated; the model annotation module uses different pre-trained models to perform preliminary annotation on the images; the RTMDet inference service module optimizes and adjusts the preliminary annotation results; the annotation output module saves and outputs the final annotation results; The multi-model incremental training system includes a digital twin layer, a model management layer, and an incremental training engine; the digital twin layer is used to mirror the model physical structure according to the models of the model annotation module to obtain multiple mirror models; the model management layer is used to perform version control, performance monitoring, and model dependency management on the multiple mirror models; the incremental training engine is used to adaptively adjust the learning rate through an online learning algorithm library and a distributed training framework.
3. A method for incremental learning of defect images under the cloud-edge collaboration of an unmanned aerial vehicle, characterized in that, The above S3: Install multiple pre-trained models required by the inference and semi-automatic annotation system on the edge side. Distribute the uploaded collected data in the first crowdsourcing method, perform pre-annotation tasks, and summarize the samples with low confidence in the automatic recognition inference results, and package and send them to the central cloud, includes: The multiple pre-trained models at least include: three models, namely, a CNN model, an improved dual Unet model, and a Logits distillation model; the first crowdsourcing method is a pipeline task implementation mechanism. Among them, the semi-automated annotation task is segmented into multiple pipeline sub-tasks; according to the number of the segmented sub-tasks, the corresponding number of pre-trained models are selected from the multiple pre-trained models for permutation and combination to obtain multiple candidate model combinations, and the pipeline tasks are respectively executed. The optimal model combination is selected according to the confidence of the automatic recognition inference result, and the samples with low confidence in the recognition results corresponding to the optimal model combination are summarized, packaged, and sent to the central cloud.
4. A method for incremental learning of defective images under the cloud-edge collaboration of an unmanned aerial vehicle according to claim 3, characterized in that, According to the number of the segmented sub-tasks, the corresponding number of pre-trained models are selected from the multiple pre-trained models for permutation and combination to obtain multiple candidate model combinations, and the pipeline tasks are respectively executed. Selecting the optimal model combination according to the confidence of the automatic recognition inference result, and summarizing, packaging, and sending the samples with low confidence in the recognition results corresponding to the optimal model combination to the central cloud includes: The number of the segmented sub-tasks is three. From the multiple candidate model combinations, the first model in the combination is respectively selected to execute the task of annotating the type of the target object, the second model executes the task of annotating the quantity and position of the target object, and the third model executes the task of annotating the contour boundary of the target object. The optimal candidate model combination is selected according to the weighted sum of the accuracies of the three execution results as the confidence. The optimal candidate model combination is used as the target model combination, the dependency relationship between the models in the target model combination is established, and the first model in the target model combination is used again to execute the task of annotating the type of the target object, the second model executes the task of annotating the quantity and position of the target object, and the third model executes the task of annotating the contour boundary of the target object according to the execution result of annotating the type of the target object combined with the execution result of annotating the quantity and position of the target object, and the samples with low confidence in the recognition results corresponding to the optimal model combination are summarized, packaged, and sent to the central cloud.
5. The method for incremental learning of defective images under the cloud-edge collaboration of an unmanned aerial vehicle according to claim 4, wherein, Step S4: In the central cloud, the samples are assisted in annotation by artificial or Label Studio software, and then multiple incremental training models constructed through digital twin technology are re-trained, including: Collecting an image data set containing transmission line defects and dividing it into a training set and a test set. Preprocessing the images, adjusting the size according to different pre-trained models, and performing normalization operations to meet the input requirements of different models; at the same time, the collected images are preliminarily screened and compressed at the edge side to reduce the data transmission volume. Selecting different pre-trained models and using the interface provided by MMDETECTION to perform preliminary annotation on the images, including: Loading the pre-trained model: loading the model using the configuration file and pre-trained weights of MMDETECTION; inputting the preprocessed images into the model for inference to obtain preliminary annotation boxes; manually viewing the inference results, dragging the boxes by hand, and correcting the positions of the boxes to obtain the corrected annotation results; saving the annotation results. Retrain multiple incremental training models constructed through digital twin technology using the labeled data.
6. A method for incremental learning of defect images under UAV cloud-edge collaboration according to claim 3, characterized in that The improved double Unet model structure is an asymmetric Unet model, including UNet1 and UNet2. Among them, low-level and high-level semantic features are extracted through the encoders of UNet1 and UNet2 respectively, and low-level and high-level semantic features are extracted in the decoding layer for interaction and fusion to output the prediction result. The loss function is the cross-entropy loss function and the Dice loss function to form a combined loss function, and the formula is: ; Among them, represents the true distribution of the sample , represents the predicted distribution of the sample , represents the total number of samples; ; Among them, respectively represent the probability value of class prediction and the true label value, is the number of classes, is the total number of classes, represents the class weight; The combined loss function is expressed as: ; The Logits distillation model structure consists of three parallel network models, namely, a standard teacher network model, a student network model, and an adversarial teacher network model. Among them, the standard teacher network model uses a UNet network structure based on the attention mechanism as the backbone network, and embeds the attention mechanism into the decoding part of the UNet network. The student network model uses a lightweight UNet network structure, and the adversarial teacher network model uses a U-CliqueNet model structure, replacing the convolutional layer of the UNet with a Clique Block. The loss function consists of a hard loss and a soft loss. Among them, the hard loss function is the cross-entropy loss function, which is: ; Among them, represents the true distribution of the sample , represents the predicted distribution of the sample , represents the total number of samples; The soft loss function is :[[]]END]] ; Among them, represents the temperature parameter, respectively represent the probability distributions after high-temperature softening of the standard teacher network model, the student network model, and the adversarial teacher network model, is the total number of categories, is the divergence; The total loss function is : ; Among them, represents a hyperparameter for adjusting the weights of the hard loss and the soft loss, .
7. A method for incremental learning of defective images under the cloud-edge collaboration of an unmanned aerial vehicle according to claim 5, characterized in that, The method of separately calculating the accuracy of the incrementally trained model and comparing it with the accuracy of the original model includes: Evaluating the accuracy of each model through the mean intersection over union, and the formula is: ; Among them, represents the total number of categories, represents that the true category is and the predicted category is also the number of samples, represents that the true category is and the predicted category is the number of samples.
8. A defect image incremental learning system under the cooperation of UAV cloud and edge, which is applied to the defect image incremental learning method under the cooperation of UAV cloud and edge according to any one of claims 1 to 7, and is characterized in that, Including: Cloud-edge collaboration platform construction module: Construct a distributed cloud-edge collaborative transmission line defect identification platform, arrange an inference and semi-automatic annotation system at the edge side, and construct a multi-model incremental training system through digital twin technology on the cloud side; Data acquisition module: Use the UAV to adopt the low-power wide-area Internet of Things technology to realize the real-time acquisition of line temperature, sag, and vibration state quantities, and upload the acquired data to the edge side; First annotation module: Load multiple pre-trained models required by the inference and semi-automatic annotation system at the edge side, distribute the uploaded acquired data in the first crowdsourcing manner, execute the pre-annotation task, and summarize the samples with low confidence in the automatic recognition inference results, and package and send them to the central cloud; Second annotation module: Manually or assisted by Label Studio software to annotate samples in the central cloud, and then retrain multiple incremental training models constructed through digital twin technology; Accuracy calculation module: Calculate the accuracy of the incrementally trained model separately and compare it with the accuracy of the original model; Model update module: If the comparison result is higher than the threshold, deploy the trained model to each edge-side node to update the original model.
9. An electronic device, characterized in that, Including: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to implement the method for incremental learning of defect images under UAV cloud-edge collaboration according to any one of claims 1 to 7 when executing the instructions.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program, and the program instructs the device to execute the method for incremental learning of defect images under UAV cloud-edge collaboration according to any one of claims 1 to 7.
Citation Information
Patent Citations
Power transmission line defect identification method and system based on incremental learning technology
CN114663751A
Power transmission line body defect detection method based on cloud edge cooperation
CN117975305A
Video transformer for deepfake detection with incremental learning
US20230401824A1
Cited By
Machine vision model training method and system based on end side computing power
CN120543948A
A machine vision model training method and system based on edge computing power
CN120543948B
AI-based low-altitude economic unmanned aerial vehicle data processing system and method thereof
CN121617097A
An ai-based low-altitude economy unmanned aerial vehicle data processing system and method thereof
CN121617097B