Grouping-DINO model acceleration method combined with multi-stage progressive multi-mode distillation mechanism
By adopting a multi-stage progressive multimodal distillation mechanism, the Grounding-DINO model is transferred to a lightweight student model, which solves the problem of insufficient detection accuracy in power line inspection and realizes efficient detection on low-computing-power devices.
Patent Information
- Application Number
- CN202511671123.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-14
- Publication Date
- 2025-12-23
AI Technical Summary
Existing power inspection and detection technologies struggle to balance lightweight design and accuracy. Traditional visual inspection models lack cross-modal fusion capabilities, and Transformer-based multimodal models cannot be directly deployed on computing-constrained devices. Furthermore, they are not optimized for the characteristics of power scenarios, resulting in insufficient detection accuracy.
A multi-stage progressive multimodal distillation mechanism is adopted to transfer the Grounding-DINO model to a lightweight student model. By constructing a dedicated power equipment dataset, designing a lightweight student model architecture, and gradually transferring knowledge during the three-stage distillation process, including visual feature alignment, cross-modal fusion, and task fine-tuning, the model is optimized to adapt to power inspection scenarios.
It significantly improves the detection accuracy of defects in power equipment and small targets, meets the low computing power requirements of power inspection equipment, achieves a synergistic improvement in lightweight design and accuracy, and adapts to the complex environment and technical terminology of power scenarios.
Smart Images

Figure CN121189376A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the technical field of power equipment inspection data processing methods, specifically a Grounding-DINO model acceleration method that combines a multi-stage progressive multimodal distillation mechanism. Background Technology
[0002] As power systems upgrade towards intelligence and automation, power line inspection, as a core component ensuring the safe and stable operation of the power grid, increasingly relies on power equipment target detection technologies. Power line inspection scenarios face complex environments such as transmission lines and substations, requiring not only the identification of large power equipment such as poles and insulators, but also the accurate detection of small target defects such as bolt corrosion and broken conductor strands. Furthermore, it necessitates compatibility with computing-constrained devices such as drone edge terminals to achieve real-time detection and data transmission. This places stringent demands on both the lightweight nature and accuracy of the detection models.
[0003] Current target detection technologies in the power line inspection field are mainly based on two types of methods: one is the traditional visual detection model, which achieves target recognition by manually designing feature extraction operators combined with classifiers. However, this type of method has poor adaptability to occlusion of power equipment and changes in lighting, and it is difficult to cope with the detection requirements guided by text commands. The other type is the multimodal detection model based on the Transformer architecture (such as Grounding-DINO), which has stronger scene adaptability and command response capability through cross-modal fusion of visual and text features. However, this type of model has a large number of parameters and high computational complexity, with extremely high hardware computing power and memory requirements, making it difficult to deploy on power line inspection edge devices with computing power typically not exceeding 20 TOPS and memory not exceeding 8GB. To adapt to the edge deployment requirements, some existing technologies adopt simple model pruning or single distillation strategies to achieve lightweighting, but they often only focus on reducing model parameters without optimizing for the characteristics of power scene.
[0004] The core flaws of existing methods make it difficult to meet the actual needs of power grid inspection accuracy: Traditional visual inspection models lack cross-modal fusion capabilities, cannot effectively utilize text instructions to guide detection, and have weak ability to identify small targets and occluded targets in power scenarios, resulting in high false negative and false positive rates. While Transformer-based multimodal models offer better performance, they cannot be directly deployed due to limitations in computing power. Simple lightweight modifications to existing models, without establishing a systematic knowledge transfer mechanism, significantly reduce detection accuracy while compressing parameters, especially in the detection of small target defects. Furthermore, existing methods have not been adapted and optimized for the hierarchical structure, occlusion characteristics, and technical terminology of power equipment, resulting in poor cross-modal alignment and further reducing the accuracy of inspection. This makes it difficult to meet the core requirement of timely and accurate identification of equipment defects in power grid inspection, increasing the risk of power grid safety hazards and the cost of manual re-inspection, thus requiring urgent solutions. Summary of the Invention
[0005] To address the technical problem of reduced inspection accuracy caused by the lack of adaptation and optimization for the layered structure, obstruction characteristics, and technical terminology of power equipment, this invention provides a Grounding-DINO model acceleration method that combines a multi-stage progressive multimodal distillation mechanism. This invention can effectively improve the accuracy of inspection result judgment.
[0006] To achieve the above objectives, the present invention provides the following technical solution: An accelerated method for the Grounding-DINO model, incorporating a multi-stage progressive multimodal distillation mechanism, includes the following steps: S1. Construct a detection dataset adapted to power equipment inspection scenarios; S2. Design a lightweight student model that meets the computational power limitations of processing the detection dataset; S3. Using Grounding-DINO as the teacher model, knowledge from the teacher model is transferred to the lightweight student model through a three-stage progressive multimodal distillation mechanism. S4. Optimize the lightweight student model trained using the detection dataset before deployment, and then deploy the optimized lightweight student model to the power inspection equipment.
[0007] As a further aspect of the present invention: the lightweight student model is based on YOLOv4 and is improved to be lightweight, and satisfies the following: the number of parameters does not exceed a set parameter threshold, the amount of computation does not exceed a set computation threshold, and the input resolution is a set resolution. The set parameter threshold, the set computation threshold, and the set resolution all need to be adapted to the computing power limitations of the power inspection equipment.
[0008] As a further aspect of the present invention, the sub-steps of step S2 are as follows: S21: The feature extraction module is designed with a ShuffleCSP structure. It balances computational load and detection accuracy through a dual-branch fusion design and lightweight built-in units. S22. Design a feature fusion module, adopt the GroupFuse mechanism, assign differentiated weights to feature maps at different stages and perform element-by-element fusion; S23. Design a multi-scale detection head, consisting of three feature maps of different scales, which are used to detect electrical equipment or defects at each scale. S24. Configure the text encoder, using BERT-tiny, whose parameter count is a set ratio of the BERT-base parameter count, and the set ratio must meet the lightweight requirements.
[0009] As a further aspect of the present invention, the sub-steps of step S3 are as follows: S31. First stage: Freeze the backbone network and cross-modal fusion module of the teacher model, and train only the feature extraction layer of the lightweight student model; S32, Second stage: Unfreeze the text encoder and cross-modal fusion module of the teacher model, and train the text branch and power-specific cross-modal attention layer of the lightweight student model; S33, Third Stage: Jointly train the detection head and all other modules of the lightweight student model, and introduce hard label supervision for power equipment.
[0010] As a further aspect of the present invention, the first stage processing in step S31 includes the following sub-steps: S311. Add the MSE loss function between the corresponding Layer 4 layers of the teacher model and the lightweight student model to achieve visual feature alignment. S312. Set training parameters: set the temperature parameter to the first set temperature value, set the feature distillation loss function hyperparameter to the first set weight value, set the learning rate to the first set learning rate, and set the batch size to the first set batch size. The set values of each parameter are determined according to the visual feature alignment requirements. S313. Set the number of training rounds based on the training parameters in step S313; S314. Before inputting the image into the lightweight student model, crop the image using the tower prediction model to focus on the core area of the transmission line to reduce invalid calculations.
[0011] As a further aspect of the present invention, the second stage processing in step S32 includes the following sub-steps: S321. Design a joint loss function, which includes visual feature distillation loss, text feature distillation loss and prediction score distillation loss. The parameter settings for each distillation loss are determined according to the cross-modal fusion requirements. S322. Integrate the SE attention mechanism and the ECA attention mechanism into the lightweight student model. The channel compression ratio of the SE attention mechanism is set to a set compression ratio, and the kernel size of the ECA attention mechanism is set to a set kernel size. The corresponding settings need to be adapted to the feature extraction requirements of power equipment. S323. Set training parameters: Set the learning rate to the second set learning rate and the batch size to the second set batch size. The setting values of each training parameter are determined according to the model training efficiency requirements. S324. Combine the training parameters in step S323 to set the number of training rounds.
[0012] As a further aspect of the present invention, the third stage of the processing in step S33 includes the following sub-steps: S331. Design a joint loss function, which includes visual feature distillation loss, text feature distillation loss, prediction score distillation loss and task-specific loss. The task-specific loss includes classification loss and localization loss, and the corresponding set values are determined according to the task fine-tuning accuracy requirements. S332. Introduce an adaptive NMS threshold to optimize the detection effect of small targets in power equipment; S333. Set training parameters: Set the learning rate to the third set learning rate and the batch size to the third set batch size. The setting values of each training parameter are determined according to the model convergence requirements. S334. Combine the training parameters in step S333 to set the number of training rounds.
[0013] As a further aspect of the present invention, the sub-steps of step S4 are as follows: S41. Quantize the lightweight student model using a preset precision to reduce memory usage and computational load. The preset precision needs to be compatible with the hardware support capabilities of the power inspection equipment. S42. Configure a dynamic batch processing strategy to adjust the batch size according to the complexity of the input image; S43. Enable feature map caching mechanism to reuse the underlying visual features of the model to reduce redundant calculations; S44. For power equipment inspection tasks, optimize the NMS parameters and classification threshold of the model; S45. Deploy the optimized model to the power inspection equipment, which includes a drone edge terminal. The computing power of the power inspection equipment shall not exceed a set computing power threshold and the memory shall not exceed a set memory threshold. All thresholds shall meet the real-time operation requirements of the model.
[0014] As a further aspect of the present invention, the objective of step S44 is to improve the recall rate of small target detection in power equipment inspection scenarios, wherein the small targets include insulators, bolts, and broken strands of wires.
[0015] As a further aspect of the present invention, the sub-steps of step S1 are as follows: S11. Collect power equipment inspection images in the scenarios of transmission lines and substations, including a set number of images without abnormalities and a set number of images with defects. The defects include broken conductor strands, damaged insulators, and corroded bolts. The images must include towers, conductors, and insulators. S12. Label the collected images, covering the categories of power equipment and equipment defects; S13. Construct a power terminology dictionary, which includes terms such as broken conductor strands, damaged insulators, and corroded bolts. S14. Data augmentation is performed on the labeled dataset by using random flipping, random rotation, random scaling, cutout operations, and StyleGAN to generate simulated images. S15. Divide the enhanced dataset into a training set and a validation set according to a set ratio, wherein the set ratio is determined according to the model training requirements.
[0016] Compared with the prior art, the beneficial effects of the present invention are: This invention takes a closed-loop design throughout the entire process as its core. Through the orderly connection of "building a dedicated power dataset → lightweight model design → three-stage progressive multimodal distillation → deployment optimization", it forms a complete system that specifically addresses the core technical pain points in power inspection scenarios. It has outstanding technical advantages and rigorous logic.
[0017] First, based on data, a dedicated dataset adapted to power inspection scenarios is constructed to provide accurate and practical support for subsequent model training, ensuring that the training data can fully cover power equipment types, defect characteristics, and scenario environments, thus guaranteeing the effectiveness of model learning from the source.
[0018] Next, a lightweight student model was designed specifically to address the key constraint of limited computing power of edge devices commonly used in power inspection. This design is not simply a reduction in model size, but rather a targeted architectural optimization based on the computing power capacity of edge devices. This effectively avoids the problem of mismatch between the general model and the computing power of power inspection equipment caused by the large number of parameters and the complexity of calculations, thus clearing the hardware adaptation obstacles for the actual deployment of the model.
[0019] Then, a three-stage progressive multimodal distillation mechanism is adopted. By systematically transferring the core capabilities of the teacher model in stages, the effective transfer of visual features is achieved first, then the cross-modal fusion capability is deepened, and finally the task adaptation capability is transferred, progressing step by step with a focus on key points. This staged transfer mode not only successfully achieves lightweight compression of the model architecture, but also effectively compensates for the accuracy loss problem that is prone to occur in the traditional lightweighting process by incorporating power scenario-specific optimization strategies, including professional optimization to enhance semantic understanding, attention enhancement to focus on key features, and post-processing optimization to adapt to target characteristics. This achieves the synergistic advancement of lightweighting and accuracy assurance.
[0020] Finally, pre-deployment optimization further adapts the model to the operating environment of edge devices, ensuring stable performance in real-world application scenarios. This invention achieves its core technical objectives: while significantly improving the model's lightweight nature and fully meeting the deployment requirements of low-computing-power power inspection equipment, it also significantly improves the detection accuracy of defects in power equipment and small targets. It possesses strong feasibility for implementation and ensures the reliability of detection results, perfectly solving the technical problem of reduced inspection accuracy caused by the lack of adaptation and optimization for the layered structure, occlusion characteristics, and technical terminology of power equipment. Attached Figure Description
[0021] Figure 1 This is a flowchart of the method of the present invention. Detailed Implementation
[0022] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0023] Please see Figure 1 This embodiment, based on the technical solution defined in the claims and combined with the actual needs of power inspection scenarios, elaborates on the specific implementation details of each step. It should be specifically noted that the "specific values" (such as data volume, parameter values, thresholds, etc.) mentioned in the following embodiments are all exemplary settings used to clearly illustrate the technical solution.
[0024] I. Constructing a dedicated detection dataset adapted to power equipment inspection scenarios To address the mismatch between existing general datasets and power line inspection scenarios, a high-quality dataset suitable for multimodal distillation training is generated through a complete "data acquisition-annotation-augmentation-partitioning" process. The specific implementation is as follows: 1.1 Power Line Inspection Image Acquisition Data collection scenarios: Focusing on two core power inspection scenarios: transmission lines (including towers, conductors, and insulators) and substations (including switchgear and transformers), covering different weather conditions such as sunny, cloudy, and foggy days, as well as lighting environments such as daytime and evening, to ensure the diversity of data scenarios.
[0025] Data volume setting: Collect a set number of images, specifically 1000 in this embodiment, including 750 images without abnormalities for learning normal equipment characteristics and 250 images with defects for learning fault characteristics; the defect types include high-frequency defects in power inspection such as broken conductor strands, broken insulators, rusted bolts, and tilted towers, and the image resolution is uniformly adjusted to 1920×1080.
[0026] 1.2 Image Annotation Labeling tool: The LabelStudio labeling platform is used, which supports dual-dimensional labeling of "object detection + defect classification".
[0027] Labeling rules: Equipment labeling: Draw bounding boxes for core power equipment (such as poles, conductors, insulators, bolts) in each image and label them with category labels (such as "pole - straight tower" and "insulator - suspension insulator"); Defect labeling: For images containing defects, draw bounding boxes in the defect areas and label the defect type (such as "conductor - broken strand" and "bolt - corrosion"), ensuring that the overlap between the bounding box and the defect area is ≥90% to avoid labeling offset.
[0028] 1.3 Constructing a dictionary of electrical terminology Dictionary purpose: To address the problem of insufficient understanding of power industry terminology by general text encoders and enhance the accuracy of cross-modal fusion.
[0029] The dictionary includes high-frequency command terms and equipment / defect names for power inspection scenarios. In this embodiment, it specifically includes more than 30 professional terms such as "conductor strand breakage detection", "insulator damage identification", "bolt corrosion location", and "tower tilt judgment". Each term corresponds to at least 5 associated images in the labeled dataset, forming a "term-image" mapping relationship.
[0030] 1.4 Dataset Augmentation The purpose of this enhancement is to expand the dataset size, improve the model's generalization ability, and avoid overfitting during training.
[0031] Enhancement Strategies: Basic Geometric Transformations: Perform random flipping (e.g., 50% probability of horizontal flipping, 20% probability of vertical flipping) and random rotation (e.g., random angle within the range of -15° to 15°) and random scaling (e.g., scaling factor of 0.8 to 1.2 times) on the labeled images; Pixel-Level Enhancement: Employ Cutout operations (e.g., occlusion area ≤ 10% of the total image area, simulating scenarios where leaves or wires obstruct equipment) and brightness / contrast adjustments (e.g., brightness ±20%, contrast ±15%, simulating different lighting conditions); Simulation Data Supplementation: Generate 100 highly realistic defect images (e.g., insulator damage of varying degrees, broken wire strands) using StyleGAN to supplement the scarce defect samples in the real data.
[0032] 1.5 Dataset Partitioning Divide the data into sets: ensure the representativeness of the training set and the objectivity of the validation set, and avoid data distribution bias.
[0033] Division ratio: The enhanced dataset is divided into a training set and a validation set according to a set ratio (9:1 in this embodiment). The training set is used for learning model parameters, and the validation set is used for real-time evaluation of model performance and adjustment of training parameters based on the validation results.
[0034] II. Design a lightweight student model that meets the computing power limitations of power inspection equipment. To address the characteristics of power line inspection equipment (such as drone edge terminals and portable inspection terminals) with "low computing power and small memory," lightweight improvements were made to the YOLOv4 architecture to ensure efficient inference of the model under limited hardware resources. The specific implementation is as follows: 2.1 Setting Core Constraints for the Model Computing power adaptation target: number of parameters ≤ set parameter threshold (5.5M in this embodiment), amount of computing power (in FLOPs) ≤ set computing threshold (12.8G in this embodiment), adapting to edge devices with computing power ≤ 20TOPS and memory ≤ 8GB.
[0035] Input resolution: Set to the set resolution (640×640 in this embodiment) to balance image detail preservation and computational load control. This avoids the loss of small targets (such as bolts) due to low resolution, and also avoids the increase in inference time due to high resolution.
[0036] 2.2 Feature Extraction Module Design (ShuffleCSP Structure) Structural improvements: Based on CSPNet (Cross-Stage Partially Connected Network), an optimization is introduced, including a Shuffle operation (channel shuffling). Specifically: Dual-branch design: The input feature map is divided into a "main branch" (70% of the total, processed by convolution, batch normalization, and activation function) and an "auxiliary branch" (30% of the total, processed only by lightweight convolution) to reduce redundant computation; Lightweight built-in units: The convolutional layer uses 3×3 depthwise separable convolution (replacing the traditional 3×3 standard convolution), and the activation function uses LeakyReLU (replacing ReLU to improve the expressive power of low feature response regions).
[0037] 2.3 Feature Fusion Module Design (GroupFuse Mechanism) Fusion Logic: Considering the "layered structure" of power equipment (environment → towers → equipment → defects), differentiated weight fusion is performed on feature maps at different stages (shallow features: edges, textures; deep features: semantics, categories): Group Weight Allocation: Feature maps are divided into three groups according to "shallow-medium-deep," and assigned weights of 0.3, 0.5, and 0.2 respectively (set values in this embodiment, which can be adjusted according to equipment type); Element-by-Element Fusion: Element-by-element addition is performed on the weighted feature maps to enhance the network's ability to simultaneously perceive "small targets (such as defects) and large equipment (such as towers)."
[0038] 2.4 Multi-scale detection head design Scale Division: To address the significant size variations of equipment in power line inspection scenarios, three scale detection heads are designed, each corresponding to different target sizes: Large-scale detection head (output feature map size: 80×80): responsible for detecting large targets, such as power poles (height ≥ 1 / 4 of image height) and substation switchgear; Medium-scale detection head (output feature map size: 40×40): responsible for detecting medium-sized targets, such as insulator strings (length approximately 1 / 8~1 / 4 of image height) and conductor segments; Small-scale detection head (output feature map size: 20×20): responsible for detecting small targets, such as bolts (diameter ≤ 1 / 20 of image width) and broken conductor strands (length ≤ 1 / 15 of image width).
[0039] Detection head structure: Each detection head contains a "convolutional layer + prediction layer". The prediction layer outputs the bounding box coordinates, class probability, and confidence of the target. The small-scale detection head adds an extra convolutional layer (to improve the expression dimension of small target features).
[0040] 2.5 Text Encoder Configuration (BERT-tiny) Selection criteria: To reduce the computational complexity of text processing, BERT-tiny is used as the text encoder, and its parameter size is a set ratio of BERT-base (1 / 8 in this embodiment, approximately 4M).
[0041] Pre-training and fine-tuning: First, BERT-tiny is pre-trained using general text corpora (such as Wikipedia), and then fine-tuned using the previously constructed "electrical terminology dictionary" to enable the encoder to accurately understand professional instructions such as "broken conductor strand" and "damaged insulator", ensuring that the cross-modal alignment accuracy between text and image is ≥88%.
[0042] III. Implementation of a three-stage progressive multimodal distillation mechanism Using Grounding-DINO as the teacher model (which has strong cross-modal detection capabilities but is computationally intensive), a three-stage progressive distillation process—"visual feature alignment → cross-modal fusion → task fine-tuning"—is employed to transfer knowledge from the teacher model to a lightweight student model. The specific implementation is as follows: 3.1 First Stage - Visual Feature Alignment Distillation (Focusing on "Visual Feature Transfer") The core objective of this stage is to enable student models to learn the visual feature extraction capabilities of teacher models, laying the foundation for subsequent cross-modal fusion.
[0043] 3.1.1 Control Model State Freeze the teacher model: Fix the parameters of the backbone network (ResNet-50) of Grounding-DINO and the cross-modal fusion module (bidirectional cross attention layer), and only unfreeze the feature extraction module of the student model (i.e. the ShuffleCSP structure designed in step S21) for training.
[0044] 3.1.2 Achieving Feature Alignment Loss function design: An MSE (mean squared error) loss function is added between the "Layer 4" layers (both deep visual feature output layers) of the teacher and student models. The formula is: L MSE =(1 / N)∑(F T i -F S i ) 2 F T i For the i-th feature map output by Layer 4 of the teacher model, F S i The i-th feature map output by Layer 4 of the student model, where N is the number of feature maps. Loss weight setting: The hyperparameter of the feature distillation loss function (corresponding to the "first set weight value" in the claim) is set to 0.9 to ensure that the student model prioritizes aligning with the visual features of the teacher model.
[0045] 3.1.3 Training Parameter Configuration Temperature parameter (first set temperature value): set to 4 (to smooth feature distribution and avoid local optima); Learning rate (first set learning rate): set to 0.0008 (cosine annealing learning rate scheduling, decreasing by 10% every 5 rounds); Batch size (first set batch size): set to 32 (to adapt to the GPU memory of the training device; in this embodiment, a single RTX3090 GPU is used); Training rounds (first training rounds): set to 20 rounds. After each round of training, the "visual feature similarity" is evaluated using a validation set (cosine similarity calculation, target ≥0.85). If the target is met, the next stage begins.
[0046] 3.1.4 Image Preprocessing Optimization (Tower Prediction and Cropping) Cropping logic: Before inputting the image into the student model, a pre-trained "tower prediction model" (built on MobileNetV3, with ≤1M parameters) is called to identify and crop the tower regions in the image, retaining "the tower and its surrounding 20% area" (removing invalid backgrounds such as the sky and ground). This reduces the computational cost of invalid backgrounds by approximately 35%, improving the student model's efficiency in feature extraction from the core areas of power equipment.
[0047] 3.2 Second Stage - Cross-modal Fusion Distillation (Focusing on "Text-Image Alignment") The core objective of this stage is to enable student models to learn the cross-modal fusion capabilities of teacher models, achieving "text-guided visual detection" (such as inputting "detect broken strands in wires," and the model accurately locating the broken strand area).
[0048] 3.2.1 Adjusting the model state Unfreeze the teacher model: Unfreeze the Grounding-DINO text encoder (BERT-base) and cross-modal fusion module to participate in the distillation process.
[0049] Training student model module: Unfreeze the text branch (BERT-tiny) of the student model with the "electric-specific cross-modal attention layer".
[0050] 3.2.2 Design the joint loss function The loss function is composed as follows: the hyperparameters of the visual feature distillation loss are set to the second set weight value, the hyperparameters of the text feature distillation loss are set to the third set weight value, the prediction score distillation loss adopts KL divergence, and the temperature parameter is set to the second set temperature value.
[0051] Visual feature distillation loss L vis The weight is set to 0.9 to continue the first stage and ensure that visual features do not degrade.
[0052] Text feature distillation loss L text We used MSE loss to align the output features of the text encoders of the teacher and student models, with a weight of 0.7.
[0053] Predicted fractional distillation loss L KL KL divergence (relative entropy) was used to align the predicted scores of the teacher model and the student model for "text-image matching degree". The temperature parameter was set to 4 and the weight was set to 0.4.
[0054] Total loss formula: L total =0.9L vis +0.7L text +0.4L KL .
[0055] 3.2.3 Integration of Attention Mechanisms SE Attention Mechanism (Channel Attention): An SE layer is added after the feature fusion module of the student model. The channel compression ratio is set to 16. The channel response to "defect features" is enhanced through "squeeze-excitation" operation.
[0056] ECA attention mechanism (spatial attention): An ECA layer is added in front of the small-scale detection head, and the kernel size is set to 3. The spatial localization capability of small targets (such as bolts) is enhanced through local cross-channel interaction. 3.2.4 Training Parameter Configuration The second learning rate is set to 0.0004, which is lower than that in the first stage, to avoid parameter oscillation.
[0057] Second batch size setting: Keep it at 32.
[0058] Training rounds: Set to 30 rounds. After each round of training, the "cross-modal detection accuracy" is evaluated using a validation set. If the accuracy is achieved, the process proceeds to the next stage.
[0059] 3.3 Third Stage - Task Fine-tuning and Distillation (Focusing on "Power Scenarios Adaptation") The core objective of this stage is to adapt student models to specific power line inspection tasks through "hard label supervision + task-specific loss," thereby further improving detection accuracy.
[0060] 3.3.1 Introducing hard label supervision Label source: Use hard labels for "equipment-defect" (such as "insulator-damaged" and "conductor-broken strand") instead of the "soft labels for teacher models" in the first two stages.
[0061] Supervision method: Hard label constraints are applied to the detection head output of the student model, that is, the category predicted by the model must be consistent with the labeled category, and the IoU (Intersection over Union) of the bounding box coordinates with the labeled box must be ≥0.5.
[0062] 3.3.2 Design a multi-dimensional joint loss function Loss function composition L vis Visual feature distillation loss: weight reduced to 0.7.
[0063] Text feature distillation loss L text The weight is reduced to 0.5.
[0064] Predicted fractional distillation loss L KL Set the KL divergence temperature parameter to 2 and the weight to 0.3.
[0065] Task-specific loss L loc This includes classification loss (cross-entropy loss, weight 0.4, constraining class prediction) and localization loss (GIoU loss, weight 0.6, constraining bounding box coordinates).
[0066] Total loss formula: L total =0.7L vis +0.5L text +0.3L KL +0.4L cls +0.6L loc .
[0067] 3.3.3 NMS Optimization (Adaptive Threshold) Optimization logic: Considering the characteristic of power equipment with "dense small targets" (such as multiple bolts on an insulator string), the fixed NMS threshold is abandoned (the traditional threshold of 0.5 easily leads to the accidental deletion of small targets), and an "adaptive threshold" is adopted: For large targets (such as towers): set the NMS threshold to 0.6 (allowing the detection boxes of larger overlapping areas to be retained).
[0068] For small targets (such as bolts or broken strands): set the NMS threshold to 0.3 (to reduce the false deletion of small target detection boxes).
[0069] 3.3.4 Configuring Training Parameters Learning rate: Set to 0.0002. Batch size: Keep it at 32. Training epochs: Set to 50 epochs. After training, evaluate the "overall detection performance" using the validation set. If the performance meets the standard, the distillation training is complete.
[0070] IV. Optimization and Equipment Deployment Before Model Deployment This section addresses the hardware limitations of power line inspection edge devices by employing a combination of quantization, dynamic scheduling, and dedicated optimization to ensure the model achieves both real-time detection and high accuracy. The specific implementation is as follows: 4.1 Model Quantization Quantification objective: Reduce model memory usage and computational load, while avoiding excessive loss of accuracy.
[0071] Quantization method: FP16 half-precision quantization (preset precision) is used to replace FP32 full precision during training. Quantization range: FP16 quantization is performed on the weights of the convolutional and fully connected layers of the student model, and the activation values retain FP16 precision. Precision compensation: The quantized model is fine-tuned through "Quantization-Aware Training" (QAT) to ensure that the precision loss after quantization is ≤3%.
[0072] 4.2 Dynamic Batch Processing Strategy Configuration Scheduling logic: Dynamically adjust the batch size based on the "complexity" of the input image (such as the number of devices and background complexity) to avoid "wasting computing power for simple images" or "memory overflow for complex images" caused by a fixed batch size: Simple images (such as a single tower with no obstructions): batch size is set to 4; Complex images (such as multiple devices in a substation with conductors obstructing insulators): batch size is set to 1.
[0073] Complexity assessment: It is achieved through "image feature complexity scoring" (based on edge detection and target quantity statistics), with a assessment time of ≤1ms, which does not affect real-time performance.
[0074] 4.3 Feature Map Caching Mechanism Enable caching logic: In response to the characteristics of "high similarity of consecutive frame images" in power line inspection (such as consecutive images taken by drones flying along the line), reuse the "low-level visual features" (such as edges and textures, accounting for 60% of the feature extraction calculation) of the previous frame image: Caching timing: Update the low-level feature cache every 3 frames of images (to avoid cache expiration and resulting decrease in accuracy); Cache storage: Store the low-level features in the device's DDR memory (read speed is 80% faster than recalculation).
[0075] 4.4 Optimization of Specific Parameters for Power Scenarios Optimization Objective: To further adjust model parameters and improve practical performance for high-frequency power line inspection tasks; Specific optimizations: NMS parameters: Based on the adaptive threshold in step S33, further fine-tune the small target threshold (e.g., adjust the bolt detection threshold from 0.3 to 0.25) to ensure a false negative rate of ≤5%; Classification threshold: Adjust the confidence threshold for defect categories from 0.5 to 0.4 (to reduce false negatives of "low-confidence but real defects"), and reduce false positives through "post-processing filtering" (e.g., defect area ≥ 5×5 pixels).
[0076] 4.5 Model Deployment and Equipment Adaptation Deployment equipment: power inspection edge equipment, specifically "UAV edge terminal" in this embodiment (corresponding to "setting computing power threshold" and "setting memory threshold" in the claims), with hardware parameters of: computing power 18 TOPS, memory 6GB, and storage 64GB.
[0077] Deployment tools: TensorRT (NVIDIA Inference Acceleration Engine) is used to optimize the model and generate .engine inference files adapted to the device.
[0078] Deployment testing: Real-time testing: Inference time for a single frame image ≤ 60ms (meets the real-time detection requirements of 15fps or higher); Stability testing: Continuous operation for 2 hours (simulating drone inspection time), with no model crashes, no memory leaks, and detection accuracy fluctuations ≤ 2%.
[0079] Final result: The deployed model can realize the entire process of "real-time detection during drone flight → defect location marking → data back transmission".
[0080] This invention addresses three core issues in power grid inspection scenarios—namely, model computational power mismatch, low accuracy for small targets, and poor scenario adaptability—through a comprehensive process of "dataset specialization → model lightweighting → distillation and gradual evolution → deployment optimization." All specific values in the embodiments are exemplary settings and can be adjusted in practical applications based on the hardware parameters of the power grid inspection equipment and the accuracy requirements of the inspection task.
[0081] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
Claims
1. A Grounding-DINO model acceleration method combining a multi-stage progressive multimodal distillation mechanism, characterized in that, Includes the following steps: S1. Construct a detection dataset adapted to power equipment inspection scenarios; S2. Design a lightweight student model that meets the computational power limitations of processing the detection dataset; S3. Using Grounding-DINO as the teacher model, knowledge from the teacher model is transferred to the lightweight student model through a three-stage progressive multimodal distillation mechanism. S4. Optimize the lightweight student model trained using the detection dataset before deployment, and then deploy the optimized lightweight student model to the power inspection equipment.
2. The Grounding-DINO model acceleration method combining a multi-stage progressive multimodal distillation mechanism as described in claim 1, characterized in that, The lightweight student model is based on YOLOv4 and improved to be lightweight, and meets the following requirements: the number of parameters does not exceed the set parameter threshold, the amount of computation does not exceed the set computation threshold, and the input resolution is the set resolution. The set parameter threshold, set computation threshold and set resolution must be adapted to the computing power limit of the power inspection equipment.
3. The Grounding-DINO model acceleration method based on a multi-stage progressive multimodal distillation mechanism as described in claim 2, characterized in that, The sub-steps of step S2 are as follows: S21: The feature extraction module is designed with a ShuffleCSP structure. It balances computational load and detection accuracy through a dual-branch fusion design and lightweight built-in units. S22. Design a feature fusion module, adopt the GroupFuse mechanism, assign differentiated weights to feature maps at different stages and perform element-by-element fusion; S23. Design a multi-scale detection head, consisting of three feature maps of different scales, which are used to detect electrical equipment or defects at each scale. S24. Configure the text encoder, using BERT-tiny, whose parameter count is a set ratio of the BERT-base parameter count, and the set ratio must meet the lightweight requirements.
4. A Grounding-DINO model acceleration method combining a multi-stage progressive multimodal distillation mechanism according to any one of claims 1-3, characterized in that, The sub-steps of step S3 are as follows: S31. First stage: Freeze the backbone network and cross-modal fusion module of the teacher model, and train only the feature extraction layer of the lightweight student model; S32, Second stage: Unfreeze the text encoder and cross-modal fusion module of the teacher model, and train the text branch and power-specific cross-modal attention layer of the lightweight student model; S33, Third Stage: Jointly train the detection head and all other modules of the lightweight student model, and introduce hard label supervision for power equipment.
5. The Grounding-DINO model acceleration method according to claim 4, which combines a multi-stage progressive multimodal distillation mechanism, is characterized in that... The first stage of processing in step S31 includes the following sub-steps: S311. Add the MSE loss function between the corresponding Layer 4 layers of the teacher model and the lightweight student model to achieve visual feature alignment. S312. Set training parameters: set the temperature parameter to the first set temperature value, set the feature distillation loss function hyperparameter to the first set weight value, set the learning rate to the first set learning rate, and set the batch size to the first set batch size. The set values of each parameter are determined according to the visual feature alignment requirements. S313. Set the number of training rounds based on the training parameters in step S313; S314. Before inputting the image into the lightweight student model, crop the image using the tower prediction model to focus on the core area of the transmission line to reduce invalid calculations.
6. The Grounding-DINO model acceleration method according to claim 5, which combines a multi-stage progressive multimodal distillation mechanism, is characterized in that... The second stage of processing in step S32 includes the following sub-steps: S321. Design a joint loss function, which includes visual feature distillation loss, text feature distillation loss and prediction score distillation loss. The parameter settings for each distillation loss are determined according to the cross-modal fusion requirements. S322. Integrate the SE attention mechanism and the ECA attention mechanism into the lightweight student model. The channel compression ratio of the SE attention mechanism is set to a set compression ratio, and the kernel size of the ECA attention mechanism is set to a set kernel size. The corresponding settings need to be adapted to the feature extraction requirements of power equipment. S323. Set training parameters: Set the learning rate to the second set learning rate and the batch size to the second set batch size. The setting values of each training parameter are determined according to the model training efficiency requirements. S324. Combine the training parameters in step S323 to set the number of training rounds.
7. The Grounding-DINO model acceleration method based on a multi-stage progressive multimodal distillation mechanism as described in claim 6, characterized in that, The third stage of processing in step S33 includes the following sub-steps: S331. Design a joint loss function, which includes visual feature distillation loss, text feature distillation loss, prediction score distillation loss and task-specific loss. The task-specific loss includes classification loss and localization loss, and the corresponding set values are determined according to the task fine-tuning accuracy requirements. S332. Introduce an adaptive NMS threshold to optimize the detection effect of small targets in power equipment; S333. Set training parameters: Set the learning rate to the third set learning rate and the batch size to the third set batch size. The setting values of each training parameter are determined according to the model convergence requirements. S334. Combine the training parameters in step S333 to set the number of training rounds.
8. The Grounding-DINO model acceleration method according to claim 7, which combines a multi-stage progressive multimodal distillation mechanism, is characterized in that... The sub-steps of step S4 are as follows: S41. Quantize the lightweight student model using a preset precision to reduce memory usage and computational load. The preset precision needs to be compatible with the hardware support capabilities of the power inspection equipment. S42. Configure a dynamic batch processing strategy to adjust the batch size according to the complexity of the input image; S43. Enable feature map caching mechanism to reuse the underlying visual features of the model to reduce redundant calculations; S44. For power equipment inspection tasks, optimize the NMS parameters and classification threshold of the model; S45. Deploy the optimized model to the power inspection equipment, which includes a drone edge terminal. The computing power of the power inspection equipment shall not exceed a set computing power threshold and the memory shall not exceed a set memory threshold. All thresholds shall meet the real-time operation requirements of the model.
9. The Grounding-DINO model acceleration method according to claim 8, which combines a multi-stage progressive multimodal distillation mechanism, is characterized in that... The objective of step S44 is to improve the recall rate of small target detection in power equipment inspection scenarios. Small targets include insulators, bolts, and broken strands of conductors.
10. The Grounding-DINO model acceleration method according to claim 8, which combines a multi-stage progressive multimodal distillation mechanism, is characterized in that... The sub-steps of step S1 are as follows: S11. Collect power equipment inspection images in the scenarios of transmission lines and substations, including a set number of images without abnormalities and a set number of images with defects. The defects include broken conductor strands, damaged insulators, and corroded bolts. The images must include towers, conductors, and insulators. S12. Label the collected images, covering the categories of power equipment and equipment defects; S13. Construct a power terminology dictionary, which includes terms such as broken conductor strands, damaged insulators, and corroded bolts. S14. Data augmentation is performed on the labeled dataset by using random flipping, random rotation, random scaling, cutout operations, and StyleGAN to generate simulated images. S15. Divide the enhanced dataset into a training set and a validation set according to a set ratio, wherein the set ratio is determined according to the model training requirements.
Citation Information
Cited By
Electric power inspection adaptive method and system based on cloud edge collaboration and hierarchical skill migration
CN122024161A