Lightweight target detection model integrating attention mechanism, optimization feature fusion method and self-supervised learning
By integrating the CBAM attention module, FPN feature fusion structure and self-supervised learning in the YOLOv7-tiny model, and combining channel pruning and distillation technology, the efficient and accurate detection capabilities of the lightweight object detection model are achieved, solving the accuracy and efficiency of power insulator defect detection under drone shooting.
Patent Information
- Application Number
- CN202510136326.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-07
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-02-07
AI Technical Summary
The prior art is difficult to efficiently and accurately detect power insulator defects under wide and variable perspectives taken by drones, resulting in low detection accuracy and large model complexity.
Based on YOLOv7-tiny, the CBAM attention module and FPN feature fusion structure are integrated, and self-supervised learning is used for backbone network pre-training, combining channel pruning and channel distillation technology to achieve lightweight and efficient model.
It significantly improves the model's defect detection capability for small goals and complex backgrounds, reduces the model parameter quantity and calculation overhead, and is suitable for real-time deployment on edge devices, meeting the needs of efficient and accurate power insulator defect detection.
Smart Images

Figure CN120070858A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular, to a lightweight object detection model integrating an attention mechanism, an optimized feature fusion method, and self-supervised learning. Background Art
[0002] With the development of the power industry, the length and complexity of transmission lines have been increasing continuously. As a key component, insulators undertake the dual functions of electrical insulation and mechanical support. However, insulators are long-term exposed to harsh environments and are easily affected by natural factors such as lightning strikes, pollution, and severe cold, resulting in insulator defects, such as insulator damage and flashover. These defects may cause the transmission line to contact the tower pole or other lines, resulting in power outages or even large-scale power failure accidents. Therefore, it is crucial to regularly inspect the insulators on transmission lines.
[0003] Traditional insulator inspections mainly rely on manual inspections, which have problems such as low efficiency, high cost, and poor safety. With the development of unmanned aerial vehicle (UAV) technology, it has become possible to use UAVs to collect images of insulators on high-voltage transmission lines. However, the wide-angle and variable perspectives captured by UAVs bring challenges such as small object detection and complex backgrounds, making it difficult to meet the actual requirements in terms of detection accuracy and efficiency. Deep learning, especially object detection algorithms, has become a research hotspot in this context. However, existing methods still have problems such as low detection accuracy and large model complexity. Therefore, it is of great practical significance to develop an efficient, lightweight, and high-precision insulator defect detection model. Summary of the Invention
[0004] The purpose of the present invention is to provide a lightweight object detection model integrating an attention mechanism, an optimized feature fusion method, and self-supervised learning, which is specifically used for the efficient and accurate detection of power insulator defects. By improving on the basis of YOLOv7-tiny, integrating the CBAM attention module and the FPN feature fusion structure, and using self-supervised learning for pre-training of the backbone network, further combined with channel pruning and channel distillation techniques, the lightweight and high-efficiency of the model are realized, which is suitable for deployment on edge devices such as UAVs to meet the needs of real-time detection.
[0005] To achieve the above purpose, the present invention provides a lightweight object detection model integrating an attention mechanism, an optimized feature fusion method, and self-supervised learning, including a backbone network Backbone, a feature fusion module Neck, a detection head Detection Head, a lightweight module, and a self-supervised learning module;
[0006] The backbone is the FasterNet architecture, and the Convolutional Block Attention Module (CBAM) is integrated into some of the pointwise convolutional (PConv) layers and depthwise separable convolutional (PWConv) layers in the architecture to enhance the attention ability to channel and spatial features;
[0007] The feature fusion module introduces the Feature Pyramid Network (FPN) structure and combines it with the OD-SlimNeck module to optimize the fusion effect of multi-scale features;
[0008] The Detection Head uses the Decoupled Head. The Decoupled Head includes independent classification and localization branches, which are used to handle the prediction of target categories and the prediction of target positions respectively;
[0009] The lightweight module reduces the number of model parameters through channel pruning and channel distillation techniques;
[0010] The self-supervised learning module pre-trains the backbone using the contrastive learning method.
[0011] Preferably, the CBAM module includes two sub-modules: channel attention and spatial attention, which are used to enhance the channel information and spatial information of the feature map respectively. The mathematical expression of the channel attention mechanism is:
[0012] M c =σ(W 2 ·δ(W 1 ·AvgPool(X))+W 2 ·δ(W 1 ·MaxPool(X)));
[0013] where M c is the channel attention map, with a size of (C, 1, 1), representing the importance of each channel; X represents the input feature map, with a size of (C, H, W), where C is the number of channels, H is the height, and W is the width; AvgPool(·) represents global average pooling of the input feature map in the spatial dimension, with an output size of (C, 1, 1); MaxPool(·) represents global max pooling of the input feature map in the spatial dimension, with an output size of (C, 1, 1); W 1 and W 2 are learnable weight matrices, corresponding to the fully connected layer weight and convolutional layer weight of channel attention respectively, determining the weighting method in the channel dimension; δ(·) is the ReLU activation function; σ(·) is the Sigmoid function, with an output range of (0, 1);
[0014]
[0015] where Y cRepresents the output feature map of the features after channel attention weighting, with the same size as X; Is an element-wise multiplication operation, that is, after broadcasting Mc to the same spatial dimension as X, it performs a per-channel product with X;
[0016] The mathematical expression of the spatial attention mechanism is:
[0017] M s = σ(Conv 7×7 ([AvgPool(Y c ) ; MaxPool(Y c )])) ;
[0018] Among them, AvgPool(Y c ) represents performing a per-pixel average operation on Y c in the channel dimension, with the output size of (1, H, W), and MaxPool(Y c ) represents performing a per-pixel maximum operation on Y c in the channel dimension, with the output size of (1, H, W); [AvgPool(Y c ) ; MaxPool(Y c )] represents concatenating the average pooling result and the maximum pooling result in the channel dimension, with the size of (2, H, W); Conv 7×7 (·) represents a 7×7 convolution operation for extracting the spatial attention information of the concatenated features; M s represents the spatial attention map, with the size of (1, H, W), indicating the importance degree of each position in the spatial dimension;
[0019] The mathematical expression of the overall attention mechanism Y:
[0020]
[0021] Among them, Y represents the final output feature map after comprehensively applying channel attention and spatial attention, with the size of (C, H, W).
[0022] Preferably, the feature fusion module combines multi-scale feature maps and uses global and local features to enhance the representation ability of the feature maps. The FPN structure fuses feature maps of different scales layer by layer through a top-down path and lateral connections. Among them, the top-down path performs feature fusion:
[0023]
[0024] Among them, is the fused feature, obtained at level l; P l is the original feature map, Denote the feature map that has been fused in the previous layer; UpSample(·) represents the upsampling operation, which increases the resolution of the high-level features and then adds them element-wise to the low-level features. l represents the index of the feature map, and L represents the index of the highest layer;
[0025] Horizontal connection: For each After adjusting the number of channels with a 1×1 convolution, add it element-wise to the original P l Element-wise addition.
[0026] Preferably, the channel pruning technique of the lightweight module reduces the number of model parameters and computational overhead by evaluating the scaling factors of the Batch Normalization layer and pruning the channels with scaling factors close to zero;
[0027] The channel distillation technique of the lightweight module uses the original YOLOv7 model as the teacher model to perform channel distillation on the pruned student model, transfers knowledge on the samples correctly predicted by the teacher model, and the distillation loss function
[0028]
[0029] Among them, w ij is the weight; it is used to measure the contribution degree of different channels or different samples to the loss; z i and z j are the two-view feature representations extracted by using FasterNet for enhancement, τ is the temperature parameter, which controls the degree of softening distribution; z i and z j represent the feature vectors obtained from different enhanced views of the same image, generally of size (d, 1) or (1, d). sim(z i ,z j ) represents the similarity function between vectors, defined as the cosine similarity; 2N is the total number of sample views, with N samples, and each sample generates two views; l [k≠i] is the indicator function, which takes 1 when k≠i and 0 otherwise, and is used to exclude the influence of the same view itself; and are the channel outputs of the student model and the teacher model respectively; ||·|| refers to the Euclidean distance or other norm distances, and here it is the square of the second norm ||·|| 2 ; is the channel distillation loss, which improves the performance of the student model by minimizing the difference between the outputs of the student model and the teacher model;
[0030] The overall loss function is
[0031]
[0032] Among them, λ and β are weight parameters that control the contribution of each loss to the overall training; is the supervised loss of the object detection task itself, such as including bounding box regression loss and classification loss; is the self-supervised contrast loss SimCLR part; is the channel distillation loss.
[0033] Preferably, the self-supervised learning module generates different views by applying random cropping, color jittering, and rotation data augmentation operations to the input image. Among them, the self-supervised contrast loss function SimCLR is calculated as follows:
[0034]
[0035] The feature representation of the backbone network is optimized using the self-supervised contrast loss function, making different views of the same image close in the feature space and the features of different images far away.
[0036] Preferably, the lightweight module further includes a model fine-tuning step and is used to perform fine-tuning training on the pruned model.
[0037] Preferably, the decoupled detection head adopts a hybrid channel strategy, sharing weights partially, and is used for classification prediction and localization prediction.
[0038] An application method of a lightweight object detection model integrating attention mechanism, optimized feature fusion method and self-supervised learning includes the following steps:
[0039] S1. Data collection: Use a drone to collect images of the electrical insulators on the high-voltage transmission line;
[0040] S2. Data preprocessing: Perform normalization processing on the collected images, including cropping, scaling, and enhancement;
[0041] S3. Model pre-training: Use the self-supervised learning method to pre-train the backbone network on a large amount of unlabeled data;
[0042] S4. Model integration: Integrate the pre-trained backbone network with the feature fusion module and the decoupled detection head to construct a complete object detection model;
[0043] S5. Model training: Use the labeled data to perform supervised training on the integrated model to optimize the detection performance;
[0044] S6. Model compression: Reduce the number of model parameters through channel pruning and channel distillation techniques to achieve model lightweight;
[0045] S7. Model Deployment: Deploy the compressed model onto the drone to achieve real-time defect detection of power insulators.
[0046] Preferably, in S5, a hybrid loss function is adopted during the training process, combining supervised detection loss and self-supervised contrast loss.
[0047] Preferably, in S6, a global pruning threshold is set to prune all channels with scaling factors lower than this threshold, and the model performance is restored through fine-tuning training.
[0048] Therefore, the present invention adopts the above lightweight object detection model integrating an attention mechanism, an optimized feature fusion method, and self-supervised learning, and the beneficial effects are as follows:
[0049] (1) By integrating the CBAM attention module and the FPN feature fusion structure, the present invention improves the model's detection ability for defects under small targets and complex backgrounds.
[0050] (2) Through channel pruning and channel distillation techniques, the present invention significantly reduces the model's parameter quantity and computational overhead, making it suitable for real-time deployment on edge devices.
[0051] (3) By the self-supervised learning method, the present invention enhances the feature representation ability of the backbone network and improves the model's generalization performance.
[0052] (4) While maintaining high detection accuracy, the lightweight model in the present invention significantly shortens the inference time, meeting the requirements of real-time detection. Description of the Drawings
[0053] Figure 1 is the overall architecture schematic diagram of an embodiment of the lightweight object detection model integrating an attention mechanism, an optimized feature fusion method, and self-supervised learning of the present invention;
[0054] Figure 2 is the schematic diagram of the FasterNet convolutional block structure of the integrated CBAM attention module in an embodiment of the lightweight object detection model integrating an attention mechanism, an optimized feature fusion method, and self-supervised learning of the present invention;
[0055] Figure 3 is the schematic diagram of the FPN feature fusion structure in an embodiment of the lightweight object detection model integrating an attention mechanism, an optimized feature fusion method, and self-supervised learning of the present invention;
[0056] Figure 4 is the schematic diagram of the self-supervised learning (SimCLR) pre-training process in an embodiment of the lightweight object detection model integrating an attention mechanism, an optimized feature fusion method, and self-supervised learning of the present invention;
[0057] Figure 5 It is a schematic diagram of the channel pruning process of an embodiment of a lightweight object detection model integrating an attention mechanism, an optimized feature fusion method, and self-supervised learning in the present invention;
[0058] Figure 6 It is a schematic diagram of the channel distillation process of an embodiment of a lightweight object detection model integrating an attention mechanism, an optimized feature fusion method, and self-supervised learning in the present invention. Detailed implementation manners
[0059] The technical solutions of the present invention will be further described below with reference to the accompanying drawings and embodiments.
[0060] Unless otherwise defined, the technical terms or scientific terms used in the present invention shall have the ordinary meanings understood by those of ordinary skill in the field to which the present invention belongs.
[0061] Embodiment
[0062] As Figure 1 shown, a lightweight object detection model integrating an attention mechanism, an optimized feature fusion method, and self-supervised learning includes a backbone network (Backbone), a feature fusion module (Neck), a detection head (Detection Head), a lightweight module, and a self-supervised learning module. The specific model implementation is as follows:
[0063] 1. Model architecture design
[0064] The object detection model of the present invention is improved based on YOLOv7-tiny as follows:
[0065] Backbone network (Backbone): It is replaced with the FasterNet architecture, and a convolutional block attention module (CBAM) is integrated after each PConv and PWConv convolutional layer to enhance the attention ability to channel and spatial features.
[0066] Neck part: The feature pyramid network (FPN) structure is introduced and combined with the OD-SlimNeck module to optimize the fusion effect of multi-scale features and improve the accuracy of small object detection.
[0067] Detection head (Detection Head): The decoupled head is adopted to process the classification and localization tasks separately, improving the detection accuracy.
[0068] Lightweight technology: Through channel pruning and channel distillation technologies, the number of model parameters is reduced, realizing the lightweight of the model, which is suitable for deployment on edge devices.
[0069] 2. Integration of attention mechanism
[0070] As Figure 2 shown, after each PConv and PWConv layer in FastNet, an integrated attention mechanism CBAM module is incorporated, which specifically includes a channel attention module and a spatial attention module. Their mathematical expressions are as follows:
[0071] Mathematical expression of the channel attention mechanism:
[0072] M c = σ(W 2 ·δ(W 1 ·AvgPool(X)) + W 2 ·δ(W 1 ·MaxPool(X)));
[0073] Among them, M c is the channel attention map, with a size of (C, 1, 1), representing the importance of each channel; X represents the input feature map, with a size of (C, H, W), where C is the number of channels, H is the height, and W is the width; AvgPool(·) represents performing global average pooling on the input feature map in the spatial dimension, with an output size of (C, 1, 1); MaxPool(·) represents performing global max pooling on the input feature map in the spatial dimension, with an output size of (C, 1, 1); W 1 and W 2 are learnable weight matrices, corresponding to the fully connected layer weight and convolutional layer weight of the channel attention respectively, determining the weighting method in the channel dimension; δ(·) is the ReLU activation function; σ(·) is the Sigmoid function, with an output range of (0, 1);
[0074]
[0075] Among them, Y c represents the feature output feature map after being weighted by the channel attention, with the same size as X; is the element-wise multiplication operation, that is, after broadcasting M c to the same spatial dimension as X, performing element-wise multiplication with X;
[0076] Mathematical expression of the spatial attention mechanism:
[0077] M s = σ(Conv 7×7 ([AvgPool(Y c ) ; MaxPool(Y c )]));
[0078] Among them, AvgPool(Y c ) represents performing average pooling on Y in the channel dimensionc Perform a per-pixel averaging operation, with the output size being (1, H, W). MaxPool(Y c ) represents performing a per-pixel maximum operation on Y along the channel dimension c Perform a per-pixel maximum operation, with the output size being (1, H, W); [AvgPool(Y c ); MaxPool(Y c )] means concatenating the average pooling result and the maximum pooling result along the channel dimension, with the size being (2, H, W); Conv 7×7 (·) represents a 7×7 convolutional operation used to extract the spatial attention information of the concatenated features; M s represents the spatial attention map, with the size being (1, H, W), indicating the importance degree of each position in the spatial dimension;
[0079] The mathematical expression of the overall attention mechanism Y:
[0080]
[0081] where Y represents the final output feature map after comprehensively applying channel attention and spatial attention, with the size being (C, H, W).
[0082] 3. Optimization of the Feature Pyramid Network (FPN) method
[0083] As Figure 3 shown, introduce the FPN structure in the Neck part and combine it with the OD-SlimNeck module. The specific process is as follows:
[0084] (1) First, perform feature fusion along the top-down path:
[0085]
[0086] where is the fused feature, obtained at level l; P l is the original feature map, represents the already fused feature map of the previous layer; UpSample(·) represents the upsampling operation, which increases the resolution of the high-level feature and then adds it element-wise to the low-level feature. l represents the index of the feature map, and L represents the index of the highest layer;
[0087] (2) Lateral connection: After adjusting the number of channels of each using a 1×1 convolution, add it element-wise to the original P l .
[0088] 4. Introduction of self-supervised learning
[0089] As Figure 4As shown in the figure, self-supervised pre-training of FasterNet is performed using self-supervised contrastive loss (SimCLR), and the specific steps are as follows:
[0090] (1) Data augmentation: Apply augmentation operations such as random cropping, color jittering, and rotation to each input image to generate two different views.
[0091] (2) Feature extraction: Use FasterNet to extract the feature representations z i and z j of the two augmented views.
[0092] (3) Self-supervised contrastive loss calculation (SimCLR): Calculate the similarity of positive sample pairs and the similarity of negative sample pairs, and optimize the contrastive loss.
[0093]
[0094] Among them, z i and z j are the feature representations of the two augmented views extracted using FasterNet, τ is the temperature parameter that controls the degree of softening distribution; z i and z j represent the feature vectors obtained from different augmented views of the same image, generally of size (d, 1) or (1, d), sim(z i ,z j ) represents the similarity function between vectors, defined as cosine similarity; 2N is the total number of sample views, with N samples, and two views are generated for each sample; l [k≠i] is the indicator function, which takes 1 when k≠i and 0 otherwise, used to exclude the influence of the same view itself. The feature representation of the backbone network is optimized using the self-supervised contrastive loss function, making the different views of the same image close in the feature space and the features of different images far away.
[0095] 5. Lightweight technology
[0096] As Figures 5 - 6 shown, the lightweighting of the model is achieved through channel pruning and channel distillation, and the specific steps are as follows:
[0097] Channel Pruning:
[0098] (1) Perform sparse training on the model optimized by the attention mechanism and feature fusion, and use the scaling factor of the BatchNormalization layer as the pruning basis.
[0099] (2) Set a global threshold to prune the channels with scaling factors close to zero, reducing the number of model parameters.
[0100] (3) Restore the model accuracy through fine-tuning.
[0101]
[0102] Among them, w ij is the weight, which is used to measure the contribution degree of different channels or different samples to the loss. and are the channel outputs of the student model and the teacher model respectively; ||·|| refers to the Euclidean distance or other norm distances, and here it is mostly the square of the second norm ||·|| 2 . is the Channel Distillation Loss, which improves the performance of the student model by minimizing the difference between the outputs of the student model and the teacher model.
[0103] The overall loss function is
[0104]
[0105] Among them, λ and β are weight parameters; they control the contribution of each loss to the overall training. is the supervised loss (Detection Loss) of the object detection task itself, such as including the box regression loss and the classification loss. is the self-supervised contrast loss (SimCLR) part. is the channel distillation loss.
[0106] 6. Training process
[0107] (1) Self-supervised pre-training stage
[0108] Use the SimCLR method to pre-train FasterNet+CBAM on a large amount of unlabeled data to optimize the feature representation.
[0109] (2) Transfer learning and fine-tuning stage
[0110] Integrate the pre-trained FasterNet+CBAM into the OD-YOLOV7-tiny model and optimize the feature fusion in combination with FPN.
[0111] Use the labeled data to perform supervised fine-tuning on the entire detection model to optimize the detection performance.
[0112] (3) Lightweight stage
[0113] Apply channel pruning to the fine-tuned model to reduce the number of model parameters.
[0114] Further improve the detection accuracy of the pruned model through the channel distillation method.
[0115] (4) Final model deployment
[0116] Deploy the lightweight OD-YOLOv7-tiny model to edge devices such as drones to achieve real-time detection of power insulator defects.
[0117] Implementation effect of the algorithm model
[0118] By integrating the attention mechanism, optimizing the feature fusion method, and introducing self-supervised learning, the present invention significantly improves the performance of the object detection model in power insulator defect detection. The specific manifestations are as follows:
[0119] Improved detection accuracy: By integrating the CBAM attention module and the FPN feature fusion structure, the detection ability of the model for defects under small targets and complex backgrounds is improved.
[0120] Model lightweight: Through channel pruning and channel distillation technologies, the number of model parameters and computational overhead are significantly reduced, making it suitable for real-time deployment on edge devices.
[0121] Enhanced feature representation: The self-supervised learning method improves the feature representation ability of the backbone network and enhances the generalization performance of the model.
[0122] Real-time guarantee: While maintaining high detection accuracy, the lightweight model significantly shortens the inference time, meeting the requirements of real-time detection.
[0123] Through experiments on the power insulator defect dataset, the effectiveness of the model of the present invention is verified. The experimental results show that after integrating the attention mechanism and optimizing the feature fusion method, the detection accuracy of the model is significantly improved. Through channel pruning and channel distillation, the number of model parameters is reduced by more than 50%, and the inference speed is increased by about 37%. While maintaining high detection accuracy, lightweight and high efficiency are achieved, making it suitable for real-time deployment on edge devices.
[0124] Therefore, the lightweight object detection model of the present invention that integrates the above attention mechanism, optimizes the feature fusion method, and self-supervised learning is particularly suitable for the efficient and accurate detection of power insulator defects. Through multiple improvements on the basis of YOLOv7-tiny, the organic combination of high performance and lightweight of the model is achieved, and it has broad application prospects.
[0125] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that they can still modify or equivalently replace the technical solutions of the present invention, and these modifications or equivalent replacements do not enable the modified technical solutions to deviate from the spirit and scope of the technical solutions of the present invention.
Claims
1. A lightweight target detection model integrating attention mechanism, optimized feature fusion method and self-supervised learning, characterized by: It includes the backbone network Backbone, feature fusion module Neck, detection head Detection Head, lightweight module and self-supervised learning module; The backbone network is the FasterNet architecture, and the convolutional block attention module CBAM is integrated in some of the convolutional PConv layers and point-by-point convolutional PWConv layers in the architecture to enhance the attention ability of channel and spatial features; The feature fusion module introduces the feature pyramid network FPN structure and combines it with the OD-SlimNeck module to optimize the fusion effect of multi-scale features; The detection head is a decoupled detection head, which includes independent classification branches and positioning branches, which are used to process the prediction of target categories and target positions respectively. Lightweight module, which reduces the number of model parameters through channel pruning and channel distillation technology; The self-supervised learning module uses contrastive learning method to pre-train the backbone network.
2. According to claim 1, a lightweight target detection model integrating attention mechanism, optimized feature fusion method and self-supervised learning is characterized in that: The CBAM module includes two sub-modules: channel attention and spatial attention, which are used to enhance the channel information and spatial information of the feature map respectively. The mathematical expression of the channel attention mechanism is: M c =σ(W2·δ(W1·AvgPool(X))+W2·δ(W1·MaxPool(X))); Among them, M c is the channel attention map, with a size of (C, 1, 1), representing the importance of each channel; X represents the input feature map, with a size of (C, H, W), where C is the number of channels, H is the channel height, and W is the channel width; AvgPool(·) means global average pooling of the input feature map in the spatial dimension, with an output size of (C, 1, 1); MaxPool(·) means global maximum pooling of the input feature map in the spatial dimension, with an output size of (C, 1, 1); W1 and W2 are learnable weight matrices, corresponding to the fully connected layer weights and convolutional layer weights of the channel attention, respectively, and determine the weighting method in the channel dimension; δ(·) is the ReLU activation function; σ(·) is the Sigmoid function, with an output range of (0, 1); Among them, Y c It is the feature output feature map after channel attention weighting, and its size is the same as that of X; For the element-by-element multiplication operation, M c After propagating to the same spatial dimension as X, perform channel-wise multiplication with X; The mathematical expression of the spatial attention mechanism is: M s =σ(Conv 7×7 ([AvgPool(Y c );MaxPool(Y c )])); Among them, AvgPool(Y c ) represents the channel dimension of Y c Perform pixel-by-pixel averaging, the output size is (1, H, W), MaxPool (Y c ) represents the channel dimension of Y c Perform pixel-by-pixel maximum operation, and the output size is (1, H, W); [AvgPool(Y c );MaxPool(Y c )] means concatenating the average pooling result and the maximum pooling result in the channel dimension, with a size of (2, H, W); Conv 7×7 (·) represents a 7×7 convolution operation, which is used to extract the spatial attention information of the concatenated features; M s Representative attention map, size is (1, H, W), indicating the importance of each position in the spatial dimension; The mathematical expression of the total attention mechanism Y is: Among them, Y represents the final output feature map after comprehensive application of channel attention and spatial attention, and its size is (C, H, W).
3. According to claim 2, a lightweight target detection model integrating attention mechanism, optimized feature fusion method and self-supervised learning is characterized in that: The feature fusion module combines multi-scale feature maps and uses global and local features to enhance the representation ability of feature maps. The FPN structure fuses feature maps of different scales layer by layer through top-down paths and lateral connections. The top-down path performs feature fusion: in, is the fusion feature, obtained in level l; P l is the original feature map, Indicates the fused feature map of the previous layer; UpSample(·) indicates the upsampling operation, which increases the resolution of high-level features and then adds them element-by-element with low-level features. l indicates the index of the feature map, and L indicates the highest layer index. Horizontal connection: For each After 1×1 convolution to adjust the number of channels, it is different from the original P l Add element-wise.
4. According to claim 3, a lightweight target detection model integrating attention mechanism, optimized feature fusion method and self-supervised learning is characterized in that: The channel pruning technology of the lightweight module reduces the number of model parameters and computational overhead by evaluating the scaling factor of the BatchNormalization layer and pruning channels with scaling factors close to zero; The channel distillation technology of the lightweight module uses the original YOLOv7 model as the teacher model, performs channel distillation on the pruned student model, and transfers knowledge on the samples correctly predicted by the teacher model. The distillation loss function Among them, w ij is the weight; it is used to measure the contribution of different channels or different samples to the loss; τ is the temperature parameter, which controls the degree of softening distribution; z i and z j It represents the feature vector obtained from different enhanced views of the same image, and its size is generally (d, 1) or (1, d). i ,z j ) represents the similarity function between vectors, which is defined as cosine similarity; 2N is the total number of sample views, and each sample in N samples generates two views; l [k≠i] is an indicator function, which takes the value 1 when k≠i and takes the value 0 otherwise, and is used to exclude the influence of the same view itself; and are the channel outputs of the student model and the teacher model respectively; ||·|| refers to the Euclidean distance or other norm distance, which is the square of the second norm here ||·|| 2 ; Channel distillation loss is used to improve the performance of the student model by minimizing the difference between the output of the student model and the teacher model. Overall loss function for Among them, λ and β are weight parameters, which control the contribution of each loss to the overall training; It is the supervised loss of the target detection task itself, including bounding box regression loss and classification loss; It is the self-supervised contrast loss SimCLR part; is the channel distillation loss.
5. According to claim 4, a lightweight target detection model integrating attention mechanism, optimized feature fusion method and self-supervised learning is characterized in that: The self-supervised learning module generates different views by applying random cropping, color jittering, and rotation data augmentation operations to the input image, where the self-supervised contrast loss function SimCLR is calculated: The feature representation of the backbone network is optimized using a self-supervised contrastive loss function, so that different views of the same image are close in the feature space and features of different images are far apart.
6. A lightweight target detection model integrating attention mechanism, optimized feature fusion method and self-supervised learning according to claim 5, characterized in that: The lightweight module also includes a model fine-tuning step and is used to fine-tune the pruned model.
7. A lightweight target detection model integrating attention mechanism, optimized feature fusion method and self-supervised learning according to claim 6, characterized in that: The decoupled detection head adopts a hybrid channel strategy, partially sharing weights for classification prediction and localization prediction.
8. An application method of a lightweight target detection model integrating attention mechanism, optimized feature fusion method and self-supervised learning, characterized in that: The following steps are involved: S1. Data collection: Use drones to collect images of power insulators on high-voltage transmission lines; S2, data preprocessing: standardize the collected images, including cropping, scaling and enhancement; S3, model pre-training: self-supervised learning method is used to pre-train the backbone network on large-scale unlabeled data; S4, model integration: Integrate the pre-trained backbone network with the feature fusion module and the decoupled detection head to build a complete object detection model; S5. Model training: Use labeled data to conduct supervised training on the integrated model to optimize detection performance; S6. Model compression: Through channel pruning and channel distillation technology, the number of model parameters is reduced to achieve lightweight model; S7. Model deployment: Deploy the compressed model to the drone to achieve real-time power insulator defect detection.
9. The application method of a lightweight target detection model integrating attention mechanism, optimized feature fusion method and self-supervised learning according to claim 8, characterized in that: In S5, the training process adopts a hybrid loss function that combines supervised detection loss and self-supervised contrastive loss.
10. The application method of a lightweight target detection model integrating attention mechanism, optimized feature fusion method and self-supervised learning according to claim 9, characterized in that: In S6, a global pruning threshold is set, all channels with scaling factors below the threshold are pruned, and the model performance is restored through fine-tuning training.
Citation Information
Patent Citations
SAR target identification method based on self-supervised learning and knowledge distillation
CN116912585A
Electrical equipment defect detection method and system based on self-supervised learning
CN117671587A
Lightweight multi-task small target detection algorithm based on adaptive pyramid and multi-stage path aggregation
CN120070857A
Deep-learning-based method for small target detection in unmanned aerial vehicle scenario
WO2024108857A1
Cited By
Method and system for big data acquisition and annotation
CN120707987A
Self-adaptive lightweight dense network mineral classification system and method
CN120976641A
Target lightweight detection method and system based on attention feature enhancement
CN121147496A
Lightweight neural network model construction method
CN121638342A