Power transmission component image noise reduction system based on improved CLIP

By improving the CLIP model and the ViT-CNN collaborative encoding and decoding architecture, the clarity and recognizability issues of the transmission component image denoising system under noise interference were solved, efficient denoising and key information retention were achieved, and the operation and maintenance efficiency and intelligent inspection capabilities of the transmission network were improved.

CN120612249APending Publication Date: 2025-09-09国网山西省电力有限公司阳泉供电分公司
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510759837.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-09-09

AI Technical Summary

Technical Problem

The existing transmission component image denoising system based on improved CLIP is susceptible to noise interference during the acquisition, transmission and storage processes, which affects the image clarity and recognizability. Traditional methods are not effective in removing complex noise and have difficulty retaining key information.

Method used

An image feature extraction algorithm based on the improved CLIP model is adopted, combined with the semantic embedding modules of ViT and CLIP. Through non-overlapping block processing, multi-head self-attention mechanism and convolutional neural network, a ViT-CNN collaborative encoding and decoding architecture is constructed to perform image denoising.

Benefits of technology

It significantly improves image clarity and reliability, effectively removes complex noise, retains key details and structural information of power transmission components, and supports intelligent inspection and fault diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure BDA0005439935450000031
    Figure BDA0005439935450000031
  • Figure BDA0005439935450000055
    Figure BDA0005439935450000055
Patent Text Reader

Abstract

A power transmission component image noise reduction system based on an improved CLIP comprises a power transmission component image acquisition module, a preprocessing module, a feature extraction module, a feature quality evaluation module, a noise reduction module and an output module, the power transmission component image acquisition module is used for acquiring a power transmission component image, and the preprocessing module is used for preprocessing the acquired original image and outputting the preprocessed image to the output module. The feature extraction module is used for carrying out feature extraction on the preprocessed image, the feature quality evaluation module is used for evaluating feature quality, the noise reduction module is used for inputting the extracted features into a pre-trained noise reduction model for noise reduction processing, and the output module is used for outputting the denoised power transmission part image. According to the power transmission component image noise reduction system based on the improved CLIP, an image feature extraction algorithm based on an improved CLIP model is proposed to perform feature extraction on a power transmission component image, and an image noise reduction algorithm based on ViT-CNN collaborative coding and decoding is proposed to perform noise reduction processing on the power transmission component image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of power system image feature extraction and image noise reduction, and in particular to an improved CLIP-based power transmission component image noise reduction system. Background Art

[0002] Power system image feature extraction technology is a technology that analyzes image information of key components of transmission lines and substation equipment. It aims to solve the problems of low recognition accuracy and blurred target boundaries in complex backgrounds caused by traditional image feature extraction methods. In a transmission component image denoising system based on improved CLIP, power system image feature extraction technology combines the image block segmentation capability of ViT with the semantic embedding module of improved CLIP to achieve distortion-invariant feature extraction of transmission component images, thereby improving the robustness and interpretability of feature representation under complex working conditions.

[0003] Power system image denoising technology is a key image processing technology used to improve the quality of power transmission images, suppress noise interference, and enhance target recognition capabilities. It aims to address the problems of traditional image denoising algorithms' insufficient ability to preserve high-frequency details and unstable denoising effects in strong interference environments. In a transmission component image denoising system based on improved CLIP, power system image denoising technology combines a cross-modal feature alignment mechanism with an image semantics-guided reconstruction strategy, which can effectively enhance the model's ability to distinguish noise and target boundaries. By introducing a semantically guided visual restoration mechanism, it achieves high-precision restoration of noisy images and structural preservation of component details, breaking through the application bottleneck of traditional image denoising technology in power system scenarios.

[0004] However, the existing power transmission component image denoising system based on improved CLIP has the following problems: the system is susceptible to various noise interferences during the image acquisition, transmission and storage process, which affects the image clarity and recognizability. In addition, although the traditional power transmission component image denoising method has certain effects, it has limitations, such as poor effect on complex noise removal and easy blurring of image edges and details. Traditional methods are difficult to effectively deal with complex noise situations and perform poorly in retaining key information. Summary of the Invention

[0005] The present invention aims to provide a power transmission component image denoising system based on an improved CLIP, so as to address the problems of the existing power transmission component image denoising system based on an improved CLIP proposed in the background art, namely, that the system is susceptible to various noise interferences during image acquisition, transmission and storage, thereby affecting image clarity and recognizability; and that traditional power transmission component image denoising methods, while effective to a certain extent, have limitations, such as poor complex noise removal effect and easy blurring of image edges and details. Traditional methods are difficult to effectively cope with complex noise situations and perform poorly in retaining key information.

[0006] To achieve the above objectives, the present invention provides the following technical solutions: a power transmission component image denoising system based on an improved CLIP model, comprising a power transmission component image acquisition module, a preprocessing module, a feature extraction module, a feature quality assessment module, a denoising module and an output module, characterized in that: the power transmission component image acquisition module is used to acquire high-quality original image data of the power transmission equipment in real time, providing standardized, high-fidelity input image data for subsequent denoising and diagnosis processes; the preprocessing module is used for image normalization and discretization processing, normalizing the pixel values ​​of the image and relatively discretizing the normalized image to highlight the image features; the feature extraction module proposes image features based on the improved CLIP model The extraction algorithm uses the improved CLIP model to extract features from the preprocessed image to obtain dense features with distortion invariance and content relevance; the feature quality assessment module is used to quantitatively analyze the reliability and effectiveness of the feature map output by the feature extraction module to ensure that the denoising model can perform image restoration based on high-quality features; the denoising module proposes an image denoising algorithm based on ViT-CNN collaborative encoding and decoding to efficiently remove complex noise in the transmission component image, and ensures denoising of the transmission component image by constructing an encoder and decoder collaborative processing architecture; the output module is used to receive the denoised image output by the decoder in the denoising module and output the final denoised transmission component image.

[0007] Preferably, the power transmission component image acquisition module efficiently and stably acquires the power transmission component image data and performs quality control, thereby ensuring that high-definition, low-distortion original data input is provided for subsequent noise reduction processing.

[0008] Preferably, the preprocessing module performs normalization and relative discretization preprocessing operations on the original image by setting the grayscale levels to 32, 64, and 128 respectively, which can effectively improve the quality of the input image and lay a solid foundation for subsequent feature extraction and noise reduction processing.

[0009] Preferably, the feature extraction module proposes an image feature extraction algorithm based on the improved CLIP model, uses the improved CLIP model to extract features from the preprocessed image, divides the image into small blocks of 16×16 pixels, and uses Vision Transformer (ViT) to process each image block, combined with the semantic feature extraction module of CLIP to extract dense features with distortion invariance and content relevance.

[0010] Preferably, the image feature extraction algorithm based on the improved CLIP model is specifically as follows: First, the input image is divided into non-overlapping blocks using the improved ViT to ensure that the subsequent feature extraction module can accurately capture key area information and enhance the algorithm's detail retention capability. Assuming the input power transmission component image size is H×W, the image block size is defined as M×N. The total number of image blocks after segmentation is calculated as follows:

[0011]

[0012] Among them, H and W represent the height and width of the image respectively, M and N represent the height and width of the image block respectively, It is expressed as rounded down, and the edge part is filled with zero to ensure the integrity of the image block. The set of image blocks after segmentation is {P1,P2,…,P K},P1,P2,…,P K They are represented as independent image block data, and each image block is independently input into the subsequent module. Then, the image block is globally encoded by the improved ViT model to capture the spatial dependency of the transmission component image, providing a basis for the subsequent semantic feature fusion. i Flattened to a vector Among them, X i Represents the input tensor of the i-th image block, i represents the number of image blocks, C represents the number of channels, Represented as a real number field and adding positional encoding Among them, E i Denote the position encoding vector of the i-th image block, D denotes the embedding dimension, and the encoding input is obtained. The specific formula is expressed as:

[0013] Z i =Linear(X i )+E i

[0014] Among them, Z i It is represented as the embedded feature vector, and Linear is represented as a linear projection operation. The image block is flattened and mapped to a high-dimensional embedding space. The feature transformation is performed through a multi-head self-attention mechanism and a feedforward network. The specific formula is expressed as follows:

[0015] Z′ i =MSA(Z i )+Z i

[0016] Z” i =FFN(Z i )+Z

[0017] Among them, Z' iIt is represented as the feature vector after MSA processing, and MSA is represented as the multi-head self-attention mechanism function, Z” i It is represented as the final ViT feature vector after FFN processing. FFN is represented as a feedforward network, which contains two fully connected layers and activation functions. The output ViT global feature matrix is:

[0018] F ViT ={Z”1,Z”2,…,Z” K}

[0019] Among them, F ViT Represented as the ViT global feature matrix, which contains the feature vectors of all image blocks, Z”1, Z”2,…, Z” K are represented as the final ViT feature vectors after independent FFN processing. Secondly, an improved CLIP image encoder is constructed to extract the semantic features of the image block, enhance the understanding of the content relevance of the transmission component image, improve the feature robustness, and convert the image block P i The image encoder of the improved CLIP model is input, and semantic features are extracted through the convolution layer and attention mechanism. The specific formula is expressed as:

[0020] S i = CLIP-Encoder(P i )

[0021] in, It is represented as the semantic feature vector of the i-th image block, L is the semantic feature dimension, and CLIP-Encoder is represented as the improved CLIP image encoder function, which consists of a convolutional neural network and an attention mechanism to extract semantic features from image blocks. The specific formula for outputting the semantic feature matrix is:

[0022] F CLIP ={S1,S2,…,S K}

[0023] Among them, F CLIP Represented as the output semantic feature matrix, S1,S2,…,S K It is represented as the semantic feature vector of each independent image block. Then, the ViT global feature and CLIP semantic feature are fused through dynamic weight allocation. Combining the advantages of both, dense features with distortion invariance and content relevance are generated. The ViT feature weight α and CLIP feature weight β are defined to satisfy α+β=1. The features of each image block are weightedly fused. The specific formula is expressed as:

[0024] F Fusion,i =α·F ViT,i +β·F CLIP,i

[0025] Among them, α and β are determined by experimental optimization, F Fusion,i Expressed as the fusion feature vector of the i-th image block, F ViT,i It is represented as the global feature vector extracted by the improved ViT of the i-th image block, F CLIP,i It is represented as the semantic feature vector extracted by the improved CLIP of the i-th image block. Secondly, the fused features are mapped to a high-dimensional space to generate a dense feature representation suitable for the task of denoising the image of the power transmission component, which provides input for the subsequent denoising module. Finally, the fused features are nonlinearly transformed through the fully connected layer. The specific formula is expressed as:

[0026] F Output,i =σ(ω·F Fusion,i +b)

[0027] in, Expressed as a weight matrix, is represented as the bias term, σ is represented as the activation function ReLU, R is represented as the target feature dimension, F Output,i The final output dense feature matrix of the i-th image block can be directly input into the denoising module for noise suppression.

[0028] Preferably, the feature quality assessment module ensures that the features extracted during the noise reduction process have high information density, semantic relevance and anti-interference ability through a comprehensive quality assessment from information theory, semantic alignment, adversarial verification to physical rules, and can provide high-precision structured data for subsequent fault diagnosis.

[0029] Preferably, the denoising module proposes an image denoising algorithm based on ViT-CNN collaborative encoding and decoding. By constructing an encoder based on ViT and a decoder based on a convolutional neural network (CNN), the extracted features are input into a pre-trained denoising model for denoising processing, and a high-resolution denoised image is gradually restored.

[0030] Preferably, the image denoising algorithm based on ViT-CNN collaborative encoding and decoding is as follows: First, the feature map obtained after processing by the feature extraction module is converted into a serialized feature map. Where N is the number of blocks, D is the block embedding dimension, It is expressed as a real number field, and a multi-head self-attention mechanism is used to globally model the features. The specific formula is:

[0031]

[0032] Among them, MHA is represented by the multi-head self-attention mechanism function, and Concat is represented by the multi-head splicing function, which is used to splice the outputs of multiple heads along the feature dimension and integrate the attention information of different subspaces. Head1,…,Head hIt is represented by the output of each independent multi-head, which is obtained by weighted summation of the value matrix through the attention weight, h is represented by the number of multi-heads, W O Represented as a learnable output transformation matrix, Head i It is represented as the output of the i-th head, and Softmax is represented as a normalization function. The attention score of each row is normalized to generate a probability distribution representing the attention weight of each position to other positions. It is represented as the query matrix of the i-th head, which represents the attention requirement of each position on the global feature. It is represented as the key matrix of the i-th head, representing the features of each position as the candidate to be paid attention to, Expressed as a scaling factor to prevent the Softmax gradient from disappearing due to excessive dot product values. It is expressed as the key matrix of the i-th head. Then, a feedforward neural network is constructed to perform nonlinear transformation on the features. The specific formula is expressed as:

[0033] FFN(X)=GELU(XW1+b1)W2+b2

[0034] Among them, FFN represents the output result of the feedforward neural network, GELU represents the nonlinear activation function, W1 represents the weight matrix of the first fully connected layer, b1 represents the bias term of the first fully connected layer, W2 represents the weight matrix of the second fully connected layer, b2 represents the bias term of the second fully connected layer. Secondly, the low-dimensional features output by the ViT encoder are enhanced by the multi-layer perceptron to improve the expression ability of the features, and random regularization is introduced to eliminate overfitting and ensure the robustness of the algorithm in complex noise environments. The features output by the ViT encoder are Enter the MLP module, the specific formula is expressed as:

[0035] H1=Dropout(GELU(F ViT W a +b a )),H2=Dropout(H1W b +b b )

[0036] Among them, H1 represents the intermediate feature after the first fully connected layer, GELU activation and Dropout, Dropout represents the random inactivation regularization function, W a It is represented as the weight matrix used to map the input features to the intermediate hidden layer, W b Denoted as the weight matrix used to map the intermediate hidden layer back to the target dimension, b a It represents the offset of adjusting the linear transformation, H2 represents the intermediate feature after the second fully connected layer, GELU activation and Dropout, and b bIt is expressed as the offset used to adjust the final output. Then, the CNN decoder is used to gradually restore the high-resolution image. The multi-scale features of the encoder are fused through the jump connection mechanism to ensure that the key details of the transmission component image, including edges and textures, are preserved during the denoising process. The decoder consists of five CNNUp modules, each of which contains transposed convolution upsampling and feature splicing. The specific formula is expressed as:

[0037] F up =DeConv 2×2 (F in ), stride=2, kernel size=2×2

[0038] F cat =Concat(F up ,F skip )

[0039] Among them, F in Denoted as input feature map, F up Represented as the output feature map after upsampling, DeConv 2×2 It is represented as a transposed convolution operation with a kernel size of 2×2, F cat It is represented as the concatenated feature map, Concat is represented as the channel concatenation operation, F skip It is represented as the lateral output of the corresponding level of the encoder. Secondly, the specific formula of the features after splicing by CNNBlock is expressed as:

[0040] F out =ReLU(BN(Conv 3×3 (F cat )))

[0041] Among them, F out Represents the output feature map of the current CNNBlock block, ReLU represents the nonlinear activation function, BN represents the normalization function, which standardizes the feature map of the convolution output to accelerate training and improve the robustness of the model. 3×3 It is represented as a standard convolution operation with a kernel size of 3×3. Finally, the ViT-CNN collaborative encoding and decoding mechanism is used to integrate global semantic information and local detail features, and finally output a high-definition denoised image, which provides a reliable basis for the status monitoring of power transmission components. The feature F enhanced by MLP is converted into MLP Input CNN decoder, upsample step by step and fuse encoder features. The specific formula is expressed as:

[0042]

[0043] in, Represented as the decoded feature map output by the i-th level CNNUp decoding block, CNNUp (i)It is represented as the operation function of the i-th level CNNUp decoding block, including upsampling, jump splicing and CNNBlock processing. Represented as the feature map output by the i-1th level CNNUp decoding block, as the input of the i-th level, It is represented as the lateral feature map output by the i-th level CNNDown module on the encoder side, which is passed to the i-th level of the decoder through jump connection. The final decoded output is activated by Sigmoid to generate the denoised image. The specific formula is expressed as:

[0044]

[0045] Among them, I denoised is the output image after denoising, σ is the Sigmoid activation function, Conv 1×1 Represented as a standard convolution operation with a kernel size of 1×1, It is represented as the output feature map of the 5th layer (i.e. the last layer) of the decoder, which contains high-resolution feature information restored after ViT-CNN collaborative encoding and decoding. It combines the global semantics of ViT encoding with the local details of CNN decoding. The algorithm can effectively remove complex noise while maintaining the structural integrity of the image, providing a reliable data foundation for intelligent inspection of power equipment.

[0046] Preferably, the output module ensures that the denoised transmission component image meets industrial standards in detail retention, semantic accuracy and system compatibility through core mechanisms of detail enhancement, format compatibility and abnormal fault tolerance, while forming a closed-loop process of "evaluation-feedback-optimization" to continuously improve the robustness of the system in complex scenarios.

[0047] Compared with the prior art, the present invention has the following beneficial effects:

[0048] 1. Feature extraction module An image feature extraction algorithm based on the improved CLIP model is proposed. First, the feature image is divided into non-overlapping blocks through the improved visual Transformer, so that the detail information in each area can be retained to the greatest extent, avoiding the problem of dilution and loss of edge details in traditional whole-image encoding, thus laying a solid foundation for subsequent refined noise reduction; secondly, the improved ViT model is applied to each image block, and the multi-head self-attention mechanism is used to capture the global spatial dependencies between image blocks, so that the understanding of the overall structure and key features of the power transmission components is more comprehensive and hierarchical, so that the real texture and noise components can be more accurately distinguished during the noise reduction process; at the same time, the improved CLIP image encoder further extracts the high-dimensional semantic features of the image blocks by fusing the local feature extraction capability of the convolutional neural network and the global context attention capability of the attention mechanism, thereby strengthening the robust representation of the semantic information of the material, shape and surface defects of the power transmission components; then, through the dynamic weight The redistribution strategy weightedly fuses the ViT global features with the CLIP semantic features, taking into account both distortion invariance and content relevance, ensuring stable capture of low-frequency structures while highlighting high-frequency details, giving the fused features a natural ability to suppress noise. Next, the fused features are mapped to a higher-dimensional dense representation space using nonlinear mapping, achieving multi-scale and multi-angle fusion of feature representations, providing more discriminative input for the subsequent denoising module. Finally, the deep nonlinear transformation based on the fused features not only further compresses noise interference but also improves the accuracy of detail recovery. When processing images of actual power transmission components, the system can effectively remove noise including Gaussian noise and impulse noise, while maximally retaining and enhancing the subtle texture and edge information on the surface of the transmission components, thereby significantly improving the clarity and reliability of the denoised images, providing high-quality visual input for downstream tasks such as intelligent inspection and defect detection, effectively reducing false detection and missed detection rates, and improving the safety and efficiency of power transmission network operation and maintenance.

[0049] 2. The denoising module proposes an image denoising algorithm based on ViT-CNN collaborative encoding and decoding. This algorithm fully utilizes the advantages of Transformer in modeling global semantic information and the ability of CNN in extracting local structural features by deeply integrating ViT and CNN, thereby realizing collaborative denoising of global and local features. Specifically, the ViT encoder first divides the image into multiple image blocks, and models the global dependency between the image blocks through a multi-head self-attention mechanism, capturing deeper semantic connections at the feature expression level, and enhancing the model's ability to distinguish between noise features and target features. At the same time, the feedforward neural network and multi-layer perceptron are introduced to perform nonlinear transformation and multi-layer enhancement on the low-dimensional features output by the Transformer. Combined with the Dropout regularization strategy, it effectively alleviates the overfitting problem that may occur in the model training of high-noise samples, and can improve the generalization ability of the algorithm in different transmission component image scenarios. Then, the CNN decoder adopts a step-by-step upsampling and jump-connection mechanism to fuse the lateral dimensions of each scale on the encoder side. The ViT-CNN collaborative mechanism structurally integrates the global modeling capability of Transformer and the local reconstruction capability of CNN, while removing background noise, retains the key information of surface texture, boundary contour and structural changes of transmission components, improving the clarity and readability of the image. The high-definition denoised image can be directly used for downstream tasks such as transmission component status recognition, surface defect detection and target abnormality diagnosis, providing stable and reliable visual data support for power inspection. In summary, the integration of this algorithm into the improved CLIP transmission component image denoising system not only improves the quality and detail reproduction capability of the denoised image, but also enhances the system's adaptability to complex noise environments and its support for downstream intelligent detection tasks, providing a reliable image data foundation for intelligent inspection and online monitoring. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 It is a structural schematic diagram of the present invention; DETAILED DESCRIPTION

[0051] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0052] See also Figure 1 The present invention provides a power transmission component image denoising system based on improved CLIP, comprising a power transmission component image acquisition module, a pre-processing module, a feature extraction module, a feature quality assessment module, a denoising module and an output module, characterized in that: the power transmission component image acquisition module is used to acquire high-quality original image data of the power transmission equipment in real time, providing standardized and high-fidelity input image data for subsequent denoising and diagnosis processes; the pre-processing module is used for image normalization and discretization processing, normalizing the pixel values ​​of the image, and relatively discretizing the normalized image to highlight the image features; the feature extraction module proposes an image feature extraction algorithm based on the improved CLIP model, utilizing The improved CLIP model is used to extract features from the preprocessed image to obtain dense features with distortion invariance and content relevance; the feature quality assessment module is used to quantitatively analyze the reliability and effectiveness of the feature map output by the feature extraction module to ensure that the denoising model can perform image restoration based on high-quality features; the denoising module proposes an image denoising algorithm based on ViT-CNN collaborative encoding and decoding to efficiently remove complex noise in the transmission component image. By constructing a collaborative processing architecture of encoder and decoder, the denoising processing of the transmission component image is ensured; the output module is used to receive the denoised image output by the decoder in the denoising module and output the final denoised transmission component image.

[0053] See Figure 1 Furthermore, the transmission component image acquisition module efficiently and stably acquires transmission component image data and performs quality control, ensuring that high-definition, low-distortion raw data input is provided for subsequent noise reduction processing.

[0054] See Figure 1 Furthermore, the preprocessing module performs normalization and relative discretization preprocessing operations on the original image by setting the grayscale levels to 32, 64, and 128 respectively, which can effectively improve the quality of the input image and lay a solid foundation for subsequent feature extraction and noise reduction processing.

[0055] See Figure 1 Furthermore, the feature extraction module proposes an image feature extraction algorithm based on the improved CLIP model, uses the improved CLIP model to extract features from the preprocessed image, divides the image into small blocks of 16×16 pixels, and uses Vision Transformer (ViT) to process each image block. Combined with the semantic feature extraction module of CLIP, dense features with distortion invariance and content relevance are extracted.

[0056] See Figure 1Furthermore, the image feature extraction algorithm based on the improved CLIP model is specifically as follows: First, the input image is segmented into non-overlapping blocks using the improved ViT to ensure that the subsequent feature extraction module can accurately capture key area information and enhance the algorithm's detail retention capability. Assuming the input image size of the power transmission component is H×W and the image block size is defined as M×N, the total number of image blocks after segmentation is calculated as follows:

[0057]

[0058] Among them, H and W represent the height and width of the image respectively, M and N represent the height and width of the image block respectively, It is expressed as rounded down, and the edge part is filled with zero to ensure the integrity of the image block. The set of image blocks after segmentation is {P1,P2,…,P K},P1,P2,…,P K They are represented as independent image block data, and each image block is independently input into the subsequent module. Then, the image block is globally encoded by the improved ViT model to capture the spatial dependency of the transmission component image, providing a basis for the subsequent semantic feature fusion. i Flattened to a vector Among them, X i Represents the input tensor of the i-th image block, i represents the number of image blocks, C represents the number of channels, Represented as a real number field and adding positional encoding Among them, E i Denote the position encoding vector of the i-th image block, D denotes the embedding dimension, and the encoding input is obtained. The specific formula is expressed as:

[0059] Z i =Linear(X i )+E i

[0060] Among them, Z i It is represented as the embedded feature vector, and Linear is represented as a linear projection operation. The image block is flattened and mapped to a high-dimensional embedding space. The feature transformation is performed through a multi-head self-attention mechanism and a feedforward network. The specific formula is expressed as follows:

[0061] Z′ i =MSA(Z i )+Z i

[0062] Z” i =FFN(Z i )+Z i

[0063] Among them, Z' iIt is represented as the feature vector after MSA processing, and MSA is represented as the multi-head self-attention mechanism function, Z” i It is represented as the final ViT feature vector after FFN processing. FFN is represented as a feedforward network, which contains two fully connected layers and activation functions. The output ViT global feature matrix is:

[0064] F ViT ={Z”1,Z”2,…,Z” K}

[0065] Among them, F ViT Represented as the ViT global feature matrix, which contains the feature vectors of all image blocks, Z”1, Z”2,…, Z” K are represented as the final ViT feature vectors after independent FFN processing. Secondly, an improved CLIP image encoder is constructed to extract the semantic features of the image block, enhance the understanding of the content relevance of the transmission component image, improve the feature robustness, and convert the image block P i The image encoder of the improved CLIP model is input, and semantic features are extracted through the convolution layer and attention mechanism. The specific formula is expressed as:

[0066] S i = CLIP-Encoder(P i )

[0067] in, It is represented as the semantic feature vector of the i-th image block, L is the semantic feature dimension, and CLIP-Encoder is represented as the improved CLIP image encoder function, which consists of a convolutional neural network and an attention mechanism to extract semantic features from image blocks. The specific formula for outputting the semantic feature matrix is:

[0068] F CLIP ={S1,S2,…,S K}

[0069] Among them, F CLIP Represented as the output semantic feature matrix, S1,S2,…,S K It is represented as the semantic feature vector of each independent image block. Then, the ViT global feature and CLIP semantic feature are fused through dynamic weight allocation. Combining the advantages of both, dense features with distortion invariance and content relevance are generated. The ViT feature weight α and CLIP feature weight β are defined to satisfy α+β=1. The features of each image block are weightedly fused. The specific formula is expressed as:

[0070] F Fusion,i =α·F ViT,i +β·F CLIP,i

[0071] Among them, α and β are determined by experimental optimization, F Fusion,i Expressed as the fusion feature vector of the i-th image block, F ViT,i It is represented as the global feature vector extracted by the improved ViT of the i-th image block, F CLIP,i It is represented as the semantic feature vector extracted by the improved CLIP of the i-th image block. Secondly, the fused features are mapped to a high-dimensional space to generate a dense feature representation suitable for the task of denoising the image of the power transmission component, which provides input for the subsequent denoising module. Finally, the fused features are nonlinearly transformed through the fully connected layer. The specific formula is expressed as:

[0072] F Output,i =σ(ω·F Fusion,i +b)

[0073] in, Expressed as a weight matrix, is represented as the bias term, σ is represented as the activation function ReLU, R is represented as the target feature dimension, F Output,i The final output dense feature matrix of the i-th image block can be directly input into the denoising module for noise suppression.

[0074] See Figure 1 Furthermore, the feature quality assessment module ensures that the features extracted in the denoising process have high information density, semantic relevance and anti-interference ability through a comprehensive quality assessment from information theory, semantic alignment, adversarial verification to physical rules, and can provide high-precision structured data for subsequent fault diagnosis.

[0075] See Figure 1 Furthermore, the denoising module proposes an image denoising algorithm based on ViT-CNN collaborative encoding and decoding. By constructing an encoder based on ViT and a decoder based on convolutional neural network CNN, the extracted features are input into a pre-trained denoising model for denoising processing, and high-resolution denoised images are gradually restored.

[0076] See Figure 1 ,Furthermore, the image denoising algorithm based on ViT-CNN collaborative encoding and decoding is as follows: First, the feature map obtained after processing by the feature extraction module is converted into a serialized feature Where N is the number of blocks, D is the block embedding dimension, It is expressed as a real number field, and a multi-head self-attention mechanism is used to globally model the features. The specific formula is:

[0077]

[0078] Among them, MHA is represented by the multi-head self-attention mechanism function, and Concat is represented by the multi-head splicing function, which is used to splice the outputs of multiple heads along the feature dimension and integrate the attention information of different subspaces. Head1,…,Head h It is represented by the output of each independent multi-head, which is obtained by weighted summation of the value matrix through the attention weight, h is represented by the number of multi-heads, W O Represented as a learnable output transformation matrix, Head i It is represented as the output of the i-th head, and Softmax is represented as a normalization function. The attention score of each row is normalized to generate a probability distribution representing the attention weight of each position to other positions. It is represented as the query matrix of the i-th head, which represents the attention requirement of each position on the global feature. It is represented as the key matrix of the i-th head, representing the features of each position as the candidate to be paid attention to, Expressed as a scaling factor to prevent the Softmax gradient from disappearing due to excessive dot product values. It is expressed as the key matrix of the i-th head. Then, a feedforward neural network is constructed to perform nonlinear transformation on the features. The specific formula is expressed as:

[0079] FFN(X)=GELU(XW1+b1)W2+b2

[0080] Among them, FFN represents the output result of the feedforward neural network, GELU represents the nonlinear activation function, W1 represents the weight matrix of the first fully connected layer, b1 represents the bias term of the first fully connected layer, W2 represents the weight matrix of the second fully connected layer, b2 represents the bias term of the second fully connected layer. Secondly, the low-dimensional features output by the ViT encoder are enhanced by the multi-layer perceptron to improve the expression ability of the features, and random regularization is introduced to eliminate overfitting and ensure the robustness of the algorithm in complex noise environments. The features output by the ViT encoder are Enter the MLP module, the specific formula is expressed as:

[0081] H1=Dropout(GELU(F ViT W a +b a )),H2=Dropout(H1W b +b b )

[0082] Among them, H1 represents the intermediate feature after the first fully connected layer, GELU activation and Dropout, Dropout represents the random inactivation regularization function, W a It is represented as the weight matrix used to map the input features to the intermediate hidden layer, W bDenoted as the weight matrix used to map the intermediate hidden layer back to the target dimension, b a It represents the offset of adjusting the linear transformation, H2 represents the intermediate feature after the second fully connected layer, GELU activation and Dropout, and b b It is expressed as the offset used to adjust the final output. Then, the CNN decoder is used to gradually restore the high-resolution image. The multi-scale features of the encoder are fused through the jump connection mechanism to ensure that the key details of the transmission component image, including edges and textures, are preserved during the denoising process. The decoder consists of five CNNUp modules, each of which contains transposed convolution upsampling and feature splicing. The specific formula is expressed as:

[0083] F up =DeConv 2×2 (F in ), stride=2, kernel size=2×2

[0084] F cat =Concat(F up ,F skip )

[0085] Among them, F in Denoted as input feature map, F up Represented as the output feature map after upsampling, DeConv 2×2 It is represented as a transposed convolution operation with a kernel size of 2×2, F cat It is represented as the concatenated feature map, Concat is represented as the channel concatenation operation, F skip It is represented as the lateral output of the corresponding level of the encoder. Secondly, the specific formula of the features after splicing by CNNBlock is expressed as:

[0086] F out =ReLU(BN(Conv 3×3 (F cat )))

[0087] Among them, F out Represents the output feature map of the current CNNBlock block, ReLU represents the nonlinear activation function, BN represents the normalization function, which standardizes the feature map of the convolution output to accelerate training and improve the robustness of the model. 3×3 It is represented as a standard convolution operation with a kernel size of 3×3. Finally, the ViT-CNN collaborative encoding and decoding mechanism is used to integrate global semantic information and local detail features, and finally output a high-definition denoised image, which provides a reliable basis for the status monitoring of power transmission components. The feature F enhanced by MLP is converted into MLP Input CNN decoder, upsample step by step and fuse encoder features. The specific formula is expressed as:

[0088]

[0089] in, Represented as the decoded feature map output by the i-th level CNNUp decoding block, CNNUp (i) It is represented as the operation function of the i-th level CNNUp decoding block, including upsampling, jump splicing and CNNBlock processing. Represented as the feature map output by the i-1th level CNNUp decoding block, as the input of the i-th level, It is represented as the lateral feature map output by the i-th level CNNDown module on the encoder side, which is passed to the i-th level of the decoder through jump connection. The final decoded output is activated by Sigmoid to generate the denoised image. The specific formula is expressed as:

[0090]

[0091] Among them, I denoised is the output image after denoising, σ is the Sigmoid activation function, Conv 1×1 Represented as a standard convolution operation with a kernel size of 1×1, It is represented as the output feature map of the 5th layer (i.e. the last layer) of the decoder, which contains high-resolution feature information restored after ViT-CNN collaborative encoding and decoding. It combines the global semantics of ViT encoding with the local details of CNN decoding. The algorithm can effectively remove complex noise while maintaining the structural integrity of the image, providing a reliable data foundation for intelligent inspection of power equipment.

[0092] See Figure 1 Furthermore, the output module ensures that the denoised transmission component images meet industrial standards in terms of detail retention, semantic accuracy and system compatibility through the core mechanisms of detail enhancement, format compatibility and abnormal fault tolerance, while forming a closed-loop processing of "evaluation-feedback-optimization" to continuously improve the robustness of the system in complex scenarios.

[0093] In specific use, first, the transmission component image acquisition module is used to collect high-quality original image data of the transmission equipment in real time, providing standardized, high-fidelity input image data for subsequent noise reduction and diagnosis processes; second, the preprocessing module is used for image normalization and discretization, normalizing the pixel values ​​of the image and relatively discretizing the normalized image to highlight the image features; then, the feature extraction module proposes an image feature extraction algorithm based on the improved CLIP model, and uses the improved CLIP model to extract features from the preprocessed image to obtain dense features with distortion invariance and content correlation; second, the feature quality assessment module is used to quantitatively analyze the reliability and effectiveness of the feature map output by the feature extraction module to ensure that the denoising model can perform image restoration based on high-quality features; then, the denoising module proposes an image denoising algorithm based on ViT-CNN collaborative encoding and decoding to efficiently remove complex noise in the transmission component image. By constructing an encoder and decoder collaborative processing architecture, the denoising of the transmission component image is ensured; finally, the output module is used to receive the denoised image output by the decoder in the denoising module and output the final denoised transmission component image.

[0094] Although the present invention has been described in detail with reference to the aforementioned embodiments, it is still possible for those skilled in the art to modify the technical solutions described in the aforementioned embodiments, or to make equivalent substitutions for some of the technical features therein. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A power transmission component image denoising system based on improved CLIP, comprising a power transmission component image acquisition module, a preprocessing module, a feature extraction module, a feature quality assessment module, a denoising module and an output module, characterized in that: The power transmission component image acquisition module is used to collect high-quality original image data of the power transmission equipment in real time, providing standardized, high-fidelity input image data for subsequent noise reduction and diagnosis processes. The preprocessing module is used for image normalization and discretization, normalizing the pixel values ​​of the image and relatively discretizing the normalized image to highlight the image features. The feature extraction module proposes an image feature extraction algorithm based on the improved CLIP model, using the improved CLIP model to extract features from the preprocessed image to obtain dense features with distortion invariance and content relevance. The feature quality assessment module is used to quantitatively analyze the reliability and effectiveness of the feature map output by the feature extraction module to ensure that the noise reduction model can perform image restoration based on high-quality features. The denoising module proposes an image denoising algorithm based on ViT-CNN collaborative encoding and decoding to efficiently remove complex noise from images of power transmission components. By building a collaborative processing architecture of encoders and decoders, it ensures denoising of images of power transmission components. The output module is used to receive the denoised image output by the decoder in the noise reduction module and output a final denoised power transmission component image.

2. The power transmission component image denoising system based on improved CLIP according to claim 1, characterized in that: The power transmission component image acquisition module efficiently and stably acquires power transmission component image data and performs quality control, thereby ensuring that high-definition, low-distortion original data input is provided for subsequent noise reduction processing.

3. The power transmission component image denoising system based on improved CLIP according to claim 1, characterized in that: The preprocessing module performs normalization and relative discretization preprocessing operations on the original image by setting the grayscale levels to 32, 64, and 128 respectively, which can effectively improve the quality of the input image and lay a solid foundation for subsequent feature extraction and noise reduction processing.

4. The power transmission component image denoising system based on improved CLIP according to claim 1, characterized in that: The feature extraction module proposes an image feature extraction algorithm based on the improved CLIP model, uses the improved CLIP model to extract features from the preprocessed image, divides the image into small blocks of 16×16 pixels, and uses the Vision Transformer (ViT) to process each image block. Combined with the semantic feature extraction module of CLIP, dense features with distortion invariance and content relevance are extracted.

5. The power transmission component image denoising system based on improved CLIP according to claim 4, characterized in that: Based on the improved CLIP model, the image feature extraction algorithm first divides the transmission component image into non-overlapping blocks of preset size and uses zero padding on the edges to ensure the integrity of the blocks. Subsequently, through the improved ViT model, after adding position encoding to each image block, it passes through the multi-head self-attention and feedforward network in sequence to extract global spatial dependency features and form a ViT feature matrix. Then, an improved CLIP image encoder is constructed, and semantic features are extracted for the same image block through convolution and attention mechanism to obtain the CLIP feature matrix. The algorithm then proportionally fuses the two types of features based on dynamic weights, taking into account detail preservation and semantic relevance, and generates a distortion-invariant dense representation. Finally, the fused features are mapped to a higher dimension and nonlinearly transformed through a fully connected layer activated by ReLU to output the final feature vector suitable for the denoising module. The image feature extraction algorithm based on the improved CLIP model has improved detail capture and semantic understanding, and the feature robustness and denoising performance have been significantly enhanced.

6. The power transmission component image denoising system based on improved CLIP according to claim 1, characterized in that: The feature quality assessment module ensures that the features extracted during the noise reduction process have high information density, semantic relevance and anti-interference ability through a comprehensive quality assessment from information theory, semantic alignment, adversarial verification to physical rules, and can provide high-precision structured data for subsequent fault diagnosis.

7. The power transmission component image denoising system based on improved CLIP according to claim 1, characterized in that: The denoising module proposes an image denoising algorithm based on ViT-CNN collaborative encoding and decoding. By constructing an encoder based on ViT and a decoder based on a convolutional neural network (CNN), the extracted features are input into a pre-trained denoising model for denoising, and high-resolution denoised images are gradually restored.

8. The power transmission component image denoising system based on improved CLIP according to claim 7, characterized in that: The image denoising algorithm based on ViT-CNN collaborative encoding and decoding first serializes the feature maps output by the feature extraction module into blocks, and then uses the ViT encoder to implement global semantic modeling using multi-head self-attention and feedforward networks to capture long-range dependencies and generate a unified global feature representation. Subsequently, the ViT output features are nonlinearly enhanced using a multi-layer perceptron, and random dropout regularization is introduced to improve the feature representation capability and the robustness of the algorithm in complex noise environments. The decoder consists of a five-level CNNUp module, which fuses multi-scale features through transposed convolution upsampling and jump connection to ensure the recovery of key details of edges and textures. Each level of convolution block contains standard convolution, batch normalization and ReLU activation, which can accelerate convergence and improve model stability. Finally, the denoised image is output after 1×1 convolution and Sigmoid activation, realizing the collaborative fusion of global semantics and local details, providing high-quality and structurally complete image data for intelligent inspection of power equipment.

9. The power transmission component image denoising system based on improved CLIP according to claim 1, characterized in that: The output module ensures that the denoised images of power transmission components meet industrial standards in terms of detail retention, semantic accuracy, and system compatibility through core mechanisms such as detail enhancement, format compatibility, and exception tolerance. It also forms a closed-loop "evaluation-feedback-optimization" process to continuously improve the system's robustness in complex scenarios.

Citation Information

Cited By

  • Image pollution value monitoring method and device

    CN121305217A