A power transmission image compression quality evaluation method based on multi-channel fusion
By employing a multi-channel fusion-based transmission image compression quality assessment method, which combines visual Transformer and multi-scale semantic feature enhancement network, the shortcomings of existing image quality assessment and compression control methods are addressed, enabling efficient and accurate image quality assessment and compression optimization in a transmission panoramic platform.
Patent Information
- Application Number
- CN202511261388.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-05
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2045-09-05
AI Technical Summary
Existing image quality assessment and compression control methods are difficult to meet the requirements of efficiency and reliability in power transmission panoramic platforms. They cannot effectively combine global and local features, lack semantic guidance mechanisms, resulting in low accuracy of compressed image quality assessment, and are prone to losing local structure and edge information during multi-scale processing and fusion.
A multi-channel fusion transmission image compression quality assessment method is adopted. Through a multi-channel feature processing module, a multi-scale attention fusion module, and an adaptive quality assessment module, a visual Transformer, a multi-scale semantic feature enhancement network, and a CNN network are used, combined with self-attention and cross-scale attention mechanisms to generate spatial position offset and attention weight maps, and dynamically adjust compression parameters to optimize image quality.
It achieves improved performance in adaptive image quality assessment under different compression ratio scenarios, enhances the ability to perceive image compression distortion, improves the accuracy and robustness of image quality assessment, takes into account both overall and local quality, and optimizes compression efficiency.
Smart Images

Figure CN120747717B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of smart power grids, and relates to a power transmission image compression quality evaluation method based on multi-channel fusion. BACKGROUND
[0002] The operation monitoring of power transmission lines increasingly relies on multi-source visual data such as images and videos. Due to the complex environment, large spatial span and frequent patrol tasks of power transmission lines, the platform needs to collect, transmit and analyze massive image data every day to ensure the safe and stable operation of the power system. Under this background, how to efficiently compress and accurately evaluate the quality of these images has become one of the key technologies for the intelligent development of the power transmission system.
[0003] Power transmission line images are usually collected by devices such as helicopters, drones and ground shooting devices. The images not only contain target regions such as towers, wires and insulators, but also contain a large amount of redundant backgrounds such as sky and vegetation. In the intelligent management and control of the power transmission panoramic platform, the clarity and texture details of the images are crucial for the monitoring of power transmission lines and power transformation equipment. If the images are directly compressed and transmitted, it is easy to cause distortion and blur in the key regions, affecting subsequent target recognition and defect detection. Moreover, when facing massive data processing and high-resolution image monitoring, today's systems still have bottlenecks in storage capacity and real-time processing capability. To ensure effective monitoring, compressed images need to maintain high quality while reducing data volume, so that operation and maintenance personnel can quickly and accurately identify equipment abnormalities.
[0004] Image quality evaluation plays an important role in the field of image compression. However, in the application of image adaptive compression, existing quality evaluation techniques have the following shortcomings: first, traditional methods only rely on local pixel features or single semantic features, making it difficult to fully represent the quality information of adaptive compression images; second, there is a lack of semantic guidance mechanism, existing methods fail to fully utilize global semantic information, only focusing on pixel-level distortion, making it difficult to adapt to adaptive compression scenarios; third, there is insufficient interaction of multi-scale features, existing methods fail to effectively combine global and local features, making it difficult to accurately analyze the interaction of different levels of features after adaptive compression. In addition, most methods assign the same weight to local regions, failing to dynamically adjust the contribution of key regions in overall quality evaluation. Furthermore, in the task of image adaptive compression quality evaluation, existing methods often rely on a single type of feature (such as pixel domain or frequency domain information), failing to fully exploit the complementarity between multiple features, resulting in insufficient perception of compression quality in complex image scenarios. At the same time, existing methods mostly use simple feature concatenation or weighted fusion strategies, lacking deep modeling of multi-scale and location information, making it difficult to effectively represent the quality changes of images at different compression levels. In addition, local structures and edge information in compressed images are easily lost in the multi-scale processing and fusion process, further reducing the accuracy of quality evaluation.
[0005] In summary, the existing image quality evaluation and compression control method is still difficult to meet the efficiency and reliability requirements of image intelligent management in the power transmission panoramic platform. Therefore, there is an urgent need for an image quality evaluation and compression optimization method combining local quality guidance and adaptive compression regulation mechanism to realize high compression rate, strong perceptual consistency and deployable image processing capability, serving actual scenes such as power line inspection, remote monitoring and fault warning. SUMMARY
[0006] The technical scheme of the present application is used to solve the problem of how to enhance the adaptive evaluation performance of image quality under different compression rate scenarios.
[0007] The present application solves the above technical problems by the following technical scheme:
[0008] The present application provides a power transmission image compression quality evaluation method based on multi-channel fusion, comprising:
[0009] S1, pre-processing the input compressed image and reference image;
[0010] S2, feature extraction of the image from three different dimensions of global semantic features, multi-scale semantic features and local detail features through a multi-channel feature processing module, generating spatial position offsets through fusion of visual Transformer and multi-scale semantic feature enhancement network branches, guiding adaptive adjustment of CNN network branch features by deformable convolution, and aligning the features in spatial dimensions by an effective fusion module to obtain unified features after fusion;
[0011] S3, using a multi-scale attention fusion module, introducing self-attention and cross-scale attention mechanisms to strengthen the perception of distorted areas from global and local, improving the discriminant ability of the model to fused features, and then accurately evaluating the image compression quality;
[0012] S4, generating a quality score through an adaptive quality evaluation module, using a collaborative structure of a prediction branch and a spatial attention branch, generating an image block score map through grouped convolution, and introducing a resolution-based dynamic convolution path selection mechanism to generate an attention weight map, weighting and integrating the score map using the attention weight map, realizing fine perception of local quality differences of the image, and effectively improving the accuracy and robustness of image quality prediction;
[0013] S5, closed-loop optimization is performed by using the compression strategy optimization module, taking the local and overall quality scores output by the adaptive quality evaluation module as input, dividing the image blocks, determining the quality sensitive area according to the local and global mean difference, dynamically adjusting the compression parameters according to the quality sensitive area, introducing a joint optimization target to balance the quality and compression efficiency, updating the parameters for image compression, repeatedly evaluating the scores until the stop condition is met, and outputting the optimized compressed image, taking into account the overall and local quality, and improving the compression efficiency.
[0014] The method of the application proposes a multi-channel feature processing module (MCFPM), which effectively alleviates the spatial offset problem between multi-channel features by spatial alignment and effective fusion of multi-channel features; the module can guide different channel features to achieve dynamic alignment in space through the learned offset, reduce the information loss caused by spatial misalignment in the multi-channel feature fusion process, enhance the expression ability of the fused features, and provide more accurate multi-channel feature representation for subsequent quality evaluation. A multi-scale attention fusion module (MSAFM) is proposed, which can extract and fuse multi-channel features from different scales. After the multi-scale features are preliminarily processed by convolution operation, the corresponding attention weight map is generated by introducing channel attention and spatial attention mechanisms, and the feature information of different scales is further weighted and fused to strengthen the attention to the key compression distortion area and improve the accuracy and robustness of image quality evaluation. An adaptive quality assessment module (AQAM) is proposed, which constructs a multi-branch fusion structure, uses multi-scale convolution and grouped convolution to capture feature information of different scales and local regions of the image, and combines a dynamic weight generation mechanism to adaptively adjust the weight according to the image features and weightedly fuse the multi-channel fused features, thereby accurately evaluating the image compression quality.
[0015] Further, the multi-channel feature processing module includes three branches, namely a visual Transformer, a multi-scale semantic feature enhancement network and a CNN network; the working process is as follows:
[0016] 1) input a pair of reference images and distorted images into the three branches of the multi-channel feature processing module, wherein the visual Transformer branch divides the input images into multiple image blocks, converts them into feature vectors through linear projection and adds class tokens, then enters the Transformer block to capture long-distance dependency relationship and optimize feature representation, and is reshaped into reference feature maps and distorted feature maps; the CNN network branch uses ResNet to extract shallow feature maps to obtain reference feature maps and distorted feature maps; the output reference feature maps and distorted feature maps of the multi-scale semantic feature enhancement network branch are used for multi-channel fusion;
[0017] 2) An offset fusion module (Offset Fusion) is designed in the multi-channel feature processing module. The Offset Fusion module introduces a dual-channel offset estimation mechanism, respectively based on the global features of the visual Transformer and the local multi-scale difference features of the multi-scale semantic feature enhancement network to estimate the offset, and dynamically fuses the two types of offset information through a learnable weight fusion strategy to generate their own spatial position offsets, which are used to adjust the two-dimensional coordinate offset of the sampling position of the deformable convolution kernel;
[0018] 3) The features of the three branch channels after offset guidance and alignment will be input into the respective fusion modules, and after the output of the fusion modules, the features will be spliced to obtain the fused unified features.
[0019] Further, the calculation formula of the spatial position offset is as follows:
[0020] (1)
[0021] wherein, is a learnable weight parameter for regulating the contribution degree of the visual Transformer and the multi-scale semantic feature enhancement network to the fusion offset, represents the visual Transformer feature, represents the multi-scale semantic feature;
[0022] The formula of the deformable convolution is as follows:
[0023] (2)
[0024] wherein, represents the value of the output feature map at position , represents the position on the output feature map, is the input feature map, is the total number of sampling points of the convolution kernel, is the weight of the kth position of the convolution kernel, represents the kth sampling position, is the learnable offset of the kth sampling position, provided by the spatial position offset .
[0025] Further, the fused unified features are represented as follows:
[0026] (3)
[0027] wherein, represents the splicing operation of the feature maps from the three branches in the channel dimension to form the fused unified features , for quality assessment of subsequent modules; denote the feature maps of the visual Transformer branch, denote the feature maps of the multi-scale semantic feature enhancement network, denote the feature maps of the CNN network.
[0028] Further, the workflow of the multi-scale semantic feature enhancement network is as follows:
[0029] 1) input the reference image and the distorted image to the shared weight lightweight backbone network to extract multi-scale semantic features;
[0030] 2) adopt the gated local pooling strategy, and select the distortion-related features through the gated convolution before pooling;
[0031] 3) the mask features are generated through window average pooling and linear dimension reduction layer , i=1, 2, 3, 4, 5;
[0032] 4) multi-scale relative position encoding is introduced in the gated local pooling, and the relative position information is embedded in the local features through the learnable position vector, enhancing the network's modeling ability for spatial structure changes;
[0033] 5) for each level of feature , input into the self-attention module to model the global context for the local features and improve the discriminative ability of the features;
[0034] 6) a depth-guided cross-scale attention module is introduced, which takes different scale features as input, constructs a cross-scale attention matrix, and realizes efficient flow and adaptive fusion of feature information through cross-scale feature query and dynamic weighting;
[0035] 7) the enhanced feature groups output by the depth-guided cross-scale attention module are further input into the scale-aware pooling module to realize efficient integration and global feature description of multi-scale features.
[0036] Further, the gated convolution is represented as:
[0037] (6)
[0038] wherein, is the mask feature, wherein is a sigmoid activation function that limits the mask value to [0, 1], is a bottleneck convolution block, is a concatenation operation; concatenation result of the distorted image feature and the reference image feature in the channel dimension, element-wise absolute difference of the distorted image feature and the reference image feature.
[0039] Further, the working process of the multi-scale attention fusion module is as follows:
[0040] 1) The multi-channel feature processing module outputs the fused uniform feature Input, using gated local pooling, selecting features related to distortion through gated convolution, generating feature through window average pooling and linear dimension reduction layer;
[0041] 2) Introduce a learnable position encoding module to model the position perception in the spatial dimension of the mask feature output by the gated local pooling. The position encoding is obtained through training and learning, and is added element by element with the original feature map, so as to embed explicit position information;
[0042] 3) Perform ReLU activation on the fused feature to introduce non-linear change; use a 3x3 convolution to extract local context information and perform normalization through batch normalization; access the ReLU activation layer to generate enhanced feature representation for subsequent attention fusion process;
[0043] 4) In the attention module part, use parallel weighted fusion strategy for self-attention and cross-scale attention, design weight generation module, and fuse the adaptive fusion weight generated by the structure edge difference of the guided network to improve the perception sensitivity to image quality degradation areas.
[0044] Further, the working process of the weight generation module is as follows: use the structural difference map between the reference image and the distorted image as additional input, use the Sobel operator to obtain the structural difference map of the reference image and the distorted image, concatenate it with the fusion feature, and then adjust the channel through 1x1 convolution, stabilize the feature through batch normalization, and obtain the self-attention weight , cross-scale attention weight , .
[0045] Further, the adaptive quality assessment module is composed of a prediction branch and a spatial attention branch.
[0046] In the prediction branch, the residual fusion feature enters a 3x3 lightweight group convolution, calculates the score of each pixel in the feature map, and generates an image block quality score map , capturing quality-related information of different image blocks;
[0047] In the spatial attention branch, a dynamic convolution path selection mechanism based on image feature size is designed: according to the spatial resolution of the feature map... With the set threshold Make a judgment, among which, Defined as the spatial size of the input feature map The mean value to reflect the image scale, threshold Determined based on statistical analysis of image size distribution in the training set; if If the receptive field is large, a 5×5 convolutional kernel with a larger receptive field is selected to expand the receptive field and enhance the modeling ability for large-scale distortion; otherwise, a 3×3 convolutional kernel is used to maintain computational efficiency. The path selection mechanism is followed by batch normalization and a sigmoid activation function to generate an attention weight map. This is used to assign saliency weights to each image patch; subsequently, the attention weight map... Image patch quality score map Perform element-wise multiplication at the pixel level to obtain a weighted graph. The feature information is then further integrated through global average pooling to obtain a one-dimensional vector, which is then input into two fully connected networks FC1 and FC2 in sequence. The first fully connected network FC1 uses the GELU activation function, and the final output is the prediction quality score of the image.
[0048] Furthermore, the comprehensive loss function of the joint optimization objective is defined as:
[0049] (12)
[0050] in, To achieve the target quality score, The overall prediction score, For compression ratio, This is the initial compression ratio. and is the weighting coefficient, with a value ranging from [0,1], and These control the trade-off between quality deviation and compression efficiency, respectively.
[0051] The beneficial effects of this invention are as follows:
[0052] The method designs a multi-channel feature processing module (MCFPM), utilizes an Offset Fusion module, realizes dynamic alignment and effective fusion of multi-channel features by introducing a learnable offset, and fully excavates complementary information between different types of features. A multi-scale attention fusion module (MSAFM) is designed, which combines multi-scale feature extraction and attention mechanism, introduces learnable position encoding, ensures global perception ability, and enhances the sensitivity of the model to local structure changes. An adaptive quality assessment module (AQAM) is proposed, which constructs a residual connection and a multi-branch structure, uses multi-scale convolution and grouped convolution to capture features and local information of different scales, combines a dynamic weight generation mechanism, adaptively adjusts the weight according to the image features, and accurately assesses the image compression quality, so as to accurately judge the quality of the image in different compression levels in a complex scene. The above modules work together to improve the image compression distortion perception ability of the model, and effectively enhance the adaptive evaluation performance of the image quality in different compression rate scenes. BRIEF DESCRIPTION OF DRAWINGS
[0053] Figure 1 is a flowchart of the power transmission image compression quality evaluation method based on multi-channel fusion of the embodiment one of the present application;
[0054] Figure 2 is a system architecture diagram of the power transmission image compression quality evaluation method based on multi-channel fusion of the embodiment one of the present application;
[0055] Figure 3 is a flowchart of the MCFPM module in the system architecture diagram of the power transmission image compression quality evaluation method based on multi-channel fusion of the embodiment one of the present application;
[0056] Figure 4 is a structure diagram of the MSFENet in the MCFPM module in the system architecture diagram of the power transmission image compression quality evaluation method based on multi-channel fusion of the embodiment one of the present application;
[0057] Figure 5 is a GLP structure diagram in the MSFENet structure in the MCFPM module in the system architecture diagram of the power transmission image compression quality evaluation method based on multi-channel fusion of the embodiment one of the present application;
[0058] Figure 6 is a DGCSA structure diagram in the MSFENet structure in the MCFPM module in the system architecture diagram of the power transmission image compression quality evaluation method based on multi-channel fusion of the embodiment one of the present application;
[0059] Figure 7 is a MSAFM module in the system architecture diagram of the power transmission image compression quality evaluation method based on multi-channel fusion of the embodiment one of the present application;
[0060] Figure 8 is the AQAM module in the system architecture diagram of the power transmission image compression quality evaluation method based on multi-channel fusion of the embodiment one of the application;
[0061] Figure 9 The score comparison chart of the power transmission image compression quality evaluation method based on multi-channel fusion of the embodiment one of the application. DETAILED DESCRIPTION
[0062] In order to make the purpose, technical scheme and advantages of the embodiments of the application clearer, the technical scheme of the embodiments of the application will be described clearly and completely below in combination with the embodiments of the application. Obviously, the described embodiments are part of the embodiments of the application, rather than all the embodiments of the application. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the application.
[0063] The technical scheme of the application will be further described below in combination with the drawings in the specification and specific embodiments:
[0064] Embodiment one
[0065] As shown in Figure 1 and Figure 2 , the embodiment of the application provides a power transmission image compression quality evaluation method based on multi-channel fusion, comprising the following steps:
[0066] Step 1, pre-processing the input compressed image and reference image, including but not limited to normalization, cropping and other operations, so that the image data is in a suitable range and format, facilitating subsequent processing.
[0067] Step 2, through the multi-channel feature processing module, the image is extracted from three different dimensions of global semantic features, multi-scale semantic features and local detail features, respectively. Through the spatial position offset generated by the fusion of visual Transformer and multi-scale semantic feature enhancement network branch, the deformable convolution is guided to adaptively adjust the CNN network branch features, and with the help of the effective fusion module, the features are aligned in the spatial dimension to obtain the unified features after fusion.
[0068] Multi-channel feature processing module (MCFPM): In image quality assessment (IQA) tasks, global semantic features, multi-scale semantic features, and local detail features are all crucial to performance. Compared to the most commonly used CNN networks, the Vision Transformer (ViT) can acquire global semantic features, while CNN networks mainly focus on local detail features. In addition, this invention proposes a novel multi-scale feature generation network structure, named Multi-Scale Semantic Feature Enhancement Network (MSFENet), for generating multi-scale semantic features.
[0069] like Figure 3 As shown, in order to effectively integrate multi-channel features, this invention innovatively constructs a multi-channel feature processing module (MCFPM). The network structure of the multi-channel feature processing module includes three branches: a visual Transformer, a multi-scale semantic feature enhancement network, and a CNN network.
[0070] The workflow of the multi-channel feature processing module is as follows:
[0071] A pair of reference images and distorted images The input images are fed into three branches of the multi-channel feature processing module, and their feature maps are extracted. The visual Transformer branch divides the input image into multiple image patches, transforms them into feature vectors through linear projection, adds class tokens, and then feeds them into the Transformer block to capture long-range dependencies, optimize the feature representation, and finally reshape them into a reference feature map. and distortion feature map The CNN network branch uses ResNet to extract shallow feature maps, obtaining reference feature maps. and distortion feature map Reference feature map of the output of the multi-scale semantic feature enhancement network branch. and distortion feature map Used for multi-channel fusion.
[0072] Considering the spatial alignment and attention region difference between different channel features, an Offset Fusion module is designed in the multi-channel feature processing module, unlike the single offset estimator used in existing methods, the Offset Fusion module designed by the application introduces a dual-channel offset estimation mechanism, respectively based on the global features of the visual Transformer and the local multi-scale difference features of the multi-scale semantic feature enhancement network for offset estimation, and through a learnable weight fusion strategy, the two types of offset information are dynamically fused to generate their own spatial position offsets, which are used to adjust the two-dimensional coordinate offset of the sampling position of the deformable convolution kernel, wherein the calculation formula of the spatial position offset is as follows:
[0073] (1)
[0074] wherein, is a learnable weight parameter for regulating the contribution of the visual Transformer and the multi-scale semantic feature enhancement network to the fusion offset, improving the flexibility and accuracy of the fusion offset, denotes the visual Transformer feature, denotes the multi-scale semantic feature.
[0075] The Offset Fusion module designed by the application applies a multi-channel semantic driven spatial guidance mechanism to the deformable convolution (Deformable Conv) sampling strategy of the CNN network branch, effectively solving the problem that the conventional deformable convolution can only estimate the offset by the backbone feature and is difficult to focus on the compression distortion area. The deformable convolution is different from the conventional convolution, and its sampling position can be adaptively adjusted, wherein the deformable convolution formula is:
[0076] (2)
[0077] wherein, denotes the value of the output feature map at position , denotes the position on the output feature map, is the input feature map, is the total number of sampling points of the convolution kernel, is the weight of the kth position of the convolution kernel, denotes the kth sampling position, is the learnable offset of the kth sampling position, provided by the spatial position offset .
[0078] After the offset guidance and alignment, the three-channel features will be input into the respective fusion module (FusionBlock), and after the output of the FusionBlock, the features are spliced to obtain the fused unified feature representation as follows:
[0079] (3)
[0080] wherein, denotes the concatenation operation of feature maps from three branches in the channel dimension to form the fused unified feature for the quality assessment of subsequent modules; denotes the feature map of the visual Transformer branch, denotes the feature map of the multi-scale semantic feature enhancement network, denotes the feature map of the CNN network.
[0081] As shown in Figure 4 , the structure diagram of the new multi-scale semantic feature enhancement network (MSFENet) proposed by the application is shown, the reference image and the distorted image are input into the shared weight lightweight backbone network to extract multi-scale semantic features. For each scale i, i = 1, 2, 3, 4, 5, corresponding to five different receptive field features, the input features are combined, and the combination formula is as follows:
[0082] (4)
[0083] (5)
[0084] wherein, is the distorted image feature, is the reference image feature, is the concatenation operation, is the concatenation result of the distorted image feature and the reference image feature in the channel dimension, which fuses the original feature information of the two images through concatenation, is the element-wise absolute difference of the distorted image feature and the reference image feature, which is used to explicitly extract the feature difference information of the two images at the current scale.
[0085] As shown in Figure 5 , the Gated Local Pooling (GLP) strategy is adopted, and the distortion-related features are accurately selected through the gated convolution before the pooling. For the full reference task, the gated convolution is represented as:
[0086] (6)
[0087] wherein, is the mask feature, wherein is a sigmoid activation function, which limits the mask value to [0, 1], is a bottleneck convolution block, is a concatenation operation. To improve efficiency, a single-channel mask .
[0088] Mask feature Block AvgPool and linear dimension reduction layer to generate features (i=1, 2, 3, 4, 5), and D is the dimension of the reduced features. Multi-scale relative position encoding (MSRPE) is introduced in GLP, which embeds relative position information in local features through learnable position vectors, enhancing the network's ability to model spatial structure changes. Subsequently, for each level of feature , input into the self-attention module (SA) to model the global context for local features and improve the discriminability of the features.
[0089] As Figure 6 shown, to further exploit the complementarity between different scale features, the application introduces a deep guided cross-scale attention module (DGCSA). DGCSA takes different scale features (i=1, 2, 3, 4, 5) as input, constructs an inter-scale attention matrix, and realizes efficient flow and adaptive fusion of feature information through cross-scale feature query and dynamic weighting. Each scale feature can be adaptively strengthened or suppressed during the fusion process according to the information of other scale features, thereby improving the global perception ability of the overall feature. The enhanced feature group (i=1, 2, 3, 4, 5) output by the DGCSA module is further input into the scale-aware pooling module (SAP) to realize efficient integration and global feature description of multi-scale features.
[0090] Step 3, use the multi-scale attention fusion module to introduce self-attention and cross-scale attention mechanisms to strengthen the perception of distorted regions from global and local perspectives, improve the discriminability of the fusion features, and accurately assess the image compression quality.
[0091] Multi-scale attention fusion module (MSAFM): to further improve the discriminability of the fusion features, a multi-scale attention fusion module (MSAFM) is designed. The unified feature output by the multi-channel feature processing module is input into the module, and a gated local pooling (GLP) strategy is adopted to accurately select features related to distortion through gated convolution, and features are generated through window average pooling and linear dimension reduction layer. Figure 7as shown.
[0092] To enhance the spatial position perception ability of the features, the application adds a set of feature enhancement operations. Specifically, first, a learnable position encoding (Learnable Positional Encoding) module is introduced to model the position perception in the spatial dimension of the mask features output by the GLP. The position encoding is learned through training and is added to the original feature map element by element, thereby embedding explicit position information and making up for the insufficient expression of position information by the convolution structure. Then, the fused features are activated by ReLU to introduce nonlinear changes. Subsequently, a 3x3 convolution is used to extract local context information, and batch normalization (BN) is used for normalization processing. Finally, a ReLU activation layer is connected to generate enhanced feature representations for subsequent attention fusion processes.
[0093] In the attention module part, the application improves the combination mode of self-attention (SA) and cross-scale attention (CSA) and adopts a parallel weighted fusion strategy. To better serve the image quality assessment task, scaled dot-product attention is used as the basis of the attention module. Given a feature vector triple (query Q, key K, value V), the attention function first calculates the similarity between the query and the key vector, and then performs weighted summation on the value vector. Assuming Q , K , V , the attention output calculation formula is:
[0094] (7)
[0095] wherein, and represent the number of feature vectors, and represent the feature dimension.
[0096] After processing by the GLP module, different scale feature groups { } will be obtained. For the SA module, the features from other positions are aggregated to enhance , and the formula is as follows:
[0097] (8)
[0098] For the CSA module, the query, key, and value are also obtained through linear projection, and attention mechanism is applied to the features from the channel and spatial dimensions to mine key information of the features in the channel and space, and the formula is as follows:
[0099] (9)
[0100] On this basis, unlike the traditional average or maximum value fusion strategy, the application designs a weight generation (Weight Generation) module, and the fusion guidance network dynamically generates adaptive fusion weights based on structural edge differences, so as to improve the perception sensitivity to the image quality decline area. The mechanism uses the structural difference graph (such as the edge graph, the gradient residual graph) between the reference graph and the distortion graph as an additional input, uses the Sobel operator to obtain the structural difference graph of the reference graph and the distortion graph, splices the structural difference graph with the fusion feature, and then adjusts the channel through 1x1 convolution, stabilizes the feature through batch normalization (BN), and obtains the self-attention weight through Sigmoid activation , cross-scale attention weight , . The output of the self-attention module and the output of the channel-space attention module are weighted and fused, that is:
[0101] (10)
[0102] The weighted fusion strategy combines the advantages of structural perception and semantic guidance, ensures that the model can dynamically focus on the area with the greatest impact on image quality during multi-scale feature fusion, and significantly improves the scoring accuracy and human eye consistency of the subsequent adaptive quality assessment module.
[0103] Step 4, generate a quality score through the adaptive quality assessment module, adopt a prediction branch and a spatial attention branch in cooperation with a structure, generate an image block score map through grouped convolution, and introduce a resolution-based dynamic convolution path selection mechanism to generate an attention weight map, and use the attention weight map to weightedly integrate the score map, realize fine perception of local quality differences of the image, and effectively improve the accuracy and robustness of image quality prediction;
[0104] Adaptive quality assessment module (AQAM): Considering that each pixel in the deep feature map corresponds to a different image block of the input image and contains rich spatial information, the traditional spatial pooling method (such as maximum pooling and average pooling) will lose information and ignore the relationship between image blocks when obtaining the final quality score. Therefore, the application introduces an adaptive quality assessment module (Adaptive Quality Assessment Module, AQAM), which is composed of a prediction branch and a spatial attention branch in cooperation, so as to realize image quality assessment more consistent with human eye perception, as shown in Figure 8 .
[0105] input feature map , batch normalization (BN) layer and ReLU activation function, to extract local features and normalize feature distribution. Then, residual connection (Residual Add) is introduced to preserve the original feature information and facilitate gradient propagation for network training. In the prediction branch, the residual fusion features enter a 3x3 lightweight group convolution (3x3 Group) to calculate the score of each pixel in the feature map and generate a patch score map (Capturing the quality-related information of different image patches). In the spatial attention branch, to enhance the perception of different image regions for various distortions (such as compression artifacts, noise interference, and blur effects), the module designs a dynamic convolution path selection mechanism based on image feature size: according to the spatial resolution of the feature map and the set threshold , the judgment is made, where is defined as the mean (or area) of the spatial size of the input feature map to reflect the image scale, and the threshold is determined according to the statistical analysis of the image size distribution in the training set. If , a 5x5 convolution kernel with a larger receptive field is used to expand the receptive field and enhance the modeling ability for large-scale distortions; otherwise, a 3x3 convolution kernel is used to maintain computational efficiency; the above path is followed by batch normalization (BN) and Sigmoid activation function to generate an attention weight map , which is used to assign significance weights to each image patch. Subsequently, the attention weight map and the image patch quality score map perform element-wise multiplication operation at the pixel level to obtain a weighted map , and the calculation formula is as follows:
[0106] (11)
[0107] where represents the element-wise multiplication operation.
[0108] After further integrating feature information through global average pooling (GAP), a one-dimensional vector is obtained, which is input into two layers of fully connected network FC1 and FC2 in turn, where the first layer of fully connected network FC1 uses GELU activation function, and finally outputs the predicted quality score of the image.
[0109] Step 5, closed-loop optimization is performed by using the compression strategy optimization module, taking the local and overall quality scores output by the adaptive quality evaluation module as input, dividing the image blocks, determining the quality sensitive area according to the difference between the local and overall mean values, dynamically adjusting the compression parameters according to the quality sensitive area, introducing a joint optimization target to balance the quality and compression efficiency; update the parameters for image compression, repeatedly evaluate the scores until the stop condition is met, output the optimized compressed image, consider the overall and local quality, and improve the compression efficiency.
[0110] Compression strategy optimization module (CSOM): in order to realize the maximum improvement of compression efficiency under the premise of ensuring subjective quality in the process of image adaptive compression, the application designs a compression strategy optimization module (Compression Strategy Optimization Module, CSOM). The compression strategy optimization module takes the image block quality score map output by the adaptive quality evaluation module and the overall prediction score as input, dynamically adjusts the compression parameters (mainly including quantization factor, encoding mode, etc.), and forms a closed-loop quality optimization mechanism. The main steps are as follows:
[0111] Firstly, the image block quality score map is divided into multiple 8x8 small blocks, and the local mean and variance of each small block are calculated. If a local score is significantly lower than the global mean (lower than the set threshold θ, which is determined according to experimental statistics and combined with different image types), it is determined that this area is a quality sensitive area. In the subsequent compression process, the compression rate of the sensitive area is appropriately reduced to retain more texture details; the compression rate of the non-sensitive area is appropriately increased to optimize the compression ratio as a whole.
[0112] Secondly, in the overall quality control, a joint optimization target is introduced, that is, while ensuring that the quality of the compressed image is close to the target quality, the compression ratio is as stable as possible to avoid excessive compression rate fluctuation, so as to realize the optimal balance between quality and compression ratio. Specifically, the comprehensive loss function of the joint optimization target is defined as:
[0113] (12)
[0114] Wherein, is the target quality score, is the overall prediction score, is the compression ratio, is the initial compression ratio, and are weight coefficients, the value range is between [0, 1], and they respectively control the trade-off between quality deviation and compression efficiency.
[0115] In the above comprehensive loss function: the first term is the quality error term, which constrains the quality of the compressed image to be close to the target score; the second term is the compression rate change term, which constrains the deviation of the compression ratio from the initial ratio to maintain the stability of the compression rate; the design of this joint optimization goal is different from the traditional method of only minimizing a single quality loss or compression rate change. It simultaneously integrates both into a unified optimization framework, which can more comprehensively achieve the balance between quality and compression rate.
[0116] Finally, the updated compression parameters are applied to the next round of image compression, and the quality score is re-evaluated until the preset stopping condition (such as the score approaching the target and the compression rate change being stable) is met or the maximum number of iterations is reached, completing the closed-loop optimization process. The above improvement strategy not only takes into account the overall image quality, but also carefully considers the importance of local regions, improving the subjective visual experience of the final compressed image, while significantly improving the compression efficiency.
[0117] Test simulation
[0118] The present application realizes accurate perception of image compression distortion through the multi-channel feature cooperative extraction and offset guided dynamic fusion mechanism. In order to quantitatively reflect the performance of the method, the present application selects large public image quality datasets, including CSIQ, KADID-10k and LIVE datasets. The performance is compared with PSNR and DISTS. As shown in Table 1, the present application effectively improves the quality evaluation accuracy in complex scenes.
[0119] Table 1 Performance comparison of different evaluation methods on three datasets
[0120]
[0121] The multi-scale attention fusion module enhances the model's attention to key areas, and the adaptive quality evaluation module realizes comprehensive perception of local quality details through a multi-branch structure. In addition, the system introduces a closed-loop optimization mechanism, which can dynamically adjust the compression parameters according to the evaluation results, achieving the optimal balance between image quality and compression rate, and enhancing the adaptability and robustness of the system in practical applications such as image transmission and monitoring. Figure 9 It can be seen that for different compression levels of power line images, the score of the compressed image 1 is 0.9509, the score of the compressed image 2 is 0.9331, and the score of the compressed image 1 is 0.5512. The method of the present application can accurately evaluate the image score.
[0122] Example two
[0123] An electronic device includes a memory for storing a program supporting a processor to execute a power transmission image compression quality evaluation method based on multi-channel fusion in embodiment one, and the processor is configured to execute the program stored in the memory.
[0124] Embodiment three
[0125] A storage medium, a computer program is stored on the storage medium, when the computer program is run by a processor, the steps of the power transmission image compression quality evaluation method based on multi-channel fusion in embodiment one are executed.
[0126] The above embodiments are only used to illustrate the technical solutions of the present application, but not limit it; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that the technical solutions recorded in the foregoing embodiments can be modified, or some technical features can be replaced by equivalents; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for power transmission image compression quality assessment based on multi-channel fusion, characterized in that, The method comprises the following steps: S1, preprocessing the input compressed image and the reference image; S2, extracting features from the image in three different dimensions of global semantic features, multi-scale semantic features and local detail features through a multi-channel feature processing module, generating spatial position offsets through the fusion of visual Transformer and multi-scale semantic feature enhancement network branches, guiding the adaptive adjustment of CNN network branch features by deformable convolution, and aligning the features in the spatial dimension through an effective fusion module to obtain unified features after fusion; S3, using a multi-scale attention fusion module, introducing self-attention and cross-scale attention mechanisms to strengthen the perception of distorted areas globally and locally, and improving the discriminant ability of the model to fused features, and then accurately evaluating the image compression quality; S4, generating quality scores through an adaptive quality evaluation module, using a collaborative structure of a prediction branch and a spatial attention branch, generating an image block score map through grouped convolution, and introducing a resolution-based dynamic convolution path selection mechanism to generate an attention weight map, and using the attention weight map to weight and integrate the score map to realize fine perception of local quality differences of the image and effectively improve the accuracy and robustness of image quality prediction; S5, using a compression strategy optimization module for closed-loop optimization, taking the local and overall quality scores output by the adaptive quality evaluation module as input, dividing the image into small blocks, determining the quality sensitive area according to the local and global mean difference, and dynamically adjusting the compression parameters accordingly, and introducing a joint optimization target to balance quality and compression efficiency; Update the parameters for image compression, repeatedly evaluate the scores until the stop condition is met, output the optimized compressed image, and balance the overall and local quality to improve the compression efficiency. 2.The multi-channel fusion based power transmission image compression quality assessment method according to claim 1, wherein, The multi-channel feature processing module comprises three branches, namely visual Transformer, multi-scale semantic feature enhancement network and CNN network; the working process is as follows: 1) a pair of reference image and distorted image are input into three branches of the multi-channel feature processing module respectively, wherein the visual Transformer branch divides the input images into multiple image blocks, converts them into feature vectors through linear projection and adds class tokens, then enters the Transformer block to capture long-distance dependencies and optimize feature representation, and is reshaped into reference feature maps and distorted feature maps; the CNN network branch uses ResNet to extract shallow feature maps to obtain reference feature maps and distorted feature maps; the output reference feature maps and distorted feature maps of the multi-scale semantic feature enhancement network branch are used for multi-channel fusion; 2) In the multi-channel feature processing module, an Offset Fusion module is designed, which introduces a double-channel offset estimation mechanism based on the global features of visual Transformer and the local multi-scale difference features of multi-scale semantic feature enhancement network, and dynamically fuses the two types of offset information through a learnable weight fusion strategy to generate their own spatial position offsets for adjusting the two-dimensional coordinate offset of the deformable convolution kernel sampling position; 3) The features of the three branch channels after offset guidance and alignment are input into the respective fusion modules, and after the output, the features are spliced to obtain the unified features after fusion. 3.The method of claim 2, wherein, The calculation formula of the spatial position offset is as follows: (1) wherein, is a learnable weight parameter for regulating the contribution degree of the visual Transformer and the multi-scale semantic feature enhancement network to the fusion offset, denotes the visual Transformer feature, denotes the multi-scale semantic feature; The formula of the deformable convolution is as follows: (2) wherein, denotes the value of the output feature map at position , denotes a position on the output feature map, is an input feature map, is the total number of sampling points of the convolution kernel, is the weight of the k-th position of the convolution kernel, denotes the k-th sampling position, is a learnable offset for the k-th sampling position, provided by a spatial position offset . 4.The method of claim 2, wherein, The unified features after fusion are represented as follows: (3) wherein, denotes a concatenation operation on the channel dimension to form the fused unified feature from the three branches for quality assessment of the subsequent modules; denotes the feature map of the visual Transformer branch, denotes the feature map of the multi-scale semantic feature enhancement network, denotes the feature map of the CNN network. 5.The method of claim 2, wherein, The working process of the multi-scale semantic feature enhancement network is as follows: 1) input the reference image and the distorted image to a shared-weight lightweight backbone network to extract multi-scale semantic features; 2) A gated local pooling strategy is used to select the features related to distortion through gated convolution before pooling; 3) Mask features are passed through a window average pooling and linear dimension reduction layer to generate features , i = 1, 2, 3, 4, 5; 4) Multi-scale relative position encoding is introduced in the gated local pooling, which embeds relative position information in local features through learnable position vectors, enhancing the network's ability to model spatial structure changes; 5) For each level of features , input self-attention module from, to the local features to impose global context modeling, enhance the discriminant ability of the features; 6) A deep guided cross-scale attention module is introduced to different scale features For the input, an inter-scale attention matrix is constructed, and through cross-scale feature query and dynamic weighting, efficient flow and adaptive fusion of feature information are realized; 7) enhanced feature groups output by the depth-guided cross-scale attention module Further input into the scale perception pooling module to achieve efficient integration of multi-scale features and global feature description. 6.The method of claim 5, wherein, The gated convolution is represented as: (6) wherein, is a mask feature, wherein is a sigmoid activation function, limiting the mask values to [0, 1], is a bottleneck convolutional block, is a concatenation operation; is a concatenation result of the distorted image feature and the reference image feature in the channel dimension, is an element-wise absolute difference of the distorted image feature and the reference image feature. 7.The method of claim 1, wherein, The workflow of the multi-scale attention fusion module is as follows: 1) the multi-channel feature processing module outputs the fused uniform feature Input, using gated local pooling, selecting distortion-related features through gated convolution, generating features through window average pooling and linear dimension reduction layer ; 2) A learnable position encoding module is introduced to model the position perception in the spatial dimension of the mask feature output by the gated local pooling. The position encoding is learned through training and is added element by element to the original feature map to embed explicit position information; 3) The fused features are activated by ReLU to introduce non-linear changes; A 3x3 convolution is used to extract local context information, and batch normalization is used for normalization; Access the ReLU activation layer to generate enhanced feature representation for subsequent attention fusion process; 4) In the attention module part, parallel weighted fusion strategy is adopted for self-attention and cross-scale attention, and weight generation module is designed to dynamically generate adaptive fusion weight based on structural edge difference, improving the perception sensitivity to image quality degradation areas. 8.The method of claim 7, wherein, The weight generation module works as follows: the structural difference graph between the reference graph and the distortion graph is used as additional input, the structural difference graph of the reference graph and the distortion graph is obtained by using a Sobel operator, which is spliced with the fusion features, and then 1*1 convolution is used to adjust the channel, batch normalization is used to stabilize the features, and Sigmoid activation is used to obtain the self-attention weight , cross-scale attention weight , . 9.The method of claim 1, wherein, The adaptive quality assessment module is composed of a prediction branch and a spatial attention branch. In the prediction branch, the residual fusion features enter a 3x3 lightweight grouped convolution to calculate the score of each pixel in the feature map and generate a quality score map of the image block , capturing quality-related information of different image blocks; In the spatial attention branch, a dynamic convolution path selection mechanism based on image feature size is designed: according to the spatial resolution of the feature map... With the set threshold Make a judgment, among which, Defined as the spatial size of the input feature map The mean value to reflect the image scale, threshold Determined based on statistical analysis of image size distribution in the training set; if If the receptive field is large, a 5×5 convolutional kernel with a larger receptive field is selected to expand the receptive field and enhance the modeling ability for large-scale distortion; otherwise, a 3×3 convolutional kernel is used to maintain computational efficiency. The path selection mechanism is followed by batch normalization and a sigmoid activation function to generate an attention weight map. This is used to assign saliency weights to each image patch; subsequently, the attention weight map... Image patch quality score map Perform element-wise multiplication at the pixel level to obtain a weighted graph. The feature information is then further integrated through global average pooling to obtain a one-dimensional vector, which is then input into two fully connected networks FC1 and FC2 in sequence. The first fully connected network FC1 uses the GELU activation function, and the final output is the prediction quality score of the image. 10.The method of claim 1, wherein, The comprehensive loss function of the joint optimization target is defined as: (12) wherein, is the target quality score, is the overall prediction score, is the compression ratio, is the initial compression ratio, and is a weight coefficient, taking values in the range [0, 1], and they control the trade-off between quality bias and compression efficiency, respectively.
Citation Information
Patent Citations
Field intensive daylily pixel classification and picking information acquisition method
CN120544044A
Expression recognition method based on attention-modulated contextual spatial information
WO2023185243A1