Transformer-Based Multi-Task Encoder-Decoder Stereo Image Quality Assessment Method
By adopting a Transformer-based multi-task encoding-decoder network in stereo image quality evaluation, combined with a depth information-guided cross-view disparity fusion and multi-level feature feedback fusion module, the problem of difficulty in accurately simulating the characteristics of human vision systems in the prior art is solved, and a more efficient and accurate stereo image quality evaluation is achieved.
Patent Information
- Application Number
- CN202411519576.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-29
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2044-10-29
AI Technical Summary
The existing CNN-based reference-free stereoscopic image quality evaluation method is difficult to accurately simulate the characteristics of human vision systems, especially in terms of left and right view feature fusion and depth information processing.
Using a multi-task encoding-decoder network architecture based on Transformer, combining a cross-view disparity fusion transformer and a multi-level feature feedback fusion module guided by depth information, it realizes information processing from global to local and cross-view feature fusion.
It improves the accuracy and generalization ability of stereoscopic image quality evaluation, can better simulate the feedback mechanism of the human visual system, and obtain more accurate binocular features and pixel-level quality maps.
Smart Images

Figure CN119402630B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image quality assessment, and in particular to a multi-task encoding-decoding stereoscopic image quality assessment method based on Transformer. Background Art
[0002] In recent years, stereoscopic images have played a crucial role in fields such as virtual reality, medical imaging, and autonomous driving. However, during the processes of acquisition, transmission, encoding, and display of stereoscopic images, distortion will inevitably occur, affecting the visual experience of viewers and even potentially causing problems such as visual fatigue and motion sickness. Therefore, developing an objective, efficient, and accurate stereoscopic image quality evaluation method is a hot topic in the field of computer vision research.
[0003] As a technology for specific applications, stereoscopic image quality evaluation technology has broad application prospects in many fields. In virtual reality and augmented reality applications, evaluating the quality of stereoscopic images can effectively enhance the immersion of users and reduce visual fatigue; during the production of 3D movies and videos, the quality evaluation method can help creators monitor the image quality in real time to ensure that the output stereoscopic video has good visual effects; in autonomous driving and drone vision systems, evaluating the quality of stereoscopic images helps to ensure that the system can correctly understand the surrounding three-dimensional environment and improve the safety of driving and navigation. Therefore, stereoscopic image quality evaluation technology not only has important academic research value but also has great market potential in practical applications.
[0004] The research on stereoscopic image quality evaluation technology can be traced back to the late 20th century. The early research was mainly subjective evaluation methods, that is, letting viewers give visual scores to stereoscopic images. Although the perceived quality of subjective evaluation of stereoscopic images is true and reliable, it is time-consuming, laborious, and difficult to promote on a large scale, and its results are often affected by individual differences of evaluators. Therefore, subjective evaluation methods are difficult to meet the diverse application requirements. With the development of computer technology, automatic and efficient objective evaluation methods have gradually become the focus of research.
[0005] Based on the dependence on reference images, objective stereoscopic image quality evaluation methods can be divided into three categories: full reference (FR), reduced reference (RR), and no reference (NR). Because it is difficult to obtain high-quality reference images in real scenarios, the applicability of full reference and reduced reference methods is limited. The no reference stereoscopic image quality evaluation method is to evaluate the quality only based on the distorted stereoscopic image itself without any reference image, which is more general and flexible in practical applications.
[0006] Traditional no-reference evaluation methods mainly rely on manually extracted features, such as edge information, color distribution, and texture complexity, and then perform quality assessment through statistical methods or machine learning algorithms. Although these methods can predict image quality to a certain extent, they usually cannot accurately capture the complex features in images and show poor generalization. In recent years, the rise of deep learning technology has provided new ideas for no-reference stereoscopic image quality evaluation technology. Since humans are one of the receptors of stereoscopic image perception, the objective evaluation method of no-reference stereoscopic image quality should be based on the formation process of human perception. Deep learning-based methods extract more complex image features through Convolutional Neural Network (CNN), achieving a more accurate modeling of the Human Visual System (HVS), and becoming one of the mainstream methods in the field of no-reference stereoscopic image quality evaluation.
[0007] Given the excellent performance of CNN in computer vision tasks, researchers have introduced CNN into the no-reference stereoscopic image quality evaluation task. CNN-based algorithms can not only automatically extract key features in images but also achieve end-to-end overall parameter optimization. Since stereoscopic images are composed of a left-eye view and a right-eye view, to enhance the accuracy of evaluation, such algorithms usually adopt a dual-branch or multi-branch network structure to simultaneously learn the features of the left and right views, and then use methods such as connecting the left and right view features to fuse binocular features, and finally obtain the final features for quality prediction. The overall structure can be divided into three parts: feature extraction, feature fusion, and quality regression.
[0008] In 2022, Liu et al. proposed a two-stream interaction network. This network extracts monocular features through the left and right branches respectively, and asymmetric convolutional kernels are introduced in both branches to enhance the model's ability to extract local features of images. At the same time, the sum and difference information of monocular features are used as binocular features of stereoscopic images. After fusing the monocular features and binocular features, the quality score is regressed. In 2022, Hu et al. constructed a four-branch CNN evaluation method. This method extracts features from the left view, right view, binocular sum image, and binocular difference image respectively, and enhances the visual feature expression ability by weighted splicing of the four features. Finally, the quality score of the image is regressed through a fully connected layer. In 2023, Chang et al. proposed a coarse-to-fine feedback-guided network model considering the characteristics of the dominant eye. This network contains left and right branches, and in each branch, three dilated convolutions with different dilation rates are parallelly used to extract low, medium, and high spatial frequency information respectively. Then, through a binocular fusion module based on binocular disparity, the binocular fusion features corresponding to the left and right views are generated. Finally, through the information feedback guidance module, the guidance of binocular to monocular is completed between different scales.
[0009] Although the above algorithms have achieved good evaluation performance, the no-reference stereoscopic image quality assessment method based on CNN usually faces the problem of how to better simulate the characteristics of the human visual system. Specifically, when extracting the features of the left and right views, existing methods all adopt an encoder structure that extracts features from low levels to high levels, which is a process of information analysis from local to global. In fact, when given a distorted image, human observers first perform global perception and then further complete complex local perception. Secondly, the disparity information between the left and right views plays a crucial role in the process of human observers perceiving depth information and generating a sense of stereoscopy when viewing stereoscopic images. However, existing methods mainly rely on methods such as connection, fusion, addition, and subtraction to fuse the features of the left and right views, without exploring the role of disparity information; although some methods consider disparity information, they ignore the depth information provided by stereoscopic images.
[0010] To solve the above problems, the present invention proposes a multi-task encoding-decoding stereoscopic image quality assessment method based on Transformer. Summary of the Invention
[0011] The object of the present invention is to propose a multi-task encoding-decoding stereoscopic image quality assessment method based on Transformer to solve the problems raised in the background art.
[0012] To achieve the above object, the present invention adopts the following technical solutions:
[0013] The multi-task encoding-decoding stereoscopic image quality assessment method based on Transformer includes the following steps:
[0014] S1. Design a multi-task framework for evaluating image quality. The multi-task framework consists of two encoding-decoding branches. One encoding-decoding branch is used for evaluating the image quality of the left and right images, and the other encoding-decoding module is used for depth estimation;
[0015] S2. Given two input images F l and F r , use the encoder composed of VGG16 to extract multi-level monocular features, denoted as where i = 1, 2, 3;
[0016] S3. Send the multi-level features output by the encoder into the Global and Local Interaction Transformer (GLIT) module for enhancement;
[0017] S4. Input the enhanced multi-level features into the decoder for upsampling and progressive fusion operations, and output multi-level features denoted as where \(i = 1, 2, 3\);
[0018] S5. For the decoder in the encoder-decoder branch for image quality assessment, design a Depth information guided Cross-view Parallax Fusion Transformer (DCPFT) module to fuse the left and right monocular features at different levels respectively;
[0019] S6. Design a Multi-Level Feedback Fusion (MLFF) module. Input the three-level binocular features processed and output by the Depth information guided Cross-view Parallax Fusion Transformer (DCPFT) module into the Multi-Level Feedback Fusion (MLFF) module to achieve feedback fusion;
[0020] S7. After the binocular features at different levels are fused by the Multi-Level Feedback Fusion (MLFF) module, output the final quality feature \(F\). Further perform an upsampling operation to obtain a pixel-level quality map, and calculate the sum of the quality scores of all pixel points to obtain the final quality score.
[0021] Preferably, the Global and Local Interaction Transformer (GLIT) module in S3 is composed of two components: Global and Local Content Interaction Attention (GLCIA) and Gated Feed-Forward Network (GFFN). A skip connection is adopted between the two components. Taking the processing process of the Global and Local Interaction Transformer module on the encoder-decoder branch for left-view quality assessment as an example, its implementation process is expressed as:
[0022]
[0023] where \(f\) GLCLA represents the processing operation of the Global and Local Content Interaction Attention component; \(f\) GFFN represents the processing operation of the Gated Feed-Forward Network component; \(f\) ln represents the layer normalization operation; \(\oplus\) represents pixel-wise addition; represents the image feature obtained after layer normalization processing, where \(i = 1, 2, 3\); represents the image feature obtained after the Global and Local Content Interaction Attention and skip connection processing, where \(i = 1, 2, 3\); represents the graphic feature strengthened by the Global and Local Interaction Transformer module.
[0024] Preferably, the processing operations of the global and local content interaction attention (GLCIA) component specifically include the following:
[0025] Group the input features along the channel dimension to obtain two groups of features where H, W, and C represent height, width, and number of channels respectively. Input the features and into the global context branch and the local context branch respectively:
[0026] In the global context branch, replace the conventional self-attention mechanism with the efficient self-attention (ESA) mechanism. Pass through a convolutional layer composed of three 1×1 convolutions and a 3×3 depth convolution respectively to obtain a matrix After transposition, obtain and Perform the softmax function processing on along the columns, and then perform batch-wise matrix multiplication with The obtained features are multiplied batch-wise with processed by the softmax function along the rows, and finally pass through a 1×1 convolutional layer to obtain the final global context features Its implementation process is expressed as:
[0027]
[0028] where and are the softmax functions along the columns and rows respectively; represents batch-wise matrix multiplication; f conv1 represents the convolution operation with a 1×1 convolution kernel;
[0029] In the local context branch, use convolutional layers in parallel to extract local context features. Obtain enhanced features through a residual block composed of a 1×1 convolution and a 3×3 depth convolution. Then, through global average pooling (GAP), a multi-layer perceptron (MLP), and a sigmoid activation function, learn the weights and perform self-correction to obtain the local context features Its implementation process is expressed as:
[0030]
[0031]
[0032] Among them, f conv1-dwconv3-conv1 represents the convolutional layer operation composed of 1×1 convolution, 3×3 depth convolution, and 1×1 convolution; f sigmoid-MLP-GAP represents global max pooling, multi-layer perceptron, and sigmoid activation function operations; ⊙ represents element-wise multiplication;
[0033] Finally, the global context features and the local context features are merged to obtain richer image features after global and local content interaction attention The implementation process is expressed as:
[0034]
[0035] Among them, f concat represents the concatenation operation along the channel dimension.
[0036] Preferably, the processing operations of the gated feed-forward network (GFFN) component specifically include the following:
[0037] Improve the first fully connected layer and non-linear layer in the ordinary feed-forward network, and use a gated linear unit built by two linear projections (one of which is activated by the sigmoid function) to control the forward flow of feature information, enabling the network to focus on features with richer information and improving the module performance. The implementation process is expressed as:
[0038]
[0039] Among them, δ represents the sigmoid function; f dwconv3 represents the depth convolution operation with a convolution kernel of 3×3.
[0040] Preferably, the depth information-guided cross-view disparity fusion transformer (DCPFT) module is composed of a depth information-guided cross-view disparity fusion attention (Depth guided Cross-view Parallax Fusion Attention, DCPFA) component and a gated feed-forward network (Gated Feed-Forward Network, GFFN) component. During the process of fusing the left and right view features (i = 1, 2, 3) output by the decoder, it aggregates the image depth information (i = 1, 2, 3) obtained from the auxiliary task. The implementation process is expressed as:
[0041]
[0042]
[0043] Among them, f DCPFA represents the operation of the depth information-guided cross-view disparity fusion attention component; f GFFN represents the operation of the gated feed-forward network component; represents the left and right view features obtained through depth information-guided cross-view disparity fusion attention processing.
[0044] Preferably, the implementation process of the depth information-guided cross-view disparity fusion attention (DCPFA) component specifically includes the following contents:
[0045] The monocular context features extracted by the left and right branches in the image quality evaluation task pass through a convolutional layer composed of 1×1 convolution and 3×3 depth convolution to obtain matrices and respectively. To obtain the cross-view disparity information based on the left / right views, matrix and the depth information obtained from the depth estimation auxiliary task are added to obtain a query matrix with depth prior information. Matrix and the depth information obtained from the depth estimation auxiliary task are added to obtain a query matrix
[0046] Then, the key matrices and are respectively transposed to obtain matrices In matrices and matrix a Batch-wise matrix multiplication is performed, and then the softmax function is used along the epipolar line to obtain the cross-view disparity attention map Its implementation process is expressed as:
[0047]
[0048]
[0049]
[0050]
[0051] Among them, f softmax represents the softmax function along the epipolar line.
[0052] The value matrix is respectively multiplied by the obtained disparity attention map Ml→r / M r→l Perform batch - wise matrix multiplication and process it through a 1×1 convolutional layer to obtain cross - view disparity compensation information Therefore, the binocular disparity feature is obtained by the following process:
[0053]
[0054]
[0055] Finally, aggregate the binocular disparity feature and send it to the gated feed - forward network (GFFN) component to obtain the binocular feature of cross - view disparity fusion
[0056]
[0057] Preferably, the multi - level feature feedback fusion (MLFF) module is used to realize the layer - by - layer fusion of binocular features at adjacent levels, and its implementation process is expressed as:
[0058] Given adjacent binocular features and as the input features of the multi - level feature feedback fusion module. First, upsample to obtain the feature to match the resolution and channel dimension of the next - level feature ;
[0059] Then, introduce a grouped attention gate (GAG) module to fuse and The grouped attention gate (GAG) module uses the gating signal generated by the high - level feature to modulate the lower - level feature. First, perform 3×3 grouped convolution operations on the features and respectively, then use batch normalization (BN) for standardization and add them together; after the generated feature map is activated by the Relu function, it is processed by a 1×1 convolution and a batch normalization layer to obtain a single - channel feature map, and this single - channel feature map generates an attention coefficient through the sigmoid function, modulate through the attention coefficient, and then add it pixel - by - pixel with the feature to obtain the fused feature Similarly, after upsampling the feature to obtain the feature and After being processed by the same steps as above, the final image quality feature F is obtained, and its implementation process is expressed as:
[0060]
[0061]
[0062]
[0063]
[0064] where f GAG () represents the operation of the grouped attention gating module; f gconv3 represents the grouped convolution with a convolution kernel size of 3×3; f BN represents the batch normalization operation; f BN-conv1-Relu represents the activation of the Relu function, the 1×1 convolution operation, and the batch normalization layer processing; δ represents the sigmoid function.
[0065] Compared with the prior art, the present invention provides a multi-task encoding-decoder stereo image quality evaluation method based on Transformer, which has the following beneficial effects:
[0066] The present invention aims to accurately evaluate the quality of stereo images, and proposes a multi-task encoding-decoder network architecture based on Transformer, using image depth estimation as an auxiliary task; the present invention adopts the encoding-decoder network to better simulate the feedback mechanism of the human visual system, solves the problems existing in the feature extraction of the left and right views, and realizes the information processing process from global to local;
[0067] In addition, in order to prevent the semantic differences of image features between the encoder and the decoder, the multi-task encoding-decoder stereo image quality evaluation method based on Transformer proposed by the present invention combines the global interaction ability of Transformer and the local attention ability of CNN to enhance image features, introduces depth estimation as an auxiliary task, can provide higher-quality depth prior information, guides the cross-view disparity fusion based on Transformer, solves the problems existing in the feature fusion of the left and right views, and obtains more accurate binocular features;
[0068] The multi-task encoding-decoder stereo image quality evaluation method based on Transformer proposed by the present invention also uses the attention mechanism to perform feedback compensation on the multi-level binocular feature fusion, obtains the final pixel-level quality map, and thus accurately evaluates the image quality. This method provides a research scheme for the deep learning method of no-reference stereo image quality evaluation, and effectively promotes the development of visual tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0069] Figure 1 This is the overall structure diagram of the Transformer-based multi-task encoding-decoding stereo image quality assessment method mentioned in Embodiment 1 of the present invention;
[0070] Figure 2 This is the specific structure diagram of the Global and Local Interaction Transformer (GLIT) module mentioned in Embodiment 1 of the present invention;
[0071] Figure 3 This is the specific structure diagram of the Depth information guided Cross-view Parallax Fusion Transformer (DCPFT) module mentioned in Embodiment 1 of the present invention;
[0072] Figure 4 This is the specific structure diagram of the Multi-Level Feedback Fusion (MLFF) module just mentioned in Embodiment 1 of the present invention. Detailed implementation manners
[0073] To make the objectives, technical solutions, and advantages of the present invention clearer, the following further describes the embodiments of the present invention in detail.
[0074] Embodiment 1:
[0075] Please refer to Figure 1 , the present invention proposes a Transformer-based multi-task encoding-decoding stereo image quality assessment method, as Figure 1 shown, including the following content:
[0076] The present invention adopts a multi-task framework to evaluate image quality. This framework consists of two encoding-decoding branches, one for evaluating the image quality of the left and right images, and the other for depth estimation. Given two input left and right images and, first, the encoder composed of VGG16 is used to extract multi-level monocular features, marked as where i = 1, 2, 3; the multi-level features output by the encoder are respectively sent into the proposed Global and Local Interaction Transformer (GLIT) module for enhancement, and then input into the decoder for upsampling and step-by-step fusion operations, and the output multi-level features are marked as Among them, i = 1, 2, 3. In the decoder of the quality evaluation branch, a Depth information guided Cross-view Parallax Fusion Transformer (DCPFT) module is designed to fuse the left and right monocular features at different levels respectively. The binocular features at three levels output by the DCPFT are input into a Multi-Level Feedback Fusion (MLFF) module to achieve feedback compensation connection. For quality regression, as Figure 1 shown, the binocular features at different levels are fused by the MLFF into the final quality feature, and through an upsampling operation, a pixel-level quality map is obtained, and the sum of the quality scores of all pixel points is calculated to obtain the final perceived quality score. The specific details are described as follows.
[0077] Please refer to Figure 2 , the specific structure of GLIT is as Figure 2 shown. Similar to the structures of other transformers, GLIT consists of two core components: Global and Local Content Interaction Attention (GLCIA) and Gated Feed-Forward Network (GFFN), and skip connections are used between the two components. Taking the GLIT on the left view quality evaluation branch as an example, its entire process is defined as:
[0078]
[0079]
[0080]
[0081] Among them, f GLCLA represents the processing operation of the Global and Local Content Interaction Attention component; f GFFN represents the processing operation of the Gated Feed-Forward Network component; f ln represents the layer normalization operation; ⊕ represents element-wise addition; represents the image feature obtained after layer normalization processing, where i = 1, 2, 3; represents the image feature obtained after the Global and Local Content Interaction Attention and skip connection processing, where i = 1, 2, 3; represents the graphic feature obtained after being enhanced by the Global and Local Interaction Transformer module.
[0082] For GLCIA, first, for the input feature Group along the channel dimension to obtain two groups of features Among them, H, W, and C represent height, width, and the number of channels respectively. The feature and are respectively input into the global context branch and the local context branch.
[0083] In the global context branch, since the self-attention mechanism will cause the computational cost to grow quadratically with the spatial resolution of the feature map. Therefore, this module uses an efficient self-attention mechanism (ESA) with equivalent effect but only linear computational cost to replace it. Respectively pass through a convolutional layer composed of three 1×1 convolutions and 3×3 depth convolutions to obtain the matrix After transposition, it gets and Perform the softmax function processing on along the columns, and then perform batch-wise matrix multiplication with The obtained feature is multiplied batch-wise with processed by the softmax function along the rows and columns, and finally passes through a 1×1 convolutional layer to obtain the final global context feature The implementation process can be expressed as:
[0084]
[0085] Among them, and are the softmax functions along the columns and rows respectively; represents batch-wise matrix multiplication; f conv1 represents the convolution operation with a 1×1 convolution kernel.
[0086] In the local context branch, considering that the convolution operation is more suitable for capturing local features in images, this method uses convolutional layers in parallel to extract local context features. Pass through a residual block composed of 1×1 convolution and 3×3 depth convolution to obtain the enhanced feature Then, through global average pooling (GAP), multi-layer perceptron (MLP), and sigmoid activation function to learn the weights and perform self-correction to obtain the local context feature
[0087]
[0088]
[0089] Among them, f conv1-dwconv3-conv1 represents the convolutional layer operation composed of 1×1 convolution, 3×3 depthwise separable convolution, and 1×1 convolution; f sigmoid-MLP-GAP represents global max pooling, multi-layer perceptron, and sigmoid activation function operations; ⊙ represents element-wise multiplication.
[0090] Finally, the global context features and local context features are merged to obtain richer image features after global and local content interaction attention processing The implementation process is expressed as:
[0091]
[0092] Among them, f concat represents the concatenation operation along the channel dimension.
[0093] For the GFFN module, the first fully connected layer and non-linear layer in the ordinary feed-forward network are improved, and a gated linear unit built by two linear projections (one of which is activated by the sigmoid function) is used to control the forward flow of feature information, so that the network focuses on more informative features and improves the module performance. This process is implemented by the following formula:
[0094]
[0095] Among them, δ represents the sigmoid function; f dwconv3 represents the depth convolution operation with a 3×3 convolution kernel.
[0096] In summary, GLIT combines the characteristics of Transformer and CNN, parallelly extracts the global features and local features of the image, obtains higher-quality features, and also reduces the computational complexity.
[0097] In order to obtain accurate binocular features, the method of the present invention innovatively designs a depth information-guided cross-view disparity fusion transformer (DCPFT). During the fusion of the left and right view features (i = 1, 2, 3) output by the decoder, the image depth information (i = 1, 2, 3) is aggregated. It should be noted that the depth information is extracted by the auxiliary task depth estimation network, which can provide important geometric features, enhance the disparity correspondence between the left and right view features, and thus guide the acquisition of more accurate binocular features.
[0098] The structure of DCPFT is as Figure 3As shown, DCPFT also adopts two core components: Depth guided Cross-view Parallax Fusion Attention (DCPFA) and Gated Feed-Forward Network (GFFN). Its overall process is defined as:
[0099]
[0100]
[0101] Among them, f DCPFA represents the operation of the depth information-guided cross-view parallax fusion attention component; f GFFN represents the operation of the gated feed-forward network component; represents the left and right view features obtained after being processed by the depth information-guided cross-view parallax fusion attention.
[0102] Please refer to Figure 3 , the structure of DCPFA is as Figure 3 shown. The monocular context features extracted by the left and right branches in the image quality assessment task pass through a convolutional layer composed of 1×1 convolution and 3×3 depth convolution to obtain matrices and respectively. To obtain the cross-view parallax information based on the left / right views, first and the depth information obtained from the depth estimation auxiliary task are added together to obtain a query matrix with depth prior information. Then, the key matrices and are transposed to obtain Batch-wise matrix multiplication is performed between the matrices and , and then the softmax function is used along the epipolar line to obtain the cross-view parallax attention map This process is expressed as follows:
[0103]
[0104]
[0105]
[0106]
[0107] Among them, f softmax represents the softmax function along the epipolar line.
[0108] The value matrix Separate from the obtained disparity attention map M l→r / M r→l Perform Batch-wise matrix multiplication and process through a 1×1 convolutional layer to obtain cross-view disparity compensation information Therefore, the binocular disparity feature Is obtained by the following process:
[0109]
[0110]
[0111] Finally, aggregate the binocular disparity feature And send it to GFFN to obtain the binocular feature of cross-view disparity fusion
[0112]
[0113] Inspired by the hierarchical fusion processing and top-down feedback mechanism of the human brain visual cortex, this method designs a multi-level feedback fusion MLFF module to achieve the layer-by-layer fusion of binocular features at adjacent levels. The detailed structure of MLFF is as Figure 4 Shown. Given adjacent binocular features And As the input features of the MLFF module. First, perform upsampling on To obtain the feature To match the resolution and channel dimension of the next-level feature Then, a Grouped Attention Gate (GAG) module is introduced to fuse And This module uses the gating signal generated by the high-level feature to modulate the lower-level feature, thereby improving its accuracy.
[0114] Specifically, first perform 3×3 grouped convolution operations on the features And respectively, then use Batch Normalization (BN) for standardization and add them together. After the generated feature map is activated by the Relu function, it is processed by a 1×1 convolution and a BN layer to obtain a single-channel feature map. This single-channel feature map generates an attention coefficient through the sigmoid function, Modulate through the attention coefficient, and then add it to the feature Pixel by pixel to obtain the fused feature Similarly, after upsampling the feature The obtained feature And After being processed by the same steps as above, the final image quality feature F is obtained. The implementation process is as follows:
[0115]
[0116]
[0117]
[0118]
[0119] Among them, f GAG () represents the operation of the grouped attention gating module; f gconv3 represents the grouped convolution with a convolution kernel size of 3×3; f BN represents the batch normalization operation; f BN-conv1-Relu represents the activation of the Relu function, the operation of 1×1 convolution, and the processing of the batch normalization layer; δ represents the sigmoid function.
[0120] Example 2:
[0121] Based on Example 1 but with differences, the following uses a specific example to illustrate the proposed Transformer-based multi-task encoding-decoder stereo image quality assessment method of the present invention. The specific content is as follows.
[0122] 1. Benchmark datasets and evaluation metrics
[0123] The proposed Transformer-based multi-task encoding-decoder stereo image quality assessment method of the present invention is experimentally verified on four benchmark datasets. The datasets used include:
[0124] (1) LIVE 3D Phase I database: Abbreviated as the LIVE I database, it contains 20 reference stereo images and 365 distorted stereo images. All distorted stereo images are symmetrically distorted and include 5 distortion types: Fast-fading (FF), White Gaussian Noise (WN), JPEG compression (JPEG), JPEG2000 compression (JP2K), and Gaussian Blur (Gblur). The Differential Mean Opinion Score (DMOS) value is the perceptual quality label of the distorted stereo image.
[0125] (2) LIVE 3D Phase II Database: Abbreviated as LIVE II Database, it contains 8 reference stereoscopic images and 360 distorted stereoscopic images, including both symmetric and asymmetric distorted stereoscopic images, with 5 distortion types: FF, WN, JPEG, JP2K, and Gblur. The DMOS value is the perceptual quality label for the distorted stereoscopic images.
[0126] (3) Waterloo-IVC 3D Phase I Database: Abbreviated as WIVC I Database, it contains 6 reference stereoscopic images and 330 distorted stereoscopic images, including both symmetric and asymmetric distorted stereoscopic images. The distortion types include WN, Gblur, and JPEG. The Mean Opinion Score (MOS) value is the perceptual quality label for the distorted stereoscopic images.
[0127] (4) Waterloo-IVC 3D Phase II Database: Abbreviated as WIVC II Database, it contains 10 reference stereoscopic images and 460 distorted stereoscopic images, including both symmetric and asymmetric distortion types of stereoscopic images. The distortion types included are WN, Gblur, and JPEG. The MOS value is the perceptual quality label for the distorted stereoscopic images.
[0128] The present invention uses three common evaluation metrics for performance evaluation:
[0129] Pearson Linear Correlation Coefficient (PLCC), Spearman Rank Order Correlation Coefficient (SROCC), and Root Mean Square Error (RMSE). Among them, PLCC represents the linear dependence between the objective score and the subjective score, SROCC reflects the monotonicity of the prediction, and RMSE reflects the accuracy of the prediction. For PLCC and SROCC, the larger their values, the better the performance, while the trend of RMSE is opposite.
[0130] 2. Training Settings
[0131] The present invention adopts a training strategy based on image patches to achieve data augmentation. First, the stereo image is segmented into non-overlapping image patches of size 40×40, and then these non-overlapping image patches are fed into the network for training. During the training phase, the subjective scores (i.e., DMOS / MOS values) of the stereo images in the database and the depth maps are used as the labels for each image patch. In this experiment, the Pytorch framework is used to train the proposed model. The initial learning rate is set to 0.001, the batch size is set to 64, and the Euclidean loss and the Stochastic Gradient Descent (SGD) optimizer are used to adjust the network parameters.
[0132] 3. Experimental Results
[0133] Table 1 Experimental results on the benchmark datasets LIVE I and LIVE II, with the best performance in bold
[0134]
[0135]
[0136] Table 2 Experimental results on the benchmark datasets WIVC I and WIVC II, with the best performance in bold
[0137]
[0138] Tables 1 and 2 show the quantitative evaluation results of the method proposed in the present invention and some of the latest stereo image quality assessment methods on four datasets, with the optimal metrics marked in bold. As can be seen from Tables 1 and 2, the method of the present invention has achieved the highest accuracy on all four datasets, demonstrating the excellent performance of the method of the present invention. This is due to the fact that the method of the present invention adopts an encoder-decoder network, which better simulates the feedback mechanism of the human visual system and realizes the information processing process from global to local during monocular feature extraction. At the same time, it combines the global interaction ability of the Transformer and the local attention ability of the CNN to enhance the image features, preventing the semantic differences of the image features between the encoder and the decoder from being misaligned. In addition, the method of the present invention introduces depth estimation as an auxiliary task, provides depth prior information, guides the cross-view disparity fusion, and obtains more accurate binocular features. The method also performs feedback compensation fusion on the multi-level binocular features to generate a pixel-level quality map, thereby more accurately evaluating the image quality.
[0139] Table 3 SROCC metrics for single distortion types on the benchmark datasets LIVE I and LIVE II, with the best performance in bold
[0140]
[0141]
[0142] Table 4 PLCC metrics for single distortion types on the benchmark datasets LIVE I and LIVE II, with the best performance bolded
[0143] In Tables 3 and 4, the comparison results of the method proposed in the present invention with some of the latest stereoscopic image quality assessment methods for different distortion types on LIVE I and LIVE II are provided, and the optimal values are marked in bold. This better demonstrates the accuracy of the proposed method in predicting the stereoscopic image quality degradation caused by different distortion types.
[0144] Table 5 Results of ablation experiments on the network architecture, with the best performance bolded
[0145]
[0146] Table 6 Results of ablation experiments on the main body module, with the best performance bolded
[0147]
[0148] Tables 5 and 6 conduct ablation experiments on the overall network framework of the method of the present invention and the functions of each module. The experimental results in Table 5 fully show the positivity of the introduction of the encoding-decoding structure and multi-tasking for the overall network. Table 6 proves the effectiveness of each main body module for the method.
[0149] It should be noted that in this invention patent, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations.
[0150] The above is the preferred implementation manner of the present invention. It should be pointed out that for those of ordinary skill in the art, without departing from the principle described in the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. A Transformer-based multi-task encoder-decoder stereo image quality assessment method, characterized in that: The following steps are involved: S1. Design a multi-task framework for image quality assessment, wherein the multi-task framework consists of two encoder-decoder branches, one of which is used for image quality assessment of left and right images, and the other encoder-decoder module is used for depth estimation; S2, given two input images F on the left and right l and F r , the encoder composed of VGG16 is used to extract multi-level monocular features, marked as Where i = 1, 2, 3; S3, sending the multi-level monocular features output by the encoder to the global and local interactive transformer modules for enhancement; S4, input the enhanced multi-level monocular features into the decoder for upsampling and step-by-step fusion operations, and output the multi-level features marked as Where i = 1, 2, 3; S5. For the decoder in the encoder-decoder branch for image quality assessment, a depth-guided cross-view disparity fusion transformer module is designed to fuse left and right monocular features of different levels respectively. S6. Design a multi-level feature feedback fusion module, and input the three-level binocular features processed and output by the cross-view disparity fusion transformer module guided by depth information into the multi-level feature feedback fusion module to realize feedback compensation connection; S7. The binocular features at different levels are fused by the multi-level feature feedback fusion module to output the final quality feature F. Further upsampling operation is performed to obtain a pixel-level quality map. The sum of the quality scores of all pixels is calculated to obtain the final quality score.
2. The Transformer-based multi-task encoder-decoder stereo image quality assessment method according to claim 1, characterized in that: The global and local interactive transformer module in S3 is composed of two components: global and local content interactive attention and gated feedforward network, and a jump connection is used between the two components; the global and local interactive transformer module processing process on the encoder-decoder branch for left view quality assessment is expressed as follows: Among them, f GLCLA represents the processing operations of the global and local content interaction attention components; f GFFN represents the processing operation of the gated feedforward network component; f ln Representation layer normalization operation; Indicates pixel-by-pixel addition; represents the image features obtained after layer normalization, where i = 1, 2, 3; represents the image features obtained after global and local content interactive attention and skip connection processing, where i = 1, 2, 3; Represents the image features obtained after being enhanced by the global and local interactive transformer modules.
3. The Transformer-based multi-task encoder-decoder stereo image quality assessment method according to claim 2, characterized in that: The processing operations of the global and local content interaction attention components specifically include the following: For input features Grouping along the channel dimension results in two sets of features Among them, H, W, and C represent the height, width, and number of channels respectively; and Enter into the global context branch and the local context branch respectively: In the global context branch, an efficient self-attention mechanism is used instead of the conventional self-attention mechanism. Through three convolution layers consisting of 1×1 convolution and 3×3 depth convolution, we get the matrix After transposing, we get and right Softmax function is performed along the column, and then Batch-wise matrix multiplication is performed between them, and the features obtained are processed by the softmax function along the row. Perform batch-wise matrix multiplication and finally pass through a 1×1 convolutional layer to obtain the final global context feature The implementation process is expressed as: in, and They are the softmax functions along the columns and rows respectively; represents Batch-wise matrix multiplication; f conv1 Indicates a convolution operation with a convolution kernel of 1×1; In the local context branch, convolutional layers are used in parallel to extract local context features. Enhanced features are obtained through a residual block consisting of 1×1 convolution and 3×3 depth convolution. Then, the weights are learned through global maximum pooling, multi-layer perceptron and sigmoid activation function, and self-correction is performed to obtain local context features. The implementation process is expressed as: Among them, f conv1-dwconv3-conv1 represents a convolutional layer operation consisting of a 1×1 convolution, a 3×3 depthwise convolution, and a 1×1 convolution; f sigmoid-MLP-GAP Represents global maximum pooling, multi-layer perceptron and sigmoid activation function operations; ⊙ represents pixel-by-pixel multiplication; Finally, the global context feature and local context features Merge to obtain richer image features after global and local content interactive attention processing The implementation process is expressed as: Among them, f concat Represents a concatenation operation along the channel dimension.
4. The Transformer-based multi-task encoder-decoder stereo image quality assessment method according to claim 2, characterized in that: The processing operations of the gated feedforward network component specifically include the following: The first fully connected layer and nonlinear layer in the ordinary feedforward network are improved, and a gated linear unit constructed by two linear projections is used to control the forward flow of feature information, so that the network focuses on features with richer information and improves module performance. The implementation process is expressed as follows: Among them, δ represents the sigmoid function; f dwconv3 Indicates a depthwise convolution operation with a convolution kernel of 3×3.
5. The Transformer-based multi-task encoder-decoder stereo image quality assessment method according to claim 1, characterized in that: The depth information guided cross-view disparity fusion transformer module consists of a depth information guided cross-view disparity fusion attention component and a gated feedforward network component, which processes the left and right view features output by the decoder. In the fusion process, the image depth information obtained by the auxiliary task is aggregated Where i = 1, 2, 3, the implementation process is expressed as: Among them, f DCPFA represents the operation of the cross-view disparity fusion attention component guided by depth information; f GFFN represents the operation of a gated feedforward network component; represents the left and right view features obtained by cross-view disparity fusion attention processing guided by depth information; Representing the binocular features of cross-view disparity fusion.
6. The Transformer-based multi-task encoder-decoder stereo image quality assessment method according to claim 5, characterized in that: The implementation process of the depth information guided cross-view disparity fusion attention component specifically includes the following contents: Monocular context features extracted by the left and right branches in the image quality assessment task After the convolution layer consisting of 1×1 convolution and 3×3 depth convolution, the matrices are obtained respectively. and In order to obtain cross-view disparity information based on the left / right view, the matrix And the depth information obtained by the depth estimation auxiliary task Add together to get the query matrix with deep prior information The matrix And the depth information obtained by the depth estimation auxiliary task Add together to get the query matrix with deep prior information Then the key matrix and Transpose and get the matrix In the matrix and matrix Batch-wise matrix multiplication is performed between them, and then the softmax function is used along the epipolar line to obtain the disparity attention map across views. The implementation process is expressed as: Among them, f softmax represents the softmax function along the epipolar line; The value matrix Respectively with the obtained disparity attention map M l→r / M r→l Perform batch-wise matrix multiplication and process it through a 1×1 convolution layer to obtain cross-view parallax compensation information Therefore, the binocular disparity feature Obtained by the following process: Finally, the binocular disparity feature Aggregate and send to the gated feedforward network component to obtain binocular features of cross-view disparity fusion 7. The Transformer-based multi-task encoder-decoder stereo image quality assessment method according to claim 6, characterized in that: The multi-level feature feedback fusion module is used to realize the layer-by-layer fusion of binocular features at adjacent levels, and its implementation process is expressed as follows: Given adjacent binocular features As the input features of the multi-level feature feedback fusion module, first, Upsampling to get features To match the next level of features resolution and channel dimension; then, a grouped attention gating module is introduced to fuse and The grouped attention gating module uses the gating signal generated by high-level features to modulate the lower-level features. and Perform a group convolution operation with a convolution kernel size of 3×3, and then use batch normalization to standardize and add and merge; After the generated feature map is activated by the Relu function, it is processed by 1×1 convolution and batch normalization layer to obtain a single-channel feature map. The single-channel feature map generates the attention coefficient through the sigmoid function. Modulated by the attention coefficient and then combined with the feature Add pixel by pixel to get fusion features Similarly, features After upsampling, the features are obtained and After the same steps as above, the final image quality feature F is obtained, and its implementation process is expressed as: Among them, f GAG () indicates the grouped attention gating module operation; f gconv3 represents a group convolution with a kernel size of 3×3; f BN represents batch normalization operation; f BN-conv1-Relu represents Relu function activation, 1×1 convolution operation and batch normalization layer processing; δ represents the sigmoid function.
Citation Information
Patent Citations
Intelligent video fault analysis and pre-warning system of transformer substation
CN105323565A
Three-dimensional image quality evaluation method based on information exchange fusion network
CN110634130A