Stereo image quality evaluation method based on dual-frequency interactive enhancement and binocular matching

Through a stereo image quality evaluation method based on dual-frequency interaction enhancement and binocular matching, a convolutional neural network is used to simulate the binocular interaction fusion process of the human visual system, which solves the problem of the existing technology failing to effectively simulate the human visual system and achieves more accurate stereo image quality evaluation.

CN119399176BActive Publication Date: 2025-10-03TIANJIN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411557848.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-04
Publication Date
2025-10-03
Estimated Expiration
2044-11-04

AI Technical Summary

Technical Problem

Existing stereo image quality evaluation methods fail to effectively simulate the binocular interactive fusion process of the human visual system, and do not fully consider the impact of high- and low-frequency information on image quality perception, resulting in evaluation results that are inconsistent with the subjective evaluation of the human eye.

Method used

A stereo image quality evaluation method based on dual-frequency interaction enhancement and binocular matching is adopted. By converting the left and right view image blocks into the frequency domain for high and low frequency decomposition, and using convolutional neural networks for feature extraction and fusion, the binocular matching and competition mechanism of the human visual system is simulated, and the alignment and attention mechanism of binocular visual features are combined to generate an objective image quality score.

Benefits of technology

The accuracy and consistency of stereoscopic image quality evaluation are improved, which can better reflect the characteristics of human visual perception and generate results that are highly consistent with the subjective evaluation of the human eye.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119399176B_ABST
    Figure CN119399176B_ABST
Patent Text Reader

Abstract

The present invention discloses a stereo image quality assessment method based on dual-frequency interactive enhancement and binocular matching, belonging to the technical field of stereo image quality assessment. The present invention comprises the following steps: dividing the left and right views of an original stereo image into non-overlapping image blocks as input samples of a network; converting the left and right view image blocks into the frequency domain using FFT and decomposing them by high and low frequencies; restoring them using IFFT to obtain high-frequency and low-frequency image blocks of the left and right views; and then extracting primary features of the high-frequency and low-frequency image blocks of the left and right views using a set of convolution and pooling operations; interactively enhancing the high-frequency and low-frequency primary features through a dual-frequency interactive attention module, and then feeding them into a dual-frequency reorganization module to obtain high-frequency and low-frequency weighted feature maps; simultaneously feeding the left and right view image blocks into a set of convolution and pooling operations to extract full-frequency primary features, and utilizing high and low frequency information for feature enhancement; performing binocular progressive registration and competitive selection on the enhanced full-frequency features through a binocular matching fusion module to obtain binocular fusion features; further extracting features from the binocular feature fusion through two sets of convolutions, splicing them with the high-frequency and low-frequency reorganized features of the left and right views, and then passing them through a fully connected layer to obtain an objective prediction score. The present invention can better simulate the binocular fusion and competition mechanism of the human visual system, utilize the influence of different spatial frequency information on quality perception, and improve the accuracy of the objective evaluation method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of stereo image quality evaluation, and in particular to a stereo image quality evaluation method based on dual-frequency interactive enhancement and binocular matching. Background Art

[0002] As people's lifestyles and entertainment options become increasingly diverse, they are no longer satisfied with flat image experiences, but are pursuing more immersive and realistic stereoscopic images. Compared to flat (2D) images, stereoscopic images (3D) can make viewers feel the sense of three-dimensionality while displaying image content, creating an immersive atmosphere. Therefore, they are widely used in many fields such as 3D movies, 3D games, and virtual reality. However, during the acquisition, compression, and transmission of stereoscopic images, the images will be distorted due to certain unavoidable factors, thus affecting the viewer's visual experience. Therefore, it is particularly important to construct a stereoscopic image quality perception model that is highly consistent with human evaluation criteria. In addition, stereoscopic image quality assessment (SIQA) can also be used as a part of the image processing system to provide optimization guidance for stereoscopic image super-resolution, segmentation, and other fields.

[0003] Generally speaking, IQA methods can be divided into two categories: subjective and objective. Subjective IQA methods involve human observers evaluating and scoring image quality based on their subjective perceptions and various image metrics under specific experimental conditions. Because images are ultimately received by humans, subjective evaluation methods based on human observers' opinions and scores provide a more reliable and reliable assessment of image quality. These methods also have a low technical barrier to entry and are considered the most effective. While subjective evaluation methods offer the most accurate assessment of image quality, they also suffer from several drawbacks. In practice, subjective evaluation results are easily influenced by factors such as the tester's psychological state, cognitive level, and surroundings. Therefore, subjective evaluation is rarely used in real-life scenarios. Objective image quality assessment methods, on the other hand, evaluate image quality by simulating the human visual system (HVS) and establishing an evaluation model. These methods effectively address the limitations of subjective evaluation and are more practical. Objective evaluation methods are further categorized by the degree of reference image involvement: full reference (FR), reduced reference (RR), and no reference (NR). Since ideal reference images cannot be obtained in many application scenarios, objective image quality evaluation methods without reference that do not require ideal reference images are more widely used.

[0004] Currently, in the field of SIQA, researchers have largely focused on NR-type stereo image quality assessment methods, and numerous NR-SIQA methods have been proposed. Early FR-SIQA methods were improvements on the classic FR-2DIQA method. Examples include Peak Signal to Noise Ratio (PSNR) and Structural Similarity Index (SSIM). However, these methods fail to reflect the characteristics of human visual perception and do not consider the influence of binocular disparity, making them inappropriate for direct application to stereo image quality assessment. Furthermore, some NR-SIQA methods rely on traditional manual feature extraction based on image features extracted from the human visual system (HVS) or natural scene statistics (NSS). These methods have a high design threshold, requiring a high level of expertise and extensive experience. However, they also have poor generalization capabilities and struggle to adapt to diverse distortion scenarios.

[0005] In recent years, driven by deep learning technology, numerous stereo binocular vision models based on convolutional neural networks (CNNs) have emerged, further advancing the development of NR-SIQA. However, existing CNN-based SIQA methods have neglected the impact of image spatial frequency on human perception of image quality. Most methods fail to consider the interplay between high- and low-frequency information and their differential contributions to image quality perception. Based on biological knowledge and human visual psychology, the human visual system processes high- and low-frequency information in the same image differently. Extracting features from the entire frequency range can encounter the problem of uneven information distribution in the spatial domain. In reality, visual information processing first extracts the overall contour features of the visual scene based on low-frequency information, which is then used to facilitate the extraction of detailed texture features corresponding to high-frequency information. On the other hand, strong edges tend to attract more attention and enhance the perception of surrounding smooth areas. These mechanisms demonstrate the complex interplay between high- and low-frequency information.

[0006] Furthermore, because SIQA evaluates stereo images, binocular visual features are often considered. However, most methods simply concatenate left and right features and perform quality regression, without any interaction or fusion between the left and right views. While some methods achieve left-right view interaction through methods such as differential summation of left and right view features or using Siamese networks, these methods clearly cannot effectively simulate the relatively complex binocular interaction and fusion process of the human visual system. According to biological knowledge and human visual psychology, the human visual system typically first fuses matching features from the left and right views and then employs a competitive mechanism to select and suppress mismatching features. Matching features between the left and right views is a key factor in stabilizing binocular perception. Aligning binocular features facilitates the integration of binocular information, resulting in a unified and coherent visual perception. By simulating this process, stable and accurate binocular visual features can be effectively extracted.

[0007] Based on the above content, the present invention proposes a stereo image quality evaluation method based on dual-frequency interactive enhancement and binocular matching. Summary of the Invention

[0008] The purpose of the present invention is to propose a stereoscopic image quality evaluation method based on dual-frequency interactive enhancement and binocular matching to solve the problem that the existing technology does not take into account the different high and low frequency sensitivities of the human eye to images with different distortion types and different semantic information when evaluating stereoscopic image quality.

[0009] In order to achieve the above object, the present invention adopts the following technical solutions:

[0010] The stereo image quality evaluation method based on dual-frequency interactive enhancement and binocular matching includes the following steps:

[0011] Step 1: Cut the left and right view original images into 32×32 non-overlapping image blocks. Use FFT to convert the left and right view image blocks into the frequency domain and decompose them into high and low frequencies. Use IFFT to restore the high and low frequency image blocks of the left and right views.

[0012] Step 2: Train a convolutional neural network based on dual-frequency interaction enhancement and binocular matching: The convolutional neural network consists of three main subnetworks: the left view dual-frequency interaction feature enhancement subnetwork, the right view dual-frequency interaction feature enhancement subnetwork, and the binocular matching fusion subnetwork. Specifically, it includes the following:

[0013] Since the left and right view subnetworks are symmetrical, only the implementation of the left view dual-frequency interaction feature enhancement subnetwork is described here. In this subnetwork, a set of convolution and pooling operations are used to extract primary feature maps from the high- and low-frequency image patches of the left view. A dual-frequency interaction attention module is used to calculate the attention weights of the high- and low-frequency signals and interactively weight the high- and low-frequency features. Three sets of convolutions are then used to further extract features. Finally, a dual-frequency recombination module is used to obtain the high- and low-frequency weighted feature maps.

[0014] In the binocular matching fusion subnetwork, a set of convolution and pooling operations are used to extract the full-frequency primary feature maps of the image blocks of the left and right views. The high- and low-frequency information of the dual-frequency interactive attention module is used to enhance the full-frequency primary feature maps to obtain full-frequency enhanced feature maps. The enhanced full-frequency features of the left and right views are then subjected to low-frequency binocular registration and high-frequency binocular registration respectively. Finally, a binocular competition selection module is used to filter features in the spatial and channel dimensions to obtain a binocular fusion feature map, and further features are extracted through two sets of convolution.

[0015] The binocular fusion feature map is concatenated with the high- and low-frequency weighted feature maps of the left and right views and then passed through a fully connected layer to obtain the objective prediction score. During the training process, the network model's loss function uses L1 loss to measure the difference between the network's predicted image quality score and the actual image quality score, and guide model optimization.

[0016] Step 3: Use the trained network to evaluate image quality. Slice the stereo image to be evaluated and input it into the network. The quality scores of all image blocks of each stereo image output by the network are averaged to obtain the quality score of the entire image.

[0017] Preferably, the processing flow of the dual-frequency interactive attention module in step 2 specifically includes the following contents:

[0018] The low-frequency primary features are parallelized through three 1×1 convolutional layers to produce three feature maps, Q, K, and V. Features Q and K are transformed and then matrix multiplied. The resulting maps are then passed through a SoftMax layer to produce a low-frequency attention map. Similarly, the same operation is performed on high-frequency features to produce a high-frequency attention map. The low-frequency feature map V is then matrix multiplied with the high-frequency attention map to produce a low-frequency feature map enhanced with high-frequency information. The same operation is performed on the high-frequency primary features to produce a high-frequency feature map enhanced with low-frequency information.

[0019] Preferably, the processing flow of the dual-frequency reconfiguration module in step 2 specifically includes the following contents:

[0020] The low-frequency and high-frequency high-level features are differentiated and summed, and the difference and sum are cascaded and sent to the residual module ResBlock composed of two convolutional layers and one residual connection layer to further extract features. The features are then passed through a global average pooling layer GAP and a fully connected layer to obtain adaptive weights. Finally, the weights are normalized and the low-frequency and high-frequency high-level features are weighted to obtain high- and low-frequency weighted features.

[0021] Preferably, the binocular matching fusion module in step 2 is composed of a binocular progressive registration module and a binocular competition selection module, and specifically includes the following contents:

[0022] ① The binocular progressive registration module consists of a low-frequency binocular registration unit and a high-frequency binocular registration unit;

[0023] In the low-frequency binocular registration unit, the low-frequency primary features of the left and right views are concatenated with their differential features and then fed into a convolution layer with a convolution kernel size of 1×1 to obtain the low-frequency offset. The full-frequency primary features of the left and right views and the low-frequency offset are then fed into a deformable convolution to obtain the features of the left and right views at the low-frequency scale.

[0024] In the high-frequency binocular registration unit, the high-frequency primary features of the left and right views are concatenated with their differential features and then fed into a convolutional layer with a kernel size of 1×1 to obtain a high-frequency offset. The low-frequency registered features of the left and right views and the high-frequency offset are then fed into a deformable convolution to obtain the registered features of the left and right views at both high and low frequency scales.

[0025] ② The binocular competitive selection module consists of a spatial dimension selection block and a channel dimension selection block. The registration features obtained by the binocular progressive registration module are added to the full-frequency primary features as the module input;

[0026] In the spatial dimension selection block, the input left and right view features are first passed through a convolution layer with a convolution kernel size of 1×1 to obtain two sets of feature maps. The two sets of feature maps are then dimensionally transformed and matrix multiplied. After passing through the SoftMax layer, the disparity attention map is obtained. The disparity attention map is matrix multiplied with the left and right view features respectively and weighted in the spatial dimension.

[0027] In the channel dimension selection block, the weighted features in the spatial dimension are sent to the global average pooling layer to obtain the left and right view channel attention maps. The channel attention maps are then normalized and dot-multiplied with the weighted features of the left and right views in the spatial dimension. Finally, the weighted left and right view features in the channel dimension are added together to obtain the binocular fusion features.

[0028] Compared with the existing technology, the present invention provides a stereo image quality evaluation method based on dual-frequency interactive enhancement and binocular matching, which has the following beneficial effects:

[0029] This method leverages information from different spatial frequencies in an image to enhance features, simulating the varying frequency sensitivity of the human visual system's quality perception. By simulating the binocular fusion and rivalry mechanisms of the human visual system and simultaneously utilizing information from different spatial frequencies to aid binocular fusion, it effectively extracts stereoscopic features consistent with human visual quality perception. The evaluation results of this method are highly consistent with subjective human evaluation, demonstrating its significant value. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 This is a flowchart of the stereo image quality assessment method based on dual-frequency interactive enhancement and binocular matching proposed in the present invention, where LFBRU represents the low-frequency binocular progressive registration unit, HFBRU represents the high-frequency binocular progressive registration unit, BRS represents the binocular competitive selection module, CCP represents two convolutional layers superimposed on a pooling layer, and nCCP represents n cascaded CCP modules;

[0031]

[0032] Figure 2 This is an overall block diagram of the binocular progressive registration module in the binocular feature matching and fusion module mentioned in Example 1 of the present invention;

[0033] Figure 3 This is an overall block diagram of the binocular competition selection module in the binocular feature matching and fusion module mentioned in Example 1 of the present invention; DETAILED DESCRIPTION

[0034] In order to make the objectives, technical solutions and advantages of the present invention more clear, the embodiments of the present invention are described in further detail below.

[0035] Example 1

[0036] This embodiment proposes a stereo image quality evaluation method based on dual-frequency interactive enhancement and binocular matching. Figure 1 As shown, the method includes the following steps:

[0037] S101: Preprocessing and data augmentation of original stereo images for training

[0038] The left and right view original images are cut into 220 32×32 non-overlapping image blocks. The quality score of each image block is the true quality score of the original image. The left and right view image blocks are converted to the frequency domain using FFT and decomposed into high and low frequencies. The high-frequency and low-frequency image blocks of the left and right views are restored using IFFT.

[0039] S102: Training a convolutional neural network based on dual-frequency interaction enhancement and binocular matching

[0040] A convolutional neural network based on dual-frequency interaction enhancement and binocular matching is constructed. The network consists of three sub-networks: a left-view dual-frequency interaction feature enhancement sub-network, a right-view dual-frequency interaction feature enhancement sub-network, and a binocular matching fusion sub-network. Specifically, it includes the following contents:

[0041] Since the left and right view subnetworks are symmetrical, only the implementation of the left view dual-frequency interaction feature enhancement subnetwork is described here. In this subnetwork, a set of 3×3 convolutions and pooling operations are used to extract primary feature maps from the high- and low-frequency image patches of the left view. A dual-frequency interaction attention module is used to calculate the attention weights of the high- and low-frequency signals and interactively weight the high- and low-frequency features. Further features are then extracted through three sets of 3×3 convolutions. Finally, a dual-frequency reorganization module is used to obtain the high- and low-frequency weighted feature maps.

[0042] In the binocular matching fusion subnetwork, a set of convolution and pooling operations are used to extract the full-frequency primary feature maps of the image blocks of the left and right views. The high- and low-frequency information of the dual-frequency interactive attention module is used to enhance the full-frequency primary feature maps to obtain full-frequency enhanced feature maps. The enhanced full-frequency features of the left and right views are then subjected to low-frequency binocular registration and high-frequency binocular registration respectively. Finally, a binocular competition selection module is used to filter features in the spatial and channel dimensions to obtain a binocular fusion feature map, and further features are extracted through two sets of convolution.

[0043] The binocular fusion feature map is concatenated with the high- and low-frequency weighted feature maps of the left and right views and then passed through a fully connected layer to obtain the objective prediction score. During the training process, the network model's loss function uses L1 loss to measure the difference between the network's predicted image quality score and the actual image quality score, and guide model optimization.

[0044] S103: Use the trained network to evaluate stereo image quality

[0045] The stereo image to be evaluated is sliced ​​and input into the network. The quality scores of all image blocks of each stereo image output by the network are averaged to obtain the quality score of the entire image.

[0046] Example 2:

[0047] Based on Example 1, but with some differences, the solution in Example 1 is further described below with reference to specific calculation formulas and example data. Since the left and right view sub-networks are symmetrical, only the implementation of one sub-network is described here. See the following description for details:

[0048] Step 1: Collect EEG signals and preprocess them

[0049] The original left and right view images are cut into 220 32×32 non-overlapping image blocks. The quality score of each image block is the true quality score of the original image. Use FFT to convert the image block I into the frequency domain according to the following formula (1) to perform high- and low-frequency decomposition. Use IFFT to restore the high- and low-frequency image blocks:

[0050]

[0051] where z represents the frequency component of the sample and t(·; r) represents a threshold function that separates the low-frequency and high-frequency components from z according to the hyperparameter radius r.

[0052] To simplify the processing, we first consider decomposing the single-channel image into high-frequency and low-frequency images. For the three-channel images used in this paper, this operation is performed independently on each channel. Define a single-channel image as I∈N n×n , which is represented in the frequency domain as z∈C n×n , where C represents a complex number. Function z lf ,z hf =t(z;r) is defined as:

[0053]

[0054] Among them, z(i,j) represents the value of position (i,j), {c i ,c j} represents the centroid, and d(·,·) is the Euclidean distance.

[0055] Step 2: Train a convolutional neural network based on dual-frequency interaction enhancement and binocular matching

[0056] This embodiment constructs a convolutional neural network based on dual-frequency interaction enhancement and binocular matching. The network consists of three sub-networks: a left-view dual-frequency interaction feature enhancement sub-network, a right-view dual-frequency interaction feature enhancement sub-network, and a binocular matching fusion sub-network. The specific method is as follows:

[0057] Since the left and right view sub-networks are symmetrical, only the implementation of the left view dual-frequency interaction feature enhancement sub-network is introduced here.

[0058] In the left-view dual-frequency interaction feature enhancement subnetwork, high- and low-frequency image patches are fed into a set of 3×3 convolutions and pooling to extract primary features. The dual-frequency interaction attention module (DFIA) calculates the attention weights for the high- and low-frequency signals and interactively weights the high- and low-frequency features. Further features are then extracted through three sets of 3×3 convolutions. Finally, a dual-frequency reorganization module (DFR) is used to generate the high- and low-frequency weighted feature maps.

[0059] Dual frequency interactive attention module DFIA Figure 1 As shown, first the primary features After three convolution layers with 1×1 convolution kernels, two sets of feature maps are obtained. and Then the dimensions of the two sets of feature maps K and Q are transformed into Where N = H × W. Afterwards, matrix multiplication is performed on the transpose of Q and K, and a SoftMax layer is applied to calculate the spatial attention map, resulting in {Q lf ,K lf} and {Q hf ,K hf}Calculated spatial attention map Att lf With Att hf Finally, the spatial attention map Att lf (Att hf ) and feature map V hf (V lf ) to perform matrix multiplication to achieve interactive weighting of high-frequency and low-frequency features. The whole process can be expressed as:

[0060]

[0061] in, Indicates matrix multiplication, {IF hf ,IF lf} are the high-frequency and low-frequency features after interactive enhancement.

[0062] Dual frequency recombination module DFR Figure 1 As shown in Figure 1, it consists of a ResBlock, a GAP layer, and a fully connected layer in cascade. The ResBlock contains two convolutional layers and a residual. The difference and sum of the interactively weighted low-frequency and high-frequency high-level features are cascaded and sent to the DFR module to obtain the weight W hf With W lf .

[0063] Finally, after normalizing the obtained weights, the low-frequency and high-frequency high-level features are weighted. The feature weighting is expressed by the following formula:

[0064]

[0065] in, is the input feature of DFR, is the final single-view frequency reorganization feature.

[0066] In the binocular matching fusion sub-network, a set of convolution and pooling operations are used to extract the full-frequency primary feature maps of the image blocks of the left and right views. In order to highlight the contours and details for better registration, the high and low frequency information of the dual-frequency interactive attention module is used to enhance the features of the full-frequency primary feature maps, that is, using Slf With S hf Respectively with V ff Perform matrix multiplication to obtain the full frequency enhancement feature map. The specific implementation process is expressed by the following formula:

[0067]

[0068] The enhanced full-frequency features of the left and right views are then sequentially subjected to low-frequency binocular registration and high-frequency binocular registration. Finally, a binocular rivalry selection module is used to filter features in the spatial and channel dimensions to obtain a binocular fusion feature map. Further features are extracted through two sets of convolutions. The binocular feature matching and fusion module consists of a binocular progressive registration module and a binocular rivalry selection module. The specific implementation process is as follows:

[0069] ① Such as Figure 2 As shown, the binocular progressive registration module consists of a low-frequency binocular registration unit and a high-frequency binocular registration unit;

[0070] In the low-frequency binocular registration unit, the low-frequency primary features of the left and right views are concatenated with their differential features and then fed into a convolutional layer with a kernel size of 1×1 to obtain a low-frequency offset. The full-frequency primary features of the left and right views and the low-frequency offset are then fed into a deformable convolution to obtain the features of the left and right views at the low-frequency scale.

[0071] In the high-frequency binocular registration unit, the high-frequency primary features of the left and right views are concatenated with their differential features and then fed into a convolutional layer with a kernel size of 1×1 to obtain a high-frequency offset. The low-frequency registered features of the left and right views and the high-frequency offset are then fed into a deformable convolution to obtain the registered features of the left and right views at both the high and low frequency scales.

[0072] The specific implementation process is expressed by the formula:

[0073]

[0074] Among them, Conv represents the convolution with a convolution kernel size of 1×1, O lf and O hf Represents the deformable convolution offset, DeConv represents the deformable convolution operation, AF left and AF right are the features after alignment.

[0075] ② If Figure 3 As shown in the figure, the binocular competitive selection module consists of a spatial dimension selection block and a channel dimension selection block, which takes the alignment features {AF left ,AF right} and full-frequency characteristics Add up to get the module input {MF left ,MFright};

[0076] In the spatial dimension selection block, the binocular features {MF left ,MF right After a convolution layer with a convolution kernel size of 1×1, we get Then F l With F r Transpose Perform matrix multiplication and pass through the SoftMax layer to obtain the attention map Then, the attention map {AM R→L ,AM L→R} and the module input {MF left ,MF right} Perform matrix multiplication:

[0077]

[0078] in is matrix multiplication, T represents transpose, SF left and SF right It is the feature weighted by disparity attention.

[0079] In the channel dimension selection block, the obtained SF left and SF right Send it to the global average pooling layer to get the channel attention map {AM left ,AM right}, and normalize the channel attention map pixel-wise:

[0080]

[0081] Then, the channel attention map is combined with SF left and SF right Perform point multiplication to obtain the fusion feature F of the binocular view b .

[0082]

[0083] Further, F b Send it to two sets of convolution and pooling layers to extract the final binocular fusion feature F bino .

[0084] Finally, the binocular fusion feature F bino High and low frequency weighted features of left and right views and After splicing, the objective prediction score is obtained through the fully connected layer. During the training process, the network model uses the L1 loss function to measure the difference between the image quality score predicted by the network and the actual image quality score, and guide the optimization of the model.

[0085] 203: Use the trained network to evaluate image quality

[0086] The stereo image to be evaluated is sliced ​​and input into the network. The quality scores of all image blocks of each stereo image output by the network are averaged to obtain the quality score of the entire image.

[0087] Example 3:

[0088] Based on Examples 1-2, but with the difference that the feasibility of the solutions in Examples 1 and 2 is verified by combining specific experiments, as described below:

[0089] To validate the performance of the proposed method, we tested it on four datasets (LIVE I, LIVE II, WIVC I, and WIVCII). To examine the performance of the proposed method, we selected two commonly used metrics: the Spearman rank order correlation coefficient (SROCC / SRCC) and the Pearson linear correlation coefficient (PLCC). The SROCC evaluates the monotonicity of the method's predictions, while the PLCC describes the linear correlation between the predicted and true values. For both the PLCC and SROCC, larger values ​​indicate better performance.

[0090] To validate the performance of our method, we compared seven mainstream no-reference stereo image quality assessment methods on the LIVE 3D database and the Waterloo IVC 3D database. These eight CNN-based methods (by Fang, Liu, Yang, Zhou, Sim, Si, and Chang, et al.) were selected for comparison. The results are shown in Tables 1 and 2. In each column, the results of the two best-performing methods are shown in bold.

[0091] Table 1 Performance comparison of methods based on LIVE 3D image database

[0092]

[0093] Table 2 Performance comparison of methods based on Waterloo IVC 3D image database

[0094]

[0095] As can be seen from Tables 1 and 2, the performance metrics SROCC and PLCC of our method lead all compared methods on both the LIVE 3D image database and the Waterloo IVC 3D image database. This method achieves exceptional results and exhibits good generalization. The model's predictions are generally highly consistent with human subjective scores.

[0096] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A stereo image quality evaluation method based on dual-frequency interactive enhancement and binocular matching, characterized in that: The following steps are involved: Step 1: Cut the left and right view original images into 32×32 non-overlapping image blocks. Use FFT to convert the left and right view image blocks into the frequency domain and decompose them into high and low frequencies. Use IFFT to restore the high and low frequency image blocks of the left and right views. Step 2: Train a convolutional neural network based on dual-frequency interaction enhancement and binocular matching. The convolutional neural network includes three subnetworks: a left-view dual-frequency interaction feature enhancement subnetwork, a right-view dual-frequency interaction feature enhancement subnetwork, and a binocular matching fusion subnetwork. Specifically, the following contents are included: The left and right view subnetworks are symmetrical. The processing flow of the left view dual-frequency interactive feature enhancement subnetwork specifically includes the following: In the left view dual-frequency interactive feature enhancement subnetwork, a set of convolution and pooling operations are used to extract primary feature maps for the high-frequency and low-frequency image blocks of the left view. The dual-frequency interactive attention module is used to calculate the attention weights of the high-frequency and low-frequency signals and interactively weight the high- and low-frequency features. Then, three sets of convolutions are used to further extract features. Finally, a dual-frequency recombination module is used to obtain the high- and low-frequency weighted feature maps. The processing flow of the binocular matching fusion sub-network specifically includes the following contents: in the binocular matching fusion sub-network, a set of convolution and pooling operations are used to extract the full-frequency primary feature map of the image blocks of the left and right views, and the high- and low-frequency information of the dual-frequency interactive attention module is used to enhance the full-frequency primary feature map to obtain a full-frequency enhanced feature map; then, the enhanced full-frequency features of the left and right views are sequentially subjected to low-frequency binocular registration and high-frequency binocular registration; finally, a binocular competition selection module is used to perform feature screening in the spatial dimension and channel dimension to obtain a binocular fusion feature map, and two sets of convolution are used to further extract features; The binocular fusion feature map is concatenated with the high- and low-frequency weighted feature maps of the left and right views and then passed through a fully connected layer to obtain the objective prediction score. During the training process, the network model's loss function uses L1 loss to measure the difference between the network's predicted image quality score and the actual image quality score, and guides model optimization. Step 3: Use the trained network to evaluate image quality. Slice the stereo image to be evaluated and input it into the network. Average the quality scores of all image blocks of each stereo image output by the network to obtain the quality score of the entire image.

2. The stereo image quality evaluation method based on dual-frequency interactive enhancement and binocular matching according to claim 1, characterized in that: The processing flow of the dual-frequency interactive attention module in step 2 specifically includes the following: The low-frequency primary features are passed through three convolution layers with a 1×1 convolution kernel in parallel to obtain three feature maps Q, K and V. The features Q and K are transformed in dimension and then matrix multiplication is performed, and then the low-frequency attention map is obtained through the SoftMax layer; the same operation is performed on the high-frequency features to obtain the high-frequency attention map; then, the low-frequency feature map V is matrix multiplied with the high-frequency attention map to obtain the low-frequency feature map after high-frequency information enhancement; the same operation is performed on the high-frequency primary features to obtain the high-frequency feature map after low-frequency information enhancement.

3. The stereo image quality evaluation method based on dual-frequency interactive enhancement and binocular matching according to claim 1, characterized in that: The processing flow of the dual-frequency reconfiguration module in step 2 specifically includes the following: The low-frequency and high-frequency high-level features are differentiated and summed, and the difference and sum are cascaded and sent to the residual module ResBlock composed of two convolutional layers and one residual connection layer to further extract features. The features are then passed through a global average pooling layer GAP and a fully connected layer to obtain adaptive weights. Finally, the weights are normalized and the low-frequency and high-frequency high-level features are weighted to obtain high- and low-frequency weighted features.

4. The stereo image quality evaluation method based on dual-frequency interactive enhancement and binocular matching according to claim 1, characterized in that: The binocular matching fusion sub-network described in step 2 consists of a binocular progressive registration module and a binocular competition selection module, specifically including the following: The binocular progressive registration module consists of a low-frequency binocular registration unit and a high-frequency binocular registration unit; In the low-frequency binocular registration unit, the low-frequency primary features of the left and right views are concatenated with their differential features and then fed into a convolution layer with a convolution kernel size of 1×1 to obtain the low-frequency offset. The full-frequency primary features of the left and right views and the low-frequency offset are then fed into a deformable convolution to obtain the features of the left and right views at the low-frequency scale. In the high-frequency binocular registration unit, the high-frequency primary features of the left and right views are concatenated with their differential features and then fed into a convolutional layer with a convolution kernel size of 1×1 to obtain a high-frequency offset. The low-frequency registered features of the left and right views and the high-frequency offset are then fed into a deformable convolution to obtain the registered features of the left and right views at both high and low frequency scales. The binocular rivalry selection module consists of a spatial dimension selection block and a channel dimension selection block. The registration features obtained by the binocular progressive registration module are added to the full-frequency primary features as the input of the module. In the spatial dimension selection block, the input left and right view features are first passed through a convolution layer with a convolution kernel size of 1×1 to obtain two sets of feature maps. The two sets of feature maps are then dimensionally transformed and matrix multiplied. After that, they are passed through a SoftMax layer to obtain the disparity attention map. The disparity attention map is matrix multiplied with the left and right view features respectively, and weighted in the spatial dimension. In the channel dimension selection block, the weighted features in the spatial dimension are sent to the global average pooling layer to obtain the left and right view channel attention maps. The channel attention maps are then normalized and dot-multiplied with the weighted features of the left and right views in the spatial dimension. Finally, the weighted left and right view features in the channel dimension are added together to obtain the binocular fusion features.

Citation Information

Patent Citations

  • A no-reference stereo image quality evaluation method based on registration distortion representation

    CN109685772A

  • Multi-scale low-illumination binocular stereo image enhancement method integrated with low-frequency information

    CN118096561A