Method for 3D image quality assessment based on deep learning
Through deep learning-based methods, the stereo image is fused and cropped, and combined with the characteristics of the human visual system, a neural network module is designed to solve the accuracy problem of stereo image quality evaluation and achieve more accurate and stable image quality prediction.
Patent Information
- Application Number
- CN202211380117.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-05
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2042-11-05
AI Technical Summary
In the evaluation of stereo image quality, the prior art ignores the comprehensiveness of human observation images and the locality of model input images, and the unreasonable design of binocular feature fusion modules, making it difficult to accurately predict the quality score of stereo image.
Using a deep learning-based method, the left and right views of 3D images are fused to generate significance maps, cropped image patches and inputted to neural networks, extracted monocular and binocular fusion features, combined with the perceptual characteristics of the human visual system, a neural network module is designed, and finally trained network weights through the MSE loss function and Adam optimizer until they are stable.
It realizes image quality evaluation in a wider range, obtains more accurate and stable result prediction, the model processing process is closer to human eye vision, simple image preprocessing, and the cropping process is close to binocular visual effects, improving the accuracy of stereoscopic image quality evaluation.
Smart Images

Figure CN115631181B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, in particular to a method for 3D image quality assessment based on deep learning. Background Art
[0002] Stereo image quality assessment (SIQA), as a method for measuring the distortion degree of stereo images, is a research hotspot in the field of image processing currently. Stereo images usually consist of a left view, a right view, and depth information. Due to the scene differences between the left view and the right view and the complex binocular vision fusion process, stereo image quality assessment is usually more challenging than general image quality detection.
[0003] In recent years, convolutional neural networks have achieved amazing results in object detection, face recognition, and image classification. The application of convolutional neural networks in IQA has attracted the attention of many researchers. The methods all consider the differences in binocular information on three-dimensional objects and obtain better results by capturing different features to predict image quality. However, due to ignoring the comprehensiveness of human image observation and the locality of model input images, as well as the unreasonable design of the binocular feature fusion module, the existing methods cannot predict the quality scores of stereo images well. Summary of the Invention
[0004] The present invention proposes a method for 3D image quality assessment based on deep learning, which can effectively predict the quality scores of the provided left and right view pictures and can accurately predict the quality scores of the pictures.
[0005] The present invention adopts the following technical solutions.
[0006] A method for 3D image quality assessment based on deep learning includes the following steps;
[0007] Step S1: Perform fusion processing on the left and right views that make up the 3D image, and generate a saliency map using the fused picture;
[0008] Step S2: Crop the original left and right views according to the fused picture;
[0009] Step S3: After combining the patch pairs of the cropped left and right views, input them into a neural network for quality assessment to obtain monocular features and binocular fusion features, and obtain the video quality after feature regression;
[0010] Step S4: Repeat Step S2 and Step S3 until the network weights of the neural network are stable, achieving the goal of 3D picture quality assessment.
[0011] The specific steps of the said Step S1 include the following steps;
[0012] Step S11: For the left and right views Im of each 3D image left and Im right perform fusion processing. Use wavelet decomposition to process the left and right views respectively. The wavelet decomposition process is expressed by the following formula:
[0013]
[0014] where cA is the approximate coefficient of decomposition, cD is the detail coefficient, ratio is the proportional coefficient of the mother wavelet, times is the number of decomposition times, Img[times] is the signal expression of any picture after times decompositions, H(·) is the high-pass filter, L(·) is the low-pass filter, y high represents the high-frequency signal in the wavelet decomposition result, and y low is the low-frequency signal in the wavelet decomposition result;
[0015] Step S12: Combine the approximate coefficients cA and detail coefficients cD decomposed from the left and right views in pairs, and calculate the arithmetic mean to obtain the fused approximate coefficient and the detail coefficient
[0016] Step S13: Use the approximate coefficient and the detail coefficient to perform inverse wavelet transform to obtain the fused saliency image Pic sal , and the inverse wavelet transform is expressed by the following formula:
[0017]
[0018] Step S14: For the fused picture with n pixels, having a length of H and a width of W, the saliency value of the pixel at the row row and column col is denoted as Sal(P row,col ), and the generation process of the saliency value Sal(P row,col ) is expressed by the following formula:
[0019]
[0020] where m is the enumeration process of the pixels and n is the total number of pixels.
[0021] The said Step S2 specifically includes the following steps;
[0022] Step S21: Preprocess the saliency image. First, calculate the saliency value of the entire saliency image divided by the proportional coefficient Rate as the threshold threshold. The calculation formula of the regional saliency value is as follows:
[0023]
[0024] Step S22: For the left view Im left and the right view Imright , provide a set of step sequences, which are denoted by the formula Step = {step k}, k = 1, 2,... Z, where Z is a positive integer;
[0025] The values of this step sequence are unique and arranged in ascending order. Starting from the initial position, try to sequentially crop the saliency map Pic sal into patches of size 40*40 in the order of the arrangement of the step S until the saliency value Sal of the patch is greater than the threshold threshold, and crop out patch pairs at the same positions in the left and right views, and then replace the initial position with the new area;
[0026] Step S23: For the rows and columns of the left and right views, repeat the above steps until the cropping process ends.
[0027] In Step S31 and Step S32, the human visual system is the visual system of the human primary visual cortex region.
[0028] Step S31: According to the perception characteristics of the human visual system, design a monocular vision processing module through basic deep learning components. The input includes general visual features Feature, and the processing process includes batch normalization BN, activation function σ, convolutional neural network etc., denotes the k-th 2D convolutional neural network. The calculation formula of the monocular vision processing result result is as follows:
[0029]
[0030] Step S32: According to the perception characteristics of the human visual system, design a binocular fusion vision processing module through basic deep learning components, which is divided into two parts: hybrid channel attention mechanism and hybrid spatial attention mechanism; in the hybrid channel attention mechanism, use the fully connected layer FC and adaptive average and max pooling methods (AvgPool / MaxPool) to assist feature extraction; the input includes the left and right view features F left and F right , and the feature maps after single-sided extraction in the channel attention mechanism are denoted as and The result of the hybrid channel attention mechanism is calculated as follows:
[0031]
[0032]
[0033]
[0034] In the hybrid spatial attention mechanism, the spatial average and maximum pooling methods are first used to process the images in the same direction and stack them to retain more information during the sampling process. Then, convolutional layers are inserted for the left and right views to filter the stacked images, filtering out small features and capturing large features.
[0035] In the spatial attention mechanism, the feature maps after unilateral extraction are denoted as and
[0036] The result of the hybrid channel attention mechanism The calculation formula is as follows:
[0037]
[0038]
[0039]
[0040] Step S33: Perform a simple fusion on the extracted spatial features and channel features using 3D convolution, and denote the mixed result as The calculation formula is as follows:
[0041]
[0042] where stack represents the stacking operation, represents the k-th 3D convolutional neural network;
[0043] Step S34: Unfold the processed feature maps and predict the graphic quality score through a fully connected layer.
[0044] In step S31 and step S32, the human visual system is the visual system of the human primary visual cortex region.
[0045] The specific content of step S4 is as follows:
[0046] Step S41: Load the MSE loss function and Adam optimizer for the designed neural network;
[0047] Step S42: Repeat the cropping step in S2 and the network training process in S3 until the weights of the neural network reach stability.
[0048] The present invention has the following beneficial effects compared with the prior art:
[0049] 1. The present invention detects the detailed semantics of the model within a wider range, making the image quality evaluation process more objective and obtaining more accurate and stable result prediction values. Compared with previous models, the processing procedures of our model and its components are all end-to-end, less dependent on the input content, and the data preprocessing is simpler.
[0050] 2. The independent visual path proposed by the present invention is closer to the process of image acquisition, image processing, and feature extraction of the human eye. The designed primary fusion module and advanced fusion module can better simulate the interaction details of the human primary visual cortex area, so that the image quality score can be better predicted.
[0051] 3. The image preprocessing scheme proposed by the present invention includes a new saliency view generation strategy and image cropping strategy. This improved scheme makes the cropping process closer to the binocular vision effect, which can not only effectively segment the picture independently according to the laws of the human visual system, but also carry the model to effectively adjust the training results of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] The present invention will be further described in detail below in conjunction with the drawings and specific embodiments:
[0053] Attached Figure 1 is a schematic flowchart of the present invention. SPECIFIC EMBODIMENTS
[0054] As shown in the figure, a method for 3D image quality evaluation based on deep learning includes the following steps;
[0055] Step S1: Perform fusion processing on the left and right views that make up the 3D image, and use the fused picture to generate a saliency map;
[0056] Step S2: Crop the original left and right views according to the fused picture;
[0057] Step S3: After combining the patch pairs of the cropped left and right views, input them into the neural network for quality evaluation to obtain monocular features and binocular fusion features, and obtain the video quality after feature regression;
[0058] Step S4: Repeat steps S2 and S3 until the network weights of the neural network are stable, achieving the goal of 3D picture quality evaluation.
[0059] The specific steps of step S1 include the following steps;
[0060] Step S11: For the left and right views Im left and Im right of each 3D image, perform fusion processing. Wavelet decomposition is used to process the left and right views respectively. The wavelet decomposition process is expressed by the following formula:
[0061]
[0062] Among them, cA is the approximate coefficient of decomposition, cD is the detail coefficient, ratio is the scale coefficient of the mother wavelet, times is the number of decomposition times, Img[times] is the signal expression of any picture after times decomposition, H(·) is the high-pass filter, L(·) is the low-pass filter, and y high represents the high-frequency signal in the wavelet decomposition result, and y low is the low-frequency signal in the wavelet decomposition result;
[0063] Step S12: Combine the approximate coefficients cA and detail coefficients cD decomposed from the left and right views in pairs, and calculate the arithmetic mean to obtain the fused approximate coefficient and the detail coefficient
[0064] Step S13: Use the approximate coefficient and the detail coefficient to perform inverse wavelet transform to obtain the fused saliency image Pic sal , and the inverse wavelet transform is expressed by the formula as follows:
[0065]
[0066] Step S14: For the fused picture with n pixels, with the length and width being H and W respectively, the saliency value of the pixel at the row row and column col is denoted as Sal(P row,col ), and the generation process of the saliency value Sal(P row,col ) is expressed by the formula as follows:
[0067]
[0068] Among them, m is the enumeration process of pixels, and n is the total number of pixels.
[0069] The specific steps of the said Step S2 include the following steps;
[0070] Step S21: Preprocess the saliency image. First, calculate the saliency value of the entire saliency image divided by the scale coefficient Rate as the threshold threshold. The calculation formula of the regional saliency value is as follows:
[0071]
[0072] Step S22: For the left view Im left and the right view Im right , provide a set of step sequences, and the sequence is denoted by the formula as Step = {step k}, k = 1, 2,... Z, where Z is a positive integer;
[0073] The values of this step sequence are unique and arranged in ascending order. Starting from the initial position, try to sequentially crop the saliency map Pic in the arrangement order of the step size S sal into patches of size 40*40 until the saliency value Sal of the patch is greater than the threshold threshold, and crop out a pair of patches at the same position in the left and right views, then replace the initial position with the new area;
[0074] Step S23: Repeat the above steps for the rows and columns of the left and right views until the cropping process ends.
[0075] In step S31 and step S32, the human visual system is the visual system of the human primary visual cortex region.
[0076] Step S31: According to the perception characteristics of the human visual system, design a monocular vision processing module through basic deep learning components. The input includes general visual features Feature, and the processing process includes batch normalization BN, activation function σ, convolutional neural network etc., representing the kth 2D convolutional neural network. The calculation formula for the monocular vision processing result result is as follows:
[0077]
[0078] Step S32: According to the perception characteristics of the human visual system, design a binocular fusion vision processing module through basic deep learning components, which is divided into two parts: hybrid channel attention mechanism and hybrid spatial attention mechanism; in the hybrid channel attention mechanism, use the fully connected layer FC and adaptive average and max pooling methods (AvgPool / MaxPool) to assist feature extraction; the input includes the left and right view features F left and F right , and the feature maps after unilateral extraction in the channel attention mechanism are denoted as and The result of the hybrid channel attention mechanism is calculated as follows:
[0079]
[0080]
[0081]
[0082] In the hybrid spatial attention mechanism, first use the spatial average and max pooling methods to process the images in the same direction and stack them to retain more information during the sampling process; then insert convolutional layers for the left and right views to filter the stacked images to filter out small features and capture large features;
[0083] In the spatial attention mechanism, the feature map after unilateral extraction is denoted as and
[0084] The result of the hybrid channel attention mechanism The calculation formula is as follows:
[0085]
[0086]
[0087]
[0088] Step S33: Perform a simple fusion on the extracted spatial features and channel features using 3D convolution, and denote the result after mixing as The calculation formula is as follows:
[0089]
[0090] where stack represents the stacking operation, represents the k-th 3D convolutional neural network;
[0091] Step S34: Unfold the processed feature map and predict the graphic quality score through a fully connected layer.
[0092] In Step S31 and Step S32, the human visual system is the visual system of the human primary visual cortex region.
[0093] The specific content of Step S4 is as follows:
[0094] Step S41: Load the MSE loss function and Adam optimizer for the designed neural network;
[0095] Step S42: Repeat the cropping step in S2 and the network training process in S3 until the weights of the neural network reach stability.
Claims
1. A method for 3D image quality assessment based on deep learning, characterized in that: Including the following steps; Step S1: Perform fusion processing on the left and right views that make up the 3D image, and generate a saliency map using the fused image; Step S2: Crop the original left and right views according to the fused image; Step S3: After combining the patch pairs of the cropped left and right views, input them into a neural network for quality assessment to obtain monocular features and binocular fusion features, and obtain the video quality after feature regression; Step S4: Repeat Step S2 and S3 until the network weights of the neural network are stable and the goal of 3D image quality evaluation is achieved; The specific steps of Step S1 include the following steps; Step S11: For the left and right views Im left and Im right of each 3D image, perform fusion processing. The left and right views are processed separately using wavelet decomposition. The wavelet decomposition process is expressed by the following formula: Among them, cA is the approximate coefficient of decomposition, cD is the detail coefficient, ratio is the proportionality coefficient of the mother wavelet, times is the number of decomposition times, Img[times] is the signal expression of any picture after times decomposition, H(·) is the high-pass filter, L(·) is the low-pass filter, and y high represents the high-frequency signal in the wavelet decomposition result, and y low is the low-frequency signal in the wavelet decomposition result; Step S12: Combine the approximation coefficients cA and detail coefficients cD decomposed from the left and right views in pairs, and calculate the arithmetic mean to obtain the fused approximation coefficients and detail coefficients Step S13: Perform inverse wavelet transform using the approximation coefficient and the detail coefficient to obtain the fused saliency image Pic sal , and the inverse wavelet transform is expressed by the following formula: Step S14: For the fused image with n pixels, having length H and width W, the saliency value of the pixel at the row row and column col is denoted as Sal(P row,col ), and the generation process of the saliency value Sal(P row,col ) is expressed by the following formula: Where m is the enumeration process of pixels and n is the total number of pixels; The specific steps of Step S2 include the following steps; Step S21: Preprocess the saliency image. First, calculate the saliency value of the entire saliency image divided by the proportionality coefficient Rate as the threshold threshold. The formula for the regional saliency value is as follows: Step S22: For the left view Im left and the right view Im right , provide a set of step sequences, which are denoted by the formula Step = {step k}, k = 1, 2,... Z, where Z is a positive integer; The values of the step sequence are unique and arranged in ascending order; starting from the initial position, try to sequentially crop the saliency map Pic in the arrangement order of the step size S sal into patches of size 40*40 until the saliency value Sal of the patch is greater than the threshold threshold, and crop out a pair of patches at the same position in the left and right views, then replace the initial position with the new region; Step S23: For the rows and columns of the left and right views, repeat the above steps until the cropping process ends; Step S31: According to the perception characteristics of the human visual system, design a monocular vision processing module through a deep learning component. The input includes visual feature Feature, and the processing process includes batch normalization BN, activation function σ, and convolutional neural network represents the k-th 2D convolutional neural network; the calculation formula of the monocular vision processing result result is as follows: Step S32: According to the perception characteristics of the human visual system, design a binocular fusion vision processing module through a deep learning component, which is divided into two parts: a hybrid channel attention mechanism and a hybrid spatial attention mechanism; In the hybrid channel attention mechanism, a fully connected layer FC and adaptive average and max pooling methods (AvgPool / MaxPool) are used to assist in feature extraction; the input includes left and right view features F left and F right , and the feature maps after single-sided extraction in the channel attention mechanism are denoted as and The result of the hybrid channel attention mechanism is calculated as follows: In the hybrid spatial attention mechanism, first use the spatial average and maximum pooling methods to process the images in the same direction and stack them to retain more information during the sampling process; then insert convolutional layers for the left and right views to filter the stacked images to filter out small features and capture large features; In the spatial attention mechanism, the feature map after unilateral extraction is denoted as and The result of the hybrid channel attention mechanism The calculation formula is as follows: Step S33: Perform a simple fusion on the extracted spatial features and channel features using 3D convolution, and denote the mixed result as The calculation formula is as follows: where stack represents a stacking operation, represents the k-th 3D convolutional neural network; Step S34: Unfold the processed feature map and predict the graphic quality score through a fully connected layer.
2. The method for 3D image quality assessment based on deep learning according to claim 1, wherein: In Step S31 and Step S32, the human visual system refers to the visual system of the human primary visual cortex region.
3. The method for 3D image quality assessment based on deep learning according to claim 1, wherein: The specific content of Step S4 is: Step S41: Equip the designed neural network with an MSE loss function and an Adam optimizer; Step S42: Repeat the cropping step in S2 and the network training process in S3 until the weights of the neural network reach stability.