An implicit neural-guided hyper-network blind light field image quality assessment method
Through the implicit neural guided supernet blind light field image quality evaluation method, Y channel information is extracted from the light field sub-aperture array image, and the implicit neural guide network and multi-scale semantic feature extraction network are used to generate perceived quality scores, which solves the accuracy of light field image quality evaluation and improves the accuracy of light field image quality evaluation.
Patent Information
- Application Number
- CN202510339497.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2045-03-21
AI Technical Summary
The prior art is difficult to objectively and accurately evaluate the quality of light field images, especially in terms of angular resolution, spatial resolution and image distortion, which affects the optimization of light field imaging devices and algorithms.
The supernet blind light field image quality evaluation method is adopted with implicit neural guidance. By extracting target Y channel information from the light field sub-aperture array image, the implicit neural guidance network and multi-scale semantic feature extraction network are used to generate perceived quality scores, and considering the visual characteristics of the human eye.
The scoring accuracy and effect of the reference-free light field image quality evaluation is significantly improved, the visual characteristics of the human eye are fully taken into account, and the accuracy of the light field image quality evaluation is improved.
Smart Images

Figure CN120198786B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing technology, and in particular to an implicit neural-guided super-network blind light field image quality evaluation method. Background Art
[0002] Light field imaging is an advanced imaging technology that can record the distribution of light in three-dimensional space. The information it captures far exceeds that of traditional two-dimensional images, covering a wealth of spatial details. Therefore, light field images have broad application value in fields such as depth estimation, biomedical testing, three-dimensional reconstruction, and virtual reality. However, due to limitations in hardware equipment and imaging conditions, the quality of light field images is often affected by resolution, noise, and distortion, which poses a challenge to their application in practical scenarios. Therefore, studying how to objectively evaluate the quality of light field images and analyze their performance in terms of angular resolution, spatial resolution, and image distortion has become an important direction for the current development of light field technology. Accurate quality evaluation can not only measure the application applicability of images, but also provide an important basis for the optimization of light field imaging equipment and algorithms. Summary of the Invention
[0003] The purpose of the present invention is to provide an implicit neural-guided super-network blind light field image quality evaluation method based on the light field image representation form of sub-aperture images, so as to perform accurate quality evaluation of light field images.
[0004] To achieve the above object, the present invention provides the following solutions:
[0005] An implicit neural-guided hyper-network blind light field image quality assessment method, comprising:
[0006] Acquire a sub-aperture array image of the light field to be evaluated;
[0007] Extracting target Y channel information of the light field sub-aperture array image;
[0008] The target Y channel information is input into a preset quality evaluation model, and a perceptual quality evaluation result is output, wherein the quality evaluation model uses an implicit neural guided network to extract high-dimensional features of the Y channel information and obtains a perceptual quality score based on the high-dimensional features.
[0009] Optionally, extracting target Y channel information of the light field sub-aperture array image includes:
[0010] Extracting vertical, horizontal, left diagonal and right diagonal light field sub-aperture image stacks from the light field sub-aperture array image, converting the light field sub-aperture image stacks from RGB space to YUV space, and extracting the Y channel information in the YUV space.
[0011] Optionally, the quality evaluation model includes:
[0012] A high-dimensional feature extraction module, configured to input the target Y channel information into the implicit neural guidance network and output high-dimensional features of the target Y channel information;
[0013] A multi-dimensional feature extraction module is used to input the high-dimensional features into a multi-scale semantic feature extraction network and output the multi-scale features of the target Y channel information;
[0014] A perception module, configured to input the highest-dimensional feature among the multi-scale features into a hypernetwork established based on perception rules, and output connection weights and bias values;
[0015] An evaluation module is used to cascade the multi-scale features and input them into a quality prediction network, and obtain a perceptual quality score by multiplying the multi-scale features with the connection weight and the bias value.
[0016] Optionally, the high-dimensional feature extraction module inputs the target Y channel information into an implicit neural guidance network and outputs high-dimensional features of the target Y channel information, including:
[0017] Obtaining corresponding coordinates and dimensions according to the spatial resolution of the light field subaperture array image, processing the coordinates using a position encoding function and sine and cosine transforms, obtaining high-frequency codes and splicing them into the coordinates in combination with the dimensions, and obtaining spliced coordinates;
[0018] Expanding the light field sub-aperture array image through F.unfold to convert it into a feature map with n angular resolutions, performing size and dimension conversion on the feature map to obtain a converted feature map;
[0019] A local integration operation is performed on each converted feature map to generate several converted feature maps of different perspectives, which are input into the feature extraction network MLP for feature extraction. The predicted value of each pixel at different perspectives is obtained and weighted summed to output a high-dimensional feature map.
[0020] Optionally, obtaining the predicted values of each pixel at different viewing angles and performing weighted summation to output the high-dimensional feature map includes:
[0021] Determine the relative coordinate information of each pixel in the spliced coordinates, splice the coordinate information with the pixel features, and input the spliced coordinate information into the feature extraction network MLP to output the predicted value of each pixel at different viewing angles;
[0022] Calculating the area of the region using the relative coordinate information, and weighting the predicted value according to the area of the region to obtain the regional weight of each pixel at each viewing angle;
[0023] The predicted values of each pixel at different viewing angles are weighted and summed according to the regional weight, and the high-dimensional feature map is output.
[0024] Optionally, the multidimensional feature extraction module inputs the high-dimensional features into a multi-scale semantic feature extraction network and outputs the multi-scale features of the target Y channel information, including:
[0025] The high-dimensional features are processed by the dimensionality reduction convolution layer and then sent to the multi-scale semantic feature extraction network to obtain multi-scale Y channel information.
[0026] Optionally, the perception module inputs the highest dimensional feature in the multi-scale features into a hypernetwork established by the perception rules, and outputs connection weights and bias values, including:
[0027] The highest-dimensional Y channel information is processed by the maximum pooling layer for dimensionality reduction and then sent to the super network established by the perception rule. After the channel is reduced by the convolution layer, the adaptive average pooling layer is used for inter-resolution compression. After passing through the fully connected layer, the connection weights and bias values are output.
[0028] Optionally, the evaluation module cascades the multi-scale features and inputs them into a quality prediction network, and obtains the perceptual quality score by dot-multiplying the multi-scale features with the connection weights and bias values, including:
[0029] The multi-scale Y channel information is cascaded through the local distortion perception module LDA to obtain a vector containing local distortion information, which is sent to the quality prediction network. After dimensionality reduction processing through several fully connected layers and Sigmoid activation functions, the vector is multiplied with the connection weight and bias value to obtain the final perceptual quality score.
[0030] The beneficial effects of the present invention are:
[0031] This method uses representations extracted from light field subaperture image stacks (including vertical, horizontal, left diagonal, and right diagonal directions) to obtain high-dimensional feature representations using an implicit neural guidance network. Furthermore, a multi-scale semantic feature extraction network is combined to generate a perceptual quality score. This method fully considers the visual characteristics of the human eye and significantly improves the scoring accuracy and effectiveness of reference-free light field image quality assessment. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0033] Figure 1This is a flow chart of an implicit neural-guided super-network blind light field image quality assessment method according to an embodiment of the present invention;
[0034] Figure 2 This is a network framework diagram of an implicit neural-guided super-network blind light field image quality assessment method according to an embodiment of the present invention;
[0035] Figure 3 A hypernetwork framework diagram established for the perception rules of an embodiment of the present invention;
[0036] Figure 4 This is a quality prediction network framework diagram of an embodiment of the present invention. DETAILED DESCRIPTION
[0037] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0038] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0039] This embodiment provides an implicit neural guided super network blind light field image quality evaluation method, such as Figure 1 As shown, including:
[0040] Acquire a sub-aperture array image of the light field to be evaluated;
[0041] Extracting target Y channel information of the light field sub-aperture array image;
[0042] The target Y channel information is input into a preset quality evaluation model, and a perceptual quality evaluation result is output, wherein the quality evaluation model uses an implicit neural guided network to extract high-dimensional features of the Y channel information and obtains a perceptual quality score based on the high-dimensional features.
[0043] Specifically, this embodiment extracts representations from light field sub-aperture image stacks, uses an implicit neural guidance network to obtain high-dimensional feature representations, and combines a multi-scale semantic feature extraction network to ultimately generate a perceptual quality score. This fully considers the visual characteristics of the human eye and significantly improves the scoring accuracy and effectiveness of reference-free light field image quality evaluation.
[0044] Furthermore, extracting target Y channel information of the light field sub-aperture array image includes:
[0045] Extracting vertical, horizontal, left diagonal and right diagonal light field sub-aperture image stacks from the light field sub-aperture array image, converting the light field sub-aperture image stacks from RGB space to YUV space, and extracting the Y channel information in the YUV space.
[0046] Specifically, this embodiment converts the vertical, horizontal, left diagonal, and right diagonal light field sub-aperture images in a 9×9 light field sub-aperture image array from RGB space to YUV space, and extracts Y channel information of the light field sub-aperture image stack.
[0047] From the 9×9 light field subaperture image array, vertical, horizontal, left diagonal, and right diagonal light field subaperture image stacks are obtained. The light field subaperture image array dimensions are 3×u×v×h×w, where u×v is the angular resolution of the subaperture images in the light field subaperture array image, h×w is the spatial resolution of a single subaperture image in the light field subaperture array image, and 3 represents the three RGB channels. First, the vertical, horizontal, left diagonal, and right diagonal light field subaperture image stacks are converted from RGB space to YUV space (each subaperture image stack has 9 subaperture images). Then, the Y channel information of the YUV space is extracted, with the dimensions of 9×h×w.
[0048] Further, if Figure 2 As shown, the quality evaluation model includes:
[0049] A high-dimensional feature extraction module is used to input the target Y channel information into the implicit neural guidance network and output the high-dimensional features of the target Y channel information;
[0050] A multi-dimensional feature extraction module is used to input the high-dimensional features into a multi-scale semantic feature extraction network and output the multi-scale features of the target Y channel information;
[0051] A perception module, configured to input the highest-dimensional feature among the multi-scale features into a hypernetwork established based on perception rules, and output connection weights and bias values;
[0052] An evaluation module is used to cascade the multi-scale features and input them into a quality prediction network, and obtain a perceptual quality score by multiplying the multi-scale features with the connection weight and the bias value.
[0053] Furthermore, the high-dimensional feature extraction module inputs the target Y channel information into an implicit neural guidance network and outputs high-dimensional features of the target Y channel information, including:
[0054] Obtaining corresponding coordinates and dimensions according to the spatial resolution of the light field subaperture array image, processing the coordinates using a position encoding function and sine and cosine transforms, obtaining high-frequency codes and splicing them into the coordinates in combination with the dimensions, and obtaining spliced coordinates;
[0055] Expanding the light field sub-aperture array image through F.unfold to convert it into a feature map with n angular resolutions, performing size and dimension conversion on the feature map to obtain a converted feature map;
[0056] A local integration operation is performed on each converted feature map to generate several converted feature maps of different perspectives, which are input into the feature extraction network MLP for feature extraction. The predicted value of each pixel at different perspectives is obtained and weighted summed to output a high-dimensional feature map.
[0057] The process of obtaining the predicted values of each pixel at different viewing angles and performing weighted summation to output the high-dimensional feature map includes:
[0058] Determine the relative coordinate information of each pixel in the spliced coordinates, splice the coordinate information with the pixel features, and input the spliced coordinate information into the feature extraction network MLP to output the predicted value of each pixel at different viewing angles;
[0059] Calculating the area of the region using the relative coordinate information, and weighting the predicted value according to the area of the region to obtain the regional weight of each pixel at each viewing angle;
[0060] The predicted values of each pixel at different viewing angles are weighted and summed according to the regional weight, and the high-dimensional feature map is output.
[0061] Specifically, this embodiment feeds the acquired vertical, horizontal, left diagonal, and right diagonal light field sub-aperture image stack Y channel information into an implicit neural guidance network to obtain a high-dimensional feature representation of the light field sub-aperture image stack.
[0062] First, the Y channel information of the acquired light field sub-aperture image stack is fed into an implicit neural guidance network. The implicit neural guidance network's computational process begins by generating corresponding coordinates (x, y) of h×w dimensions based on the spatial resolution h×w of the light field sub-aperture image. Next, a positional encoding function is used to generate high-frequency features associated with the coordinates. The coordinate points are encoded using sine and cosine transforms. Finally, the coordinate information is concatenated with the high-frequency encoding to form coordinates of h×w×(2L+3) dimensions, where L is the number of frequency levels encoded. The light field sub-aperture image features are then expanded using F.unfold to convert them into feature maps with nine angular resolutions. This expands the size of each image from h×w to 9×(h×w), thereby converting the image dimensions from 3×u×v×h×w to 9×(h×w), where u×v is the angular resolution of the light field sub-aperture array. Next, for each sub-aperture image, a local integration operation is performed to resample the image using different disparities vx_lst and vy_lst to generate multiple images with different perspectives. The dimensions of these images are kept as 9×(h×w) and are input to the feature extraction network (MLP).
[0063] During feature extraction, the features of each pixel are concatenated with their relative coordinate information and fed into the MLP network. Assuming the input feature dimension is 9×(h×w), after MLP processing, the output feature dimension is still 9×(h×w), representing the predicted value of each pixel at different viewing angles.
[0064] The weighting process first calculates the regional weight of each pixel at each viewpoint. The area of the region is calculated from the relative coordinates, and the predicted value is weighted according to the area. The regional weight has the dimension h×w and is used to balance the contribution of the image from different viewpoints.
[0065] Finally, all predictions from different viewpoints are weighted and summed according to the region weights. This weighted summation ensures the integration of information from different viewpoints and adjusts the contribution of different regions based on their area. The final output is a feature map with a dimension of 9×h×w, and the final result is returned.
[0066] Furthermore, the multidimensional feature extraction module inputs the high-dimensional features into a multi-scale semantic feature extraction network and outputs the multi-scale features of the target Y channel information, including:
[0067] The high-dimensional features are processed by the dimensionality reduction convolution layer and then sent to the multi-scale semantic feature extraction network to obtain multi-scale Y channel information.
[0068] Specifically, this embodiment feeds the obtained high-dimensional features of the Y channel information of the light field sub-aperture image stack into a multi-scale semantic feature extraction network to obtain the multi-scale Y channel information of the light field sub-aperture image array.
[0069] The high-dimensional features of the Y channel information of the light field sub-aperture image stack are first passed through a dimensionality reduction convolution layer with a 1×1 kernel, a stride of 1, and a padding value of 0, reducing the 9×H×W high-dimensional features to 3×h×w. This is then fed into a multi-scale semantic feature extraction network. After processing by four feature extraction modules at different scales, the Y channel information of four multi-scale sub-aperture array images is obtained, with dimensions of 64×64×C1, 32×32×C2, 16×16×C3, and 8×8×C4. C1, C2, C3, and C4 represent the number of channels that increases with increasing depth of the layer.
[0070] Furthermore, the perception module inputs the highest dimensional feature in the multi-scale feature into the hypernetwork established by the perception rule, and outputs connection weights and bias values, including:
[0071] The highest-dimensional Y channel information is processed by the maximum pooling layer for dimensionality reduction and then sent to the super network established by the perception rule. After the channel is reduced by the convolution layer, the adaptive average pooling layer is used for inter-resolution compression. After passing through the fully connected layer, the connection weights and bias values are output.
[0072] Specifically, this embodiment feeds the obtained highest-dimensional light field sub-aperture image array Y channel information into the super network established by the perception rule to obtain the connection weights and bias values between the FC layers.
[0073] like Figure 3 As shown in the figure, the Y channel information of the highest-dimensional light field subaperture image array, with dimensions of 8×8×C4, is first sent to a max pooling layer with a convolution kernel size of 2, a stride of 1, and padding of 0, reducing the dimension to 7×7×C4. It is then sent to the hypernetwork established by the perceptual rules. It first passes through a convolutional network layer conv1, which first undergoes a first convolution operation using a 1×1 convolution kernel, a stride of 1, and padding of 0. Next, a second convolution operation also uses a 1×1 convolution kernel. Finally, a third convolution operation still uses a 1×1 convolution kernel. The entire process gradually reduces the number of channels to 112 while maintaining the spatial dimension unchanged. Next, an AdaptiveAvgPool2d pooling operation is used, which compresses the spatial resolution from 7×7 to 1×1, preserving the average information of each channel.
[0074] Then the connection weights and bias values are generated. The final output of the hypernetwork contains multiple feature maps in the following format: target_in_vec is the initial input feature map vector, target_fc1w and target_fc1b represent the weights and biases of the first fully connected layer, target_fc2w and target_fc2b represent the weights and biases of the second fully connected layer, target_fc3w and target_fc3b represent the weights and biases of the third fully connected layer, target_fc4w and target_fc4b represent the weights and biases of the fourth fully connected layer, and target_fc5w and target_fc5b represent the weights and biases of the last fully connected layer.
[0075] Furthermore, the evaluation module cascades the multi-scale features and inputs them into the quality prediction network, and obtains the perceptual quality score by dot-multiplying the multi-scale features with the connection weights and the bias values, including:
[0076] The multi-scale Y channel information is cascaded through the local distortion perception module LDA to obtain a vector containing local distortion information, which is sent to the quality prediction network. After dimensionality reduction processing through several fully connected layers and Sigmoid activation functions, the vector is multiplied with the connection weight and bias value to obtain the final perceptual quality score.
[0077] Specifically, this embodiment cascades the obtained multi-scale light field sub-aperture image array Y channel information and feeds it into the quality prediction network, and obtains the final perceptual quality score by performing a dot product with the connection weight and the bias value.
[0078] like Figure 4 As shown in the figure, the Y channel information of the obtained multi-scale light field sub-aperture image array first passes through the local distortion perception module (LDA). After cascading, a vector containing local distortion information is formed and sent to the quality prediction network. The network includes multiple fully connected layers (FC layers), each of which is followed by a Sigmoid activation function. Finally, through a series of linear transformations and nonlinear activation functions, the final perceptual quality score is obtained by multiplying the connection weights and bias values.
[0079] The following provides a specific evaluation example based on the implicit neural-guided super-network blind light field image quality evaluation method provided in this embodiment:
[0080] (1) Converting the vertical, horizontal, left diagonal, and right diagonal light field sub-aperture image stacks from RGB space to YUV space, and extracting Y channel information of the light field sub-aperture image stacks. In this embodiment, the dimensions of the vertical, horizontal, left diagonal, and right diagonal light field sub-aperture image stacks are 3×9×256×256. The dimensions of the extracted Y channel information of the light field sub-aperture image stacks are 9×256×256.
[0081] (2) First, the Y channel information of the light field subaperture image stack is input into the implicit neural guidance network. The image dimension is 9×256×256, where 9 represents the number of subapertures of the light field image and 256×256 is the spatial resolution of each subaperture image. In order to introduce spatial position information, the make_coord function is first used to generate coordinates to obtain a coordinate representation with a dimension of 2×256×256, which represents the x and y coordinates of each pixel. Then, the generated coordinates are position encoded, and the dimension becomes 18×256×256. This process maps the coordinates of each pixel to a higher dimension through frequency transformation, thereby obtaining richer spatial information. The dimension of the position encoding is 2L, where L is the predefined number of frequencies (here 8), and each coordinate generates a 16-dimensional (2×8=16) encoding result. Then, the coordinate information and the position encoding are spliced together to obtain a 19×256×256 feature representation, which provides richer feature input for subsequent calculations.
[0082] Next, the input image is expanded (via F.unfold) to convert the image from 9×256×256 to 81×256×256. The expansion operation converts each small area of the image into a higher-dimensional feature map, where the information of the nine sub-aperture images is combined to form a higher-dimensional feature representation. The expanded features are then concatenated with the relative coordinates again to obtain an 83×256×256 feature representation. The relative coordinates represent the relative position of each pixel and are combined with the expanded features to provide more detailed spatial information.
[0083] If cell decoding is enabled (cell_decode=True), further cell information is added. The cell position of each pixel is magnified to the spatial scale of the image, increasing the cell information dimension by 2, resulting in an 85×256×256 feature input. All features are then fed into a multi-layer perceptron (MLP) network for processing. The input dimension is 85×256×256, and after layer-by-layer linear transformation and activation functions in the MLP, the output dimension becomes 9×256×256, ultimately outputting a 9-dimensional feature vector for each pixel.
[0084] Finally, the output feature map is weighted averaged, and the feature values of different regions are weighted calculated, and the dimension is maintained at 9×256×256.
[0085] (3) The high-dimensional features of the Y channel information of the light field sub-aperture image stack obtained in step (2) are first passed through a dimensionality reduction convolution layer with a convolution kernel of 1×1, a step size of 1, and a padding value of 0 to reduce the 9×256×256 high-dimensional features to 3×256×256. Then, the features are sent to the multi-scale semantic feature extraction network of Swin Transformer V2 Tiny. After being processed by four feature extraction modules of different scales, four multi-scale sub-aperture array image Y channel information are obtained, with dimensions of 64×64×96, 32×32×192, 16×16×384 and 8×8×768, where the number of channels increases with the deepening of the layer.
[0086] (4) The highest-dimensional light field subaperture image array Y channel information obtained in step (3), whose dimension is 8×8×768, is first sent to a maximum pooling layer with a convolution kernel size of 2, a stride of 1, and a padding value of 0 to reduce the dimension to 7×7×768. Next, the feature map is processed by the conv1 module, which includes multiple 1×1 convolution layers. The first convolution reduces the number of channels from 768 to 384 and is activated by ReLU; the second convolution reduces the number of channels from 384 to 192 and is activated by ReLU; the last convolution reduces the number of channels from 192 to 112 and is activated by ReLU. The final feature map size is 112×7×7, that is, the number of channels is 112 and the spatial size remains 7×7. At this stage, the target input vector target_in_vec is extracted from res_out['target_in_vec'] and reshaped to 224×1×1, converting the LDA result vector into a tensor with a spatial dimension of 1×1. Next, the network further processes the feature maps through multiple convolutional layers, gradually reducing their spatial dimensions. Specifically, the feature map generated by the convolution of fc1w_conv is 112×224×1×1 with 112 channels, a target input size of 224, and a spatial dimension of 1×1, while the bias of fc1b_fc has a dimension of 112. The convolution of fc2w_conv reduces the number of channels from 112 to 56, resulting in an output feature map of 56×112×1×1, and a bias of 56 for fc2b_fc. Similarly, fc3w_conv reduces the number of channels from 56 to 28, resulting in an output feature map of 28×56×1×1, with a bias of 28. fc4w_conv reduces the number of channels from 28 to 14, resulting in an output feature map of 14×28×1×1, with a bias of 14. Finally, the convolutional weights generated by fc5w_fc reduce the number of channels from 14 to 1, resulting in an output feature map with dimensions of 1×14×1×1. The bias term, generated by fc5b_fc, has a dimension of 1. Each convolutional layer and the corresponding fully connected layer generate corresponding weights and biases with the following dimensions: target_fc1w is 112×224×1×1, target_fc1b is 112, target_fc2w is 56×112×1×1, target_fc2b is 56, target_fc3w is 28×56×1×1, target_fc3b is 28, target_fc4w is 14×28×1×1, target_fc4b is 14, target_fc5w is 1×14×1×1, and target_fc5b is 1. This step-by-step structure allows the model to effectively extract and adjust features from the input image, ultimately generating adjusted feature maps and associated bias terms.
[0087] (5) Before the multi-scale semantic features obtained in step (3) are fed into the quality prediction network, they are first passed through a local distortion awareness module (LDA). The dimension of LDA is gradually transformed according to the operations of different convolutional layers, pooling layers, and fully connected layers. The purpose of the LDA module is to extract local distortion-related features from different feature maps and map them to a fixed dimension through the fully connected layer. In the LDA module, the processing of each feature map is gradually transformed through convolution, pooling, and fully connected layers. First, LDA1 processes the input feature map from layer1 with a size of 1×256×56×56, using a 1×1 convolution kernel to reduce the number of channels from 256 to 16, with a convolution stride of 1 and padding of 0. The output feature map has a size of 1×16×56×56. Then, 7×7 average pooling is applied with a stride of 7 to reduce the spatial size to 1×16×8×8, which is then flattened to a one-dimensional vector of 1×1024 through a flattening operation and mapped to a 1×16 output through a fully connected layer. LDA2 processes the feature map from layer2 with a size of 1×512×28×28. It also uses a 1×1 convolution kernel to reduce the number of channels from 512 to 32. The output feature map has a size of 1×32×28×28. It then uses a 7×7 average pooling with a stride of 7 to reduce the spatial size to 1×32×4×4. After flattening to 1×512, it is mapped to 1×16. LDA3 processes the input feature map from layer3 with a size of 1×1024×14×14. It uses a 1×1 convolution kernel to reduce the number of channels from 1024 to 64. The output size is 1×64×14×14. It then uses a 7×7 average pooling with a stride of 7 to reduce the spatial size to 1×64×2×2. After flattening to 1×256, it is mapped to 1×16 through a fully connected layer. Finally, LDA4 processes the feature map from layer 4, which has a size of 1×2048×7×7. It directly applies 7×7 average pooling with a stride of 7, reducing the spatial dimensions to 1×2048×1×1. After flattening, it obtains a 1×2048 output, which is then mapped to a 1×176 output through a fully connected layer. Through this series of convolution, pooling, and fully connected operations, the LDA module gradually extracts local distortion features and maps them to the final output dimension. The output features of all LDA layers are concatenated together to form a vector containing local distortion information with a dimension of 1×224.
[0088] The vector containing local distortion information is then fed into the quality prediction network. The quality prediction network consists of the following steps: The input passes through the first fully connected layer l1, which converts the input dimension from 224 to 112. After processing with the Sigmoid activation function, the output dimension is 112. Next, the input passes through the second fully connected layer l2, mapping the input dimension from 112 to 56. After processing with the Sigmoid activation function, the output dimension is 56. The input then passes through the third fully connected layer l3, mapping the dimension from 56 to 28. It also passes through the Sigmoid activation function, and the output dimension is 28. Finally, the input passes through the fourth fully connected layer l4, which first maps 28 to 14, then passes through the Sigmoid activation function, and then passes through the fifth fully connected layer to map 14 to the final output dimension of 1. Single dimensions are removed using .squeeze(), and the final output is a quality prediction score with a dimension of 1.
[0089] The embodiments described above are merely descriptions of preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Without departing from the spirit of the present invention, various modifications and improvements made to the technical solutions of the present invention by persons skilled in the art should fall within the scope of protection defined by the claims of the present invention.
Claims
1. An implicit neural-guided hypernetwork blind light field image quality assessment method, characterized in that: include: Acquire a sub-aperture array image of the light field to be evaluated; Extracting target Y channel information of the light field sub-aperture array image; Inputting the target Y channel information into a preset quality evaluation model and outputting a perceptual quality evaluation result, wherein the quality evaluation model uses an implicit neural network to extract high-dimensional features of the target Y channel information and obtains a perceptual quality score based on the high-dimensional features; Wherein, the quality evaluation model includes: A high-dimensional feature extraction module, configured to input the target Y channel information into the implicit neural guidance network and output high-dimensional features of the target Y channel information; A multi-dimensional feature extraction module is used to input the high-dimensional features into a multi-scale semantic feature extraction network and output the multi-scale features of the target Y channel information; A perception module, configured to input the highest-dimensional feature among the multi-scale features into a hypernetwork established based on perception rules, and output connection weights and bias values; An evaluation module, configured to concatenate the multi-scale features and input them into a quality prediction network, and obtain a perceptual quality score by performing a dot product with the connection weight and the bias value; The high-dimensional feature extraction module inputs the target Y channel information into the implicit neural guidance network and outputs the high-dimensional features of the target Y channel information, including: Obtaining corresponding coordinates and dimensions according to the spatial resolution of the light field subaperture array image, processing the coordinates using a position encoding function and sine and cosine transforms, obtaining high-frequency codes and splicing them into the coordinates in combination with the dimensions, and obtaining spliced coordinates; Expanding the light field sub-aperture array image through F.unfold to convert it into a feature map with n angular resolutions, performing size and dimension conversion on the feature map to obtain a converted feature map; A local integration operation is performed on each converted feature map to generate several converted feature maps of different perspectives, which are input into the feature extraction network MLP for feature extraction. The predicted value of each pixel at different perspectives is obtained and weighted summed to output a high-dimensional feature map.
2. The implicit neural-guided super-network blind light field image quality assessment method according to claim 1 is characterized in that: Extracting target Y channel information of the light field sub-aperture array image includes: Extracting vertical, horizontal, left diagonal, and right diagonal light field sub-aperture image stacks from the light field sub-aperture array image, converting the light field sub-aperture image stacks from RGB space to YUV space, and extracting the target Y channel information in the YUV space.
3. The implicit neural-guided super-network blind light field image quality assessment method according to claim 1, characterized in that: Obtain the predicted values of each pixel at different viewing angles and perform weighted summation to output the high-dimensional feature map, including: Determine the relative coordinate information of each pixel in the spliced coordinates, splice the coordinate information with the pixel features, and input the spliced coordinate information into the feature extraction network MLP to output the predicted value of each pixel at different viewing angles; Calculating the area of the region using the relative coordinate information, and weighting the predicted value according to the area of the region to obtain the regional weight of each pixel at each viewing angle; The predicted values of each pixel at different viewing angles are weighted and summed according to the regional weight, and the high-dimensional feature map is output.
4. The implicit neural-guided super-network blind light field image quality assessment method according to claim 1, characterized in that: The multidimensional feature extraction module inputs the high-dimensional features into a multi-scale semantic feature extraction network and outputs the multi-scale features of the target Y channel information, including: The high-dimensional features are processed by the dimensionality reduction convolution layer and then sent to the multi-scale semantic feature extraction network to obtain multi-scale Y channel information.
5. The implicit neural-guided super-network blind light field image quality assessment method according to claim 1, characterized in that: The perception module inputs the highest dimensional feature in the multi-scale features into the hypernetwork established by the perception rule, and outputs connection weights and bias values, including: The highest-dimensional Y channel information is processed by the maximum pooling layer for dimensionality reduction and then sent to the super network established by the perception rule. After the channel is reduced by the convolution layer, the adaptive average pooling layer is used for inter-resolution compression. After passing through the fully connected layer, the connection weights and bias values are output.
6. The implicit neural-guided super-network blind light field image quality assessment method according to claim 1, characterized in that: The evaluation module cascades the multi-scale features and inputs them into the quality prediction network, and obtains the perceptual quality score by dot multiplication with the connection weight and the bias value, including: The multi-scale Y channel information is cascaded through the local distortion perception module LDA to obtain a vector containing local distortion information, which is sent to the quality prediction network. After dimensionality reduction processing through several fully connected layers and Sigmoid activation functions, the vector is multiplied with the connection weight and bias value to obtain the final perceptual quality score.
Citation Information
Patent Citations
Full-reference quality evaluation method for light field image
CN117911367A
Light field image full-reference quality evaluation method based on lens image
CN117934408A