Implicit nerve-guided super-network blind light field image quality evaluation method

Through the implicit neural guided hypernetwork method, high-dimensional features are extracted from the light field sub-aperture image stack, solving the problem of light field image quality evaluation, and achieving more accurate quality evaluation and higher scoring accuracy.

CN120198786AActive Publication Date: 2025-06-24ANQING NORMAL UNIV

Patent Information

Application Number
CN202510339497.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-06-24
Estimated Expiration
2045-03-21

AI Technical Summary

Technical Problem

The prior art is difficult to effectively evaluate the quality of light field images, especially in terms of angular resolution, spatial resolution and image distortion, which affects its effect in practical applications.

Method used

The hypernetwork method of implicit neural guidance is adopted to extract target Y channel information from the optical field sub-aperture image stack, and high-dimensional features are extracted using the implicit neural guidance network, and the network is extracted in combination with multi-scale semantic features to finally generate perceived quality scores.

Benefits of technology

It significantly improves the scoring accuracy and effect of image quality evaluation without reference light field, takes into account the visual characteristics of the human eye, and provides more accurate quality evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198786A_ABST
    Figure CN120198786A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image processing, in particular to an implicit nerve-guided super-network blind light field image quality evaluation method, which comprises the following steps: acquiring a to-be-evaluated light field sub-aperture array image; extracting target Y channel information of the light field sub-aperture array image; the target Y channel information is input into a preset quality evaluation model, a perception quality evaluation result is output, and the quality evaluation model uses an implicit neural guidance network to extract high-dimensional features of the Y channel information and obtains a perception quality score according to the high-dimensional features. According to the method, through a representation form extracted from a light field sub-aperture image stack, high-dimensional feature representation is obtained by using an implicit neural guidance network. In addition, a multi-scale semantic feature extraction network is combined, a perception quality score is finally generated, visual characteristics of human eyes are fully considered, and the scoring accuracy and effect of non-reference light field image quality evaluation are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and particularly to an implicit neural-guided hypernetwork blind light field image quality evaluation method. Background Art

[0002] Light field imaging is an advanced imaging technology that can record the distribution of light in three-dimensional space. The information it captures far exceeds that of traditional two-dimensional images and covers rich spatial details. Therefore, light field images have wide application values in fields such as depth estimation, biomedical detection, three-dimensional reconstruction, and virtual reality. However, due to limitations in hardware devices and imaging conditions, the quality of light field images is often affected by resolution, noise, and distortion, which poses challenges to their application effects in actual scenarios. Therefore, studying how to objectively evaluate the quality of light field images and analyze their performance in terms of angular resolution, spatial resolution, and image distortion has become an important direction for the current development of light field technology. Accurate quality evaluation can not only measure the application suitability of images but also provide an important basis for optimizing light field imaging devices and algorithms. Summary of the Invention

[0003] The purpose of the present invention is to provide an implicit neural-guided hypernetwork blind light field image quality evaluation method based on the light field image representation form of sub-aperture images to accurately evaluate the quality of light field images.

[0004] To achieve the above purpose, the present invention provides the following solutions:

[0005] An implicit neural-guided hypernetwork blind light field image quality evaluation method, comprising:

[0006] Obtaining a light field sub-aperture array image to be evaluated;

[0007] Extracting target Y-channel information of the light field sub-aperture array image;

[0008] Inputting the target Y-channel information into a preset quality evaluation model to output a perceptual quality evaluation result, wherein the quality evaluation model uses an implicit neural guidance network to extract high-dimensional features of the Y-channel information and obtains a perceptual quality score according to the high-dimensional features.

[0009] Optionally, extracting the target Y-channel information of the light field sub-aperture array image includes:

[0010] Extracting a stack of light field sub-aperture images in the vertical, horizontal, left diagonal, and right diagonal directions from the light field sub-aperture array image, converting the stack of light field sub-aperture images from the RGB space to the YUV space, and extracting the Y-channel information in the YUV space.

[0011] Optionally, the quality evaluation model includes:

[0012] A high-dimensional feature extraction module, configured to input the target Y-channel information into the implicit neural guidance network and output the high-dimensional features of the target Y-channel information;

[0013] A multi-dimensional feature extraction module, configured to input the high-dimensional features into a multi-scale semantic feature extraction network and output the multi-scale features of the target Y-channel information;

[0014] A perception module, configured to input the highest-dimensional feature in the multi-scale features into a hypernetwork established based on perception rules and output connection weights and bias values;

[0015] An evaluation module, configured to cascade the multi-scale features and input them into a quality prediction network, and obtain a perceived quality score by multiplying with the connection weights and bias values.

[0016] Optionally, the high-dimensional feature extraction module inputs the target Y-channel information into the implicit neural guidance network and outputs the high-dimensional features of the target Y-channel information, including:

[0017] Obtain the corresponding coordinates and dimensions according to the spatial resolution of the light field sub-aperture array image, process the coordinates using a position encoding function and sine and cosine transforms, obtain high-frequency encoding and splice it into the coordinates to obtain the spliced coordinates;

[0018] Unfold the light field sub-aperture array image through F.unfold, convert it into feature maps with n angular resolutions, and perform size and dimension conversion on the feature maps to obtain the converted feature maps;

[0019] Perform local integration operations on each converted feature map, generate several converted feature maps with different perspectives, and input them into the feature extraction network MLP for feature extraction, obtain the predicted values of each pixel under different perspectives and perform weighted summation, and output a high-dimensional feature map.

[0020] Optionally, obtain the predicted values of each pixel under different perspectives and perform weighted summation to output the high-dimensional feature map, including:

[0021] Determine the relative coordinate information of each pixel in the spliced coordinates, splice the coordinate information with the pixel features and input them into the feature extraction network MLP, and output the predicted values of each pixel under different perspectives;

[0022] Calculate the area of the region through the relative coordinate information, and weight the predicted values according to the area of the region to obtain the region weights of each pixel under each perspective;

[0023] The predicted values of each pixel at different viewing angles are weighted and summed according to the regional weights to output the high-dimensional feature map.

[0024] Optionally, the multi-dimensional feature extraction module inputs the high-dimensional feature into a multi-scale semantic feature extraction network to output the multi-scale features of the target Y-channel information, including:

[0025] After reducing the dimension of the high-dimensional feature through a dimensionality reduction convolutional layer, it is sent into the multi-scale semantic feature extraction network to obtain multi-scale Y-channel information.

[0026] Optionally, the perception module inputs the highest-dimensional feature in the multi-scale features into a hypernetwork established by perception rules to output connection weights and bias values, including:

[0027] After reducing the dimension of the highest-dimensional Y-channel information through a max pooling layer, it is sent into the hypernetwork established by perception rules. After reducing the number of channels through a convolutional layer, it uses an adaptive average pooling layer for intermediate resolution compression, and outputs connection weights and bias values after passing through a fully connected layer.

[0028] Optionally, the evaluation module cascades the multi-scale features and inputs them into a quality prediction network, and obtains the perceptual quality score by multiplying pointwise with the connection weights and bias values, including:

[0029] Cascade the multi-scale Y-channel information through a local distortion perception module LDA to obtain a vector containing local distortion information, send it into the quality prediction network, and after dimensionality reduction through several fully connected layers and a Sigmoid activation function, multiply pointwise with the connection weights and bias values to obtain the final perceptual quality score.

[0030] The beneficial effects of the present invention are:

[0031] The present invention obtains a high-dimensional feature representation by using an implicit neural guidance network from the representation forms extracted from the light field sub-aperture image stack (including vertical, horizontal, left diagonal, and right diagonal directions). In addition, combined with a multi-scale semantic feature extraction network, a perceptual quality score is finally generated. This method fully considers the visual characteristics of the human eye and significantly improves the scoring accuracy and effect of no-reference light field image quality evaluation. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0033] Figure 1Flowchart of an implicit neural-guided hypernetwork blind light field image quality evaluation method according to an embodiment of the present invention;

[0034] Figure 2 Network framework diagram of an implicit neural-guided hypernetwork blind light field image quality evaluation method according to an embodiment of the present invention;

[0035] Figure 3 Hypernetwork framework diagram for establishing perception rules according to an embodiment of the present invention;

[0036] Figure 4 Network framework diagram of the quality prediction network according to an embodiment of the present invention. Detailed implementation manners

[0037] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0038] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below in conjunction with the drawings and specific implementation manners.

[0039] This embodiment provides an implicit neural-guided hypernetwork blind light field image quality evaluation method, as Figure 1 shown, including:

[0040] Obtain the light field sub-aperture array image to be evaluated;

[0041] Extract the target Y-channel information of the light field sub-aperture array image;

[0042] Input the target Y-channel information into a preset quality evaluation model, and output a perceptual quality evaluation result, where the quality evaluation model uses an implicit neural guidance network to extract high-dimensional features of the Y-channel information, and obtains a perceptual quality score according to the high-dimensional features.

[0043] Specifically, in this embodiment, by extracting the representation form from the light field sub-aperture image stack, using an implicit neural guidance network to obtain a high-dimensional feature representation, and in addition, combining a multi-scale semantic feature extraction network, a perceptual quality score is finally generated, fully considering the visual characteristics of the human eye, and significantly improving the scoring accuracy and effect of the no-reference light field image quality evaluation.

[0044] Further, extracting the target Y-channel information of the light field sub-aperture array image includes:

[0045] Extract the light field sub-aperture image stacks of vertical, horizontal, left diagonal, and right diagonal from the light field sub-aperture array image, convert the light field sub-aperture image stacks from the RGB space to the YUV space, and extract the Y-channel information in the YUV space.

[0046] Specifically, in this embodiment, the light field sub-aperture images of vertical, horizontal, left diagonal, and right diagonal in the 9×9 light field sub-aperture image array are converted from the RGB space to the YUV space, and the Y-channel information of the light field sub-aperture image stacks is extracted.

[0047] From the light field sub-aperture image stacks of vertical, horizontal, left diagonal, and right diagonal in the 9×9 light field sub-aperture image array, the dimension of the light field sub-aperture image array is 3×u×v×h×w, where u×v is the angular resolution of the sub-aperture images in the light field sub-aperture array image, h×w is the spatial resolution of a single sub-aperture image in the light field sub-aperture array image, and 3 is the three RGB channels. First, the obtained light field sub-aperture image stacks of vertical, horizontal, left diagonal, and right diagonal are converted from the RGB space to the YUV space (each sub-aperture image stack has 9 sub-aperture images). Then, the Y-channel information in the YUV space is extracted, and its dimension is 9×h×w.

[0048] Furthermore, as Figure 2 shown, the quality evaluation model includes:

[0049] A high-dimensional feature extraction module for inputting the target Y-channel information into the implicit neural guidance network and outputting the high-dimensional features of the target Y-channel information;

[0050] A multi-dimensional feature extraction module for inputting the high-dimensional features into a multi-scale semantic feature extraction network and outputting the multi-scale features of the target Y-channel information;

[0051] A perception module for inputting the highest-dimensional feature in the multi-scale features into a super network established based on perception rules and outputting connection weights and bias values;

[0052] An evaluation module for cascading the multi-scale features and inputting them into a quality prediction network, and obtaining a perceived quality score by multiplying with the connection weights and bias values.

[0053] Furthermore, the high-dimensional feature extraction module inputs the target Y-channel information into the implicit neural guidance network and outputs the high-dimensional features of the target Y-channel information, including:

[0054] Obtain the corresponding coordinates and dimensions according to the spatial resolution of the light field sub-aperture array image, process the coordinates using a position encoding function and sine and cosine transforms, obtain high-frequency encoding and splice it into the coordinates to obtain the spliced coordinates;

[0055] Unfold the optical field sub-aperture array image through F.unfold to convert it into feature maps with n angular resolutions, and perform size and dimension conversion on the feature maps to obtain the converted feature maps;

[0056] Perform local integration operations on each converted feature map to generate several converted feature maps from different perspectives, and input them into the feature extraction network MLP for feature extraction. Obtain the predicted values of each pixel from different perspectives and perform weighted summation to output a high-dimensional feature map.

[0057] Among them, obtaining the predicted values of each pixel from different perspectives and performing weighted summation to output the high-dimensional feature map includes:

[0058] Determine the relative coordinate information of each pixel in the coordinates after splicing, splice the coordinate information with the pixel features and then input them into the feature extraction network MLP to output the predicted values of each pixel from different perspectives;

[0059] Calculate the area of the region through the relative coordinate information, and weight the predicted values according to the area of the region to obtain the regional weights of each pixel from each perspective;

[0060] Perform weighted summation on the predicted values of each pixel from different perspectives according to the regional weights to output the high-dimensional feature map.

[0061] Specifically, in this embodiment, the Y-channel information of the optical field sub-aperture image stack obtained vertically, horizontally, diagonally left, and diagonally right is sent into the implicit neural guidance network to obtain the high-dimensional feature representation of the optical field sub-aperture image stack.

[0062] First, the Y-channel information of the obtained light field sub-aperture image stack is fed into the implicit neural guidance network. During the specific calculation process of the implicit neural guidance network, first, coordinates (x, y) corresponding to the spatial resolution h×w of the light field sub-aperture image are generated, with a dimension of h×w. Then, a position encoding function is used to generate high-frequency features related to the coordinates. The coordinate points are encoded through sine and cosine transforms. Finally, the coordinate information is concatenated with the high-frequency encoding to form coordinates with a dimension of h×w×(2L + 3), where L is the number of frequency layers of the encoding. Then, the light field sub-aperture image features are unfolded by F.unfold and converted into feature maps with 9 angular resolutions. The size of each image is expanded from h×w to 9×(h×w), thereby converting the dimension of the image from 3×u×v×h×w to 9×(h×w), where u×v is the angular resolution of the light field sub-aperture array. Next, for each sub-aperture image, a local integration operation is performed. Different disparities vx_lst and vy_lst are used to resample the image to generate multiple images with different perspectives. The dimensions of these images remain 9×(h×w) and are input into the feature extraction network (MLP).

[0063] During the feature extraction process, the features of each pixel are concatenated with their relative coordinate information and input into the MLP network. Assuming the input feature dimension is 9×(h×w), after being processed by the MLP, the output feature dimension is still 9×(h×w), which is the predicted value of each pixel under different perspectives.

[0064] During the weighting process, first, the regional weights of each pixel under each perspective are calculated. The regional area is calculated through relative coordinates, and the predicted values are weighted according to the area. The dimension of the regional weights is h×w, which is used to balance the contributions of images under different perspectives.

[0065] Finally, all the prediction results from different perspectives are weighted and summed according to the regional weights. The weighted sum process ensures the merging of information from different perspectives and adjusts the contributions of different regions according to the regional area. Finally, a feature map with a dimension of 9×h×w is output and the final result is returned.

[0066] Furthermore, the multi-dimensional feature extraction module inputs the high-dimensional features into a multi-scale semantic feature extraction network and outputs multi-scale features of the target Y-channel information, including:

[0067] After the high-dimensional features are dimension-reduced by a dimension-reducing convolutional layer, they are fed into the multi-scale semantic feature extraction network to obtain multi-scale Y-channel information.

[0068] Specifically, in this embodiment, the high-dimensional features of the Y-channel information of the obtained light field sub-aperture image stack are fed into the multi-scale semantic feature extraction network to obtain multi-scale Y-channel information of the light field sub-aperture image array.

[0069] The high-dimensional features of the Y-channel information of the obtained light field sub-aperture image stack are first passed through a 1×1 convolutional kernel of a dimensionality reduction convolutional layer with a stride of 1 and a padding value of 0, reducing the 9×H×W high-dimensional features to 3×h×w. Then it is fed into the multi-scale semantic feature extraction network and processed by four feature extraction modules of different scales to obtain four multi-scale sub-aperture array image Y-channel information with dimensions of 64×64×C1, 32×32×C2, 16×16×C3, and 8×8×C4. Among them, C1, C2, C3, and C4 are the number of channels that increase with the deepening of the layer.

[0070] Furthermore, the perception module inputs the highest-dimensional feature in the multi-scale features into the hypernetwork for establishing perception rules and outputs the connection weights and bias values, including:

[0071] After the highest-dimensional Y-channel information is dimensionally reduced by the max pooling layer, it is fed into the hypernetwork for establishing perception rules. After the channel reduction by the convolutional layer, the adaptive average pooling layer is used for intermediate resolution compression, and the connection weights and bias values are output after passing through the fully connected layer.

[0072] Specifically, in this embodiment, the highest-dimensional light field sub-aperture image array Y-channel information obtained is fed into the hypernetwork for establishing perception rules to obtain the connection weights and bias values between the FC layers.

[0073] As Figure 3 shown, the highest-dimensional light field sub-aperture image array Y-channel information obtained, with dimensions of 8×8×C4, is first sent to a max pooling layer with a convolutional kernel size of 2, a stride of 1, and a padding value of 0 to be dimensionally reduced to 7×7×C4. Then it is fed into the hypernetwork for establishing perception rules. First, it passes through a convolutional network layer conv1 and undergoes the first convolutional operation using a 1×1 convolutional kernel with a stride of 1 and a padding of 0. Then, the second convolutional operation also uses a 1×1 convolutional kernel. Finally, the third convolutional operation still uses a 1×1 convolutional kernel. The entire process gradually reduces the number of channels to 112 while keeping the spatial dimension unchanged. Then the AdaptiveAvgPool2d pooling operation is used, which compresses the spatial resolution from 7×7 to 1×1, retaining the average information of each channel.

[0074] Next, connection weights and bias values are generated. The final output of the hypernetwork contains multiple feature maps, in the following format: target_in_vec is the initial input feature map vector, target_fc1w and target_fc1b represent the weights and biases of the first fully connected layer respectively, target_fc2w and target_fc2b represent the weights and biases of the second fully connected layer respectively, target_fc3w and target_fc3b represent the weights and biases of the third fully connected layer respectively, target_fc4w and target_fc4b represent the weights and biases of the fourth fully connected layer respectively, and target_fc5w and target_fc5b represent the weights and biases of the last fully connected layer respectively.

[0075] Furthermore, the evaluation module cascades the multi-scale features and inputs them into the quality prediction network, and obtains the perceptual quality score by multiplying with the connection weights and bias values, including:

[0076] Cascade the multi-scale Y-channel information through the local distortion perception module LDA to obtain a vector containing local distortion information, send it into the quality prediction network, and after dimensionality reduction through several fully connected layers and the Sigmoid activation function, multiply with the connection weights and bias values to obtain the final perceptual quality score.

[0077] Specifically, in this embodiment, the obtained multi-scale Y-channel information of the light field sub-aperture image array is cascaded and then sent into the quality prediction network, and the final perceptual quality score is obtained by multiplying with the connection weights and bias values.

[0078] As Figure 4 shown, the obtained multi-scale Y-channel information of the light field sub-aperture image array is first passed through the local distortion perception module (LDA), cascaded to form a vector containing local distortion information and sent into the quality prediction network. This network includes multiple fully connected layers (FC layers), and each fully connected layer is followed by a Sigmoid activation function. Finally, through a series of linear transformations and non-linear activation functions, the final perceptual quality score is obtained by multiplying with the connection weights and bias values.

[0079] The following provides a specific evaluation example based on an implicit neural-guided hypernetwork blind light field image quality evaluation method provided in this embodiment:

[0080] (1) Convert the vertical, horizontal, left diagonal and right diagonal light field sub-aperture image stacks from RGB space to YUV space, and extract the Y channel information of the light field sub-aperture image stacks. In this embodiment, the dimensions of the vertical, horizontal, left diagonal and right diagonal light field sub-aperture image stacks are 3×9×256×256. The dimension of the extracted Y channel information of the light field sub-aperture image stack is 9×256×256.

[0081] (2) First, the Y channel information of the light field sub-aperture image stack is input into the implicit neural guidance network. The image dimension is 9×256×256, where 9 represents the number of sub-apertures of the light field image and 256×256 is the spatial resolution of each sub-aperture image. In order to introduce spatial position information, the make_coord function is first used to generate coordinates to obtain a coordinate representation with a dimension of 2×256×256, which represents the x and y coordinates of each pixel. Next, the generated coordinates are position encoded, and the dimension becomes 18×256×256. This process maps the coordinates of each pixel to a higher dimension through frequency transformation, thereby obtaining richer spatial information. The dimension of the position encoding is 2L, where L is the predefined number of frequencies (here 8), and each coordinate generates a 16-dimensional (2×8=16) encoding result. Then, the coordinate information and the position encoding are spliced ​​together to obtain a 19×256×256 feature representation, which provides richer feature input for subsequent calculations.

[0082] Next, the input image is expanded (via F.unfold) to convert the image from 9×256×256 to 81×256×256. The expansion operation converts each small area of ​​the image into a higher dimensional feature map, where the information of the 9 sub-aperture images is combined to form a higher dimensional feature representation. The expanded features are then concatenated with the relative coordinates again to obtain a 83×256×256 feature representation. The relative coordinates represent the relative position of each pixel, which is combined with the expanded features to provide more detailed spatial information.

[0083] If cell decoding is enabled (cell_decode=True), cell information is further added. The cell position of each pixel is magnified to the spatial scale of the image, and the cell information dimension is increased by 2 dimensions, resulting in a feature input of 85×256×256. After that, all features are input into the multi-layer perceptron (MLP) network for processing. The input dimension is 85×256×256. After the layer-by-layer linear transformation and activation function of the MLP, the output dimension becomes 9×256×256, and finally each pixel outputs a 9-dimensional feature vector.

[0084] Finally, perform weighted average processing on the output feature map, calculate the weighted values of the feature values in different regions, and keep the dimension as 9×256×256.

[0085] (3) First, pass the high-dimensional features of the Y-channel information of the optical sub-aperture image stack obtained in step (2) through a dimensionality reduction convolutional layer with a convolutional kernel of 1×1, a stride of 1, and a padding value of 0 to reduce the 9×256×256 high-dimensional features to 3×256×256. Then send them into the multi-scale semantic feature extraction network of Swin Transformer V2 Tiny. After being processed by four feature extraction modules at different scales, four multi-scale sub-aperture array image Y-channel information are obtained, with dimensions of 64×64×96, 32×32×192, 16×16×384, and 8×8×768, where the number of channels increases as the level deepens.

[0086] (4) The Y-channel information of the highest-dimensional optical sub-aperture image array obtained in step (3), with a dimension of 8×8×768, is first sent to a max pooling layer with a convolution kernel size of 2, a stride of 1, and a padding value of 0 to be reduced in dimension to 7×7×768. Then, the feature map is processed through the conv1 module, which includes multiple 1×1 convolutional layers. The first convolution reduces the number of channels from 768 to 384 and is followed by ReLU activation; the second convolution reduces the number of channels from 384 to 192 and is again followed by ReLU activation; the last convolution reduces the number of channels from 192 to 112 and is also followed by ReLU activation. The final obtained feature map has a size of 112×7×7, that is, the number of channels is 112 and the spatial size remains 7×7. At this stage, the target input vector target_in_vec is extracted from res_out['target_in_vec'] and reshaped into 224×1×1, and the LDA result vector is converted into a tensor with a spatial dimension of 1×1. Next, the network further processes the feature map through multiple convolutional layers, gradually reducing the spatial size. Specifically, the feature map generated by the convolution of fc1w_conv has a size of 112×224×1×1, the number of channels is 112, the target input size is 224, and the spatial size is 1×1, while the bias term dimension of fc1b_fc is 112. The convolution of fc2w_conv reduces the number of channels from 112 to 56, and the dimension of the output feature map is 56×112×1×1, and the bias term dimension of fc2b_fc is 56. Similarly, the convolution of fc3w_conv reduces the number of channels from 56 to 28, and the dimension of the output feature map is 28×56×1×1, and the bias term dimension is 28. The convolution of fc4w_conv reduces the number of channels from 28 to 14, and the dimension of the output feature map is 14×28×1×1, and the bias term dimension is 14. Finally, the convolution weights generated by fc5w_fc reduce the number of channels from 14 to 1, and the dimension of the output feature map is 1×14×1×1, and the bias term is generated by fc5b_fc with a dimension of 1. Each convolutional layer and the corresponding fully connected layer generate corresponding weights and biases, which have the following dimensions respectively: target_fc1w is 112×224×1×1, target_fc1b is 112, target_fc2w is 56×112×1×1, target_fc2b is 56, target_fc3w is 28×56×1×1, target_fc3b is 28, target_fc4w is 14×28×1×1, target_fc4b is 14, target_fc5w is 1×14×1×1, target_fc5b is 1. Through this step-by-step processing structure, the model can effectively extract and adjust the features in the input image, and finally generate the adjusted feature map and related bias terms.

[0087] (5) Before feeding the multi-scale semantic features obtained in step (3) into the quality prediction network, they first pass through a Local Distortion Awareness module (LDA). The dimension of the LDA is gradually transformed according to the operations of different convolutional layers, pooling layers, and fully connected layers. The purpose of the LDA module is to extract features related to local distortion from different feature maps and map them to a fixed dimension through a fully connected layer. In the LDA module, the processing of each feature map is gradually transformed through convolutional, pooling, and fully connected layers. First, LDA1 processes the input feature map from layer1, with a size of 1×256×56×56. Using a 1×1 convolutional kernel, the number of channels is reduced from 256 to 16, the convolutional stride is 1, and the padding is 0. The size of the output feature map is 1×16×56×56. Then, 7×7 average pooling is applied with a stride of 7 to reduce the spatial size to 1×16×8×8. Next, it is flattened into a one-dimensional vector of 1×1024 through a flattening operation and mapped to an output of 1×16 through a fully connected layer. LDA2 processes the feature map from layer2, with a size of 1×512×28×28. Similarly, using a 1×1 convolutional kernel, the number of channels is reduced from 512 to 32, and the size of the output feature map is 1×32×28×28. Then, 7×7 average pooling is applied with a stride of 7 to reduce the spatial size to 1×32×4×4. After flattening to 1×512, it is mapped to 1×16. LDA3 processes the input feature map from layer3, with a size of 1×1024×14×14. The number of channels is reduced from 1024 to 64 through a 1×1 convolutional kernel, and the output size is 1×64×14×14. Then, 7×7 average pooling is applied with a stride of 7 to reduce the spatial size to 1×64×2×2. After flattening to 1×256, it is mapped to 1×16 through a fully connected layer. Finally, LDA4 processes the feature map from layer4, with a size of 1×2048×7×7. 7×7 average pooling is directly applied with a stride of 7 to reduce the spatial dimension to 1×2048×1×1. After flattening, it gets 1×2048 and is mapped to an output of 1×176 through a fully connected layer. Through this series of convolutional, pooling, and fully connected operations, the LDA module gradually extracts local distortion features and maps them to the final output dimension. The output features of all LDA layers are concatenated together to form a vector containing local distortion information, with a dimension of 1×224.

[0088] Next, the vector containing local distortion information is fed into the quality prediction network. The quality prediction network consists of: the input passes through the first fully connected layer l1, which converts the input dimension from 224 to 112. After being processed by the Sigmoid activation function, the output dimension is 112. Next, it passes through the second fully connected layer l2, which maps the input dimension from 112 to 56. After passing through the Sigmoid activation function, the output dimension is 56. Then, the input passes through the third fully connected layer l3, which maps the dimension from 56 to 28 and is processed by the Sigmoid activation function, with an output dimension of 28. Finally, the input passes through the fourth fully connected layer l4. First, it maps 28 to 14, then passes through the Sigmoid activation function, and then passes through the fifth fully connected layer to map 14 to the final output dimension of 1. After removing the single dimension through.squeeze(), the final quality prediction score is output, with a dimension of 1.

[0089] The embodiments described above are only descriptions of the preferred embodiments of the present invention and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope determined by the claims of the present invention.

Claims

1. An implicit neural-guided hypernetwork blind light field image quality assessment method, characterized in that: include: Acquire a sub-aperture array image of the light field to be evaluated; Extracting target Y channel information of the light field sub-aperture array image; The target Y channel information is input into a preset quality evaluation model, and a perceptual quality evaluation result is output, wherein the quality evaluation model uses an implicit neural guidance network to extract high-dimensional features of the Y channel information, and obtains a perceptual quality score based on the high-dimensional features.

2. The implicit neural-guided super-network blind light field image quality assessment method according to claim 1 is characterized in that: Extracting target Y channel information of the light field sub-aperture array image includes: Extracting vertical, horizontal, left diagonal and right diagonal light field sub-aperture image stacks from the light field sub-aperture array image, converting the light field sub-aperture image stacks from RGB space to YUV space, and extracting the Y channel information in the YUV space.

3. The implicit neural-guided super-network blind light field image quality assessment method according to claim 1 is characterized in that: The quality evaluation model includes: A high-dimensional feature extraction module, used for inputting the target Y channel information into the implicit neural guidance network and outputting high-dimensional features of the target Y channel information; A multi-dimensional feature extraction module, used for inputting the high-dimensional features into a multi-scale semantic feature extraction network, and outputting the multi-scale features of the target Y channel information; A perception module, used for inputting the highest dimensional feature among the multi-scale features into a hypernetwork established based on perception rules, and outputting connection weights and bias values; The evaluation module is used to cascade the multi-scale features and input them into the quality prediction network, and obtain the perceptual quality score by multiplying the multi-scale features with the connection weight and the bias value.

4. The implicit neural-guided super-network blind light field image quality assessment method according to claim 3 is characterized in that: The high-dimensional feature extraction module inputs the target Y channel information into an implicit neural guidance network and outputs high-dimensional features of the target Y channel information, including: Acquire corresponding coordinates and dimensions according to the spatial resolution of the light field sub-aperture array image, process the coordinates using a position encoding function and sine and cosine transforms, obtain high-frequency encoding and splice it into the coordinates in combination with the dimensions, and obtain the spliced ​​coordinates; Expand the light field sub-aperture array image through F.unfold to convert it into a feature map with n angular resolutions, perform size and dimension conversion on the feature map, and obtain a converted feature map; A local integration operation is performed on each converted feature map to generate several converted feature maps of different perspectives, which are input into the feature extraction network MLP for feature extraction, and the predicted value of each pixel at different perspectives is obtained and weighted summed to output a high-dimensional feature map.

5. The implicit neural-guided super-network blind light field image quality assessment method according to claim 4 is characterized in that: Obtain the predicted value of each pixel at different viewing angles and perform weighted summation to output the high-dimensional feature map, including: Determine the relative coordinate information of each pixel in the spliced ​​coordinates, splice the coordinate information with the pixel features and input them into the feature extraction network MLP, and output the predicted value of each pixel at different viewing angles; Calculating the area of ​​the region through the relative coordinate information, and weighting the predicted value according to the area of ​​the region to obtain the regional weight of each pixel at each viewing angle; The predicted values ​​of each pixel at different viewing angles are weighted and summed according to the regional weight, and the high-dimensional feature map is output.

6. The implicit neural-guided super-network blind light field image quality assessment method according to claim 3 is characterized in that: The multidimensional feature extraction module inputs the high-dimensional features into a multi-scale semantic feature extraction network and outputs the multi-scale features of the target Y channel information, including: The high-dimensional features are processed by the dimensionality reduction convolution layer and then sent to the multi-scale semantic feature extraction network to obtain multi-scale Y channel information.

7. The implicit neural-guided super-network blind light field image quality assessment method according to claim 3 is characterized in that: The perception module inputs the highest dimensional feature in the multi-scale feature into the hypernetwork established by the perception rule, and outputs connection weights and bias values, including: The highest-dimensional Y channel information is processed by the maximum pooling layer for dimensionality reduction and then sent to the hypernetwork established by the perception rule. After the channel reduction is performed by the convolutional layer, the inter-resolution compression is performed using the adaptive average pooling layer. After passing through the fully connected layer, the connection weights and bias values ​​are output.

8. The implicit neural-guided super-network blind light field image quality assessment method according to claim 3 is characterized in that: The evaluation module cascades the multi-scale features and inputs them into the quality prediction network, and obtains the perceived quality score by dot multiplication with the connection weight and the bias value, including: The multi-scale Y channel information is cascaded through the local distortion perception module LDA to obtain a vector containing local distortion information, which is sent to the quality prediction network. After dimension reduction processing through several fully connected layers and Sigmoid activation functions, it is dot-multiplied with the connection weight and bias value to obtain the final perceived quality score.

Citation Information

Patent Citations

  • No-reference light field image quality evaluation method based on high-dimensional discrete cosine transform

    CN112950592A

  • No-reference light field image quality evaluation method and system based on light field decoupling

    CN117495841A

  • ViT-based light field image quality evaluation method for carrying out spatial domain and frequency domain feature fusion

    CN117522813A

  • Full-reference quality evaluation method for light field image

    CN117911367A

  • Light field image full-reference quality evaluation method based on lens image

    CN117934408A

Cited By

  • Light field image quality evaluation method combining parallax perception and capsule network

    CN121258908A