Light field image no-reference quality evaluation method and system

Through Fourier channel attention and multi-scale fusion strategy, combined with KANs network, the problem of insufficient feature extraction and multi-scale fusion in light field image quality evaluation is solved, and high-precision reference-free quality evaluation is achieved.

CN120471854APending Publication Date: 2025-08-12江淮前沿技术协同创新中心 +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510548704.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

The existing light field image quality evaluation methods are difficult to achieve high-precision quality prediction in terms of insufficient feature extraction, poor multi-scale fusion effect and insufficient regression model performance.

Method used

The Fourier channel attention is used to extract the spatial domain characteristics of the light field image and enhance the frequency domain characteristics. Combined with the progressive multi-scale fusion strategy, the KANs network is used for reference-free quality evaluation.

Benefits of technology

It significantly improves feature extraction capabilities, comprehensively captures multi-dimensional quality information of light field images, and improves the accuracy and generalization capabilities of quality evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120471854A_ABST
    Figure CN120471854A_ABST
Patent Text Reader

Abstract

The invention provides a light field image no-reference quality evaluation method, and belongs to the field of light field image quality evaluation. The method comprises the following steps: S1, preprocessing a light field image to obtain different two-dimensional slices; s2, extracting spatial domain features of each slice by using Fourier channel attention, further obtaining and enhancing frequency domain features corresponding to each spatial domain feature, and correspondingly fusing the spatial domain features and the enhanced frequency domain features to obtain reconstruction features of each slice; s3, gradually fusing each reconstruction feature and carrying out multi-scale integration to obtain a total fusion feature; and S4, inputting the total fusion feature into a KANs network to obtain a no-reference quality score. According to the invention, the fusion of the frequency domain features and the spatial domain features is enhanced, so that the key features of the light field image can be simultaneously reserved in two dimensions; a progressive multi-scale fusion strategy is adopted, so that the features contain local details and also capture global context information; a nonlinear mapping task can be processed more efficiently by using the KANs network, and stronger generalization ability and quality prediction performance are shown.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of light field image quality evaluation, and in particular to a light field image quality evaluation method and system without reference. Background Art

[0002] With the rapid development of virtual reality and immersive multimedia technologies, light field imaging (LFI) has become a research hotspot in computer vision and image processing due to its ability to record the direction and intensity of light in three-dimensional space. By capturing rich spatial and angular information, light field technology enables highly realistic 3D reconstruction and depth estimation, and is widely used in fields such as virtual reality, panoramic displays, and medical imaging.

[0003] However, the high-dimensional nature of light field images makes them susceptible to various distortions during acquisition, compression, transmission, and display. These distortions not only degrade the integrity of the light field content, but also disrupt the user's visual experience and even pose potential risks to visual health. Therefore, effective assessment of light field image quality has become a key challenge in advancing the theoretical research and practical application of light field technology.

[0004] Early light field image quality assessment methods mainly relied on manually designed features. For example, the paper "A No-Reference Image Quality Assesment Metric by Multiple Characteristics of LightField Images" (Liang Shan, Ping An, Chunli Meng, Xinpeng Huang, Chao Yang, and Liquan Shen. IEEE Access, vol. PP, no. 99, pp. 1-1, 2019.) combines the three-dimensional features of depth maps and the two-dimensional features of sub-aperture images (SAI) to evaluate light field quality. Another example is the paper "Tensor-Oriented No-Refeience Light Field Image Quality Assessment" (Wei Zhou, Likun Shi, Zhibo Chen, Jinglin Zhang. IEEE Transactions on Image Processing, Vol. 29, pp. 4070-4084, 2020), which establishes a light field image quality model based on tensor theory. However, these traditional methods have limitations in feature selection and are unable to fully capture the complex characteristics of light fields.

[0005] In recent years, deep learning technology has made some progress in light field image quality assessment. For example, the paper "DEEBLIF: DEEP BLIND LIGHT FIELD IMAGE QUALITY ASSESSMENT BY EXTRACTING ANGLARAND SPATIAL INFORMATION" (Zhengyu Zhang, Shishun Tian, Wenbin Zou, Luce Morin, and Lu Zhang. 2022 IEEE International Conference on Image Processing, pp. 2266-2270, 2022) proposes a reference-free light field quality assessment model based on a two-stream convolutional neural network (CNN); the paper "Reduced Reference Quality Assessment of Light Field Images" (Pradip Paudyal, Federica Battisti, Marco Carli. IEEE TRANSACTIONS ON BROADCASTING, VOL. 65, NO. 1, MARCH, 2019.) indirectly reflects the light field quality by evaluating the quality of the light field depth map; and the paper "Blind light field image quality assessment based on deep meta-learning" (JIAN MA, XIAOYIN ZHANG, AND JUNBO WANG. Optics Letters, Vol. 48, No. 23, pp. 6184-6187, 2023) uses meta-learning to solve the problem of limited sample size. Despite this, existing methods still face the following problems: (1) Insufficient feature extraction: Existing methods have difficulty in extracting key features related to image quality, resulting in limited model performance; (2) Limited effect of multi-scale feature fusion: How to effectively integrate features of different scales to comprehensively evaluate the quality dimension of light field images remains a challenge; (3) Regression model limitations: Traditional fully connected networks (such as multi-layer perceptrons, MLPs) perform poorly in quality regression tasks, making it difficult to achieve high-precision quality prediction. Summary of the Invention

[0006] The technical problem to be solved by the present invention is how to solve the problems of insufficient feature extraction and poor multi-scale fusion effect in light field image quality assessment.

[0007] To solve the above technical problems, the present invention provides the following technical solution: a method for light field image quality assessment without reference, comprising the following steps:

[0008] S1: Preprocess the input light field image to obtain different two-dimensional slices;

[0009] S2: Use Fourier channel attention to extract the spatial domain features of each 2D slice, further obtain the frequency domain features corresponding to each spatial domain feature and enhance them, and fuse the spatial domain features and the enhanced frequency domain features to obtain the reconstructed features of each 2D slice;

[0010] S3: Gradually fuse the reconstructed features and perform multi-scale integration to obtain the total fusion feature;

[0011] S4: The total fusion features are input into the KANs network to obtain the final no-reference quality score.

[0012] The present invention extracts frequency domain features through Fourier transform and uses the channel attention mechanism to enhance features related to image quality. It can effectively capture representative information related to quality, thereby significantly improving the feature extraction capability. The enhanced frequency domain features are fused with spatial domain features to simultaneously retain the key features of light field images in two dimensions. In addition, the present invention adopts progressive, multi-scale fusion reconstruction features, which can make full use of complementary information from different feature maps, fuse multi-scale features under different receptive fields, and comprehensively capture the multi-dimensional quality information of light field images, which can improve the accuracy and comprehensiveness of quality assessment. Compared with the traditional MLP, the present invention uses the KANs network to more efficiently process nonlinear mapping tasks, showing stronger generalization ability and quality prediction performance.

[0013] Preferably, in step S1, the two-dimensional slice includes a sub-aperture image, a macro-pixel image, a horizontal extreme plane image and a vertical extreme plane image.

[0014] Preferably, the specific process of step S1 is:

[0015] S11: Convert the light field image format into a shape of R(U×V×H×W×C), where U and V are angular resolutions, H and W are height and width respectively, and C represents the channel;

[0016] S12: Reshape the light field image into a sub-aperture image form R(UV×H×W×C); reshape the light field image into a macro-pixel image form R(HW×U×V×C); reshape the light field image into a horizontal polar plane image form R(VW×U×H×C); reshape the light field image into a vertical polar plane image form R(UH×V×W×C).

[0017] After converting the light field image, the present invention reshapes it into a sub-aperture image form to facilitate the effective extraction of spatial features, reshapes it into a macro-pixel image form to facilitate the extraction of angular features, and reshapes it into a horizontal polar image form and a vertical polar plane image form to facilitate the extraction of spatial angular structural features, which can reduce the complexity of feature extraction.

[0018] Preferably, in step S2, the spatial domain features include spatial features, angular features, horizontal polar plane image features and vertical polar plane image features obtained by respectively extracting features of the sub-aperture image, the macro-pixel image, the horizontal polar plane image and the vertical polar plane image using a 3×3 convolutional layer with a GELU activation function.

[0019] Preferably, the specific process of step S2 is:

[0020] S21: spatial features, angle features, horizontal polar plane image features, and vertical polar plane image features are used as input features;

[0021] S22: Use fast Fourier transform to map each input feature from the spatial domain to the frequency domain, and obtain the channel descriptor corresponding to each input feature through global average pooling;

[0022] S23: Each channel descriptor passes through two convolutional layers with ReLU activation function and Sigmoid activation function to obtain the corresponding channel attention weight;

[0023] S24: The attention weight of each channel is applied to the corresponding input feature to obtain enhanced frequency domain features;

[0024] S25: Return each enhanced frequency domain feature to the corresponding input feature through residual connection, and obtain the reconstructed features of the sub-aperture image, macro-pixel image, horizontal polar plane image and vertical polar plane image respectively.

[0025] The present invention maps the image features of each two-dimensional slice from the spatial domain to the frequency domain. Through frequency domain processing, the representativeness of feature extraction can be significantly improved, and the frequency domain features can be enhanced, which can effectively capture the high-frequency distortion characteristics of the light field image, thereby effectively improving the problem that the traditional light field image quality assessment model is insufficiently sensitive to high-frequency loss. The enhanced frequency domain features are returned to the input features, that is, the spatial domain features are returned. After fusion, features containing multi-dimensional information can be obtained, which can simultaneously retain the key features of the light field image in both the spatial domain and the frequency domain.

[0026] Preferably, in step S22, the formula for mapping each input feature from the spatial domain to the frequency domain using fast Fourier transform is: Where j is an imaginary unit, I(p,q) represents the spatial domain feature with a size of M×N, (p,q) represents the coordinate information, and F(u,v) is the frequency domain feature after mapping, where u and v are the frequency indices in the horizontal and vertical directions respectively, and 0≤u <M,0≤v<N。

[0027] Preferably, the specific process of step S3 is:

[0028] S31: performing preliminary integration of features reconstructed from the sub-aperture image and the macro-pixel image, performing preliminary integration of features reconstructed from the horizontal polar plane image and the vertical polar plane image, and combining the two preliminary integrated features;

[0029] S32: Use multiple convolution kernels of different scales to process the combined features separately;

[0030] S33: Concatenate the combined features processed by convolution kernels of different scales, perform dimensionality reduction through 1×1 convolution, and output the total fusion feature F fused .

[0031] The progressive fusion adopted by the present invention, that is, first fusing two by two and then combining them, can fully utilize the complementary information from different features and improve the ability to evaluate the quality of light field images. In addition, the multi-scale fusion mechanism of the present invention can make the integrated total fusion feature contain local details while also capturing global context information, which can effectively integrate the diversity and hierarchy of multidimensional features in light field images.

[0032] Preferably, in step S32, the multiple convolution kernels of different scales include 7×7, 11×11 and 21×21 convolution kernels.

[0033] Preferably, the specific process of step S4 is: taking the total fusion feature as the input of the KANs network; using the formula Q=W·S k (F fused )+b calculates the no-reference quality score of the light field image, where Q is the no-reference quality score of the light field image, W is the weight, and S k (x) represents the B-spline basis function, and b represents the bias.

[0034] The present invention uses the KANs network to evaluate the quality of light field images and replaces the fixed activation function with a learnable B-spline function. Compared with the traditional fully connected layer, it has stronger nonlinear expression ability and good generalization ability for small sample problems, and can achieve more accurate quality prediction.

[0035] A light field image no-reference quality assessment system includes the following modules:

[0036] Preprocessing module: used to preprocess the input light field image to obtain different two-dimensional slices;

[0037] Obtaining reconstruction feature module: It is used to extract the spatial domain features of each two-dimensional slice using Fourier channel attention, further obtain the frequency domain features corresponding to each spatial domain feature and enhance them, and fuse the spatial domain features and the enhanced frequency domain features to obtain the reconstruction features of each two-dimensional slice;

[0038] Obtaining the total fusion feature module: used to gradually fuse each reconstructed feature and perform multi-scale integration to obtain the total fusion feature;

[0039] Obtaining quality score module: used to input the total fusion features into the KANs network to obtain the final no-reference quality score.

[0040] Compared with the existing technology, the advantages of the present invention are: it can improve the representativeness of feature extraction through frequency domain processing, and after enhancing the frequency domain features and then fusing them with the spatial domain features, it can retain the key features of the light field image in multiple dimensions; it adopts a progressive fusion strategy to integrate spatial, angular and structural features through multi-scale feature fusion, which can fully capture the multi-dimensional quality characteristics of the light field image; it adopts the KANs network and uses a learnable B-spline function instead of a fixed activation function to improve the prediction performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 This is a flow chart of Example 1 of the present invention;

[0042] Figure 2 Schematic diagram of preprocessing a light field image in Example 1 of the present invention;

[0043] Figure 3 This is the Fourier channel attention network diagram in Example 1 of the present invention;

[0044] Figure 4 Schematic diagram of progressive fusion in Example 1 of the present invention;

[0045] Figure 5 These are the experimental results of the method of Example 1 of the present invention conducted on different data sets. DETAILED DESCRIPTION

[0046] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0047] Example 1

[0048] like Figure 1As shown, this embodiment provides a light field image quality assessment method without reference, including the following steps:

[0049] S1: If Figure 2 As shown, the raw data, i.e., the input light field image, is preprocessed to obtain different two-dimensional slices. In this embodiment, there are four types of two-dimensional slices: sub-aperture image (2D SAI), macro pixel image (MacPi), horizontal epipolar plane image (H-EPI), and vertical epipolar plane image (V-EPI), specifically:

[0050] S11: Convert the light field image format into a shape of R(U×V×H×W×C), where U and V are angular resolutions, H and W are height and width respectively, and C represents the channel;

[0051] S12: Reshape the light field image into a sub-aperture image form R (UV×H×W×C) to effectively extract spatial features; reshape the light field image into a macro-pixel image form R (HW×U×V×C) to facilitate the extraction of angular features; reshape the light field image into a horizontal polar plane image form R (VW×U×H×C) to facilitate the extraction of spatial angular structural features; reshape the light field image into a vertical polar plane image form R (UH×V×W×C), which is also used for spatial angular structural feature extraction.

[0052] S2: If Figure 3 As shown in the figure, the spatial domain features of each two-dimensional slice are extracted using Fourier channel attention, and the frequency domain features corresponding to each spatial domain feature are further obtained and enhanced. The spatial domain features and the enhanced frequency domain features are fused to obtain the reconstructed features of each two-dimensional slice.

[0053] The spatial domain features in this embodiment are: spatial features, angular features, horizontal polar plane image features, and vertical polar plane image features obtained by extracting features of the sub-aperture image, macro-pixel image, horizontal polar plane image, and vertical polar plane image, respectively, using a 3×3 convolutional layer with a GELU activation function.

[0054] This embodiment is described by taking the spatial features extracted from the sub-aperture image as an example. The specific process of S2 is as follows:

[0055] S21: Use spatial features as input features;

[0056] S22: Use Fast Fourier Transform (FFT) to map the spatial feature x from the spatial domain to the frequency domain to obtain the frequency domain feature X, that is, X = F(x) = FFT(x). The specific mapping formula is: where \(j\) is the imaginary unit, \(I(p,q)\) represents the spatial domain feature with the size of \(M\times N\), \((p,q)\) represents the coordinate information, \(F(u,v)\) is the frequency domain feature after mapping, \(u\) and \(v\) are the frequency indices in the horizontal and vertical directions respectively, and \(0\leq u\lt M\), \(0\leq v\lt N\). And the channel descriptor \(z\) corresponding to the spatial feature is obtained through global average pooling, that is, \(z = AvgPool(x)\);

[0057] S23: The channel descriptor \(z\) passes through two convolutional layers with ReLU activation function and obtains the corresponding channel attention weight \(G\) through the Sigmoid activation function. The formula is \(G=\sigma(W_2\cdot ReLU(W_1z))\), where \(G\) is the channel attention weight, \(W_1\) and \(W_2\) represent the convolutional layer weights, and \(\sigma()\) represents the Sigmoid activation function;

[0058] S24: Apply the channel attention weight \(G\) to the spatial feature \(x\) to obtain the enhanced frequency domain feature \(X\) enhanced For enhanced representation, that is \(X\) enhanced \(=x\cdot G\);

[0059] S25: Through the residual connection, the enhanced frequency domain feature \(X\) enhanced Returns the spatial feature \(x\) to retain key information, and obtains the feature \(Y\) after reconstruction of the sub-aperture image, \(Y = x+X\) enhanced ;

[0060] Perform the above operations on the angular feature, horizontal polar plane image feature, and vertical polar plane image feature to obtain the features after reconstruction of the macro-pixel image, horizontal polar plane feature image, and vertical polar plane image.

[0061] The present invention extracts the frequency domain features through Fourier transform and uses the channel attention mechanism to enhance the features related to the image quality, which can effectively capture the representative information related to the quality, thereby significantly improving the feature extraction ability. The fusion of the enhanced frequency domain features and the spatial domain features can retain the key features of the light field image in two dimensions simultaneously.

[0062] S3: Gradually fuse each reconstructed feature and perform multi-scale integration to obtain the total fusion feature. Specifically:

[0063] S31: As Figure 4 shown, the features after reconstruction of the sub-aperture image and the macro-pixel image are concatenated along the channel dimension to obtain a preliminarily fused feature. This preliminarily fused feature is processed by two convolutional layers with dilation rates of 1 and \(3\times3\) kernels to extract deep features. The same processing is performed on the features after reconstruction of the horizontal polar plane image and the vertical polar plane image. The two preliminarily fused features are combined along the channel dimension after being processed by the convolutional layer, and the combined feature is further processed by two additional convolutional layers to extract deeper features. The progressive fusion here can make full use of the complementary information from different feature maps;

[0064] S32: Use convolution kernels of different scales, in this embodiment, 7×7, 11×11, and 21×21, to process the combined features processed by the convolution layer respectively;

[0065] S33: Concatenate the combined features processed by convolution kernels of different scales, perform dimensionality reduction through 1×1 convolution, and output the total fusion feature F fused ,The total fusion feature contains both local details and captures global context information.

[0066] The present invention adopts a progressive fusion strategy and multi-scale fusion, which can make full use of the complementary information from different features and improve the ability of light field image quality assessment. It can also make the integrated total fusion feature contain local details while also capturing global context information, and can effectively integrate the diversity and hierarchy of multidimensional features in light field images.

[0067] S4: Input the total fusion features into the KANs network to obtain the final no-reference quality score. Specifically, the total fusion features are used as the input of the KANs network; the formula Q = W·S k (F fused )+b calculates the no-reference quality score of the light field image, where Q is the no-reference quality score of the light field image, W is the weight, and S k (x) represents the B-spline basis function, and b represents the bias. The present invention adopts the KANs network and uses a learnable B-spline function instead of a fixed activation function, which can flexibly express the nonlinear relationship between features and quality scores, further improving the prediction performance.

[0068] This embodiment also conducts extensive experiments on two public datasets (Win 5-LID and SHU). In the experiments, 20% of the data is used for testing and 80% of the data is used for training. Cross-validation is performed 1000 times on each dataset, and the median of the root mean square error (RMSE), linear correlation coefficient (LCC), and Spearman rank correlation coefficient (SROCC) are calculated as objective indicators of quality assessment. The closer the SROCC and LCC are to 1 and the closer the RMSE is to 0, the better the performance. The experimental results are shown in Figure 2. Figure 5 shown.

[0069] It can be seen that compared with the other nine full-reference and no-reference quality assessment methods (SSIM, VIF, NIQE, CHEN, SINQ, BELIF, VBLIF, MDFM, Zhang), the performance of the proposed method on both datasets is better than the existing light field image quality assessment methods, verifying the effectiveness of the proposed method.

[0070] Example 2

[0071] Corresponding to Embodiment 1 of the present invention, this embodiment provides a no-reference quality evaluation system for light field images, including the following modules:

[0072] Preprocessing module: used to preprocess the input light field image to obtain different two-dimensional slices, specifically including the following units:

[0073] Light field image format conversion unit: used to convert the light field image format into a form with a shape of R(U×V×H×W×C), where U and V are angular resolutions, H and W are height and width respectively, and C represents the channel;

[0074] Light field image shaping unit: used to shape the light field image into the form of a sub-aperture image R(UV×H×W×C); shape the light field image into the form of a macro-pixel image R(HW×U×V×C); shape the light field image into the form of a horizontal epipolar plane image R(VW×U×H×C); shape the light field image into the form of a vertical epipolar plane image R(UH×V×W×C).

[0075] Reconstruction feature acquisition module: used to extract the spatial domain features of each two-dimensional slice by using Fourier channel attention, further obtain the corresponding frequency domain features of each spatial domain feature and enhance them, and fuse the spatial domain features and the enhanced frequency domain features to obtain the reconstruction features of each two-dimensional slice, specifically including the following units:

[0076] Input feature acquisition unit: used to extract the spatial feature, angular feature, horizontal epipolar plane image feature and vertical epipolar plane image feature obtained by using a 3×3 convolutional layer with a GELU activation function for the sub-aperture image, macro-pixel image, horizontal epipolar plane image and vertical epipolar plane image respectively, and use the spatial feature, angular feature, horizontal epipolar plane image feature and vertical epipolar plane image feature as input features;

[0077] Frequency domain mapping and average pooling unit: used to map each spatial domain feature from the spatial domain to the frequency domain by using the fast Fourier transform (FFT) to obtain the frequency domain feature, and the mapping formula is: where j is the imaginary unit, I(p,q) represents the spatial domain feature with a size of M×N, (p,q) represents the coordinate information, F(u,v) is the mapped frequency domain feature, u and v are the frequency indices in the horizontal and vertical directions respectively, and 0≤u<M, 0≤v<N, and the channel descriptor z corresponding to each input feature is obtained through global average pooling, that is, z = AvgPool(x);

[0078] Obtain channel attention weight unit: used to pass each channel descriptor through two layers of convolutional layers with ReLU activation function and Sigmoid activation function to obtain the corresponding channel attention weight G. The formula is G = σ(W2 ReLU(W1z)), where G is the channel attention weight, W1 and W2 are the convolutional layer weights, and σ() is the Sigmoid activation function;

[0079] Obtain enhanced frequency domain feature unit: used to apply channel attention weights to each spatial domain feature to obtain enhanced frequency domain features;

[0080] Obtaining the reconstructed feature unit: used to return the enhanced frequency domain features to the corresponding spatial domain features through residual connections to retain key information and obtain the reconstructed features of each two-dimensional slice.

[0081] Obtaining the total fusion feature module: It is used to gradually fuse the various reconstructed features and perform multi-scale integration to obtain the total fusion feature. It specifically includes the following units:

[0082] Progressive fusion unit: This unit is used to concatenate the features reconstructed from the sub-aperture image and the macro-pixel image along the channel dimension to obtain a preliminary fused feature. This preliminary fused feature is processed by two convolutional layers with a dilation rate of 1 and a 3×3 kernel to extract deep features. The same processing is performed on the features reconstructed from the horizontal and vertical polar plane images. The two preliminary fused features are processed by the convolutional layer and then combined along the channel dimension. The combined features are then processed by two additional convolutional layers to extract deeper features. The progressive fusion here can fully utilize the complementary information from different feature maps.

[0083] Multi-scale fusion unit: used to use multiple convolution kernels of different scales to process the combined features processed by the convolution layer;

[0084] Output total fusion feature unit: used to connect the combined features processed by convolution kernels of different scales in series, and perform dimensionality reduction through 1×1 convolution to output the total fusion feature F fused ,The total fusion feature contains both local details and captures global context information.

[0085] Obtaining quality score module: It is used to input the total fusion features into the KANs network to obtain the final no-reference quality score. Specifically, the total fusion features are used as the input of the KANs network; the formula Q = W·S is used. k (F fused )+b calculates the no-reference quality score of the light field image, where Q is the no-reference quality score of the light field image, W is the weight, and S k(x) represents the B-spline basis function, and b represents the bias. The present invention adopts the KANs network and uses a learnable B-spline function instead of a fixed activation function, which can flexibly express the nonlinear relationship between features and quality scores, further improving the prediction performance.

[0086] This embodiment first obtains sub-aperture images, macro-pixel images, horizontal polar plane images and vertical polar plane images that are easy to extract features through the preprocessing module, and then uses the acquisition and reconstruction feature module to perform frequency domain mapping, frequency domain feature enhancement, etc. to obtain the features of each two-dimensional slice reconstruction. These features can retain the key features of the light field image in both spatial and frequency domain dimensions. Then, the total fusion feature acquisition module performs progressive fusion and multi-scale fusion to obtain the total fusion feature. At this time, the total fusion feature effectively integrates the diversity and hierarchy of the multi-dimensional features of the light field image. Finally, by obtaining the quality score module, the B-spline function that can be learned by the KANs network is used to replace the fixed activation function, which can flexibly express the nonlinear relationship between the feature and the quality score, further improve the prediction performance, and finally obtain a reliable reference-free quality score.

[0087] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A light field image no-reference quality assessment method, characterized in that: The following steps are involved: S1: Preprocess the input light field image to obtain different two-dimensional slices; S2: Use Fourier channel attention to extract the spatial domain features of each 2D slice, further obtain the frequency domain features corresponding to each spatial domain feature and enhance them, and fuse the spatial domain features and the enhanced frequency domain features to obtain the reconstructed features of each 2D slice; S3: Gradually fuse the reconstructed features and perform multi-scale integration to obtain the total fusion feature; S4: The total fusion features are input into the KANs network to obtain the final no-reference quality score.

2. The light field image no-reference quality assessment method according to claim 1, characterized in that: In step S1, the two-dimensional slice includes a sub-aperture image, a macro-pixel image, a horizontal polar plane image, and a vertical polar plane image.

3. The light field image no-reference quality assessment method according to claim 2, characterized in that: The specific process of step S1 is: S11: Convert the light field image format into a shape of R(U×V×H×W×C), where U and V are angular resolutions, H and W are height and width respectively, and C represents the channel; S12: Reshape the light field image into a sub-aperture image form R(UV×H×W×C); reshape the light field image into a macro-pixel image form R(HW×U×V×C); reshape the light field image into a horizontal polar plane image form R(VW×U×H×C); reshape the light field image into a vertical polar plane image form R(UH×V×W×C).

4. The light field image no-reference quality assessment method according to claim 3, characterized in that: In step S2, the spatial domain features include spatial features, angular features, horizontal polar plane image features, and vertical polar plane image features obtained by extracting features of the sub-aperture image, macro-pixel image, horizontal polar plane image, and vertical polar plane image respectively using a 3×3 convolutional layer with a GELU activation function.

5. The light field image no-reference quality assessment method according to claim 4, characterized in that: The specific process of step S2 is: S21: spatial features, angle features, horizontal polar plane image features, and vertical polar plane image features are used as input features; S22: Use fast Fourier transform to map each input feature from the spatial domain to the frequency domain, and obtain the channel descriptor corresponding to each input feature through global average pooling; S23: Each channel descriptor passes through two convolutional layers with ReLU activation function and Sigmoid activation function to obtain the corresponding channel attention weight; S24: The attention weight of each channel is applied to the corresponding input feature to obtain enhanced frequency domain features; S25: Return each enhanced frequency domain feature to the corresponding input feature through residual connection, and obtain the reconstructed features of the sub-aperture image, macro-pixel image, horizontal polar plane image and vertical polar plane image respectively.

6. The light field image no-reference quality assessment method according to claim 5, characterized in that: In step S22, the formula for mapping each input feature from the spatial domain to the frequency domain using fast Fourier transform is: Where j is an imaginary unit, I(p,q) represents the spatial domain feature with a size of M×N, (p,q) represents the coordinate information, and F(u,v) is the frequency domain feature after mapping, where u and v are the frequency indices in the horizontal and vertical directions respectively, and 0≤u <M,0≤v<N。 7. The light field image no-reference quality assessment method according to claim 5, characterized in that: The specific process of step S3 is: S31: performing preliminary integration of features reconstructed from the sub-aperture image and the macro-pixel image, performing preliminary integration of features reconstructed from the horizontal polar plane image and the vertical polar plane image, and combining the two preliminary integrated features; S32: Use multiple convolution kernels of different scales to process the combined features separately; S33: Concatenate the combined features processed by convolution kernels of different scales, perform dimensionality reduction through 1×1 convolution, and output the total fusion feature F fused .

8. The light field image no-reference quality assessment method according to claim 7, characterized in that: In step S32, the multiple convolution kernels of different scales include 7×7, 11×11 and 21×21 convolution kernels.

9. The light field image no-reference quality assessment method according to claim 7, characterized in that: The specific process of step S4 is: taking the total fusion feature as the input of the KANs network; using the formula Q = W·S k (F fused )+b calculates the no-reference quality score of the light field image, where Q is the no-reference quality score of the light field image, W is the weight, and S k (x) represents the B-spline basis function, and b represents the bias.

10. A light field image no-reference quality assessment system, characterized in that: Includes the following modules: Preprocessing module: used to preprocess the input light field image to obtain different two-dimensional slices; Obtaining reconstruction feature module: It is used to extract the spatial domain features of each two-dimensional slice using Fourier channel attention, further obtain the frequency domain features corresponding to each spatial domain feature and enhance them, and fuse the spatial domain features and the enhanced frequency domain features to obtain the reconstruction features of each two-dimensional slice; Obtaining the total fusion feature module: used to gradually fuse each reconstructed feature to obtain the total fusion feature; Obtaining quality score module: used to input the total fusion features into the KANs network to obtain the final no-reference quality score.