No-reference light field image quality evaluation method based on pseudo video sequence state modeling

By combining the Mamba and Transformer hybrid architecture with the Kolmogorov-Arnold network (KAN), the problems of cross-dimensional correlation modeling and nonlinear mapping in light field image quality assessment are solved, improving the accuracy and computational efficiency of immersive experience quality assessment, and making it suitable for light field image quality assessment.

CN121811219APending Publication Date: 2026-04-07ZHEJIANG GONGSHANG UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-09
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing light field image quality assessment methods lack explicit modeling of cross-dimensional correlations, cannot fully characterize the geometric and depth consistency during smooth viewpoint transitions in light field displays, and have weak nonlinear mapping capabilities, making it difficult to simulate the nonlinear relationship between subjective perception and objective distortion. They also have high computational complexity and are difficult to adapt to high-dimensional light field data.

Method used

Employing a hybrid architecture of Mamba and Transformer, spatial, angular, and epipolar geometric features of light field pseudo-video sequences are extracted. Combining a 3D convolution module and a quality prediction module, cross-dimensional modeling and nonlinear mapping are achieved through the Mamba-Transformer hybrid architecture and Kolmogorov-Arnold network (KAN), outputting an objective quality evaluation score.

Benefits of technology

It achieves explicit modeling of cross-dimensional correlations, improves the accuracy of quality assessment related to immersive experiences, has strong nonlinear mapping capabilities, maintains high efficiency in computationally limited scenarios, and has higher evaluation accuracy than existing methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121811219A_ABST
    Figure CN121811219A_ABST
Patent Text Reader

Abstract

The invention discloses a non-reference light field image quality evaluation method based on pseudo video sequence state modeling. The method comprises the following steps: S1, preprocessing an original light field distortion image, and extracting a light field pseudo video sequence; s2, extracting and fusing features of a spatial domain, an angle domain and an epipolar geometric domain of the light field pseudo video sequence by adopting a mixed framework of Mama and Transform to obtain a domain fusion feature code; and S3, inputting the domain fusion feature code into a quality prediction module, and outputting an objective quality evaluation score. According to the method, cross-dimension correlation explicit modeling is carried out, so that the quality evaluation precision of immersive experience is improved; and the strong nonlinear mapping capability is adopted, the performance and complexity are efficiently balanced, and the evaluation accuracy is high.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of light field image quality assessment, specifically a referenceless light field image quality assessment method based on pseudo-video sequence state modeling. Background Technology

[0002] Light field imaging, as a core technology of immersive VR, can capture the intensity and direction information of light in a scene, enabling continuous synthesis of multiple perspectives and depth perception. However, light field data is prone to distortions such as resolution degradation, geometric inconsistency of viewpoints, focus blur, and artifacts during acquisition, compression, and transmission, which seriously affect the user's immersion. Therefore, light field image quality evaluation has become a key research direction.

[0003] Existing methods for evaluating the quality of light field images typically project a four-dimensional light field into two-dimensional slices representing the spatial, angular, and epipolar geometric domains. Features are extracted independently and then fused using implicit encoding or attention mechanisms. However, these methods have two major limitations: first, they lack explicit modeling of cross-dimensional correlations, failing to fully characterize the geometric and depth consistency and other viewpoint coherence characteristics during smooth viewpoint transitions in light field displays—characteristics crucial for ensuring an immersive visual experience; second, subjective score prediction relies on linear weighting matrices, making it difficult to simulate the nonlinear relationship between subjective human perception and objective distortion, such as visual masking effects and differences in contrast sensitivity.

[0004] Furthermore, existing methods mostly employ a single network architecture, which also has many limitations. For example, convolutional neural networks have limited receptive fields and cannot capture long-range dependencies between feature sequences; the computational complexity of Transformer networks increases exponentially with sequence length, making them difficult to adapt to high-dimensional light field data; while Mamba networks can efficiently model medium-range dependencies between feature sequences, their ability to capture global correlations is insufficient. Therefore, there is an urgent need for a light field image quality assessment method that balances cross-dimensional modeling, nonlinear mapping, and computational efficiency. Summary of the Invention

[0005] The purpose of this invention is to provide a no-reference light field image quality assessment method based on pseudo-video sequence state modeling, in order to solve the technical problems in existing light field image quality assessment methods such as "lack of cross-dimensional correlation modeling", "weak nonlinear mapping ability" and "difficulty in balancing local-global feature requirements and computational efficiency".

[0006] The technical solution adopted by this invention to solve its technical problem is as follows:

[0007] Step S1: Preprocess the original light field distortion image to extract the light field pseudo-video sequence;

[0008] Step S2: Using a hybrid architecture of Mamba and Transformer, the spatial domain, angular domain, and epipolar geometric domain features of the light field pseudo-video sequence are extracted and fused to obtain the domain fusion feature encoding;

[0009] Step S3: Input the domain fusion feature encoding into a quality prediction module and output an objective quality evaluation score.

[0010] The specific implementation of step S1 includes:

[0011] Step S1-1: Obtain the original light field distortion image and extract its center. The sub-aperture images are used to form a sub-aperture image array. , The spatial resolution of the sub-aperture images is typically represented by values ​​of 5 for U and V. The sub-aperture image array L is arranged in a "W-shaped" scanning order (odd columns scan first, then even columns, odd columns scan from top to bottom, even columns from bottom to top) to form a sub-aperture image sequence. This sequence contains... Individual aperture images.

[0012] Step S1-2: A sliding window of size H×W is used to continuously and non-overlap sample each sub-aperture image in the sub-aperture image sequence. Typically, H and W are set to 32. The sliding window extracts data from the sub-aperture image sequence each time. Individual aperture image patches are used to form a pseudo-video sequence of the light field. .

[0013] Furthermore, step S2 specifically includes:

[0014] Step S2-1: Convert the light field pseudo-video sequence The input is fed into a 3D convolutional module to extract local feature sequences. , Includes The size is Local feature map.

[0015] The 3D convolution module contains four convolutional layers with 1×3×3 convolutional kernels, and the convolutional operations of the last three layers are all connected to the LeakyReLU activation function.

[0016] Step S2-2: Employ a Mamba-Transformer hybrid architecture to extract local feature sequences. High-dimensional features from the spatial domain, angular domain, and outer pole geometric domain of the light field are extracted and fused to obtain the domain fusion feature encoding. ;

[0017] The described Mamba-Transformer hybrid architecture includes an intra-frame / inter-frame feature encoder and a horizontal-vertical EPI feature encoder. Each encoder is designed based on a structurally identical Mamba-Transformer module.

[0018] The intra-frame to inter-frame feature encoder adopts a cascaded architecture design, which sequentially connects the intra-frame feature module, the inter-frame feature module, and the intra-frame to inter-frame residual fusion module, respectively used to extract and fuse intra-frame high-dimensional features and inter-frame correlation features in the light field pseudo-video sequence.

[0019] The intra-frame feature module aims to mine high-dimensional intra-frame features of light field pseudo-video sequences. The specific process involves: processing local feature sequences... Enter the Mamba-Transformer module, for Feature aggregation is performed on each local feature map to generate intra-frame feature codes with spatial domain distortion awareness. , where C is the number of feature channels, which is usually 32.

[0020] The inter-frame feature module aims to extract inter-frame correlation features from light field pseudo-video sequences. These inter-frame correlation features actually reflect the high-dimensional features of the light field in the angular domain. The specific process is as follows: [The text abruptly ends here, likely due to an incomplete sentence or missing information.] Dimensional transformation ,but Includes HW= Feature maps of each angle; Enter the Mamba-Transformer module, for Feature aggregation is performed on each angle feature map to generate inter-frame feature codes with angle domain distortion awareness. .

[0021] The intra-frame-inter-frame residual fusion module first performs... Dimensional transformation Then, the initial feature sequence and Add them together to get the intra-frame inter-frame coding. ;

[0022] The horizontal-vertical EPI feature encoder adopts a cascaded architecture design, which sequentially connects the horizontal EPI feature module, the vertical EPI feature module, and the EPI residual fusion module, respectively used to extract and fuse the horizontal / vertical outer pole geometric features of the light field pseudo-video sequence.

[0023] The horizontal EPI feature module is designed to mine the horizontal outer pole geometric features of light field pseudo-video sequences; the specific process is as follows: intra-frame and inter-frame coding Dimensional transformation ,but Includes HV= Each level EPI feature map; will Enter the Mamba-Transformer module, for Each horizontal EPI feature map in the dataset is used for feature aggregation to generate a horizontal EPI feature code. .

[0024] The vertical EPI feature module is designed to mine the vertical epipolar geometric features of light field pseudo-video sequences; the specific process is as follows: [The text abruptly ends here, likely due to an incomplete sentence or missing information.] Dimensional transformation ,but ;Will Enter the Mamba-Transformer module, for Each vertical EPI feature map in the dataset is aggregated to generate a vertical EPI feature code. .

[0025] The EPI residual fusion module first performs... Dimensional transformation Then and Adding them together yields the domain fusion feature code. .

[0026] The specific processing procedure of the Mamba-Transformer module is as follows: First, input... After the normalization layer, the input is fed into a two-layer Mamba network architecture to extract the data. The short-range dependency between adjacent frames and the medium-range dependency across frames are used to obtain the medium-short-range sequence features. Then, The input is fed into a Transformer network architecture to extract global long-range correlation features and output global long-range features. Finally, a residual connection method is used to... Adding them together yields the characteristics of the mixed polymerization. ,Will The data is sequentially fed into a normalization layer, a convolutional layer, and a channel attention layer, finally outputting channel-modulated features. Finally and The final enhanced feature is obtained by adding the features together. .

[0027] The Mamba network architecture is referenced from the paper "Mamba: Linear-Time Sequence Modeling with Selective State Spaces" published by Albert Gu et al. in 2023; the Transformer network architecture is referenced from the paper "An Image is Worth 16x16Words: Transformers for Image Recognition at Scale" published by Alexey Dosovitskiy et al. in 2021; this embodiment directly uses the existing Mamba architecture and Transformer to construct the above state update and feature extraction process.

[0028] Step S3 specifically includes:

[0029] For step S2 , , Cascaded fusion for multi-domain joint coding Multi-domain joint coding The data is then fed into the quality prediction module to output an objective quality evaluation score.

[0030] The quality prediction module comprises an average pooling layer, a gated recurrent unit, and a Kolmogorov-Arnold network (KAN) layer. Specifically, it involves joint encoding of multiple domains. The feature vector is converted through an average pooling layer, input to a gated recurrent unit (GRU) to model the long-range temporal dependencies of the features, and outputs temporal fusion features. ; in Represents average pooling. Represents a gated recursive unit. It incorporates temporal fusion features. Input a Kolmogorov-Arnold network (KAN) to perform nonlinear regression instead of a traditional multilayer perceptron, and output the final quality assessment score.

[0031] The Kolmogorov-Arnold Network (KAN) described is referenced from the paper "KAN: Kolmogorov-Arnold Networks" published by Ziming Liu et al. in 2024. This embodiment directly adopts the existing Kolmogorov-Arnold Network (KAN) construction process.

[0032] The beneficial effects of this invention are as follows:

[0033] 1. Cross-dimensional correlation explicit modeling: By preserving the viewpoint continuity through light field pseudo-video sequences, and combining the Mamba-Transformer hybrid architecture, we can jointly model intra-frame high-dimensional features, inter-frame correlation features, and the outer pole geometric features of light field pseudo-video sequences, fully characterizing the "viewpoint coherence" of the light field and improving the accuracy of quality assessment related to immersive experience.

[0034] 2. Strong nonlinear mapping capability: Kolmogorov-Arnold network (KAN) is used to replace the traditional multilayer perceptron as the regression head. The complex nonlinear relationship between subjective perception and objective distortion of the human eye is simulated through a learnable function. The prediction performance is better than the scheme with linear weighted matrix and fixed activation function.

[0035] 3. Highly efficient balance between performance and complexity: The overall model uses 3D convolution to extract local features, Mamba to efficiently model mid-range dependencies, and Transformer to complete global relationships. The number of parameters (0.48M) and computational cost (2.97G FLOPs) are significantly lower than existing methods such as DeeBliF (16.66M / 8.06G) and LFACon (10M / 4.77G), making it suitable for scenarios with limited computing resources.

[0036] 4. High evaluation accuracy: On the Win5-LID dataset, the SROCC and PLCC metrics are both superior to existing advanced methods. Attached Figure Description

[0037] Figure 1 This is a schematic diagram of the main steps of the present invention;

[0038] Figure 2 This is a diagram showing the generation of the optical field pseudo-video sub-aperture sequence in this invention.

[0039] Figure 3 This is a schematic diagram of the referenceless light field image quality assessment method based on light field pseudo-video sequences of the present invention.

[0040] Figure 4 This invention relates to multi-view dimensional reshaping and sequence encoding of sub-aperture images;

[0041] Figure 5 This is a structural diagram of the Mamba module of the present invention;

[0042] Figure 6 This is a structural diagram of the Transformer module of the present invention;

[0043] Figure 7 This is a structural diagram of the Mamba-Transformer module of the present invention; Detailed Implementation

[0044] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0045] like Figure 1 As shown, a method for evaluating the quality of no-reference light field images based on pseudo-video sequence state modeling is presented.

[0046] The technical steps taken are as follows:

[0047] Step S1: Preprocess the original light field distortion image to extract the light field pseudo-video sequence;

[0048] Step S2: Using a hybrid architecture of Mamba and Transformer, the spatial domain, angular domain, and epipolar geometric domain features of the light field pseudo-video sequence are extracted and fused to obtain the domain fusion feature encoding;

[0049] Step S3: Input the domain fusion feature encoding into a quality prediction module and output an objective quality evaluation score.

[0050] The specific implementation of step S1 includes:

[0051] Step S1-1: Use the Win5-LID dataset as the experimental set. This dataset contains 220 distorted light field images of different types and degrees, and provides the average subjective score for each distorted light field image. It is randomly split into 80% of the dataset and 20% of the test set. The light field is represented as four-dimensional data U×V×H×W, where U×V is the angular resolution and H×W is the spatial resolution. The spatial dimension (details and textures within the same viewpoint) is analogous to intra-frame information in a video, while the angular dimension (geometric consistency across different viewpoints) is analogous to inter-frame dependencies in a video. Obtain the original distorted light field image and extract its center. The sub-aperture images are used to form a sub-aperture image array. , The spatial resolution of the sub-aperture images is typically represented by values ​​of 5 for U and V. The sub-aperture image array L is arranged in a "W-shaped" scanning order (odd columns scan first, then even columns, odd columns scan from top to bottom, even columns from bottom to top) to form a sub-aperture image sequence. This sequence contains... Individual aperture images.

[0052] Step S1-2: A sliding window of size H×W is used to continuously and non-overlap sample each sub-aperture image in the sub-aperture image sequence. Typically, H and W are set to 32. The sliding window extracts data from the sub-aperture image sequence each time. Individual aperture image patches are used to form a pseudo-video sequence of the light field. ,in , This represents the spatial resolution of the sub-aperture image patch. For example, in... Figure 2The top-left red boxes of all central 5×5 sub-aperture images are aggregated to obtain a single red sub-aperture block. Finally, the light field viewpoint matrix is ​​rearranged along a "W-shaped" scan path into a pseudo-video sequence of length T = U×V (as shown below), thus ensuring geometric continuity between adjacent viewpoints in the scan sequence. Several spatial angle block sequences can be generated for each scene in the light field image in this way. All intra-frame and inter-frame block sequences generated for the same scene have the same average opinion score.

[0053]

[0054] Among them, parameters , , ,in For the first A "pseudo-video frame" This indicates a "W-shaped" scanning sequence. This represents the original light field pseudo-video sequence. This is a sequence of spatial angle blocks obtained after scanning, i.e., a pseudo-video sequence of the light field. .

[0055] Step S2 specifically includes:

[0056] Figure 3 This is a schematic diagram of a no-reference light field image quality assessment method model based on light field pseudo-video sequences. The preprocessing result of step S1 is the input of this model.

[0057] Step S2-1: Convert the light field pseudo-video sequence Input into a 3D convolutional module to extract local feature sequences , Includes The size is Local feature map.

[0058] The 3D convolution module contains four convolutional layers with 1×3×3 convolutional kernels, and the convolutional operations of the last three layers are all connected to the LeakyReLU activation function.

[0059] Step S2-2: Figure 4 In this process, when different feature modules are input, the light field is first structurally decomposed and parameterized into multi-dimensional slices to match the sequence modeling requirements of the hybrid architecture. Specifically, the four-dimensional light field image is parameterized by a biplane model, represented as... ,in Indicates angular resolution. This represents spatial resolution. When the hybrid architecture is applied to the light field, the four-dimensional light field data... It is considered as a combination of four information-rich 2D slices.

[0060]

[0061] in, Corresponding intra-frame features. Corresponding inter-frame features. and EPI image features, corresponding to the horizontal and vertical directions respectively, are used to capture cross-view structural dependencies. Multi-dimensional slices are input into different feature modules. Represents the set of intra-frame feature slices. Represents a set of inter-frame feature slices. Represents a set of horizontal EPI feature slices. This represents a set of vertical EPI feature slices.

[0062] Step S2-3: Employ a Mamba-Transformer hybrid architecture to extract local feature sequences. High-dimensional features from the spatial domain, angular domain, and outer polar geometric domain of the light field are extracted and fused to obtain the domain fusion feature encoding. ;

[0063] The described Mamba-Transformer hybrid architecture includes an intra-frame / inter-frame feature encoder and a horizontal-vertical EPI feature encoder. Each encoder is designed based on a structurally identical Mamba-Transformer module.

[0064] The intra-frame / inter-frame feature encoder employs a cascaded architecture to capture intra-frame high-dimensional spatial features and inter-frame correlation features in light field pseudo-video sequences. Structurally, the encoder consists of an intra-frame feature module, an inter-frame feature module, and an intra-frame / inter-frame residual fusion module connected in series.

[0065] The intra-frame feature module aims to fully exploit the high-dimensional intra-frame spatial features of light field pseudo-video sequences. Specifically, it involves processing local feature sequences... Input the Mamba-Transformer module, and use the Mamba-Transformer module to... Feature aggregation is performed on each local feature map to generate an intra-frame feature sequence with spatial domain distortion awareness. Where C is the number of feature channels, which is usually 32.

[0066] The inter-frame feature module aims to capture the inter-frame correlation features of light field pseudo-video sequences. These inter-frame correlation features actually reflect the high-dimensional features of the light field in the angular domain. The specific process involves... After dimensional rearrangement , Includes HW= Feature maps of each angle; Enter the Mamba-Transformer module, for Each angular feature map in the dataset is aggregated to generate an inter-frame feature sequence with angular domain distortion awareness. .

[0067] The intra-frame-inter-frame residual fusion module first performs... Dimensional rearrangement Initial feature sequence and Adding them together yields the intra-frame and inter-frame codes. ;

[0068] The horizontal-vertical EPI feature encoder employs a cascaded architecture design to capture the horizontal / vertical epipolar geometric features of optical field pseudo-video sequences. Structurally, the encoder consists of a horizontal EPI feature module, a vertical EPI feature module, and an EPI residual fusion module connected in series.

[0069] The horizontal EPI feature module aims to fully exploit the horizontal outer pole geometric features of the light field pseudo-video sequence; the specific process involves intra-frame and inter-frame coding. After dimensional rearrangement , Includes HV= Each level EPI feature map; will Enter the Mamba-Transformer module, for Each horizontal EPI feature map in the dataset is aggregated to generate a horizontal EPI feature sequence. .

[0070] The vertical EPI feature module aims to fully exploit the vertical epipolar geometric features of the light field pseudo-video sequence; the specific process is as follows: After dimensional rearrangement , ;Will Enter the Mamba-Transformer module, for Each vertical EPI feature map in the dataset is aggregated to generate a vertical EPI feature sequence. .

[0071] The EPI residual fusion module first performs... Dimensional rearrangement Then the initial feature sequence and Adding them together yields the domain fusion feature encoding. .

[0072] Leveraging the Mamba network architecture's ability to model short-to-medium range epipolar slice sequences and the Transformer network architecture's ability to model long range epipolar slice sequences, this module effectively identifies the linear distribution and displacement patterns of pixels within the space-angle plane, thereby capturing depth cues and epipolar geometric consistency of the light field in the horizontal and vertical dimensions. This enables the model to perceive structural distortions in the epipolar geometric domain of the light field, thus outputting a high-dimensional feature representation that reflects the integrity of the multidimensional geometric structure of the light field.

[0073] exist Figure 7 The specific processing procedure of the Mamba-Transformer module is as follows: First, input... After the normalization layer, the input is fed into a two-layer Mamba network architecture to extract the data. The short-range dependency between adjacent frames and the medium-range dependency across frames are used to obtain the medium-short-range sequence features. Then, The input is fed into a Transformer network architecture to extract global long-range correlation features and output global long-range features. Finally, a residual connection method is used to... Adding them together yields the characteristics of the mixed polymerization. ,Will The data is sequentially fed into a normalization layer, a convolutional layer, and a channel attention layer, finally outputting channel-modulated features. Finally and The final enhanced feature is obtained by adding the features together. .

[0074] The first-layer Mamba module models the mixed initial features. With short-range dependencies between adjacent frames, the state update formula is:

[0075]

[0076] in, Indicates the first layer The hidden state of the next update. , For state-space model parameters, For activation function, , These are the parameters for the linear transformation.

[0077] The second-layer Mamba module strengthens cross-frame mid-range dependencies, outputting the second-layer... Mid-range features of the next update .

[0078] The third-layer Transformer module handles the features of the middle layer. Perform self-attention and feedforward network operations to aggregate global long-range correlation features, and fuse them through residual connections to output features. .

[0079] The Mamba network architecture (such as) Figure 6 (As shown) This is referenced from the 2023 paper "Mamba: Linear-Time Sequence Modeling with Selective State Spaces" by Albert Gu et al.; the Transformer network architecture (such as...) Figure 7 (As shown) This is based on the paper "An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale" published by Alexey Dosovitskiy et al. in 2021; this embodiment directly uses the existing Mamba architecture and Transformer to construct the above state update and feature extraction process.

[0080] Step S3 specifically includes:

[0081] Step S3-1: For the steps in S2 , , Cascading fusion is characterized by multi-domain collaboration. For multi-domain joint features Perform average pooling to convert the data into feature vectors. Input the gated recurrent unit (GRU) to model the long-range temporal dependencies of the features, and output temporal fusion features. ; in Represents average pooling. This represents a gated recursive unit.

[0082] Step S3-2: Fuse temporal features Input a Kolmogorov-Arnold network (KAN) to perform nonlinear regression instead of a traditional multilayer perceptron, and output the final quality assessment score.

[0083] The Kolmogorov-Arnold Network (KAN) described is referenced from the paper "KAN: Kolmogorov-Arnold Networks" published by Ziming Liu et al. in 2024. This embodiment directly adopts the existing Kolmogorov-Arnold Network (KAN) construction process.

[0084] Step S3-3: Train the mini-batch stochastic gradient descent optimizer for 60 epochs with a weight momentum of 0.9, a decay rate of 0.0001, and an initial learning rate of 0.0001 (decreasing by 0.1 every 30 epochs); use the mean squared error (MSELoss) as the loss function, with the following formula:

[0085]

[0086] in, This represents the batch size during training, i.e., the number of samples processed in parallel during each iteration. Indicates the index number of the sample in the current batch ( =1, 2, ..., ), This represents the objective quality score predicted by the model for the b-th light field pseudo-video sequence sample in the current batch. This represents the true subjective quality score label corresponding to the b-th sample.

[0087] During the testing phase, the prediction scores of all spatial angle block sequences in the same scene are averaged to obtain the overall average subjective score of the scene. ( (Number of spatial angle block sequences).

[0088] Table 1 compares the method of this invention with other full-reference / no-reference light field quality assessment methods on the Win5-LID dataset. The metrics are Spearman Rank Correlation Coefficient (SROCC) and Pearson Linear Correlation Coefficient (PLCC). Higher SROCC and Pearson Linear Correlation Coefficient indicate better method performance. The method of this invention outperforms the second-ranked method by 0.0212 in SROCC and 0.0131 in PLCC, demonstrating its superiority. More importantly, this invention significantly reduces the computational complexity of the model while maintaining the highest evaluation accuracy. As shown in Table 2, the second-ranked DeeBliF method has a parameter count as high as 16.66M, while this invention has only 0.48M, approximately one-thirty-fifth of the latter. This indicates that this invention does not achieve performance improvement by simply stacking parameters, but rather through an efficient hybrid architecture design, achieving accurate capture of light field features with extremely low parameter count, thus balancing excellent evaluation performance with extremely high computational efficiency, making it more suitable for practical application deployment.

[0089] Table 1

[0090]

[0091] Table 2. Model Complexity Analysis of Real Light Field Data

[0092]

[0093] To verify the effectiveness of the key modules and cascading order in this invention, three variant models were constructed for comparative experiments, and the results are shown in Table 3. Importance of EPI features (w / o EPI vs Ours): When the horizontal-vertical EPI feature encoder was removed (w / o EPI), the model's SROCC metric on the Win5-LID dataset decreased by 0.0209. This indicates that relying solely on intra-frame and inter-frame features is insufficient to fully characterize the complex geometric structure of the light field; introducing a horizontal-vertical EPI feature encoder to capture features in both horizontal and vertical directions is crucial for improving evaluation accuracy.

[0094] The rationality of the module order (w / reverse vs Ours): When the feature extraction order is reversed (w / reverse), i.e., epipolar features are extracted first and then intra-frame and inter-frame features are extracted, the performance drops significantly. This verifies that the strategy of this invention—that is, first establishing the basic spatiotemporal (space-angle) dependency through intra-frame and inter-frame modules, and then refining the cross-view geometric structure through the epipolar image feature extraction module on this basis—is a better feature learning path.

[0095] In summary, the complete architecture (Ours) proposed in this invention achieves optimal performance through reasonable module combination and cascading order.

[0096] Table 3 Performance Comparison of Different Light Field Structure Learning Modules

[0097]

[0098] Finally, it should be noted that the purpose of disclosing the embodiments is to help further understand the present invention. However, those skilled in the art will understand that various substitutions and modifications are possible without departing from the spirit and scope of the present invention and the appended claims. Therefore, the present invention should not be limited to the content disclosed in the embodiments, and the scope of protection of the present invention is defined by the claims.

Claims

1. A method for evaluating the quality of a no-reference light field image based on pseudo-video sequence state modeling, characterized in that, Includes the following steps: Step S1: Preprocess the original light field distortion image to extract the light field pseudo-video sequence; Step S2: Using a hybrid architecture of Mamba and Transformer, the spatial domain, angular domain, and epipolar geometric domain features of the light field pseudo-video sequence are extracted and fused to obtain the domain fusion feature encoding; Step S3: Input the domain fusion feature encoding into a quality prediction module and output an objective quality evaluation score.

2. The method according to claim 1, characterized in that, Step S1 is specifically implemented by the following steps: Step S1-1: Obtain the original light field distortion image and extract its center. The sub-aperture images are used to form a sub-aperture image array. , The spatial resolution of the sub-aperture images is typically represented by values ​​of 5 for U and V. The sub-aperture image array L is arranged in a "W-shaped" scanning sequence to form a sub-aperture image sequence, which contains... Individual aperture images; Step S1-2: A sliding window of size H×W is used to continuously and non-overlap sample each sub-aperture image in the sub-aperture image sequence. Typically, H and W are set to 32. The sliding window extracts data from the sub-aperture image sequence each time. Individual aperture image patches are used to form a pseudo-video sequence of the light field. .

3. The method according to claim 2, characterized in that, The "W-shaped" scanning order is implemented as follows: the odd-numbered columns are scanned first, followed by the even-numbered columns. The odd-numbered columns are scanned from top to bottom, and the even-numbered columns are scanned from bottom to top.

4. The method according to claim 1, characterized in that, Step S2 specifically includes the following steps: Step S2-1: Convert the light field pseudo-video sequence The input is fed into a 3D convolutional module to extract local feature sequences. , Includes The size is Local feature map; Step S2-2: Employ a Mamba-Transformer hybrid architecture to extract local feature sequences. High-dimensional features from the spatial domain, angular domain, and outer pole geometric domain of the light field are extracted and fused to obtain the domain fusion feature encoding. .

5. The method according to claim 4, characterized in that, Step S2 specifically includes the following steps: The 3D convolution module contains four convolutional layers with 1×3×3 convolutional kernels, wherein the convolutional operations of the last three layers are all connected to the LeakyReLU activation function; The Mamba-Transformer hybrid architecture includes an intra-frame / inter-frame feature encoder and a horizontal-vertical EPI feature encoder; each encoder is based on a structurally identical Mamba-Transformer module.

6. The method according to claim 5, characterized in that, Step S2 specifically includes the following steps: The intra-frame-inter-frame feature encoder adopts a cascaded architecture design, which sequentially connects the intra-frame feature module, the inter-frame feature module, and the intra-frame-inter-frame residual fusion module, respectively used to extract and fuse intra-frame high-dimensional features and inter-frame correlation features in the light field pseudo-video sequence. The intra-frame feature module is designed to mine intra-frame high-dimensional features of light field pseudo-video sequences; The specific process is as follows: [The local feature sequence is then processed.] Enter the Mamba-Transformer module, for Feature aggregation is performed on each local feature map to generate intra-frame feature codes with spatial domain distortion awareness. Where C is the number of feature channels; The inter-frame feature module is designed to extract inter-frame correlation features of light field pseudo-video sequences; The specific process is as follows: Dimensional transformation ,but Includes HW= Feature maps of each angle; Enter the Mamba-Transformer module, for Feature aggregation is performed on each angle feature map to generate inter-frame feature codes with angle domain distortion awareness. ; The intra-frame-inter-frame residual fusion module first performs... Dimensional transformation Then, the initial feature sequence and Add them together to get the intra-frame inter-frame coding. .

7. The method according to claim 5 or 6, characterized in that, Step S2 specifically includes the following steps: The horizontal-vertical EPI feature encoder adopts a cascaded architecture design, which sequentially connects the horizontal EPI feature module, the vertical EPI feature module, and the EPI residual fusion module, respectively used to extract and fuse the horizontal / vertical epipolar geometric features of the light field pseudo-video sequence. The horizontal EPI feature module is designed to mine the horizontal outer pole geometric features of light field pseudo-video sequences; The specific process is as follows: intra-frame inter-frame coding Dimensional transformation ,but Includes HV= Each level EPI feature map; will Enter the Mamba-Transformer module, for For each horizontal EPI feature map, feature aggregation is performed to generate a horizontal EPI feature code. ; The vertical EPI feature module is designed to mine the vertical exopolar geometric features of light field pseudo-video sequences; The specific process is as follows: Dimensional transformation ,but ;Will Enter the Mamba-Transformer module, for Each vertical EPI feature map in the dataset is aggregated to generate a vertical EPI feature code. ; The EPI residual fusion module first performs... Dimensional transformation Then and Adding them together yields the domain fusion feature code. .

8. The method according to claim 7, characterized in that, Step S2 specifically includes the following steps: The specific processing procedure of the Mamba-Transformer module is as follows: First, input... After the normalization layer, the input is fed into a two-layer Mamba network architecture to extract the data. The short-range dependency between adjacent frames and the medium-range dependency across frames are used to obtain the medium-short-range sequence features. Then, The input is fed into a Transformer network architecture to extract global long-range correlation features and output global long-range features. Finally, a residual connection method is used to... Adding them together yields the characteristics of the mixed polymerization. ,Will The data is sequentially fed into a normalization layer, a convolutional layer, and a channel attention layer, finally outputting channel-modulated features. Finally and The final enhanced feature is obtained by adding the features together. .

9. The method according to claim 8, characterized in that, Step S3 specifically includes the following steps: For step S2 , , Cascaded fusion for multi-domain joint coding Multi-domain joint coding The objective quality evaluation score is then sent to the quality prediction module. The quality prediction module includes an average pooling layer, a gated recurrent unit, and a Kolmogorov-Arnold network; the specific process is as follows: multi-domain joint encoding The features are converted into feature vectors through an average pooling layer, input to a gated recurrent unit (GRU) to model the long-range temporal dependencies of the features, and output temporal fusion features. ; in Represents average pooling. Represents a gated recursive unit; integrates temporal fusion features Input a Kolmogorov-Arnold network to perform nonlinear regression instead of a traditional multilayer perceptron, and output the final quality assessment score.