A linear attention-based degradation-aware blind light field image quality assessment method, system, storage medium and device

CN122820535APending Publication Date: 2026-09-25ANQING NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610625150.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-08
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0006]本申请提供一种基于线性注意力的退化感知盲光场图像质量评估方法、系统、存储介质及设备,旨在解决背景技术中提出的现有技术中基于卷积神经网络的方法受限于局部感受野,难以有效捕获空间与角度域的长程依赖关系,而基于视觉Transformer的模型虽能实现全局建模,却面临计算复杂度高、参数量冗余以及对微透镜异质退化适配能力不足的问题

Benefits of technology

[0017]本申请通过引入线性注意力机制,将计算复杂度从传统的二次方降低至线性级别,显著提升了处理高分辨率光场图像的效率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122820535A_ABST
    Figure CN122820535A_ABST
Patent Text Reader

Abstract

The application provides a degradation perception blind light field image quality evaluation method and system based on linear attention, a storage medium and equipment, and belongs to the field of light field image evaluation. The method generates a degradation perception weight map based on the spatial distribution characteristics of a microlens array, and uses the degradation perception weight map to weight the microlens light field image to obtain a degradation perception enhanced image. Spatial features and angle features of the degradation perception enhanced image are extracted to obtain spatial domain basic features and angle domain basic features. A linear attention mechanism is used to model long-range dependence of the feature sequence, and combined with residual connection, enhanced spatial features and angle features are obtained. The enhanced spatial features and angle features are fused, the fused features are adaptively weighted through a channel attention mechanism to obtain fused features, the fused features are globally pooled, and a multi-layer fully connected regression network is used to output a quality score of the light field image under a reference-free scene, so that efficient and accurate quality evaluation can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of light field image quality assessment technology, specifically a method, system, storage medium, and device for assessing the quality of degenerate perceptually blind light field images based on linear attention. Background Technology

[0002] Light field images, as a type of high-dimensional visual data that simultaneously records the spatial distribution and angular propagation information of light in a three-dimensional scene, not only contain the texture details of traditional two-dimensional images but also contain multi-view correlations and depth geometry. This makes them of significant application value in fields such as virtual reality (VR), augmented reality (AR), computational photography, and medical microscopy.

[0003] However, the unique four-dimensional data structure and microlens array acquisition mechanism of light field images present unique technical challenges to image quality assessment. First, due to the physical imaging characteristics of microlenses, image degradation often exhibits a non-uniform distribution, meaning that the degree of degradation varies significantly across different microlens regions. This places extremely high demands on the model's ability to perceive local degradation. Second, the perceived quality of light field images is jointly determined by the spatial fidelity within a single viewpoint and the angular consistency across multiple viewpoints. This means that the assessment model must be able to simultaneously model both spatial and angular dual-dimensional information. However, current assessment methods often struggle to simultaneously achieve accurate perception of local non-uniform degradation and complete extraction of dual-dimensional features when processing such high-dimensional data.

[0004] In no-reference (blind) evaluation scenarios, the lack of original, clear images as comparison benchmarks further increases the difficulty of evaluation. Although deep learning methods such as Convolutional Neural Networks (CNNs) and Visual Transformers (ViTs) have been introduced into this field, they still face the challenge of balancing computational power and accuracy in practical applications: CNNs are limited by local receptive fields and have difficulty capturing long-range dependencies across viewpoints; while ViT-like models can model global information, the computational complexity of their self-attention mechanism increases quadratically with sequence length, often facing huge computational overhead when processing high-resolution images of light fields, making it difficult to meet the requirements of lightweight deployment and real-time evaluation.

[0005] Therefore, this application provides a method, system, storage medium, and device for evaluating the quality of degenerate perception blind light field images based on linear attention, in order to solve the above problems. Summary of the Invention

[0006] This application provides a method, system, storage medium, and device for quality assessment of degraded perception blind light field images based on linear attention. It aims to solve the problems of existing methods based on convolutional neural networks in the background art, which are limited by local receptive fields and have difficulty effectively capturing long-range dependencies in the spatial and angular domains. While models based on visual Transformers can achieve global modeling, they face problems such as high computational complexity, redundant parameters, and insufficient adaptability to heterogeneous degradation of microlenses.

[0007] To achieve the above objectives, this application provides the following technical solution: a method for evaluating the quality of degraded perceptual blind light field images based on linear attention, the evaluation method comprising the following steps: The light field image of the microlens to be evaluated is acquired, a degradation perception weight map is generated based on the spatial distribution characteristics of the microlens array, and the degradation perception weight map is used to weight the light field image of the microlens to obtain a degradation perception enhanced image. The spatial feature extraction branch and the angular feature extraction branch are used to extract the basic spatial domain features and the basic angular domain features of the degraded perception enhancement image, respectively. The linear attention mechanism is used to model the long-range dependency of the feature sequence, and the enhanced spatial and angular features are obtained by combining residual connections. The enhanced spatial and angular features are fused, and the fused features are adaptively weighted using a channel attention mechanism to obtain fused features. The fused features are subjected to global pooling, and the quality score of the light field image is output through a multi-layer fully connected regression network.

[0008] Preferably, the generation of the degradation-aware weight map specifically involves: The microlens light field image is divided into multiple microlens sub-regions according to the periodic structure of the microlens array; Each microlens sub-region is input into a lightweight feature mapping network for local structure encoding, and the degradation response value of each sub-region is output. All degradation response values ​​are normalized to generate a degradation-aware weight map corresponding to the spatial size of the input image.

[0009] Preferably, the step of weighting the microlens light field image using the degradation-perceived weight map specifically involves: The degradation perception weight map is multiplied pixel by pixel with the microlens light field image.

[0010] Preferably, in the spatial feature extraction branch, dilated convolution is used to extract features from the image, wherein the dilation rate of the dilated convolution is set as the angular resolution parameter of the light field image to expand the receptive field and cover the structural information across the microlens region, thereby obtaining the basic features of the spatial domain.

[0011] Preferably, in the angle feature extraction branch, the viewing angle information inside the microlens sub-region is aggregated through convolution operation to extract the basic features of the angle domain.

[0012] Preferably, the method of using linear attention mechanism to model long-range dependencies of feature sequences specifically includes: The input feature map is unfolded into a feature sequence in the spatial dimension; The feature sequence is mapped to query features, key features, and value features through a linear projection layer; By introducing kernel functions to transform query features and key features, the self-attention matrix operation is transformed into a computational form with linear complexity. A recursive update mechanism is introduced to progressively accumulate and model the features in the sequence, resulting in an enhanced feature sequence. The enhanced feature sequence is restored to the feature map size and then added to the input feature map using residuals.

[0013] Preferably, the adaptive weighting of the fused features through the channel attention mechanism specifically involves: Global average pooling is performed on the fused features to generate channel description vectors; The channel description vector is input into the multilayer perceptron to generate the weight coefficients for each channel; The weighting coefficients are multiplied channel by channel with the fused features to achieve adaptive recalibration of the features.

[0014] A degradation-perceived blind light field image quality assessment system based on linear attention, used to perform the aforementioned degradation-perceived blind light field image assessment method based on linear attention, the assessment system comprising: The degradation-aware weighting module is used to receive the microlens light field image to be evaluated, generate a degradation-aware weight map based on the spatial distribution characteristics of the microlens array, and use the degradation-aware weight map to weight the microlens light field image to obtain a degradation-aware enhanced image. A dual-branch feature extraction module is used to receive the degraded perception enhancement image and extract the corresponding spatial domain basic features and angular domain basic features using spatial feature extraction branch and angular feature extraction branch. The linear attention enhancement module is used to receive the basic features of the spatial domain and the basic features of the angular domain, use the linear attention mechanism to perform long-range dependency modeling on the feature sequence, and combine it with residual connections to obtain the enhanced spatial features and angular features. The feature fusion and weighting module is used to fuse the enhanced spatial features and angular features, and to adaptively weight the fused features through a channel attention mechanism to obtain fused features; The quality regression module is used to perform global pooling processing on the fused features and output a quality score of the light field image through a multi-layer fully connected regression network.

[0015] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-described method for assessing the quality of degenerate perceptual blind light field images based on linear attention.

[0016] An electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the above-described method for assessing the quality of degenerate perceptual blind light field images based on linear attention.

[0017] This application introduces a linear attention mechanism, which reduces the computational complexity from the traditional quadratic level to the linear level, significantly improving the efficiency of processing high-resolution light field images.

[0018] This application enhances the model's sensitivity to local image degradation and improves evaluation accuracy by designing a degradation-aware weighting module that utilizes prior structural information from the microlens array to generate a weight map.

[0019] This application uses spatial and angular feature calculations to extract two-dimensional spatial features and multi-view angular features of light field images, and makes a comprehensive judgment through a fusion mechanism, which is more in line with the physical characteristics of light field data. Attached Figure Description

[0020] Figure 1 This is a flowchart illustrating a method for assessing the quality of degraded perception blind light field images based on linear attention. Figure 2 This is a detailed flowchart of a method for evaluating the quality of degraded perception blind light field images based on linear attention. Detailed Implementation

[0021] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0022] This embodiment provides a method for quality assessment of degraded perceptual blind light field images based on linear attention, such as... Figure 1 and Figure 2 As shown, this evaluation method specifically includes the following steps: S1. Obtain the light field image of the microlens to be evaluated, generate a degradation perception weight map based on the spatial distribution characteristics of the microlens array, and use the degradation perception weight map to perform weighted processing on the light field image of the microlens to obtain a degradation perception enhanced image.

[0023] Specifically, generating the degradation-aware weight map involves: The microlens light field image is divided into multiple microlens sub-regions according to the periodic structure of the microlens array; Each microlens sub-region is input into a lightweight feature mapping network for local structure encoding, and the degradation response value of each sub-region is output. All degradation response values ​​are normalized to generate a degradation-aware weight map corresponding to the spatial size of the input image.

[0024] The weighting process of the microlens light field image using the degradation perception weight map is specifically as follows: The degradation perception weight map is multiplied pixel by pixel with the microlens light field image.

[0025] In one embodiment, the microlens light field image L to be evaluated is acquired. The input image L is a single-channel grayscale image with dimensions set to B×1×224×224, where B represents the batch size, the spatial resolution is 224×224, and the angular resolution parameter of the light field image is set to 7×7. Since the light field image is formed by acquiring a microlens array, its spatial structure has periodic characteristics. Therefore, based on the arrangement pattern of the microlens array, the input image L is regularly divided in both the vertical and horizontal directions according to a sampling period of 7 pixels. Thus, the input image is divided into 32×32 microlens sub-regions, each sub-region corresponding to a local viewpoint unit with a size of 7×7, represented as B×1×7×7 in the batch dimension.

[0026] Each microlens sub-region is input into a lightweight feature mapping network. In this process, local structure encoding is performed independently for each sub-region to extract structural information and degradation response features of that region. This mapping process is performed independently for each sub-region, and the output is the corresponding degradation response value, which characterizes the health or degradation degree of that region.

[0027] Normalization was applied to constrain all degradation response values ​​to a range of 0 to 1, where a larger value indicates a more complete structure and lower degree of degradation, while a smaller value indicates more severe degradation. Based on the response results of all sub-regions, they were rearranged according to their spatial relationship in the original image to construct a coarse-grained health weight matrix with a size of B×1×32×32.

[0028] The health weight matrix is ​​spatially expanded so that the weight value corresponding to each microlens sub-region is spatially extended to a corresponding 7×7 pixel range, thereby restoring a degradation-perceived weight map of the same size as the input image, with dimensions of B×1×224×224. During this process, the weight value of each microlens unit is consistently copied to its corresponding local pixel region to ensure that the weight distribution is strictly aligned with the microlens structure.

[0029] The original input image L is multiplied pixel-by-pixel with the degradation-aware weight map and then weighted and fused to obtain the degradation-aware enhanced image I, which maintains a size of B×1×224×224. This weighting process achieves adaptive modulation of different microlens regions, suppressing the response of severely degraded regions while enhancing the expression of structurally intact regions. This provides a more stable data input with quality prior information for subsequent extraction of spatial and angular features.

[0030] S2. Using the spatial feature extraction branch and the angle feature extraction branch, respectively extract the spatial domain basic features Fs and the angle domain basic features Fa of the degraded perception enhancement image I, specifically: The degraded perception-enhanced image I is input into the spatial feature extraction branch and the angular feature extraction branch respectively for feature extraction.

[0031] The spatial feature extraction branch is used to extract structural information related to local texture, edge continuity, and spatial correlation across lens regions in the light field image. It employs a convolution operation with a dilation parameter to extract features from the image; where the dilation rate is set to the angular resolution. This expands the receptive field and enhances the ability to model translens structures. After convolution and nonlinear activation, the basic spatial features Fs are obtained. The angular feature extraction branch is used to extract the arrangement relationship between different viewpoints and viewpoint consistency information. Based on the microlens array structure, the degraded perception enhancement image I is rearranged, and the original spatial arrangement is converted into an angular dimension unfolded form, so that the information of different viewpoints can be displayed on the channel. The angular domain features are extracted through convolution operations to obtain the basic angular domain features Fa.

[0032] In one embodiment, the spatial feature extraction branch uses a 2D convolutional layer with 1 input channel and 64 output channels to initially encode image I. This convolutional layer has a kernel size of 3×3, a dilation rate of 7, a stride of 1, and padding of 7. Since the dilation rate is consistent with the angular resolution, this convolution expands the receptive field while maintaining the input spatial resolution, allowing the feature response at each spatial location to be simultaneously associated with information from adjacent microlens regions. After this convolutional layer and subsequent nonlinear activation, the intermediate feature size output by the spatial branch remains B×64×224×224. Then, a 2D convolutional layer with 64 input channels and 128 output channels is used to compress and aggregate these intermediate features. This convolutional layer has a kernel size of 7×7, a stride of 7, and padding of 0, thereby downsampling the feature map according to the microlens cycle, ultimately obtaining the spatial feature Fs, with a size of B×128×32×32.

[0033] The angular feature extraction branch directly models the viewpoint organization corresponding to the microlens array: the input image I is considered as a two-dimensional sampling plane formed by a regular arrangement of 32×32 microlens sub-regions, each sub-region corresponding to a 7×7 local viewpoint unit. The angular branch uses a two-dimensional convolutional layer with 1 input channel and 128 output channels for feature extraction. The kernel size of this convolutional layer is 7×7, the stride is 7, and the padding is 0. This setting ensures that the receptive range of the convolutional kernel is strictly aligned with a single microlens sub-region, and the convolution operation can directly aggregate the viewpoint sampling information within a sub-region and encode the local viewpoint structure into angular domain features. After this convolutional layer and nonlinear activation, the angular branch outputs angular features Fa, which also have dimensions of B×128×32×32.

[0034] Through the above processing, the output Fs of the spatial feature extraction branch and the output Fa of the angle feature extraction branch maintain the same number of channels and spatial scale, both being B×128×32×32, providing a unified feature input for the subsequent long-range dependency modeling module. This step achieves separate representations of the spatial fidelity and angular consistency of the light field image, enabling subsequent networks to further model and fuse the two types of information at the same feature scale.

[0035] S3. Use the linear attention mechanism to model the long-range dependency of the feature sequence, and combine it with residual connections to obtain the enhanced spatial and angular features.

[0036] Specifically, the method of using linear attention mechanism to model long-range dependencies of feature sequences includes: The input feature map is unfolded into a feature sequence in the spatial dimension; The feature sequence is mapped to query features, key features, and value features through a linear projection layer; By introducing kernel functions to transform query features and key features, the self-attention matrix operation is transformed into a computational form with linear complexity. A recursive update mechanism is introduced to progressively accumulate and model the features in the sequence, resulting in an enhanced feature sequence. The enhanced feature sequence is restored to the feature map size and then added to the input feature map using residuals.

[0037] In one embodiment, the spatial domain basic feature Fs and the angular domain basic feature Fa are respectively subjected to spatial dimension expansion processing. Each 32×32 spatial position is rearranged into a sequence of length 1024 in a fixed order, while keeping the channel dimension unchanged, thereby obtaining spatial sequence features and angular sequence features respectively. The dimension of each sequence feature is B×1024×128.

[0038] For each sequence feature, query features, key features, and value features are generated through three independent linear mapping layers. The query features describe the response requirements at the current spatial location, the key features characterize the association strength of global locations, and the value features carry specific semantic information. Based on this, a linearly complex attention mechanism is introduced. By mapping the query features and key features position-by-position and performing global cumulative computation, the output feature at each position can establish associations with all positions globally, thus achieving effective modeling of long-distance dependencies. This process avoids the traditional quadratic complexity of global attention computation, which causes computational complexity to increase linearly with sequence length.

[0039] In the sequence modeling process, a recursive state update mechanism is introduced, in which the state representation at the current moment is determined by the cumulative state at the previous moment and the current input, thereby realizing the gradual memorization and enhancement of sequence information, enabling the model to effectively capture long-range dependencies and structural consistency information across microlens regions and across viewpoints in the light field image.

[0040] After completing sequence-level feature modeling, the output sequence features are restored to a two-dimensional feature map structure according to their original spatial arrangement, and their size is restored to B×128×32×32. Then, the restored features are added element-wise with the input features of the corresponding branches to obtain the enhanced spatial features Hs and angular features Ha, respectively. This further enhances the global modeling capability while preserving the original local structural information, providing a more discriminative feature representation for subsequent feature fusion and quality regression.

[0041] S4. The enhanced spatial features Hs and angular features Ha are fused together, and the fused features are adaptively weighted using a channel attention mechanism to obtain the fused features.

[0042] Specifically, the adaptive weighting of the fused features through a channel attention mechanism involves: Global average pooling is performed on the fused features to generate channel description vectors; The channel description vector is input into the multilayer perceptron to generate the weight coefficients for each channel; The weighting coefficients are multiplied channel by channel with the fused features to achieve adaptive recalibration of the features.

[0043] In one embodiment, spatial features Hs and angular features Ha are concatenated along the channel dimension to obtain a fused feature. Its size is expanded from the original single-branch feature dimension to B×256×32×32, thereby jointly expressing spatial information and angular information in the same feature space.

[0044] A 1×1 two-dimensional convolution is used to transform and interact with the fused features. This operation is used to linearly combine and redistribute information from feature channels from different sources, thereby achieving effective fusion of cross-branch features to obtain the intermediate fused feature Ff, whose size is restored to B×128×32×32. In this process, the 1×1 convolution does not change the spatial structure, but only reconstructs the channel dimension, allowing spatial and angular information to fully interact at the channel level.

[0045] A global average pooling operation is performed on the fused feature Ff to compress the spatial dimension into a channel vector with a dimension of B×128, which is used to represent the response intensity and importance distribution of each feature channel in the overall image. Subsequently, the channel vector is input into a nonlinear mapping structure consisting of two fully connected mapping layers. The first layer is used to perform inter-channel information compression and nonlinear transformation, and the second layer is used to restore the channel dimension and generate the corresponding channel weight vector with a dimension of B×128.

[0046] The obtained channel weight vector is applied to each channel of the fused feature Ff, and each channel is adaptively recalibrated to enhance the feature response that makes an important contribution to quality discrimination, while suppressing redundant or interfering features. Finally, the fused enhanced feature Fout is obtained, with its size maintained at B×128×32×32.

[0047] Through the above processing, adaptive fusion and selective enhancement of spatial and angular information are achieved, which improves the discriminative ability and stability of feature representation and provides a more reliable input feature basis for subsequent quality regression.

[0048] S5 performs global pooling on the fused feature Fout and outputs a quality score for the light field image through a multi-layer fully connected regression network.

[0049] In one embodiment, the fused feature Fout is subjected to global average pooling and compressed in a spatial dimension of 32×32 to obtain a global feature vector with a dimension of B×128, which is used to characterize the overall spatial structure and angular consistency information of the light field image.

[0050] The global feature vector is input into a multi-layer fully connected regression network for mapping and prediction. This regression network consists of multiple fully connected layers and non-linear activation functions, and achieves the mapping from high-dimensional features to the quality score space through layer-by-layer feature transformation.

[0051] The output scalar Q is used as the quality score of the input light field image to characterize the overall perceived quality of the light field image, thus achieving end-to-end quality evaluation under no-reference conditions.

[0052] This embodiment also provides a degradation-aware blind light field image quality assessment system based on linear attention, used to perform the above-described degradation-aware blind light field image assessment method based on linear attention. The assessment system includes: The degradation-aware weighting module is used to receive the microlens light field image to be evaluated, generate a degradation-aware weight map based on the spatial distribution characteristics of the microlens array, and use the degradation-aware weight map to weight the microlens light field image to obtain a degradation-aware enhanced image. A dual-branch feature extraction module is used to receive the degraded perception enhancement image and extract the corresponding spatial domain basic features and angular domain basic features using spatial feature extraction branch and angular feature extraction branch. The linear attention enhancement module is used to receive the basic features of the spatial domain and the basic features of the angular domain, use the linear attention mechanism to perform long-range dependency modeling on the feature sequence, and combine it with residual connections to obtain the enhanced spatial features and angular features. The feature fusion and weighting module is used to fuse the enhanced spatial features and angular features, and to adaptively weight the fused features through a channel attention mechanism to obtain fused features; The quality regression module is used to perform global pooling processing on the fused features and output a quality score of the light field image through a multi-layer fully connected regression network.

[0053] This embodiment addresses the shortcomings in heterogeneous degradation adaptation, spatial-angular dual-dimensional modeling, and lightweight global modeling. First, a weight map is generated using a microlens degradation modulation module to perform degradation-aware weighting on the input light field image, suppressing interference from damaged areas. Second, a spatial and angular dual-branch feature extraction architecture is constructed to extract spatial fidelity and angular consistency features from the light field image, respectively. Subsequently, an RWKV linear attention module is used to efficiently model long-range dependencies of the dual-branch features. Finally, through feature fusion, channel attention weighting, and fully connected regression, a light field image quality score for a referenceless scene is output. This approach achieves efficient and accurate quality assessment without a reference image, making it suitable for fields with high requirements for light field image quality, such as virtual reality, augmented reality, computational photography, and medical imaging.

[0054] Secondly, this embodiment provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the aforementioned method for evaluating the quality of degenerate perceptually blind light field images based on linear attention.

[0055] Specifically, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0056] In this embodiment, the computer-readable storage medium may include, but is not limited to: portable computer disk, hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof.

[0057] Preferably, the computer-readable storage medium is a non-volatile storage medium, such as a solid-state drive (SSD) or flash memory.

[0058] In addition, this embodiment also provides an electronic device, which may be a server, a desktop computer, a laptop computer, or an embedded image processing device. The electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the aforementioned method for assessing the quality of degenerate perceptually blind light field images based on linear attention.

[0059] The processor can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. In this embodiment, the processor is mainly used to perform convolution operations, linear attention matrix operations, and regression calculations for fully connected layers.

[0060] Memory, as a computer-readable storage medium, is used to store software programs, computer-executable programs, and data. Memory can primarily include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a given function (such as an image processing driver, a quality assessment algorithm library, etc.); the data storage area may store data created based on terminal usage (such as a light field image dataset to be evaluated, model weight files, etc.). Furthermore, memory may include high-speed random access memory and non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device.

[0061] It should be noted that the electronic device in this embodiment can be a standalone computing device or part of a cloud server cluster. When applied in the cloud, the electronic device receives light field images uploaded by the terminal via the network, completes quality assessment, and then feeds back the scoring results to the terminal.

[0062] The above description is merely a preferred embodiment of this application, but the scope of protection of this application is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in this application, based on the technical solution and concept of this application, should be included within the scope of protection of this application.

Claims

1. A method for quality assessment of degenerate perceptually blind light field images based on linear attention, characterized in that, The evaluation method includes the following steps: The light field image of the microlens to be evaluated is acquired, a degradation perception weight map is generated based on the spatial distribution characteristics of the microlens array, and the degradation perception weight map is used to weight the light field image of the microlens to obtain a degradation perception enhanced image. The spatial feature extraction branch and the angular feature extraction branch are used to extract the basic spatial domain features and the basic angular domain features of the degraded perception enhancement image, respectively. The linear attention mechanism is used to model the long-range dependency of the feature sequence, and the enhanced spatial and angular features are obtained by combining residual connections. The enhanced spatial and angular features are fused, and the fused features are adaptively weighted using a channel attention mechanism to obtain fused features. The fused features are subjected to global pooling, and the quality score of the light field image is output through a multi-layer fully connected regression network.

2. The method for assessing the quality of degenerate perceptually blind light field images based on linear attention according to claim 1, characterized in that, The generation of the degradation-aware weight map is specifically as follows: The microlens light field image is divided into multiple microlens sub-regions according to the periodic structure of the microlens array; Each microlens sub-region is input into a lightweight feature mapping network for local structure encoding, and the degradation response value of each sub-region is output. All degradation response values ​​are normalized to generate a degradation-aware weight map corresponding to the spatial size of the input image.

3. The method for assessing the quality of degenerate perceptually blind light field images based on linear attention according to claim 1, characterized in that, The weighting process of the microlens light field image using the degradation perception weight map is specifically as follows: The degradation perception weight map is multiplied pixel by pixel with the microlens light field image.

4. The method for assessing the quality of degenerate perceptually blind light field images based on linear attention according to claim 1, characterized in that, In the spatial feature extraction branch, dilated convolution is used to extract features from the image. The dilation rate of the dilated convolution is set as the angular resolution parameter of the light field image to expand the receptive field and cover the structural information across the microlens region, thereby obtaining the basic features in the spatial domain.

5. The method for assessing the quality of degenerate perceptually blind light field images based on linear attention according to claim 1, characterized in that, In the angle feature extraction branch, the viewpoint information inside the microlens sub-region is aggregated through convolution operations to extract the basic features of the angle domain.

6. The method for quality assessment of degenerate perceptually blind light field images based on linear attention according to claim 1, characterized in that, The method of using linear attention mechanism to model long-range dependencies of feature sequences specifically involves: The input feature map is unfolded into a feature sequence in the spatial dimension; The feature sequence is mapped to query features, key features, and value features through a linear projection layer; By introducing kernel functions to transform query features and key features, the self-attention matrix operation is transformed into a computational form with linear complexity. A recursive update mechanism is introduced to progressively accumulate and model the features in the sequence, resulting in an enhanced feature sequence. The enhanced feature sequence is restored to the feature map size and then added to the input feature map using residuals.

7. The method for assessing the quality of degenerate perceptually blind light field images based on linear attention according to claim 1, characterized in that, The adaptive weighting of the fused features through a channel attention mechanism is specifically as follows: Global average pooling is performed on the fused features to generate channel description vectors; The channel description vector is input into the multilayer perceptron to generate the weight coefficients for each channel; The weighting coefficients are multiplied channel by channel with the fused features to achieve adaptive recalibration of the features.

8. A degradation-perceived blind light field image quality assessment system based on linear attention, used to execute the degradation-perceived blind light field image assessment method based on linear attention as described in any one of claims 1-7, characterized in that, The evaluation system includes: The degradation-aware weighting module is used to receive the microlens light field image to be evaluated, generate a degradation-aware weight map based on the spatial distribution characteristics of the microlens array, and use the degradation-aware weight map to weight the microlens light field image to obtain a degradation-aware enhanced image. A dual-branch feature extraction module is used to receive the degraded perception enhancement image and extract the corresponding spatial domain basic features and angular domain basic features using spatial feature extraction branch and angular feature extraction branch. The linear attention enhancement module is used to receive the basic features of the spatial domain and the basic features of the angular domain, use the linear attention mechanism to perform long-range dependency modeling on the feature sequence, and combine it with residual connections to obtain the enhanced spatial features and angular features. The feature fusion and weighting module is used to fuse the enhanced spatial features and angular features, and to adaptively weight the fused features through a channel attention mechanism to obtain fused features; The quality regression module is used to perform global pooling processing on the fused features and output a quality score of the light field image through a multi-layer fully connected regression network.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the degradation perception blind light field image quality assessment method based on linear attention as described in any one of claims 1-7.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the degradation perception blind light field image quality assessment method based on linear attention as described in any one of claims 1-7.