Image quality evaluation method based on dynamic weight feature extraction and three-dimensional fusion under meta-learning framework

Through the dynamic weight feature extraction and three-dimensional fusion method under the meta-learning framework, the problem that the feature weight relationship in the prior art is not considered is solved, and more accurate and adaptable image quality evaluation is achieved, which is suitable for the evaluation of multiple distorted images.

CN120544012APending Publication Date: 2025-08-26NANJING NORMAL UNIVERSITY
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510717215.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

The existing reference-free image quality evaluation algorithm fails to effectively consider the weight relationship between features during feature fusion, resulting in poor robustness and difficulty in adapting to the distortion fluctuation laws in different scenarios, affecting the evaluation effect.

Method used

The dynamic weight feature extraction and three-dimensional fusion method under the meta-learning framework are adopted to filter significant areas through gradient analysis, and multi-receptive fields and dynamic coefficient modules are designed, combined with three-dimensional convolution and improved Vision Transformer network for feature fusion and semantic analysis, dynamically adjust feature weights, and enrich feature expression methods.

Benefits of technology

It improves the accuracy and adaptability of image quality evaluation, reduces the dependence on reference images, and achieves a more objective and rapid evaluation process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120544012A_ABST
    Figure CN120544012A_ABST
Patent Text Reader

Abstract

The invention discloses an image quality evaluation method based on dynamic weight feature extraction and three-dimensional fusion under a meta-learning framework, and the method comprises the steps: designing a multi-receptive-field and dynamic parameter module in an image quality evaluation model, extracting feature information under different views, and enriching and weighing multi-receptive-field features; screening out a salient region by using image gradient information outside the image quality evaluation model, designing a distortion classification model based on the salient region, and providing shared features for the image quality evaluation model; using random cutting and three-dimensional convolution to deeply fuse the multi-receptive-field features and the shared features to obtain deep fusion features; shartcut Vision Transform is proposed to perform semantic analysis on the fusion features, the association between the overall features and quality fluctuation is summarized, and the quality score of the image is obtained through linear mapping. The multi-distortion image quality evaluation method can accurately evaluate the quality of the multi-distortion image, can effectively process a large amount of data, and has the advantages of being objective, rapid and high in usability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of image quality assessment, and specifically relates to an image quality assessment method based on dynamic weight feature extraction and three-dimensional fusion under a meta-learning framework. Background Art

[0002] Image quality assessment algorithms, a key technology in computer vision, evaluate image quality based on human visual characteristics and are often used as screening solutions in a range of visual tasks. No-reference image quality assessment algorithms, which can evaluate distorted images without requiring a reference image, are more practical and have therefore attracted increasing research and attention. Early no-reference image quality assessment algorithms described visual quality by leveraging various image features. However, the limited number of features and simple analysis methods resulted in poor adaptability and limited application to diverse datasets.

[0003] The rise of deep learning has provided new insights into computer vision. With the help of some classic neural convolutional networks, many image quality assessment algorithms have achieved new improvements. From the perspective of classic neural convolutional networks and the network hierarchy of these methods, the richer the network layer, the more feature information is extracted, and the more accurate the predicted quality score after analysis. To further enrich the feature information and representation in network models, some methods have designed multi-scale, interactive, and fusion schemes within the network. Taking the common Feature Pyramid Network (FPN) as an example, this model can effectively capture multi-level features with different receptive fields, thereby extracting semantic and detailed information. This shows that diverse receptive field variations can effectively improve the model's performance. However, these multi-scale methods use relatively simple methods to fuse different features, mostly concatenating channel features or adding matrix elements. The fused features are then processed through common convolution and pooling. In reality, different features have different impacts on the overall quality assessment, and the weight relationship between features is not fixed. Therefore, simple fusion methods do not give much consideration to the weight relationship between features. As a result, the fused features cannot generalize distortion fluctuations in different scenarios, and the robustness of the features is relatively poor, which affects the image quality assessment results. Summary of the Invention

[0004] Purpose of the invention: In order to overcome the shortcomings of the existing technology, an image quality evaluation method based on dynamic weight feature extraction and three-dimensional fusion under a meta-learning framework is provided, which aims to enable the model to accurately evaluate the quality of various distorted images, and has the advantages of objectivity, speed, and strong usability, providing a scientific basis for image quality evaluation methods.

[0005] Technical Solution: To achieve the above objectives, the present invention provides an image quality assessment method based on dynamic weight feature extraction and three-dimensional fusion under a meta-learning framework, comprising the following steps:

[0006] S1: Perform gradient analysis on the overall image information and local information. The gradient analysis results combine horizontal and vertical information changes, use the information difference between local and global regions to design a saliency score calculation method, and filter out salient areas of the image based on the saliency score;

[0007] S2: Use the salient regions obtained in step S1 as the input of the constructed distortion classification network to extract shared features;

[0008] S3: Extract multiple receptive field features of the entire image through the constructed multi-receptive field and dynamic coefficient module, and add dynamic coefficients between different network layers of the module to balance the weight proportion of features under different receptive fields;

[0009] S4: Design a feature fusion module to randomly crop the multi-receptive field features obtained in step S3 and the shared features obtained in step S2, and splice them into new square features according to the second dimension. Evenly segment the spliced ​​overall features, and use three-dimensional convolution through the feature fusion module to perform spatial information fusion to obtain deep fusion features.

[0010] S5: The improved Vision Transformer network is used as an analysis module to perform semantic analysis and distortion information analysis on the deep fusion features, and finally obtain the image quality score through linear mapping.

[0011] Furthermore, the method for screening the salient areas in step S1 includes:

[0012] Design a sliding window to take small blocks of the image; use the Sobel operator to calculate the grayscale gradient information of each image block in the x and y directions;

[0013] After the pixel gradient information of each image block is processed by Gaussian filtering and generalized Pareto distribution, the overall salient information distribution of the local image block is obtained;

[0014] The mean of the salient information of the image block is used to represent the standard of the overall image gradient change. The corresponding saliency score is obtained according to the difference between the salient information distribution of each image block and the mean, and the image block with the largest score is regarded as the salient area.

[0015] Furthermore, the significance score in step S1 is calculated using the following formula:

[0016]

[0017] Among them, x irepresents the gradient feature matrix of the i-th image block, and E represents the identity matrix.

[0018] Furthermore, the construction and operation of the distortion classification network in step S2 includes:

[0019] Design a residual network with inter-layer feature interaction. Use the residual network to extract features and perform semantic analysis on salient regions. Use the distortion type of the distorted image as a label and use cross entropy to calculate the loss, thereby performing gradient descent. After the residual network reaches convergence, freeze the network parameters.

[0020] When testing new images, the feature information processed by the third residual module of the residual network is retained and provided as shared knowledge to subsequent image quality evaluation tasks.

[0021] Furthermore, in step S3, convolution kernels of different sizes are designed to achieve feature extraction under different receptive fields. Each receptive field branch network is composed of three dense blocks and three transition blocks. Each dense block is composed of multiple feature reuse layers, wherein each feature reuse layer is composed of two convolutional layers, two normalization layers, and two ReLU activation functions. Each feature reuse layer concatenates the input features and output features of this layer as the input of the next layer, completing feature reuse between layers and allowing features at different layers to interact only under the same receptive field before fusion.

[0022] Each transition block is used to integrate the shallow features obtained by the previous dense block, while compressing the feature information, highlighting the role of key features, and providing a more refined input for the next dense block.

[0023] Furthermore, in step S3, a dynamic coefficient is added in each convolution process of each receptive field, and the dynamic coefficient is updated with gradient descent and back propagation during network training, so as to set a flexible weight relationship for different receptive field features under the same hierarchical relationship.

[0024] Furthermore, in step S4, the channel features of the multiple multi-receptive field features and the shared features are clipped by the screening frame and then spliced ​​according to the second dimension to obtain a new feature, and the feature is evenly divided according to the channel dimension and then dimensionally expanded to obtain a three-dimensional spatial feature;

[0025] The three-dimensional spatial features are processed by three-dimensional convolution to obtain features; the channel dimension and sequence dimension are compressed to obtain fused features.

[0026] Furthermore, in step S5, the improvements to the Vision Transformer network include:

[0027] Vision Transformer consists of 6 internal modules 1 and 2 alternatingly;

[0028] Design a feature reuse strategy in every two sub-modules, compress the reused features and add them to the new features as feature input;

[0029] The multi-head attention mechanism in every two sub-modules is improved to multi-feature fusion attention, which performs bitwise mapping based on the position and semantic information of the initial features and reused features to obtain the fused features and position weight information.

[0030] Furthermore, in step S5, the Vision Transformer is used to perform semantic analysis on the obtained overall features, and by adding a feature reuse strategy to the Vision Transformer, the network's reuse and analysis of long-distance information is further improved, specifically including: the output features obtained after processing by the internal module 1 are subjected to dimensionality reduction processing to obtain the main features, the output features obtained after processing by the internal module 2 are subjected to dimensionality reduction processing to obtain the reused features, and then the main features and the reused features are re-spliced ​​into new features as the input of the next internal module 1.

[0031] Furthermore, in step S5, the multi-head attention mechanism in the module with the overall feature as input is improved to a multi-feature fusion attention mechanism, specifically including: the features before being input to the multi-head attention mechanism are divided into position information features, main features and reused features, and more efficient feature information is obtained after multi-layer perceptron and normalization processing, position information features, main features f main and reuse feature f re After being compressed into one-dimensional features, they are input into the multi-feature fusion attention mechanism. In the multi-feature fusion attention mechanism, the three features are used to obtain their corresponding Q, K, and V features. The weight information after matrix multiplication of the respective Q and K features is mapped to the V feature after softmax normalization, so as to obtain the weight relationship and semantic information between different features. The main features and reused features are fused after such an attention mechanism, which can obtain richer semantic information relationships and help the overall network more accurately describe the laws of image quality fluctuations and distortion changes;

[0032] After the multi-feature fusion attention mechanism processes the position information features, main features, and reused features, they are concatenated to form a new overall feature for subsequent feature reuse operations. The resulting semantic information is linearly mapped to obtain the final quality score. The process of the multi-feature fusion attention mechanism can be expressed as follows:

[0033]

[0034] The present invention utilizes the differences in feature extraction of the model under different receptive fields to design a multi-receptive field feature extraction module. At the same time, in order to further flexibly adjust the feature weight relationship between different receptive fields, dynamic parameters are added to the module to meet the real-time adjustment of feature weights during training, and ultimately generate richer feature information within the model.

[0035] The present invention utilizes the commonalities between different visual tasks to provide more features for the image quality assessment model. Specifically, it obtains shared features from the distortion classification task, including different analyses and expressions of distortion information, and provides more distortion feature information from an external perspective of the model.

[0036] From the perspective of multi-dimensional feature processing, the present invention randomly crops two-dimensional features and splices them according to the row and column dimensions of the two-dimensional features. The spliced ​​features are deeply processed and fused using three-dimensional convolution to obtain fused features.

[0037] The semantic analysis module designed in this paper utilizes a multi-head attention mechanism and semantic reuse for short-term connections, accurately analyzing feature semantic information and deriving the final image quality score through linear mapping. This invention boasts high accuracy and strong adaptability, providing a feasible solution and theoretical basis for the research of image quality assessment methods.

[0038] When the present invention uses a network model to complete the image quality assessment task, the effectiveness and richness of the feature information largely determine the performance of the model. In order to further enrich the features in the training process, the present invention captures more information from the inside and outside of the image quality assessment model. Inside the model, multi-receptive field features are captured, and at the same time, dynamic coefficients are added to flexibly adjust the weight relationship between features. Outside the model, a distortion classification model is used to provide shared features. In order to fuse different features, random cropping and three-dimensional convolution are designed for deep fusion to obtain fused features. In order to further analyze the semantic relationship of the fused features, a Shortcut Vision Transformer is proposed, and the obtained semantic information is mapped to obtain the image quality score.

[0039] Beneficial effects: Compared with the prior art, the present invention has the following advantages:

[0040] 1. Compared with the subjective image quality evaluation method that is time-consuming, unstable and relies on manual experience, the present invention has the advantages of high efficiency and real-time.

[0041] 2. Compared with the full-reference image quality evaluation method and the semi-reference image quality evaluation method, the present invention does not require a reference image, reduces the cost of obtaining a reference image, and has a wider range of application scenarios.

[0042] 3. Compared with the existing image quality evaluation method that selects a single feature for training, the present invention designs feature analysis under multiple receptive fields, enriching the feature expression method.

[0043] 4. Compared with the existing single image quality assessment task, the present invention uses distortion classification as an auxiliary task to obtain shared features, further enriching the feature source of the image quality assessment model and providing more effective information.

[0044] 5. Compared with common fusion processing methods such as splicing and convolution, this paper starts from the perspective of multi-dimensional processing of features and uses spatial random cropping and three-dimensional convolution to fuse features.

[0045] 6. Compared with the existing Transformer-based semantic analysis method, the idea of ​​feature short connection is applied to semantic analysis, realizing semantic short connection and interaction. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Figure 1 is a flow chart of the method of the present invention;

[0047] Figure 2 Schematic diagram of the operation of feature extraction and dynamic parameters for different sizes;

[0048] Figure 3 This is a diagram of the analysis process of the multi-feature fusion attention mechanism in Shortcut Vision Transformer;

[0049] Figure 4 The following diagram shows the network gradient change with and without the Shortcut Vision Transformer. To further compare the gradient change process and training time, an overhead view of the gradient change is shown, and the loss change and training time are recorded at the corresponding locations. DETAILED DESCRIPTION

[0050] The present invention is further illustrated below with reference to the accompanying drawings and specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and are not used to limit the scope of the present invention. After reading the present invention, modifications of various equivalent forms of the present invention made by those skilled in the art all fall within the scope defined by the claims attached to this application.

[0051] Example 1:

[0052] This embodiment provides an image quality evaluation method based on dynamic weight feature extraction and three-dimensional fusion under a meta-learning framework. Figure 1 As shown, it includes the following steps:

[0053] S1: Perform gradient analysis on the overall image information and local information. The gradient analysis results combine horizontal and vertical information changes, use the information difference between local and global regions to design a saliency score calculation method, and filter out salient areas of the image based on the saliency score;

[0054] S2: Design a distortion classification network as an auxiliary task. Taking the salient regions obtained in step S1 as input, a residual network is used to extract features from the salient regions. The linear layer analyzes the distortion information in the features and assigns the distortion type. During the forward propagation, the 14*14 features in the third module of the residual network are retained to provide shared features for the main task (image quality assessment).

[0055] S3: Design a multi-receptive field and dynamic coefficient module for image quality assessment tasks. This module is used to extract features from multiple receptive fields across the entire image. Dynamic coefficients are added between different network layers in the module to balance the weights of features from different receptive fields, thereby optimizing the representation of features from multiple receptive fields.

[0056] S4: Design a feature fusion module to randomly crop the multi-receptive field features obtained in step S3 and the shared features obtained in step S2, and splice them into new square features along the second dimension. The spliced ​​overall features are evenly segmented and spatial information is fused using 3D convolution. The resulting fused features are a combination of multi-receptive field features and shared knowledge.

[0057] S5: To perform semantic analysis and further distortion information analysis on the fused features in step S4, the improved Vision Transformer network is used as the analysis module. To further enhance the reuse and analysis of long-range features by the Vision Transformer, a feature reuse strategy is designed in each of the two submodules. The reused features are compressed and added to the new features as feature input.

[0058] S6: To better analyze the reused features, the multi-head attention mechanism in every two submodules is improved to multi-feature fusion attention. Bit-by-bit mapping is performed based on the position and semantic information of the initial and reused features to obtain the fused features and position weight information. The overall Vision Transformer is improved to the Shortcut Vision Transformer. Finally, the image quality score is obtained through linear mapping.

[0059] In step S1:

[0060] The salient region screening method is as follows: a sliding window with a size of 224*224 and a step size of 112 is designed to select small blocks of the image; the grayscale gradient information of each image block in the x and y directions is calculated using the Sobel operator; to amplify and enhance the information difference between the local image block and the overall image, the pixel gradient information of each image block is processed by Gaussian filtering and generalized Pareto distribution to obtain the overall salient information distribution of the local image block; the mean of the salient information of the image block is used to represent the standard of the overall image gradient change, and the corresponding salient score is obtained based on the difference between the salient information distribution of each image block and the mean. The image block with the largest score is selected as the salient region.

[0061] The significance score is calculated using the following formula:

[0062]

[0063] Among them, x i represents the gradient feature matrix of the i-th image block, and E represents the identity matrix.

[0064] In step S2:

[0065] An auxiliary task for image distortion classification is designed. This task takes the salient regions from step S1 as input, and constructs a residual network with inter-layer feature interaction. This network is used to extract features and perform semantic analysis on the salient regions. The distortion type of the distorted image is used as a label, and the loss is calculated using cross-entropy, allowing for gradient descent. After the residual network reaches convergence, the network parameters are frozen. Because the distortion classification task emphasizes capturing detailed information about the distortion to distinguish between different types of distortion, the extracted features represent distortion patterns differently from those used in the primary task (image quality assessment). When testing new images, the feature information processed by the third residual module of the residual network is retained. This feature information contains rich distortion and semantic information, and is provided as shared knowledge to the subsequent image quality assessment task, providing more local feature information for the features in the primary task.

[0066] In step S3:

[0067] like Figure 2As shown in the figure, convolution kernels of different sizes are designed to realize feature extraction under different receptive fields. Different from the effectiveness of pyramid features in analyzing multi-scale features at the same network depth, step S3 realizes the analysis process of different receptive fields between networks of different depths, that is, the extraction process of each receptive field corresponds to a receptive field branch network, each receptive field branch network is composed of three dense blocks and three transition blocks, each dense block is composed of multiple feature reuse layers, each feature reuse layer is composed of two convolution layers, two normalization layers, and two Relu activation functions alternately, and each feature reuse layer will splice the input features and output features of this layer as the input of the next layer, completing the feature reuse between layers, and allowing features at different levels to interact only under the same receptive field before fusion.

[0068] Each transition block integrates shallow features obtained from the previous dense block, compressing feature information and highlighting the role of key features, providing more refined input for the next dense block. Furthermore, a dynamic coefficient is added to each convolution process for each receptive field. This coefficient is updated during network training with gradient descent and backpropagation. This provides a flexible weight relationship for features in different receptive fields within the same hierarchical relationship, making the analysis process more balanced across different receptive fields. Convolution operations in different receptive field paths can better adapt to changes in input data, enhancing the model's ability to express complex features.

[0069] In step S4:

[0070] The channel features (14*14) of the three multi-receptive field features (1024*14*14) and the shared features (1024*14*14) are cropped with a 7*7 filter frame and then concatenated along the second dimension to obtain a new 1024*14*14 feature. This feature is then equally divided into 16 64*14*14 features along the channel dimension and then dimensionally expanded to obtain a 16*64*14*14 three-dimensional spatial feature. The three-dimensional spatial feature undergoes three-dimensional convolution to obtain a 32*24*14*14 feature. The channel and sequence dimensions are compressed to obtain a 768*14*14 fused feature. This fused feature combines multi-receptive field features and shared features while meeting the input requirements of subsequent modules, enhancing the model's overall ability to integrate and analyze different features.

[0071] In step S5:

[0072] The Vision Transformer performs semantic analysis on the resulting overall features. By incorporating a feature reuse strategy into the Vision Transformer, the network's ability to reuse and analyze long-range information is further enhanced. Specifically, the Vision Transformer consists of six alternating internal modules 1 and 2. The output features (size 197*768) obtained by internal module 1 undergo dimensionality reduction to obtain primary features (size 197*680). The output features (size 197*768) obtained by internal module 2 undergo dimensionality reduction to obtain reused features (size 197*88). The primary and reused features are then concatenated into new features (size 197*768) as input to the next internal module 1. Therefore, feature information is compressed in each module, allowing semantics at different levels to interact. Long-range semantic relationships also incorporate semantic information from different levels, ensuring the richness of overall semantic information and enabling more accurate and comprehensive analysis of image quality fluctuation patterns.

[0073] In step S6:

[0074] like Figure 3 As shown in the figure, in order to better analyze the overall features after reorganization, the multi-head attention mechanism in the module with the overall features as input is improved to a multi-feature fusion attention mechanism. The specific details are as follows: the features before input to the multi-head attention mechanism are divided into position information features, main features and reused features. After the multi-layer perceptron (MLP) and normalization (Norm), more efficient feature information is obtained. The position information features, main features f main and reuse feature f reAfter being compressed into one-dimensional features, they are input into the multi-feature fusion attention mechanism. In this multi-feature fusion attention mechanism, the three features are converted into corresponding Q, K, and V features. The weight information of each Q and K feature after matrix multiplication is normalized by softmax and then mapped to the V feature. This method obtains the weight relationship and semantic information between different features. The main features and reused features are fused after this attention mechanism, which can obtain richer semantic information relationships, helping the overall network to more accurately describe the laws of image quality fluctuations and distortion changes. The position information features, main features, and reused features processed by the multi-feature fusion attention mechanism are spliced ​​to form a new overall feature for subsequent feature reuse operations. Due to the feature reuse and residual connection of the Vision Transformer after steps S5 and S6, it is improved to the Shortcut Vision Transformer. The semantic information obtained by the Shortcut Vision Transformer is linearly mapped to obtain the final quality score. The process of the multi-feature fusion attention mechanism can be expressed as follows:

[0075]

[0076] Where, and Respectively represent f main and f re The transpose of the corresponding Q feature; and Respectively represent f main and f re The corresponding K features; and Respectively represent f main and f re Corresponding V features; Represents the dimension of K features.

[0077] The above steps S1 to S6 can be summarized as follows:

[0078] Within the image quality assessment model, a multi-receptive field and dynamic parameter module is designed to extract feature information under different fields of view, enrich and weigh the multi-receptive field features; outside the image quality assessment model, image gradient information is used to screen out salient areas, and a distortion classification model based on salient areas is designed to provide shared features for the image quality assessment model; random cropping and three-dimensional convolution are used to deeply fuse multi-receptive field features and shared features to obtain deeply fused features; a ShortcutVision Transformer is proposed to perform semantic analysis on the fused features, summarize the correlation between overall features and quality fluctuations, and obtain the image quality score through linear mapping.

[0079] Example 2:

[0080] Based on the method in Example 1, in order to verify the effectiveness and practical effect of the method of the present invention, this example applies the method of the present invention and other methods to the LIVEC dataset. The specific comparison results are shown in Table 1. The optimal performance values ​​corresponding to different evaluation indicators are highlighted in bold:

[0081] Table 1 - Performance comparison results on the LIVEC database

[0082]

[0083] It can be seen from the experimental data in Table 1 that the method of the present invention can accurately evaluate the distorted image and has a high consistency with the subjective evaluation results of the human eye.

[0084] In this example, in order to verify the effect of the Shortcut Vision Transformer proposed in the present invention, Figure 4 The figure shows the network gradient change diagram with and without the Shortcut Vision Transformer. At the same time, in order to further compare the gradient change process and training time, the top view of the gradient change is shown, and the loss change and training time are recorded at the corresponding position. Figure 4 It can be intuitively observed that after using Shortcut Vision Transformer, the training convergence speed of the model is significantly faster, the final converged Loss value is smaller, and the efficiency of the model is significantly improved.

Claims

1. An image quality assessment method based on dynamic weight feature extraction and three-dimensional fusion under a meta-learning framework, characterized by: The steps include: S1: Perform gradient analysis on the overall image information and local information. The gradient analysis results combine horizontal and vertical information changes, use the information difference between local and global regions to design a saliency score calculation method, and filter out salient areas of the image based on the saliency score; S2: Use the salient regions obtained in step S1 as the input of the constructed distortion classification network to extract shared features; S3: Extract multiple receptive field features of the entire image through the constructed multi-receptive field and dynamic coefficient module, and add dynamic coefficients between different network layers of the module to balance the weight proportion of features under different receptive fields; S4: Design a feature fusion module to randomly crop the multi-receptive field features obtained in step S3 and the shared features obtained in step S2, and splice them into new square features according to the second dimension. Evenly segment the spliced ​​overall features, and use three-dimensional convolution through the feature fusion module to perform spatial information fusion to obtain deep fusion features. S5: The improved Vision Transformer network is used as an analysis module to perform semantic analysis and distortion information analysis on the deep fusion features, and finally obtain the image quality score through linear mapping.

2. The image quality assessment method based on dynamic weight feature extraction and three-dimensional fusion under a meta-learning framework according to claim 1 is characterized in that: The method for screening the salient areas in step S1 includes: Design a sliding window to take small blocks of the image; use the Sobel operator to calculate the grayscale gradient information of each image block in the x and y directions; After the pixel gradient information of each image block is processed by Gaussian filtering and generalized Pareto distribution, the overall salient information distribution of the local image block is obtained; The mean of the salient information of the image block is used to represent the standard of the overall image gradient change. The corresponding saliency score is obtained according to the difference between the salient information distribution of each image block and the mean, and the image block with the largest score is regarded as the salient area.

3. The image quality assessment method based on dynamic weight feature extraction and three-dimensional fusion under a meta-learning framework according to claim 2 is characterized in that: The significance score in step S1 is calculated using the following formula: Among them, x i represents the gradient feature matrix of the i-th image block, and E represents the identity matrix.

4. The image quality assessment method based on dynamic weight feature extraction and three-dimensional fusion under a meta-learning framework according to claim 1 is characterized in that: The construction and operation of the distortion classification network in step S2 includes: Design a residual network with inter-layer feature interaction. Use the residual network to extract features and perform semantic analysis on salient regions. Use the distortion type of the distorted image as a label and use cross entropy to calculate the loss, thereby performing gradient descent. After the residual network reaches convergence, freeze the network parameters. When testing new images, the feature information processed by the third residual module of the residual network is retained and provided as shared knowledge to subsequent image quality evaluation tasks.

5. The image quality assessment method based on dynamic weight feature extraction and three-dimensional fusion under a meta-learning framework according to claim 1 is characterized in that: In step S3, convolution kernels of different sizes are designed to achieve feature extraction under different receptive fields. Each receptive field branch network is composed of three dense blocks and three transition blocks. Each dense block is composed of multiple feature reuse layers, wherein each feature reuse layer is composed of two convolutional layers, two normalization layers, and two Relu activation functions. Each feature reuse layer concatenates the input features and output features of this layer as the input of the next layer, completing feature reuse between layers and allowing features at different layers to interact only under the same receptive field before fusion. Each transition block is used to integrate the shallow features obtained by the previous dense block, while compressing the feature information, highlighting the role of key features, and providing a more refined input for the next dense block.

6. The image quality assessment method based on dynamic weight feature extraction and three-dimensional fusion under a meta-learning framework according to claim 5 is characterized in that: In step S3, a dynamic coefficient is added in each convolution process of each receptive field. The dynamic coefficient is updated with gradient descent and back propagation during network training, and a flexible weight relationship is set for different receptive field features under the same hierarchical relationship.

7. The image quality assessment method based on dynamic weight feature extraction and three-dimensional fusion under a meta-learning framework according to claim 4 is characterized in that: In step S4, after the channel features of the multiple receptive field features and the shared features are clipped by the screening frame, they are spliced ​​according to the second dimension to obtain a new feature, and the feature is evenly divided according to the channel dimension and then dimensionally expanded to obtain a three-dimensional spatial feature; The three-dimensional spatial features are processed by three-dimensional convolution to obtain features; the channel dimension and sequence dimension are compressed to obtain fused features.

8. The image quality assessment method based on dynamic weight feature extraction and three-dimensional fusion under a meta-learning framework according to claim 1 is characterized in that: In step S5, the improvements to the Vision Transformer network include: Vision Transformer consists of 6 internal modules 1 and 2 alternatingly; Design a feature reuse strategy in every two sub-modules, compress the reused features and add them to the new features as feature input; The multi-head attention mechanism in every two sub-modules is improved to multi-feature fusion attention, which performs bitwise mapping based on the position and semantic information of the initial features and reused features to obtain the fused features and position weight information.

9. The image quality assessment method based on dynamic weight feature extraction and three-dimensional fusion under a meta-learning framework according to claim 8, characterized in that: In step S5, the Vision Transformer is used to perform semantic analysis on the obtained overall features. By adding a feature reuse strategy to the Vision Transformer, the network's reuse and analysis of long-distance information is further improved, specifically including: the output features obtained after processing by the internal module 1 are subjected to dimensionality reduction processing to obtain the main features, the output features obtained after processing by the internal module 2 are subjected to dimensionality reduction processing to obtain the reused features, and then the main features and the reused features are re-spliced ​​into new features as the input of the next internal module 1.

10. The image quality assessment method based on dynamic weight feature extraction and three-dimensional fusion under a meta-learning framework according to claim 1, characterized in that: In the step S5, the multi-head attention mechanism in the module with the overall feature as input is improved to a multi-feature fusion attention mechanism, specifically including: the features before being input to the multi-head attention mechanism are divided into position information features, main features and reused features, and more efficient feature information is obtained after multi-layer perceptron and normalization processing, position information features, main features f main and reuse feature f re After being compressed into one-dimensional features, they are input into the multi-feature fusion attention mechanism. In the multi-feature fusion attention mechanism, the three features are used to obtain their corresponding Q, K, and V features. The weight information after matrix multiplication of the respective Q and K features is mapped to the V feature after softmax normalization, so as to obtain the weight relationship and semantic information between different features. The main features and reused features are fused after such an attention mechanism, which can obtain richer semantic information relationships and help the overall network more accurately describe the laws of image quality fluctuations and distortion changes; The position information features, main features, and reused features processed by the multi-feature fusion attention mechanism are spliced ​​together to form a new overall feature for subsequent feature reuse operations. The final quality score is obtained by linearly mapping the obtained semantic information.

Citation Information

Cited By

  • Method and system for evaluating multi-dimensional ability of Chinese large language model

    CN121681392A

  • Multi-mode-based salient object detection method and device and related medium

    CN121861030A