No-reference super-resolution image quality evaluation method based on semantic feature enhancement

By adopting a multi-scale network model, gated pooling module, external attention mechanism and cross-scale attention mechanism in the reference-free super-resolution image quality evaluation method, the limitations of existing methods in evaluating super-resolution image quality are solved, and more accurate and reliable image quality evaluation is achieved.

CN120198407APending Publication Date: 2025-06-24CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510348014.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

Existing reference-free super-resolution image quality evaluation methods have limitations in evaluating super-resolution image quality based on generative adversarial networks, especially due to inconsistencies between human vision systems and machine vision, resulting in limited accuracy and reliability of evaluation results.

Method used

The reference-free super-resolution image quality evaluation method based on semantic feature enhancement is adopted, and the multi-scale features of the image are extracted through a multi-scale network model, the feature enhancement is performed using the gated pooling module and the external attention mechanism, and the attention map between the features is calculated through the cross-scale attention mechanism, and the prediction of the quality score is finally achieved using a multi-layer perceptron.

Benefits of technology

It improves the accuracy and reliability of image quality evaluation, can more effectively capture complex and mixed image distortion in super-resolution images, enhances the generalization ability of the model, and provides better image quality evaluation results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198407A_ABST
    Figure CN120198407A_ABST
Patent Text Reader

Abstract

The invention relates to a no-reference super-resolution image quality evaluation method based on semantic feature enhancement, and belongs to the technical field of image processing, and the method comprises the following steps: S1, extracting multi-scale features of a super-resolution image to be scored through a multi-scale network model, and obtaining fine-grained and coarse-grained image quality features; s2, adjusting the multi-scale features to be the same as the highest-layer features in space size by using a gating pooling module; s3, performing feature enhancement on the pooled features by using an external attention mechanism; s4, utilizing a cross-scale attention mechanism to calculate an attention map of low-level features guided by high-level features among the multi-scale features after feature enhancement, and enabling the high-level semantic features to guide the network to pay more attention to a semantically important local distortion region; and S5, completing mapping between the image features and the quality scores thereof by using a multi-layer perceptron (MLP) to obtain a final quality prediction score.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image processing, and relates to a no-reference super-resolution image quality evaluation method based on semantic feature enhancement. Background Art

[0002] In today's digital age, images have become an important carrier for people's information interaction. News media tend to report current affairs intuitively by combining videos and images, and people also tend to share and obtain information through images on social platforms such as WeChat and Weibo. Image clarity not only affects people's visual experience, but also affects information transmission and expression. Therefore, the image super-resolution reconstruction technology aimed at improving image resolution and enhancing users' visual experience has become a research hotspot in the field of computer vision and image processing.

[0003] In the past three decades, a large number of image super-resolution reconstruction algorithms have been proposed, such as interpolation methods and deep learning-based methods, etc., while there are relatively few related performance evaluation studies. In practical applications, this will further affect people's extraction of information contained in images and the analysis and processing of related data. Therefore, how to efficiently and accurately evaluate the performance of image super-resolution algorithms has become an urgent problem to be solved.

[0004] Image quality evaluation has universal application value in the field of image processing. It can objectively compare the performance advantages and disadvantages of different image algorithms and provide certain guidance for algorithm improvement and optimization. Researchers have conducted in-depth studies on image quality evaluation algorithms and proposed a variety of image quality evaluation indicators, which can effectively evaluate the distortion degree of images. According to different evaluation means, image quality evaluation can be divided into two categories: subjective quality evaluation and objective quality evaluation. Subjective evaluation methods refer to the quality evaluation of images by testers through some devices. The evaluation results are reliable and there are no technical obstacles. However, subjective evaluation methods require multiple repeated experiments, and have many disadvantages such as time-consuming, costly, and difficult to achieve real-time quality evaluation. Therefore, a rich variety of evaluation models have been developed for image objective quality evaluation, which can be roughly divided into three categories of image quality evaluation models: full-reference, partial-reference, and no-reference, according to different reference information.

[0005] Numerous researchers have made important contributions to the evaluation of image quality, proposing a large number of image quality evaluation methods. However, most of the existing image quality evaluation algorithms are aimed at general natural image distortions and are not applicable to super-resolution images. This is because the distortion of super-resolution images is different from the degradation of natural images. It is the loss of image details introduced by the super-resolution reconstruction algorithm, namely blurred edges and incompatible textures, such as jagged distortions, ringing artifacts, checkerboard artifacts, and artificial texture distortions. In addition, considering that it is difficult to obtain reference information in actual super-resolution problems, it is of greater practical significance to develop a no-reference type quality assessment (NR-SRIQA) method for super-resolution images.

[0006] In recent years, the rapid development of generative adversarial networks has greatly promoted the progress of super-resolution reconstruction technology, making the generated images more excellent in visual perception. However, with the wide application of super-resolution algorithms based on generative adversarial networks in various fields, it has been found that the previously proposed NR-SRIQA methods have certain limitations in evaluating the quality of such images. This limitation stems from the inconsistency between the human visual system and machine vision. Most users expect rich and realistic details, while machines focus on distinguishing the misalignment between degraded images and original-quality images. Specifically, the restored images usually have unrealistic and lifelike textures, meeting the perception of the human eye but deviating from the prior knowledge used by deep learning models for image quality evaluation. To solve the above problems, existing methods usually use a combination of global and local representations (i.e., multi-scale features) to evaluate the image quality at different granularities to achieve better performance. However, most of the above methods adopt simple linear fusion of multi-scale features, ignoring the possible complex relationships and interactions between multi-scale features, thus limiting the accuracy and reliability of the evaluation results. Summary of the Invention

[0007] In view of this, the purpose of the present invention is to provide a no-reference super-resolution image quality evaluation method based on semantic feature enhancement.

[0008] To achieve the above purpose, the present invention provides the following technical solutions:

[0009] A no-reference super-resolution image quality evaluation method based on semantic feature enhancement, comprising the following steps:

[0010] S1: Use a multi-scale network model to extract multi-scale features of the super-resolution image to be scored, and obtain fine-grained and coarse-grained image quality features;

[0011] S2: Use a gated pooling module to adjust the multi-scale features to the same spatial size as the highest-level features;

[0012] S3: Use an external attention mechanism to enhance the features after pooling; the external attention mechanism uses two external and learnable shared memory units to capture the potential relationships of the entire dataset, thereby enhancing the generalization ability of the model;

[0013] S4: Use a cross-scale attention mechanism to calculate the attention map of the low-level features guided by the high-level features among the multi-scale features after feature enhancement, enabling the high-level semantic features to guide the network to pay more attention to the semantically important local distortion regions;

[0014] S5: Use a multi-layer perceptron (MLP) to complete the mapping between the image features and their quality scores to obtain the final quality prediction score.

[0015] Furthermore, in step S1, a swin transformer is used as the multi-scale network model to extract the multi-scale features of the super-resolution image to be scored. The specific steps are as follows:

[0016] First, the input image is divided into small image patches and an initial feature map is obtained through a linear embedding layer; in the first stage, the window multi-head self-attention mechanism and the multi-layer perceptron are used for processing to extract the features of the input image; starting from the second stage, downsampling is performed, adjacent patches are merged to halve the resolution of the feature map and double the number of channels, and then the window self-attention mechanism, the shifted window multi-head self-attention mechanism, and the multi-layer perceptron are continued to be used to extract features on the new feature map; as the stage progresses, the resolution of the feature map continuously decreases and the number of channels continuously increases, thereby extracting the features of different scales of the input image.

[0017] Furthermore, the gated pooling module is composed of gated convolution and global average pooling;

[0018] The multi-scale features of the input image are first processed by gated convolution, while retaining the spatial hierarchical structure of the features, and integrating the features of different scales; the gated convolution first uses a one-dimensional convolutional layer to generate features of two branches; one branch uses a non-linear activation function GELU to generate a feature map with the same spatial size as the input features, and the other branch uses three convolutional layers to generate a single-channel gated weight, and uses the sigmid activation function to limit the weight value between 0 and 1; finally, the activated feature map is multiplied by the gated weight to obtain the final feature output;

[0019] Then the features processed by gated convolution are fed into the global average pooling module, thereby adjusting the multi-scale features to the same spatial size as the highest-level features.

[0020] Furthermore, the external attention mechanism implicitly captures the feature correlations in the entire dataset using two linear layers and two normalization layers, enhancing the generalization ability of the model. The specific steps are as follows:

[0021] First, the input features are mapped to an externally learnable key memory through a linear layer, and the affinity between the input features and the key memory is calculated to obtain an attention map.

[0022] Then, this attention map is multiplied by another externally learnable value memory through a linear layer to generate the final feature representation.

[0023] These two memories share parameters across the entire dataset, are independent of individual samples, and play a role in regularization.

[0024] Furthermore, in step S4, the cross-attention mechanism uses high-level semantic features as clustering centers, high-level features as queries, and low-level features to form (key, value) pairs, enabling high-level semantic features to guide the network to pay more attention to semantically important local distortion regions; the attention weights generated between adjacent scales represent the similarity between low-level and high-level features; the attention map aggregates the "value" features through weighted summation; the calculation process of the cross-scale attention mechanism is expressed as:

[0025]

[0026] where i = 1, 2, 3, n, representing each stage of the multi-scale network model, F i represents the multi-scale features after pooling and feature enhancement, represents the final output obtained by gradually applying the cross-scale attention mechanism; The features in are calculated using the attention map A from to F i ; the attention weights in A represent and F i the similarity between.

[0027] Furthermore, the multi-layer perceptron is established by three fully connected layers to map the connected visual feature vectors to the perceptual quality score. They perform a linear transformation on the input features through weights and biases, and then usually introduce non-linearity through a non-linear activation function, enabling the network to learn complex function mappings and finally output a predicted quality score of a one-dimensional unit.

[0028] The beneficial effects of the present invention are as follows: First, a multi-scale network model is used to extract multi-scale features of the input image to capture image quality features at coarse-grained and fine-grained levels, which is more helpful for quantifying complex and mixed image distortions in super-resolution images. Second, a gated pooling module is used to process the extracted multi-scale features, so as to adjust the multi-scale features to the same spatial size as the highest-level features, facilitating the reduction of the computational complexity of subsequent feature interactions. Compared with a simple average pooling operation, this pooling method combined with gated convolution has performed deep learning on the features through gated convolution before performing global average pooling, so that it can more effectively retain the spatial hierarchical structure of the features and avoid the possibility of fusing features within a local window. After the pooling operation, an external attention mechanism is used to enhance the features of the pooled features. By using two external and learnable shared memory units to capture the potential relationships of the entire dataset, the generalization ability of the model is enhanced. In addition, the present invention applies a cross-scale attention mechanism between the enhanced multi-scale features to calculate the attention map of the low-level features guided by the high-level features, so that the high-level semantic features can guide the network to pay more attention to the semantically important local distortion regions. Finally, a multi-layer perceptron is used to complete the mapping relationship between the quality score and the image features.

[0029] Other advantages, objectives, and features of the present invention will be described to some extent in the subsequent specification, and to some extent, will be obvious to those skilled in the art based on the study of the following text, or can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be realized and obtained through the following specification. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be described in detail preferably with reference to the accompanying drawings, where:

[0031] Figure 1 Schematic diagram of the no-reference super-resolution image quality evaluation model based on semantic feature enhancement of the present invention;

[0032] Figure 2 Schematic diagram of the gated pooling module of the present invention;

[0033] Figure 3 Schematic diagram of the external attention mechanism of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0034] The following describes the embodiments of the present invention through specific examples. Those skilled in the art can easily understand the other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0035] It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. Therefore, only the components related to the present invention are shown in the diagrams, rather than being drawn according to the number, shape, and size of the components in actual implementation. The type, quantity, and ratio of each component in actual implementation can be arbitrarily changed, and the layout type of its components may also be more complex.

[0036] In the following description, a large number of details are explored to provide a more thorough explanation of the embodiments of the present invention. However, it is obvious to those skilled in the art that the embodiments of the present invention can be implemented without these specific details. In other embodiments, well-known structures and devices are shown in the form of block diagrams rather than in detail to avoid making the embodiments of the present invention difficult to understand.

[0037] Embodiment 1:

[0038] The present invention provides a no-reference super-resolution image quality evaluation method based on semantic feature enhancement, which is characterized in that it includes the following steps:

[0039] S1: First, the image x to be processed is passed through a multi-scale network model, the swin transformer, to extract multi-scale features, thereby obtaining fine-grained and coarse-grained image quality features.

[0040] S2: The obtained multi-scale features are sent into a gated pooling module to adjust the multi-scale features to the same spatial size as the highest-level features;

[0041] The described gated pooling module is composed of gated convolution and global average pooling. Specifically, when processing the multi-scale features of the input image, gated convolution operation is first adopted. This operation method plays a crucial role. It can help the model not only fully retain the spatial hierarchical structure of the features but also efficiently integrate features of different scales. The gated convolution mentioned here initially uses a one-dimensional convolutional layer to generate features of two branches. Among these two branches, one branch uses the non-linear activation function GELU to generate a feature map with the same spatial size as the input features. The other branch generates the gating weights through three convolutional layers with convolutional kernel sizes of 1, 3, and 1 respectively. It should be noted that in order to limit the value of the gating weights within the range of 0 to 1, the sigmid activation function is used to achieve this goal. After completing the above steps, multiplying the activated feature map by the gating weights can obtain the final feature output.

[0042] Next, the features processed by gated convolution are passed to the global average pooling module. In this module, the main purpose is to adjust the spatial size of the multi-scale features to be the same as that of the highest-level features. This has two-fold significance. On the one hand, it can effectively reduce the computational complexity, and on the other hand, it can make the subsequent processing operations of multi-scale feature interaction more convenient. Compared with the conventional average pooling operation, this pooling method that combines enhanced gated convolution has used gated convolution to deeply learn the features before performing global average pooling. Therefore, it can more efficiently retain the spatial hierarchical structure of the features, thus avoiding the problems that may occur in the process of simple average pooling. Simple average pooling usually forcibly fuses different features within a local window, and this processing method is likely to cause the loss of some important information, while the current method that combines gated convolution can effectively avoid such problems.

[0043] Generally speaking, through this two-stage processing flow, the gated pooling module in step S2 not only realizes the unification of the spatial dimensions of multi-scale features but also enhances the feature expression ability, providing richer and more accurate feature information for subsequent image quality evaluation. This design enables the model to better understand and evaluate the distortion of the image when processing super-resolution images, thereby improving the accuracy and reliability of image quality evaluation.

[0044] S3: Use the external attention mechanism to enhance the features after pooling, thereby enhancing the generalization ability of the model. Different from the traditional self-attention mechanism that only focuses on the feature relationships within a single sample, the external attention mechanism adopted in the present invention can transcend the limitations of individual samples and implicitly capture the feature associations at the level of the entire dataset. This global feature association capture ability enables the model to better understand and adapt to the internal connections between different samples, and thus shows more excellent generalization performance when facing diverse input data. Specifically, the implementation of the external attention mechanism is relatively simple and efficient, and it can be completed by only two linear layers and two normalization layers. Compared with the self-attention mechanism, it has achieved a significant reduction in both computational complexity and memory occupancy.

[0045] S4: After obtaining the multi-scale features with enhanced features, use the high-level semantic features as the clustering centers to guide the network to pay more attention to the locally distorted regions that are semantically important, thereby allowing information to propagate between different layers and achieving the purpose of intra-layer feature interaction. Since the query feature Q in the self-attention mechanism naturally plays a guiding role when calculating the output, the cross-attention mechanism here uses the high-level features as queries and the low-level features as (key, value) pairs, so that the high-level semantic features can guide the network to pay more attention to the locally distorted regions that are semantically important. Specifically, the attention weights generated between adjacent scales represent the similarity between the low-level features and the high-level features, thereby highlighting the features with more active semantics in the high-level features. This attention map aggregates the "value" features in V through weighted summation. Therefore, we can describe the high-level semantic features as clustering centers, and the clustering centers can aggregate the low-level features with greater semantic significance.

[0046] S5: Use the multi-layer perceptron MLP to complete the mapping between the image features and their quality scores, thereby obtaining the final quality prediction score. The multi-layer perceptron consists of three fully connected layers (also called linear layers) to establish the mapping relationship from the connected visual feature vectors to the perceptual quality scores. They perform linear transformations on the input features through weights and biases, and then usually introduce non-linearity through a non-linear activation function (such as ReLU), enabling the network to learn complex function mappings and finally output a one-dimensional unit prediction quality score.

[0047] Embodiment 2:

[0048] Please refer to Figure 1 、 Figure 2 and Figure 3 , in this embodiment, a no-reference super-resolution image quality evaluation method based on semantic feature enhancement is provided, including:

[0049] Step 1: The image to be processed is subjected to feature extraction through the multi-scale network model swin transformer to obtain image quality features of different granularities. The network model extracts multi-scale features through its hierarchical structure. First, the input image is divided into small image patches and an initial feature map is obtained through a linear embedding layer. In the first stage, the window multi-head self-attention mechanism and multi-layer perceptron are used to process and extract the features of the input image. Starting from the second stage, downsampling is performed. Adjacent patches are merged to halve the resolution of the feature map and double the number of channels, and then the window self-attention mechanism, shifted window multi-head self-attention mechanism, and multi-layer perceptron are continued to be used to extract features on the new feature map. As the stage progresses, the resolution of the feature map continuously decreases and the number of channels continuously increases, thereby extracting features of different scales of the input image;

[0050] Step 2: The obtained multi-scale features are fed into the gated pooling module to adjust the multi-scale features to the same spatial size as the highest-level features. The gated pooling module is a composite structure that combines two mechanisms: gated convolution and global average pooling;

[0051] Step 2.1: The image features are first processed by gated convolution. This processing method helps the model to effectively integrate features of different scales while retaining the spatial hierarchical structure of the features. The gated convolution first uses a one-dimensional convolutional layer to generate features of two branches. These two branches respectively undertake different functions and tasks and cooperate together to complete the optimization and enhancement of the features. In the first branch, the nonlinear activation function GELU is used to activate the features. The GELU function is widely used in the field of deep learning due to its good nonlinear characteristics and the ability to smooth the features. Through the activation of GELU, the generated feature map has the same spatial size as the input features. At the same time, the second branch generates corresponding weights through three convolutional layers. In order to ensure that these weight values are within a reasonable range and can effectively regulate the expression of the features, the sigmoid activation function is particularly used to perform the final limit processing on the weights, thereby limiting the value of the gated weights within the range of 0 to 1.

[0052] Step 2.2: The features processed by gated convolution are fed into the global average pooling module to adjust the multi-scale features to the same spatial size as the highest-level features, so as to reduce the computational complexity of subsequent processing operations of multi-scale feature interaction. Compared with the conventional average pooling operation, this pooling method combined with enhanced gated convolution has performed deep learning on the features through gated convolution before performing global average pooling, so that it can more effectively retain the spatial hierarchical structure of the features. This method avoids the possible loss of important information caused by the feature fusion within the local window in simple average pooling.

[0053] Step 3: Feed the pooled features into an external attention mechanism for feature enhancement. The self-attention mechanism updates the features at each position by calculating the weighted sum of the features using the pairwise affinities of all positions, so as to capture the long-range dependencies in a single sample and plays an increasingly important role in the deep feature representation of visual tasks. However, the self-attention mechanism has quadratic complexity and ignores the potential correlations between different samples. Therefore, the present invention chooses to use an external attention mechanism to enhance the image features.

[0054] The working principle of the external attention mechanism is realized by introducing two external, small, and learnable shared memories. Specifically, first, the input features are mapped to an external learnable key memory through a linear layer, and the affinity between the input features and the key memory is calculated to obtain an attention map. Then, this attention map is multiplied by another linear layer and the external learnable value memory to generate the final feature representation. These two memories share parameters across the entire dataset and are independent of individual samples, playing a powerful regularization role and improving the generalization ability of the attention mechanism. Since the number of elements in the memory is much smaller than the number of elements in the input features, the external attention mechanism has a linear computational complexity, which gives it a significant advantage in computational efficiency. By using the external attention mechanism, the network model can implicitly capture the feature correlations in the entire dataset, thereby enhancing the generalization ability of the model;

[0055] Step 4: Apply a cross-scale attention mechanism to the multi-scale features after feature enhancement to calculate the attention map of the low-level features guided by the high-level features, so that the high-level semantic features can guide the network to pay more attention to the semantically important local distortion regions;

[0056] The cross-scale attention mechanism mentioned above refers to using the self-attention mechanism between adjacent multi-scale features. The self-attention mechanism allows each element in the input sequence to consider the information of the entire sequence during the calculation process, so as to capture the long-range dependencies within the sequence. Specifically, the self-attention mechanism generates three matrices: query (Q), key (K), and value (V) through linear transformation of the input features, and then obtains the attention weights by calculating the similarity or matching degree between the query and the key. These weights are then used to weight-sum the value matrix to obtain the final output feature representation. Since the query feature Q in the self-attention mechanism naturally plays a guiding role in calculating the output, here the cross-attention mechanism uses the high-level features as the query and the low-level features as the (key, value) pair, so that the high-level semantic features can guide the network to pay more attention to the semantically important local distortion regions. The formula is as follows:

[0057]

[0058] In the above formula, i = 1, 2, 3, ..., n, representing each stage of the multi-scale network model. F i represents the multi-scale features after pooling and feature enhancement, and represents the final output obtained by gradually applying the cross-scale attention mechanism. The features in i are calculated using the attention map A from to F i . The attention weights in A represent the similarity between and F i , thus highlighting the semantically more active features in F i . It should be noted that here represents the highest-level multi-scale features. In this process, each feature vector in generates an attention map, which is aggregated with the "value" features in the lower-level multi-scale features. Therefore, the high-level semantic features can be described as clustering centers. That is to say, by using the cross-scale attention mechanism between different features, low-level features with more semantics can be aggregated, enabling the high-level semantic features to guide the network to pay more attention to the semantically important local distortion regions;

[0059] Step 5: Use a multi-layer perceptron composed of three fully connected layers to establish a mapping relationship from the connected visual feature vectors to the perceived quality scores. They perform a linear transformation on the input features through weights and biases, and then usually introduce non-linearity through a non-linear activation function (such as ReLU), enabling the network to learn complex function mappings and finally output a one-dimensional unit prediction quality score.

[0060] In the above embodiments, the reference to "this embodiment" in the specification means that the specific features, structures, or characteristics described in connection with the embodiments are included in at least some embodiments, but not necessarily all embodiments. Multiple occurrences of "this embodiment" do not necessarily all refer to the same embodiment.

[0061] In the above embodiments, although the present invention has been described in connection with specific embodiments of the present invention, many substitutions, modifications, and variations of these embodiments will be apparent to those of ordinary skill in the art based on the foregoing description. For example, other storage structures (e.g., dynamic RAM (DRAM)) can be used in the embodiments discussed. The embodiments of the present invention are intended to cover all such substitutions, modifications, and variations that fall within the broad scope of the appended claims.

[0062] This embodiment also provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements any one of the methods in this embodiment.

[0063] ​​​​This embodiment also provides an electronic terminal, including: a processor and a memory;

[0064] The memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory, so that the terminal executes any one of the methods in this embodiment.

[0065] For the computer-readable storage medium in this embodiment, those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to the computer program. The foregoing computer program can be stored in a computer-readable storage medium. When this program is executed, it executes the steps including the above method embodiments; and the foregoing storage medium includes: various media such as ROM, RAM, magnetic disk or optical disc that can store program codes.

[0066] The electronic terminal provided in this embodiment includes a processor, a memory, a transceiver, and a communication interface. The memory and the communication interface are connected to the processor and the transceiver and complete communication with each other. The memory is used to store a computer program, and the communication interface is used for communication. The processor and the transceiver are used to run the computer program to make the electronic terminal execute each step of the above method.

[0067] In this embodiment, the memory may include a random access memory (Random Access Memory, abbreviated as RAM), and may also include a non-volatile memory, such as at least one disk memory.

[0068] The above-mentioned processor may be a general-purpose processor, including a central processing unit (Central Processing Unit, abbreviated as CPU), a network processor (Network Processor, abbreviated as NP), etc.; it may also be a digital signal processor (Digital Signal Processing, abbreviated as DSP), an application specific integrated circuit (Application SpecificIntegrated Circuit, abbreviated as ASIC), a field programmable gate array (Field-Programmable Gate Array, abbreviated as FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0069] The present invention can be used in many general-purpose or special-purpose computing system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on.

[0070] The present invention may be described in the general context of computer-executable instructions, such as program modules, executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform particular tasks or implement particular abstract data types. The present invention may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communications network. In a distributed computing environment, program modules may be located in both local and remote computer storage media including storage devices.

[0071] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention may be modified or equivalently replaced without departing from the spirit and scope of the technical solutions, and they should all be covered by the scope of the claims of the present invention.

Claims

1. A reference-free super-resolution image quality assessment method based on semantic feature enhancement, characterized in that: The following steps are involved: S1: Use a multi-scale network model to extract multi-scale features of the super-resolution image to be scored, and obtain fine-grained and coarse-grained image quality features; S2: Use the gated pooling module to adjust the multi-scale features to the same spatial size as the highest-level features; S3: Enhance the pooled features using an external attention mechanism that uses two external, learnable shared memory units to capture the potential relationships of the entire dataset, thereby enhancing the generalization ability of the model. S4: Use the cross-scale attention mechanism to calculate the attention map of low-level features guided by high-level features between multi-scale features after feature enhancement, so that high-level semantic features can guide the network to pay more attention to semantically important local distortion areas; S5: Use the multi-layer perceptron MLP to complete the mapping between image features and their quality scores to obtain the final quality prediction score.

2. The method for non-reference super-resolution image quality assessment based on semantic feature enhancement according to claim 1, characterized in that: In step S1, the swin transformer is used as a multi-scale network model to extract the multi-scale features of the super-resolution image to be scored. The specific steps are as follows: First, the input image is divided into small image patches and the initial feature map is obtained through a linear embedding layer. In the first stage, the features of the input image are extracted using the windowed multi-head self-attention mechanism and multi-layer perceptron processing. Downsampling is performed from the second stage, and adjacent patches are merged to halve the feature map resolution and double the number of channels. Then, the window self-attention mechanism, shifted window multi-head self-attention mechanism, and multi-layer perceptron are used to extract features on the new feature map. As the stages progress, the resolution of the feature map decreases and the number of channels increases, thereby extracting features of different scales of the input image.

3. The method for non-reference super-resolution image quality assessment based on semantic feature enhancement according to claim 1, characterized in that: The gated pooling module is composed of gated convolution and global average pooling; The multi-scale features of the input image are first processed by gated convolution, which integrates features of different scales while preserving the spatial hierarchy of the features. The gated convolution first generates two branch features using a one-dimensional convolution layer. One branch uses the nonlinear activation function GELU to generate a feature map of the same size as the input feature space. The other branch uses three convolutional layers to generate a single-channel gating weight and uses the sigmid activation function to limit the weight value between 0 and 1. Finally, the activated feature map is multiplied by the gating weight to obtain the final feature output. The features processed by gated convolution are then fed into the global average pooling module to adjust the multi-scale features to the same spatial size as the highest-level features.

4. The method for non-reference super-resolution image quality assessment based on semantic feature enhancement according to claim 1, characterized in that: The external attention mechanism uses two linear layers and two normalization layers to implicitly capture feature associations in the entire dataset and enhance the generalization ability of the model. The specific steps are as follows: First, the input features are mapped to an external learnable key memory through a linear layer, and the affinity between the input features and the key memory is calculated to obtain an attention map; This attention map is then multiplied by an external learnable value memory through another linear layer to generate the final feature representation; These two memories share parameters across the entire dataset, independent of individual samples, and act as a regularizer.

5. The method for non-reference super-resolution image quality assessment based on semantic feature enhancement according to claim 1, characterized in that: In step S4, the cross-attention mechanism uses high-level semantic features as cluster centers, high-level features as queries, and low-level features as (key, value) pairs, so that high-level semantic features guide the network to pay more attention to semantically important local distortion areas; the attention weights generated between adjacent scales represent the similarity between low-level features and high-level features; the attention map aggregates the "value" features together through weighted sum; the calculation process of the cross-scale attention mechanism is expressed as: Among them, i = 1, 2, 3, n, represents each stage of the multi-scale network model, F i Represents the multi-scale features after pooling and feature enhancement, represents the final output obtained by gradually applying the cross-scale attention mechanism; The feature is to use to F i The attention weight in A represents and F i The similarities between.

6. The method for non-reference super-resolution image quality assessment based on semantic feature enhancement according to claim 1, characterized in that: The multilayer perceptron is composed of three fully connected layers to establish a mapping relationship from the connected visual feature vector to the perceptual quality score. They linearly transform the input features through weights and biases, and then usually introduce nonlinearity through a nonlinear activation function, so that the network can learn complex function mappings and finally output a one-dimensional unit of predicted quality score.