No-reference image quality evaluation method and device and computer equipment

Through discrete wavelet transformation and frequency-scale dual-path feature coding, combined with graph attention network and Transformer, the problem of insufficient global information modeling of traditional methods is solved, and efficient and accurate evaluation of complex image distortions is achieved.

CN120374624AActive Publication Date: 2025-07-25BEIJING ZHISHENG VISION TECH CO LTD
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202510866787.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-26
Publication Date
2025-07-25
Estimated Expiration
2045-06-26

AI Technical Summary

Technical Problem

The traditional reference-free image quality evaluation method has limited global information modeling capabilities and is difficult to effectively capture complex distortion patterns of images. The local receptive field limitation of deep learning methods leads to insufficient global distortion capture.

Method used

Discrete wavelet transformation is used to perform multi-scale layer decomposition, heterogeneous graphs are constructed through frequency-scale dual-path feature encoding, and global feature extraction is performed in combination with graph attention network and Transformer, and mass scores are generated using linear regression network.

Benefits of technology

The perceived accuracy and robustness of complex image distortion patterns are improved, and efficient and accurate reference-free image quality evaluation is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120374624A_ABST
    Figure CN120374624A_ABST
Patent Text Reader

Abstract

The invention is suitable for the technical field of image quality evaluation, and provides a no-reference image quality evaluation method and device and computer equipment, and the method comprises the steps: carrying out the multi-scale layer decomposition of a target image through discrete wavelet transform, and extracting statistical features for each sub-band; converting the multi-scale and multi-band sub-band statistical features into structured coding frequency features and coding scale features through scale frequency dual-channel feature coding; defining a frequency node and a scale node, connecting the frequency node and the scale node to form a heterogeneous graph, and fusing according to the attention weight of each node to form a graph embedding vector; mapping the image embedding vector to a Transform input dimension, and generating a global feature vector representing the overall quality of the image based on an output result of the Transform; and the global feature vector is mapped into an image quality score by using a linear regression network, so that the capturing and modeling capabilities of global information are effectively improved, and the recognition capability of global structural distortion is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of image quality assessment, and particularly relates to a no-reference image quality assessment method, apparatus, and computer device. Background Art

[0002] Image Quality Assessment (IQA) is an important research direction in the fields of computer vision and image processing. In traditional image quality assessment methods, no-reference image quality assessment (NR-IQA) has attracted wide attention due to its characteristic of not relying on the original distortion-free image and showing higher practical value.

[0003] Traditional NR-IQA methods mainly rely on manually designed statistical features and perceptual models, such as models based on Natural Scene Statistics (NSS). These methods analyze the statistical distribution of images in the spatial or frequency domain, construct quality value evaluation indicators (such as mean, standard deviation, entropy, energy, etc.), and aim to evaluate image quality by detecting the deviation of statistical characteristics caused by distortions such as compression, noise, and blur. However, traditional methods have obvious limitations: First, traditional methods rely on predefined manual features and are difficult to fully and effectively characterize the complex distortion patterns presented by images at multiple scales and frequency bands. For example, simple statistics (such as mean, variance, etc.) often cannot comprehensively capture the subtle changes in image details and textures. Second, many methods overly focus on local statistical characteristics and lack the ability to model the global structure and cross-scale information of images, resulting in a large deviation between the evaluation results and human eye subjective perception when processing high-resolution images or complex scenes. Finally, models based on manual features usually have good adaptability to specific distortion types, but have poor generalization ability when faced with the superposition of multiple distortion types or actual complex application scenarios, and often require feature engineering adjustments for different scenarios, lacking a unified and robust evaluation framework.

[0004] To overcome the above deficiencies, deep learning methods, especially Convolutional Neural Networks (CNNs), have been introduced for automatic feature learning and have made progress in image processing tasks. However, existing CNN-based methods still have insufficient ability to model the global distortion structure in image quality assessment due to their inherent local receptive field limitations. Summary of the Invention

[0005] Embodiments of this application provide a no-reference image quality assessment method, apparatus, and computer device, which can effectively improve the ability to capture and model global information and enhance the ability to identify global structural distortions.

[0006] In a first aspect, an embodiment of the present application provides a no-reference image quality assessment method, including: Perform multi-scale layer decomposition on the target image through discrete wavelet transform. Each scale layer includes a low-frequency sub-band and high-frequency sub-bands. Statistical features are extracted for each sub-band. Through scale-frequency dual-channel feature encoding, convert the statistical features of each sub-band into structured encoded frequency features and encoded scale features; define frequency nodes and scale nodes according to the above encoded frequency features and the above encoded scale features, connect the above frequency nodes and the above scale nodes to form a heterogeneous graph, and use a graph attention network to calculate the attention weights of each node and fuse them to form a graph embedding vector. Map the above graph embedding vector to the input dimension of the Transformer to obtain a global feature vector representing the overall quality of the image output by the Transformer; the positional encoding of the above Transformer fuses frequency information and scale information, and the self-attention mechanism fuses the scale weight matrix. Use a linear regression network to map the above global feature vector to an image quality score.

[0007] Exemplarily, each scale layer includes one low-frequency sub-band and three high-frequency sub-bands; the above low-frequency sub-band is the LL sub-band, which contains the overall structure information of the image; the above high-frequency sub-bands include the LH sub-band, the HL sub-band, and the HH sub-band. The above LH sub-band reflects the edge features in the horizontal direction of the image, the above HL sub-band reflects the edge features in the vertical direction of the image, and the above HH sub-band contains the texture and detail information in the diagonal direction of the image.

[0008] Exemplarily, extracting statistical features for each sub-band includes: For each sub-band, extract four statistical features: mean μ, standard deviation σ, energy E, and entropy value H.

[0009] Exemplarily, converting the statistical features of each sub-band into structured encoded frequency features and encoded scale features includes: Construct a frequency feature vector for the sub-band based on all the statistical features of each sub-band, stack the frequency feature vectors of all sub-bands to obtain a frequency feature sequence, and encode the above frequency feature sequence to obtain encoded frequency features. Perform global average pooling on the low-frequency sub-band and high-frequency sub-bands of each scale layer respectively to obtain low-frequency features and high-frequency features, splice the above low-frequency features and the above high-frequency features to obtain the scale features of each scale layer, and encode the scale features of each scale layer to obtain encoded scale features.

[0010] Exemplarily, defining frequency nodes and scale nodes according to encoded frequency features and encoded scale features includes: Define the encoded frequency features of each high-frequency subband as frequency nodes; Define the encoded scale features of each scale layer as scale nodes.

[0011] Exemplarily, connecting the above frequency nodes and the above scale nodes forms a heterogeneous graph, including: Connect the scale nodes of adjacent scale layers to form scale-scale edges; Connect the frequency nodes and scale nodes within the same scale layer to form frequency-scale edges; There is no explicit connection between different frequency nodes or a weight less than the threshold is assigned.

[0012] Exemplarily, mapping the above graph embedding vector to the input dimension of the Transformer, including: Divide the above graph embedding vector into multiple small patches, and convert each patch into a node embedding vector with a fixed dimension through linear mapping.

[0013] Exemplarily, the positional encoding of the Transformer fuses frequency information and scale information, including: Generate positional encoding according to the spatial position, frequency information, and scale information of each patch, and add the positional encoding to the above node embedding vector to obtain the fused input vector; The above self-attention mechanism fuses the scale weight matrix, including: Introduce a scale matrix into the self-attention calculation formula. The above scale matrix is calculated according to the scale information of each patch, and the attention weight is positively correlated with the similarity of the scale or the similarity of the frequency band.

[0014] In a second aspect, an embodiment of the present application provides a no-reference image quality evaluation device, including: A wavelet transform feature extraction module, configured to perform multi-scale layer decomposition on a target image through discrete wavelet transform. Each scale layer includes a low-frequency subband and a high-frequency subband, and statistical features are extracted for each subband; A frequency-scale encoding and graph structure fusion module, configured to convert the statistical features of each subband into structured encoded frequency features and encoded scale features through scale-frequency dual-path feature encoding; define frequency nodes and scale nodes according to the above encoded frequency features and the above encoded scale features, connect the above frequency nodes and the above scale nodes to form a heterogeneous graph, and use a graph attention network to calculate the attention weights of each node and fuse them to form a graph embedding vector; The Transformer global modeling module is used to map the above graph embedding vectors to the input dimension of the Transformer, and obtain a global feature vector representing the overall quality of the image output by the Transformer; the position encoding of the above Transformer fuses frequency information and scale information, and the self-attention mechanism fuses the scale weight matrix; The quality evaluation and regression module is used to map the above global feature vector into an image quality score by using a linear regression network.

[0015] In a third aspect, an embodiment of the present application provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the method described in any one of the above first aspects is implemented.

[0016] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the method described in any one of the above first aspects can be implemented.

[0017] In a fifth aspect, an embodiment of the present application provides a computer program product. When the computer program product runs on a computer device, the computer device is enabled to execute the method described in any one of the above first aspects.

[0018] The beneficial effects of the embodiments of the present application compared with the prior art are: The present application provides a no-reference image quality evaluation method based on discrete wavelet transform, frequency-scale dual-channel coding and graph structure fusion, which is mainly used for automatic detection and evaluation of image quality. This method overcomes the problems of insufficient local feature expression, limited global information modeling ability and limited generalization ability in traditional statistical characteristic-based and manually designed perceptual models, and at the same time solves the problem of insufficient capture of global distortion caused by the local receptive field limitation of the CNN in deep learning methods. By utilizing the frequency decomposition ability of discrete wavelet transform, the global modeling ability of the Transformer, and the perception ability of the graph attention network for the non-linear interaction between frequency and scale, a frequency-scale dual-channel structure is constructed, and a graph structure attention fusion mechanism is introduced to perform in-depth joint modeling on the statistical features of the image at different frequency bands and different scales, effectively improving the perception accuracy and robustness of complex distortion patterns of the image, so as to achieve efficient and accurate no-reference image quality evaluation.

[0019] It can be understood that the beneficial effects of the above second to fifth aspects can refer to the relevant descriptions in the above first aspect, and will not be elaborated here. Description of the Drawings

[0020] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0021] Figure 1 is a schematic flowchart of a no-reference image quality assessment method provided by an embodiment of the present application; Figure 2 is a schematic flowchart of step 102 provided by an embodiment of the present application; Figure 3 is a schematic flowchart of generating a global feature vector based on a Transformer provided by an embodiment of the present application; Figure 4 is a schematic diagram of a Transformer architecture provided by an embodiment of the present application; Figure 5 is a structural block diagram of a no-reference image quality assessment device provided by an embodiment of the present application; Figure 6 is a schematic diagram of the structure of a computer device provided by an embodiment of the present application. Detailed implementation manners

[0022] In the following description, specific details such as specific system architectures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.

[0023] It should be understood that when used in the specification and the appended claims of the present application, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0024] It should also be understood that the term "and / or" used in the specification and the appended claims of the present application refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0025] As used in the specification and the appended claims of this application, the term "if" may be construed contextually as "when" or "once" or "in response to determining" or "in response to detecting". Similarly, the phrases "if determined" or "if [the described condition or event] is detected" may be construed contextually to mean "once determined" or "in response to determining" or "once [the described condition or event] is detected" or "in response to detecting [the described condition or event]".

[0026] In addition, in the description of the specification and the appended claims of this application, the terms "first", "second", "third", etc. are only used for differential description and should not be construed as indicating or implying relative importance.

[0027] Reference to "one embodiment" or "some embodiments" or the like described in the specification of this application means that a particular feature, structure, or characteristic described in connection with that embodiment is included in one or more embodiments of this application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "comprising", "including", "having", and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.

[0028] The technical solutions in the embodiments of this application will be described in detail below.

[0029] Figure 1 is a schematic flowchart of a no-reference image quality assessment method provided by an embodiment of this application. By way of example and not limitation, this method can be applied to a computer device. As Figure 1 shown, this method includes: Step 101, perform multi-scale layer decomposition on the target image through discrete wavelet transform, and extract statistical features for each sub-band.

[0030] In this step, discrete wavelet transform is used to perform multi-scale multi-band decomposition on the target image, so as to capture the feature information of the target image at different scales and frequencies. Among them, the statistical features include statistical features such as mean, standard deviation, energy, and entropy, which are used to characterize the content distribution of the image at different frequency levels.

[0031] In one embodiment, performing multi-scale layer decomposition on the target image through discrete wavelet transform includes: performing multi-scale layer multi-band decomposition on the target image through discrete wavelet transform, and each scale layer includes a low-frequency sub-band and a high-frequency sub-band, that is, sub-bands of different frequencies. Exemplarily, the target image after preprocessing can be The multi-level decomposition is performed using the discrete wavelet transform according to the following formula: (1) where j is the scale parameter, representing the scale level in the multi-scale layer decomposition; k is the two-dimensional translation parameter, representing the position offset of the wavelet function in the image spatial domain; l is the direction parameter, identifying different sub-band types; represents the wavelet function after translation and scaling at scale level j; represents the corresponding wavelet coefficient.

[0032] Through the above formula, j scale levels can be decomposed, and each scale level is decomposed into four sub-bands: the LL sub-band (containing the overall structure information of the image), LH (reflecting the edge features in the horizontal direction of the image), HL (reflecting the edge features in the vertical direction of the image), and HH (containing the texture and detail information in the diagonal direction of the image). Among them, the LL sub-band is the low-frequency sub-band, and the LH, HL, and HH sub-bands are high-frequency sub-bands.

[0033] In this step, the discrete wavelet transform can be used to analyze the image both spatially (i.e., multi-scale) and in terms of frequency. The low-frequency component (LL sub-band) reflects the global structure and basic luminance information of the image, while the high-frequency components (LH, HL, HH sub-bands) contain the edge, texture, and local detail information of the image.

[0034] Exemplarily, for each sub-band, the following four statistical features can be extracted: the mean μ, the standard deviation σ, the energy E, and the entropy value H. These statistical features can quantify the local luminance distribution, contrast, texture complexity, and detail information of the image. The specific calculation formulas are as follows: Mean μ: (2) Standard deviation σ: (3) Energy E: (4) Entropy H: (5) where, is the coefficient of a certain sub-band at pixel , is the probability distribution after normalizing the coefficient histogram, and M and N are the number of rows and columns of the sub-band.

[0035] In one embodiment, before performing the discrete wavelet transform, the target image can be preprocessed, and the preprocessing includes at least one of the following: radiometric correction, geometric correction, normalization, and noise removal.

[0036] The following takes geometric correction and noise removal as examples to illustrate image preprocessing. During the image acquisition process, geometric distortion may occur due to sensor movement, platform jitter, or lens distortion. Through geometric correction, the positions of each pixel in the image can be corrected to the true spatial coordinates. To ensure that the pixels in each band or each frame of the image accurately correspond to the actual physical positions, this example uses geometric correction methods to register and resample the image.

[0037] (6) Suppose is the position of a certain pixel in the original image, and the new position is obtained after correction by the projective transformation f. Then, Gaussian blur filtering is used for noise removal processing, and the specific processing is as follows: (7) In formulas (7) and (8), i and j are input vectors; represents the pixel intensity value of the image after correction by the projective transformation at the coordinate ; is the Gaussian kernel, which is used to convolve and smooth the image. The Gaussian kernel is defined as follows: (8) In this formula, σ represents the standard deviation of the Gaussian kernel.

[0038] In this step 101, through wavelet transform, the target image is decomposed into multiple subbands with multiple scales and multiple frequency bands, and in each subband, statistical features such as mean, standard deviation, energy, and entropy are extracted to characterize the local structure, texture change, and edge details of the image. The formed feature vector can not only retain spatial locality but also enhance the frequency representation ability, providing more refined and structure-aware input data support for the subsequent image quality evaluation step.

[0039] In step 102, through scale-frequency dual-path feature encoding, the statistical features of each subband are converted into structured encoded frequency features and encoded scale features. Based on this, frequency nodes and scale nodes are defined, an heterogeneous graph is formed by connecting the frequency nodes and scale nodes, and the attention weights of each node are calculated using a graph attention network and fused to form a graph embedding vector.

[0040] To better extract and organize the multi-scale and multi-frequency information generated by wavelet transform, this step introduces dual-path hierarchical encoding for feature vector generation. In this step, through the idea of separate modeling, the frequency information and scale information are regarded as two different levels of input signals respectively, and two independent neural paths are used for feature encoding. Finally, a graph attention network (Graph Attention Network, GAT) is used for feature fusion, so as to construct a clear-structured and semantic-rich embedding vector for downstream modules to use.

[0041] Figure 2 It is a schematic flow diagram of step 102 provided by an embodiment of the present application. As Figure 2 shown, the process includes: Step 201, scale-frequency dual-channel feature encoding.

[0042] In the present application, the image is decomposed into multiple frequency components (sub-bands) through the discrete wavelet transform of step 101, and statistical features (such as mean μ, standard deviation σ, energy E, and entropy value H) are extracted from each sub-band respectively. These statistical features describe the local brightness distribution, contrast, texture complexity, and details of the image. In step 201, through the idea of separate modeling, the frequency feature and the scale feature are regarded as two different levels of input signals respectively, and two independent neural channels are used for feature encoding.

[0043] In one embodiment, the frequency feature encoding includes: constructing a frequency feature vector (such as mean, standard deviation, energy, entropy) based on the statistical features of each sub-band and encoding. This process specifically includes the following steps: 1) Construct the frequency feature vector of each sub-band based on all the statistical features of each sub-band.

[0044] Exemplarily, for each sub-band s and each sub-layer l, calculate its mean μ, standard deviation σ, energy E, and entropy value H to obtain a feature vector containing 4 elements: (9) The calculation formulas of the mean μ, standard deviation σ, energy E, and entropy value H are specifically shown in formulas (2) to (5) above and will not be elaborated here.

[0045] 2) Stack the frequency feature vectors of all sub-bands to obtain a frequency feature sequence, and then encode the frequency feature sequence.

[0046] Exemplarily, stack the statistical vectors of all sub-bands into a frequency feature sequence (N is the total number of sub-bands), and then use a one-dimensional convolutional neural network 1D CNN for in-channel modeling to achieve encoding: (10) where Conv1D represents a one-dimensional convolutional operation.

[0047] In one embodiment, the scale feature encoding includes: performing global average pooling on the low-frequency sub-band and the high-frequency sub-band of each scale layer respectively to obtain a low-frequency feature and a high-frequency feature, concatenating the low-frequency feature and the high-frequency feature to obtain the scale feature of each scale layer, and encoding the scale feature of each scale layer.

[0048] Exemplarily, to obtain the scale information carried by the multi-layer wavelet decomposition, global average pooling is performed on the high-frequency sub-bands and low-frequency sub-bands of each layer respectively to form a richer scale structure representation vector, as shown in the following formula: (11) where represents the set of wavelet high-frequency sub-bands of the l-th layer (such as LH, HL, HH sub-bands), represents the low-frequency sub-band of the l-th layer (such as LL sub-band), and GAP represents the global average pooling operation.

[0049] Then, the low-frequency features and high-frequency features obtained by global average pooling are concatenated, and a one-dimensional convolutional neural network is used to encode the concatenated features. The concatenation result and the encoding result are shown in the following formulas respectively: (12) where, is the concatenated feature of the high-frequency feature and the low-frequency feature after global average pooling of the -th layer, is the encoding result of the concatenated features from layer 1 to layer L .

[0050] Through the feature encoding method provided by the embodiments of the present application, while compressing the feature dimension, the structural relationship between sub-bands is retained, and a structured input is provided for the construction of frequency / scale nodes of the subsequent heterogeneous graph. It should be noted that, in addition to the one-dimensional convolutional neural network, other methods such as fully connected layers, self-attention mechanisms, or recurrent neural networks can also be used to implement the encoding of frequency feature sequences or scale features.

[0051] Step 202, construct a heterogeneous graph according to the encoded frequency features and encoded scale features.

[0052] To accurately model the complementary relationship between high-frequency details and global structures, the present application designs a node connection strategy based on a heterogeneous graph. First, the encoded frequency features of the high-frequency sub-bands (such as LH, HL, HH sub-bands) of each scale layer are defined as frequency nodes, and the encoded scale features of each scale layer are defined as scale nodes. Based on this definition, frequency nodes are responsible for capturing the local edges and textures of the image, and scale nodes are responsible for describing the global structures and brightness distributions of the corresponding levels.

[0053] Based on the above-defined frequency nodes and scale nodes, during the graph structure generation process, there will be three types of edges connecting different types of nodes: 1) Frequency-frequency edge (F-F): Reflects the semantic similarity between different frequency sub-bands. The edge weights can be constructed by calculating the cosine similarity between each frequency node through the following formula: (13) Among them, f i and f j Used to represent different frequency nodes.

[0054] 2) Scale-scale edge (SS): reflects the hierarchical correlation between different scale layers, and uses a fixed hierarchical structure to connect adjacent scale layers (such as the lth layer and the l+1th layer).

[0055] 3) Frequency-scale edge (FS): connects the frequency nodes and corresponding scale nodes in the same scale layer. This edge reflects the structural relationship between the subband and its scale layer.

[0056] For example, in the process of graph structure generation, the following three types of connection principles can be followed: a) strong frequency-scale connection in the same layer, that is, each frequency node is connected to the scale node of the layer to which it belongs, so that the edge and texture features are directly associated with the semantic structure of the layer; b) weak connection across scales, that is, only weak edges are established between scale nodes of adjacent layers to reduce cross-resolution interference and maintain semantic hierarchy; c) weak processing of nodes of the same type, that is, frequency-frequency or scale-scale nodes are not explicitly connected or given very small weights (i.e., weights less than the threshold or approaching 0), highlighting the key interaction between frequency and scale. Finally, the constructed heterogeneous graph and its node features are input into the graph attention network, and the neighbor information is adaptively aggregated through edge weights to achieve bidirectional fusion of frequency and scale semantics. The output graph embedding has both local details and global structure, laying the foundation for the subsequent global modeling of Transformer.

[0057] Step 203: A graph embedding vector is formed based on the attention weights of each node.

[0058] In order to achieve a more refined fusion of frequency features and scale features, an embodiment of the present application introduces a graph attention mechanism. By constructing a heterogeneous graph of frequency nodes and scale nodes, each node establishes a relationship with at least one other node, and the graph attention mechanism can dynamically adjust the influence of each node. The specific steps are as follows: 1) Apply the graph attention mechanism to calculate the attention weights between nodes.

[0059] Use the graph attention mechanism to learn the nodes in the heterogeneous graph constructed above and dynamically calculate the attention weights between nodes. The calculation formula is as follows: (14) in, is the set of neighbor nodes of node i (i.e. other nodes connected to the node), W is the shared weight matrix, σ is the nonlinear activation function, and h j is the original feature vector of neighbor node j. is the attention weight between node i and node j, calculated through the self-attention mechanism: (15) where d is the feature dimension, is the attention vector.

[0060] Through the self-attention mechanism of GAT, the mutual relationships and importance between nodes can be modeled more precisely.

[0061] 2) Update the features of each node based on the attention weights and fuse them to obtain the graph embedding vector.

[0062] After being processed by the GAT layer, the features of each node will be updated with weights, and finally a new graph embedding vector will be obtained : (16) The above is the fused feature representation, which contains the information of frequency and scale features; are the initial features of each node.

[0063] Finally, the feature vector fused by the graph attention mechanism will be passed as input to the Transformer for global modeling. Based on this, the subsequent Transformer will be able to process high-level image representations and further improve the performance of the model.

[0064] Step 103, map the graph embedding vector to the input dimension of the Transformer to obtain a global feature vector representing the overall quality of the image output by the Transformer.

[0065] This step mainly generates a global feature vector representing the overall quality of the image based on the Transformer. The Transformer is a deep learning model architecture used for natural language processing and other sequence-to-sequence tasks. This architecture introduces the self-attention mechanism, making it perform excellently in processing sequence data. The Transformer usually includes the following important components: self-attention mechanism, position encoding, encoder, residual connection, and layer normalization.

[0066] Figure 3 is a schematic flow diagram of generating a global feature vector based on the Transformer provided by an embodiment of the present application. As Figure 3 shown, this flow includes: Step 301: Linearly map the graph embedding vectors to the input dimension of the Transformer, and perform encoding by fusing spatial position, frequency information, and scale information.

[0067] To enable the Transformer to process image data, first divide the graph embedding vectors output in Step 102 according to image patches, and convert them into node embedding vectors of a fixed dimension through linear mapping.

[0068] Let the input local feature vector be , and convert it into a node embedding vector through linear mapping. The formula is: (17) where is the mapping matrix, is the bias term, represents the number of patches.

[0069] To enable the Transformer model to distinguish the spatial position, frequency information, and scale information of different patches, a position encoding method that fuses frequency information and scale information is introduced. Specifically: for each node embedding vector of a patch, add the position encoding , where can be jointly composed of the spatial position, frequency index and the scale factor : (18) where is the standard sine-cosine position encoding; and are the encodings of frequency and scale information after linear or non-linear mapping respectively, , is the adjustment coefficient.

[0070] Then, add the position encoding to the node embedding vector to obtain the fused input vector sequence: (19) Step 302: Introduce a scale weight matrix into the standard self-attention mechanism to enhance the semantic association between nodes of the same scale.

[0071] The traditional self-attention mechanism can calculate the correlation between each node, but it cannot distinguish patches from different scales or frequency bands. In this application, an additional scale weight is introduced into the self-attention mechanism to encourage higher association between the same scales. For this purpose, an improved Transformer encoder layer is used for the input sequence Perform global modeling and utilize the self-attention mechanism to capture the long-range dependencies and cross-scale correlations between different Patches.

[0072] The traditional self-attention calculation formula is: (20) Among them, Q, K, and V are three input representation vectors. Q represents the query vector, K represents the key vector, V represents the value vector, and D is the calculation parameter.

[0073] After introducing the scale matrix (i.e., the cross-scale weight S) in this application, the attention score is updated to: (21) Among them, S can be calculated from the scale information of each Patch, ensuring that Patches from similar scales or frequency bands obtain higher attention weights, that is, the attention weight is positively correlated with the similarity of the scale or the similarity of the frequency band.

[0074] Step 303: Stack multiple layers of Transformer encoders to output a global feature sequence, and perform average pooling on this global feature sequence to generate a global feature vector that can represent the quality of the entire image.

[0075] Stack multiple layers of Transformer encoders (each layer contains cross-scale self-attention and a feed-forward network), and finally generate a fused global feature sequence. Perform average pooling on the sequence output by the encoder to obtain the global quality feature vector of the entire image : (22) Among them represents the features output by the Transformer encoder. This global feature will be used as the input for subsequent quality evaluation.

[0076] So far, the description of the Figure 3 shown process is completed. It should be noted that in addition to Figure 3 the method of using average pooling to generate the global feature vector, other methods such as max pooling and self-attention aggregation can also be used based on the global feature sequence output by the Transformer. Relatively speaking, average pooling can treat all nodes equally, with high computational efficiency, and is suitable for the case where the importance of nodes is evenly distributed; max pooling can capture the most significant distortion features and suppress background noise, and is suitable for the case where local severe distortion dominates the quality evaluation; self-attention aggregation can dynamically learn the node importance weights and focus on the key distortion regions, and is suitable for the case where differential attention is required for complex distortion distributions.

[0077] Figure 4 is a schematic diagram of a Transformer architecture provided by an embodiment of this application. AsFigure 4 As shown are multi-scale and multi-band features extracted through wavelet transform; Patch Embedding refers to dividing an image into patches and converting each patch into a vector of a fixed dimension; Positional Encoding fuses spatial, frequency, and scale information so that each node carries complete position-frequency-scale information; Cross-Scale Self-Attention uses an improved self-attention mechanism to capture the relationships between patches of different scales and different frequency bands globally; Global feature aggregation obtains the global quality feature representation F of the entire image through global pooling, providing a stable and accurate input for subsequent quality regression. Transformer global modeling fully combines the multi-scale and multi-band features extracted by wavelet transform, and realizes global information integration through the cross-scale self-attention mechanism, providing a powerful global feature modeling ability for no-reference image quality assessment. In some embodiments, the multi-scale and multi-band features extracted by wavelet transform in step 101 can also be subjected to similar processing and used as the input of the Transformer. As a relatively simplified implementation, this method can still improve the ability to capture and model global information to a certain extent compared with the prior art, and enhance the ability to identify global structural distortions.

[0078] Step 104, use a linear regression network to map the global feature vector to an image quality score.

[0079] Exemplarily, the mapping formula adopted in this step is as follows: (23) Wherein, is the predicted image quality score; is the regression weight; is the bias term. In addition to the formula (23) provided in this example, other linear ones that can realize the mapping of the image quality score can also be adopted In one embodiment, to make the predicted quality score as close as possible to the true quality score , the mean square error (MSE) can be used as the loss function. At the same time, the Adam optimization algorithm can be used to update the regression network parameters during the training process to ensure the minimization of the loss function.

[0080] So far, the completion of the Figure 1Description of the shown process. It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not imply the sequence of execution. The execution sequence of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0081] Based on Figure 1 the shown process, the present application provides a no-reference image quality assessment method based on discrete wavelet transform, frequency-scale dual-channel coding, and graph structure fusion, mainly used for automatic detection and assessment of the quality of images (including ordinary images and hyperspectral images). This method overcomes the problems of insufficient local feature expression, limited global information modeling ability, and limited generalization ability in traditional methods based on statistical characteristics and manually designed perceptual models. At the same time, it solves the problem of insufficient capture of global distortion caused by the local receptive field limitation of the CNN in deep learning methods. By leveraging the frequency decomposition ability of discrete wavelet transform, the global modeling ability of Transformer, and the perception ability of graph attention network for the non-linear interaction between frequency and scale, a frequency-scale dual-channel structure is constructed, and a graph structure attention fusion mechanism is introduced to perform in-depth joint modeling of the statistical features of the image at different frequency bands and scales, effectively improving the perception accuracy and robustness of complex distortion patterns in the image, thus achieving efficient and accurate no-reference image quality assessment.

[0082] Specifically, the main innovations of the present application are as follows: 1. Introduction of a multi-channel frequency-scale coding mechanism: The present application decouples and models the frequency features obtained after wavelet transform and the multi-scale structural features, constructs a frequency channel and a scale channel, and respectively extracts the expression information of the image at different sub-bands and different resolutions, effectively improving the fine-grainedness and structural sensitivity of feature representation.

[0083] 2. Construction of a frequency-scale heterogeneous graph structure and introduction of a graph attention fusion mechanism: The present application uses the frequency channel and the scale channel as heterogeneous graph nodes, realizes the relationship modeling and feature fusion between nodes through a graph neural network, and adaptively captures the non-linear interaction between different frequency components and scale structures, breaking through the limitations of traditional feature stitching and weighted average fusion methods.

[0084] 3. Improvement of the cross-scale modeling ability of the Transformer model: The present application designs a Patch embedding strategy that combines frequency position coding and scale-aware position coding, and introduces a cross-scale self-attention mechanism in the Transformer module, enabling the model to simultaneously focus on the context information of the image at multiple frequencies and scales, and enhancing the ability to identify global structural distortions.

[0085] 4. Establish an end-to-end reference-free image quality evaluation system: The model proposed in this application does not rely on the original distortion-free image as a reference and can directly score the quality of the input image under various distortion conditions such as compression, blur, and noise, with stronger practicability, generality, and deployment value.

[0086] 5. Universal for ordinary image and hyperspectral image quality evaluation scenarios: The designed frequency-scale joint structure has strong migration ability and can be extended to a variety of image types, including but not limited to natural images, remote sensing images, and hyperspectral images, laying a foundation for subsequent cross-modal image perception.

[0087] Corresponding to the reference-free image quality evaluation method described in the above embodiments, Figure 5 The structural block diagram of a reference-free image quality evaluation device provided by an embodiment of the present application is shown. For the convenience of description, only the parts related to the embodiments of the present application are shown.

[0088] Referring to Figure 5 , the device includes: A wavelet transform feature extraction module 501, configured to perform multi-scale layer decomposition on a target image through discrete wavelet transform. Each scale layer includes a low-frequency subband and a high-frequency subband, and statistical features are extracted for each subband; A frequency-scale encoding and graph structure fusion module 502, configured to convert the statistical features of each subband into structured encoded frequency features and encoded scale features through scale-frequency dual-channel feature encoding; define frequency nodes and scale nodes according to the encoded frequency features and encoded scale features, connect the frequency nodes and scale nodes to form a heterogeneous graph, and use a graph attention network to calculate the attention weights of each node and fuse them to form a graph embedding vector; A Transformer global modeling module 503, configured to map the graph embedding vector to the Transformer input dimension to obtain a global feature vector representing the overall quality of the image output by the Transformer; the position encoding of the Transformer fuses frequency information and scale information, and the self-attention mechanism fuses the scale weight matrix; A quality evaluation and regression module 504, configured to map the global feature vector to an image quality score using a linear regression network.

[0089] Furthermore, the device may further include an image preprocessing module (not shown in the figure), configured to perform preprocessing such as geometric correction and noise removal on the target image, and the wavelet transform feature extraction module 501 performs wavelet decomposition on the preprocessed target image.

[0090] It should be noted that for the information interaction, execution process, etc. among the above-mentioned devices / units / modules, since they are based on the same concept as the method embodiments of this application, for their specific functions and the technical effects brought about, reference can be specifically made to the method embodiment part, and details will not be elaborated here.

[0091] To facilitate the understanding of the solution provided by this application, a detailed description will be given below in conjunction with an embodiment: 1. Image preprocessing: This module preprocesses the image using methods of geometric correction and noise removal.

[0092] 1.1 Geometric correction: Assume is the position of a certain pixel in the original image, and the new position is obtained after correction through the projective transformation f:

[0093] (6) 1.2 Noise removal: Use a Gaussian kernel to perform convolution smoothing on the image. After the convolution operation, the smoothed image is obtained as the input for subsequent processing: (7) 2. Wavelet transform feature extraction: This module uses discrete wavelet transform to perform multi-layer decomposition on the preprocessed image to capture the feature information of the image at different scales and frequency bands. In this embodiment, 3-layer wavelet decomposition is adopted, and 4 sub-bands are generated for each layer: LL, LH, HL, HH. Subsequently, statistical features are calculated for each sub-band respectively, and are concatenated into a final feature vector in a fixed order.

[0094] 2.1 Wavelet decomposition: Use a wavelet basis to perform 3-layer decomposition on the image. The decomposition formula can be expressed as: (1) where is the translated and scaled wavelet function at scale j, is the corresponding wavelet coefficient. After decomposition, the following are obtained: The first layer:

[0095] The second layer: Perform sequential decomposition on to obtain

[0096] The third layer: Perform sequential decomposition on to obtain

[0097] 2.2 Sub-band feature extraction: For each sub-band , the following statistical features are extracted: Mean: (2) Standard Deviation: (3) Energy: (4) Entropy: (5) 3. Frequency-Scale Dual-Pathway Feature Encoding and Graph Structure Fusion: In this module, the frequency-domain features and scale information are regarded as input signals at two different levels, and two independent neural pathways are used for feature encoding. Then, a graph attention network is used for feature fusion, so as to construct an embedding vector with clear structure and rich semantics for the downstream module to use.

[0098] 3.1 Frequency-Domain Feature Extraction and Encoding Calculate frequency-domain features: For each sub-band s and each sub-layer l, calculate its mean, standard deviation, energy, and entropy to obtain a feature vector containing 4 elements: (9) Stack frequency-domain features: Stack the statistical vectors of all sub-bands into a frequency-domain feature sequence (N is the total number of sub-bands), and then use a one-dimensional convolutional neural network (1D CNN) for in-channel modeling: (10) 3.2 Scale Feature Extraction and Encoding Perform global average pooling on the high-frequency sub-bands and low-frequency sub-bands of each layer respectively to form a more abundant scale structure representation vector: (11) Then perform the operations of concatenation and encoding on the features after global pooling of the high-frequency features: (12) 3.3 Sub-Band Semantics-Guided Heterogeneous Graph Construction Heterogeneous graph node design: The frequency nodes come from the high-frequency sub-band feature vectors in each row of F freq and are responsible for capturing the local edges and textures of the image; the scale nodes come from the scale feature vectors in each row of F scale and are responsible for describing the global structure and brightness distribution of the corresponding level.

[0099] The edge connection rules are designed as follows: 1) Strong frequency-scale connections in the same layer, that is, each frequency node is connected to the scale node of the layer to which it belongs, so that the edge and texture features are directly associated with the semantic structure of the layer; 2) Weak connections across scales, that is, only weak edges are established between scale nodes in adjacent layers to reduce cross-resolution interference and maintain semantic hierarchy; 3) Weak processing of nodes of the same type, that is, frequency-frequency or scale-scale nodes are not explicitly connected or given extremely small weights, highlighting the key interaction between frequency and scale.

[0100] 3.4 Graph Attention Mechanism Feature Fusion: Apply graph attention mechanism: Use graph attention mechanism to learn the nodes in the graph. The attention weights between nodes are dynamically calculated by GAT. The formula is as follows: (14) in, is the attention weight between node i and node j, which is calculated by the self-attention mechanism: (15) Through the self-attention mechanism of GAT, the relationships and importance between nodes can be modeled more accurately.

[0101] Feature fusion: After being processed by the GAT layer, the features of each node will be weighted and updated, and finally a new graph embedding vector will be obtained. : (16) Above It is the fused feature representation, which contains information about frequency and scale features.

[0102] The constructed heterogeneous graph and its node features are input into the Graph Attention Network (GAT), and neighbor information is adaptively aggregated through edge weights to achieve a bidirectional fusion of frequency and scale semantics. The output graph embedding has both local details and global structure, laying the foundation for the subsequent Transformer global modeling.

[0103] 4. Transformer global modeling module: This module embeds the graph into a vector It is converted into a vector sequence of fixed dimension and globally modeled through specially designed frequency position encoding and cross-scale self-attention mechanism, finally generating a global feature vector describing the quality of the entire image.

[0104] 4.1PatchEmbedding: Embedding Graphs into Vectors Divided into multiple local patches, each patch has a Transformed into an embedding vector after linear mapping : (17) 4.2 Frequency position encoding: To preserve the spatial position information of each Patch and the inherent frequency and scale information in wavelet decomposition, position encoding is added to each embedding vector: (18) Finally, the fused input sequence is obtained: (19) 4.3 Cross-scale self-attention mechanism: The improved Transformer encoder is used to globally model the input sequence. The self-attention calculation formula with the scale weight matrix S added is: (21) 4.4 Transformer encoder and global feature output: After passing through multiple layers of Transformer encoders, average pooling is performed on the outputs of all Patches to obtain the global feature vector : (22) 5. Quality evaluation and regression module: This module uses a simple linear regression network to map the global feature vector output by the Transformer to an image quality score. Specifically, a single-layer linear mapping is used to convert F into a quality score : (23) Loss function: During the training process, the mean squared error (MSE) is used as the loss function, where B is the number of samples: (24) So far, the description of this embodiment is completed. As a simplified implementation method, the output of step 2, that is, the multi-scale and multi-band features extracted by wavelet transform , after being similarly processed, can be used as the input of the Transformer (that is, converting the multi-scale and multi-band features extracted by wavelet transform into a vector sequence with a fixed dimension), so as to simplify the implementation steps, improve the calculation speed, and at the same time, compared with the prior art, still improve the ability to capture and model global information to a certain extent and enhance the ability to identify global structural distortions.

[0105] Those skilled in the art can clearly understand that, for the convenience and simplicity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be assigned to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of this application. The specific working process of the units and modules in the above system can refer to the corresponding process in the foregoing method embodiment and will not be elaborated herein.

[0106] An embodiment of this application also provides a computer device, which includes: at least one processor, a memory, and a computer program stored in the memory and executable on the at least one processor. When the processor executes the computer program, it implements the steps in any of the foregoing method embodiments.

[0107] An embodiment of this application also provides a computer-readable storage medium storing a computer program, which when executed by a processor can implement the steps in each of the foregoing method embodiments.

[0108] An embodiment of this application provides a computer program product, which when running on a computer device enables the computer device to execute the steps in each of the foregoing method embodiments.

[0109] Figure 6 It is a schematic structural diagram of a computer device provided by an embodiment of this application. As Figure 6 shown, the computer device of this embodiment includes: at least one processor 60 ( Figure 6 only one is shown in the figure), a memory 61, and a computer program 62 stored in the memory 61 and executable on the at least one processor 60. When the processor 60 executes the computer program 62, it implements the steps in any of the foregoing visual programming method embodiments.

[0110] The computer device may include, but is not limited to, a processor 60 and a memory 61. Those skilled in the art can understand, Figure 6The above are only examples of computer devices and do not limit computer devices. There may be more or fewer components than shown in the figure, or some components may be combined, or different components. For example, input / output devices, network access devices, etc. may also be included.

[0111] The so-called processor 60 may be a central processing unit (CPU), and the processor 60 may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.

[0112] In some embodiments, the memory 61 may be an internal storage unit of the computer device, such as the hard disk or memory of the computer device. In other embodiments, the memory 61 may also be an external storage device of the computer device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the computer device. Further, the memory 61 may also include both the internal storage unit and the external storage device of the computer device. The memory 61 is used to store an operating system, application programs, a boot loader, data, and other programs, such as the program code of the computer program. The memory 61 may also be used to temporarily store data that has been output or will be output.

[0113] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-described embodiment methods of this application, a computer program can be used to instruct the relevant hardware to complete. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-described method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can at least include: any entity or device that can carry the computer program code to the device / computer equipment, recording medium, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk, or an optical disc, etc. In some jurisdictions, according to legislation and patent practice, the computer-readable medium cannot be an electrical carrier signal and a telecommunication signal.

[0114] In the above embodiments, the descriptions of the various embodiments have their own focuses. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0115] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0116] In the embodiments provided in this application, it should be understood that the disclosed device / computer equipment and method can be implemented in other ways. For example, the device / computer equipment embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical, mechanical, or other form.

[0117] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or may be distributed over multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0118] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.

Claims

1. A no-reference image quality assessment method, characterized in that Including: Perform multi-scale layer decomposition on the target image through discrete wavelet transform. Each scale layer includes a low-frequency subband and high-frequency subbands, and extract statistical features for each subband; Through scale-frequency dual-pathway feature encoding, convert the statistical features of each subband into structured encoded frequency features and encoded scale features; define frequency nodes and scale nodes according to the encoded frequency features and the encoded scale features, connect the frequency nodes and the scale nodes to form a heterogeneous graph, and use a graph attention network to calculate the attention weights of each node and fuse them to form a graph embedding vector; Map the graph embedding vector to the input dimension of the Transformer to obtain a global feature vector representing the overall quality of the image output by the Transformer; the positional encoding of the Transformer fuses frequency information and scale information, and the self-attention mechanism fuses a scale weight matrix; Use a linear regression network to map the global feature vector to an image quality score.

2. The method according to claim 1, characterized in that, Each scale layer includes one low-frequency subband and three high-frequency subbands; the low-frequency subband is the LL subband, which contains the overall structural information of the image; the high-frequency subbands include the LH subband, the HL subband, and the HH subband. The LH subband reflects the edge features in the horizontal direction of the image, the HL subband reflects the edge features in the vertical direction of the image, and the HH subband contains the texture and detail information in the diagonal direction of the image.

3. The method according to claim 1, wherein The extracting statistical features for each subband includes: For each subband, extract four statistical features: mean μ, standard deviation σ, energy E, and entropy value H.

4. The method according to claim 1, characterized in that The converting the statistical features of each subband into structured encoded frequency features and encoded scale features includes: Construct a frequency feature vector for each subband based on all the statistical features of the subband, stack the frequency feature vectors of all subbands to obtain a frequency feature sequence, and encode the frequency feature sequence to obtain encoded frequency features; Perform global average pooling on the low-frequency subband and high-frequency subbands of each scale layer respectively to obtain low-frequency features and high-frequency features, concatenate the low-frequency features and the high-frequency features to obtain the scale features of each scale layer, and encode the scale features of each scale layer to obtain encoded scale features.

5. The method according to claim 4, wherein The defining frequency nodes and scale nodes according to the encoded frequency features and the encoded scale features includes: Define the encoded frequency features of each high-frequency subband as frequency nodes; Define the encoded scale features of each scale layer as scale nodes.

6. The method according to claim 4, wherein The connecting the frequency nodes and the scale nodes to form a heterogeneous graph includes: Connect the scale nodes of adjacent scale layers to form scale-scale edges; Connect the frequency nodes and scale nodes within the same scale layer to form frequency-scale edges; Do not explicitly connect different frequency nodes or assign weights less than a threshold.

7. The method according to claim 1, characterized in that The mapping the graph embedding vector to the input dimension of the Transformer includes: Divide the graph embedding vector into multiple small patches Patch, and convert each Patch into a node embedding vector with a fixed dimension through a linear mapping.

8. The method according to claim 7, wherein The positional encoding of the Transformer fusing frequency information and scale information includes: Generate position encodings based on the spatial position, frequency information, and scale information of each patch, and add the position encodings to the node embedding vectors to obtain fused input vectors; The self-attention mechanism fuses the scale weight matrix, including: Adding a scale matrix to the self-attention calculation formula, where the scale matrix is calculated based on the scale information of each Patch, and the attention weights are positively correlated with the similarity of scales or the similarity of frequency bands.

9. An image quality evaluation device without a reference image, characterized in that Including: A wavelet transform feature extraction module for performing multi-scale layer decomposition on the target image through discrete wavelet transform. Each scale layer includes a low-frequency sub-band and a high-frequency sub-band, and statistical features are extracted for each sub-band; A frequency-scale encoding and graph structure fusion module for converting the statistical features of each sub-band into structured encoded frequency features and encoded scale features through scale-frequency dual-pathway feature encoding; defining frequency nodes and scale nodes according to the encoded frequency features and the encoded scale features, connecting the frequency nodes and the scale nodes to form a heterogeneous graph, and using a graph attention network to calculate the attention weights of each node and fuse them into a graph embedding vector; A Transformer global modeling module for mapping the graph embedding vector to the Transformer input dimension to obtain a global feature vector representing the overall quality of the image output by the Transformer; the position encoding of the Transformer fuses frequency information and scale information, and the self-attention mechanism fuses the scale weight matrix; A quality evaluation and regression module for mapping the global feature vector to an image quality score using a linear regression network.

10. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Non-reference image quality evaluation method based on multi-domain distortion learning

    CN116823794A

  • High-low frequency feature enhanced super-resolution image quality evaluation method and system

    CN118096750A

  • Image restoration system fusing bidirectional perception Transform and frequency analysis strategy

    CN118429228A

  • Multi-focus image fusion method based on multilayer semantics and multi-scale self-attention

    CN119251062A

  • Stereo image quality evaluation method based on dual-frequency interaction enhancement and binocular matching

    CN119399176A