Image quality evaluation method and device
By combining the Mamba model and hypernetworks, the problem of capturing global dependencies and local features in image quality assessment is solved, achieving more accurate image quality assessment and reducing the deviation between assessment results and human subjective perception.
Patent Information
- Application Number
- CN202511653054.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-12
- Publication Date
- 2026-02-06
AI Technical Summary
Existing image quality assessment methods cannot effectively capture global dependencies and local features of images, resulting in a large discrepancy between the assessment results and human subjective perception.
The Mamba model is used for semantic feature extraction. Multi-scale content features are obtained through a bidirectional window scanning strategy. The weight parameters and bias parameters of the target network are generated using a hypernetwork. The target network is then evaluated in conjunction with the image quality prediction network.
It achieves global dependency modeling, local distortion capture, and adaptive quality prediction, which improves the accuracy of image quality assessment and reduces the deviation between assessment results and human subjective perception.
Smart Images

Figure CN121481978A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing technology, and more specifically, to an image quality assessment method and an image quality assessment device. Background Technology
[0002] Image quality assessment aims to quantify the subjective perception of image quality by the human eye through computational models. Based on whether an original reference image is required, image quality assessment is divided into three categories: full-reference (requires the original image), semi-reference (requires some original features), and no-reference (does not require the original image). Among these, no-reference assessment has become a research hotspot due to its suitability for real-world scenarios (such as situations where distorted images exist independently).
[0003] Current image quality assessment uses convolutional neural networks (CNNs). However, CNNs cannot effectively capture global dependencies in images (such as the relationship between "character outline integrity" and "facial local noise"), resulting in a lack of global quality perception. At the same time, the local feature bias of CNNs prevents them from utilizing both global and local information simultaneously, ultimately leading to a significant deviation between the assessment results and human subjective perception.
[0004] Therefore, existing image quality assessment methods suffer from inaccurate assessments and significant discrepancies between the assessment results and human subjective perception. Summary of the Invention
[0005] The purpose of this invention is to provide an image quality assessment method and apparatus, which improves the problems of inaccurate image quality assessment and significant deviation between the assessment results and human subjective perception in the prior art.
[0006] In a first aspect, this application provides an image quality assessment method, comprising the following steps: Acquire the image to be evaluated; The Mamba model is used to extract semantic features from the image to be evaluated, resulting in multi-scale content features, which include local distortion detail information and global content information. Based on the multi-scale content features, a hypernetwork is used to generate the weight parameters and bias parameters of the target network; Based on the multi-scale content features, the weight parameters and bias parameters of the target network, the image quality results are obtained by using a preset image quality prediction target network.
[0007] In this embodiment of the application, the step of using the Mamba model to extract semantic features from the image to be evaluated to obtain multi-scale content features includes: The image to be evaluated is scanned using a bidirectional window scanning strategy to obtain a pixel block sequence; The Mamba model is used to extract features from the pixel block sequence to obtain multi-scale content features.
[0008] In this embodiment of the application, the step of scanning the image to be evaluated based on a bidirectional window scanning strategy to obtain a pixel block sequence includes: The non-overlapping window of the image to be evaluated is divided to obtain multiple sub-window images; Perform bidirectional horizontal scanning on each sub-window image to obtain the horizontal scanning results for each sub-window; With the goal of eliminating spatial gaps between adjacent windows, feature aggregation is performed on the horizontal scanning results of adjacent sub-windows to obtain an aggregated window sequence. The aggregated window sequence is subjected to bidirectional vertical scanning to obtain a pixel block sequence.
[0009] In this embodiment of the application, the Mamba model includes a Mamba block, which includes a feature preprocessing submodule, an SSM feature extraction submodule, and a feature aggregation submodule; The Mamba model is used to extract features from the pixel block sequence to obtain multi-scale content features, including: The feature preprocessing submodule preprocesses the pixel block sequence to obtain an initial pixel block sequence; The SSM feature extraction submodule extracts features from the initial pixel block sequence to obtain a bidirectional global feature sequence; The feature aggregation submodule performs convolution operations at different scales on the bidirectional global feature sequence to obtain feature maps at multiple scales. The feature aggregation submodule performs global average pooling on the feature maps at each scale to obtain feature vectors with multiple dimensions. The feature aggregation submodule concatenates the feature vectors of the multiple dimensions to obtain multi-scale content features.
[0010] In this embodiment of the application, the horizontal scanning results of each sub-window include the forward sequence and the reverse sequence of each sub-window; The step of performing feature aggregation on the horizontal scanning results of adjacent sub-windows with the goal of eliminating spatial gaps between adjacent windows, to obtain an aggregated window sequence, includes: The forward sequence of the current sub-window is concatenated with the reverse sequence of the next sub-window to obtain the aggregated window sequence.
[0011] In this embodiment of the application, the supernetwork generating target network includes a preliminary processing submodule, a weight generation branch submodule, and a bias generation branch submodule; The step of generating the weight parameters and bias parameters of the target network using a hypernetwork based on the multi-scale content features includes: The preliminary processing submodule compresses the multi-scale content features to obtain a compressed feature vector; The weight generation branch submodule performs a convolution operation on the compressed feature vector and converts the output of the convolution operation into a weight vector to obtain the weight parameters of the target network. The bias generation branch submodule performs global average pooling and full connection on the compressed feature vector to obtain the bias parameters of the target network.
[0012] In this embodiment of the application, the preset image quality prediction target network includes a dynamic weighted fully connected layer submodule and a score normalization submodule; The process of evaluating image quality results using a pre-set image quality prediction target network based on the multi-scale content features, the weight parameters and bias parameters of the target network, includes: The multi-scale content features, the weight parameters and bias parameters of the target network are input into the dynamic weight fully connected layer submodule to obtain the original image evaluation score; The original image evaluation score is normalized by the score normalization submodule to obtain the image quality result.
[0013] A second aspect of this application provides an image quality assessment apparatus, comprising: The acquisition module is used to acquire the image to be evaluated; The extraction module is used to extract semantic features from the image to be evaluated using the Mamba model to obtain multi-scale content features, which include local distortion detail information and global content information. The generation module is used to generate the weight parameters and bias parameters of the target network based on the multi-scale content features using a hypernetwork. The evaluation module is used to evaluate the image quality results using a preset image quality prediction target network based on the multi-scale content features, the weight parameters and bias parameters of the target network.
[0014] The embodiments of the present invention have at least the following advantages or beneficial effects: This invention provides an image quality assessment method and apparatus. The method involves acquiring an image to be assessed; extracting semantic features from the image using a Mamba model to obtain multi-scale content features, including local distortion details and global content information; generating weight and bias parameters for a target network using a hypernetwork based on the multi-scale content features; and evaluating the image quality using a pre-set image quality prediction target network based on the multi-scale content features, the weight and bias parameters of the target network. Feature extraction using the Mamba model effectively captures global image dependencies, and the combination with an adaptive hypernetwork dynamically determines quality perception rules. Finally, image quality assessment is performed using the image quality prediction target network. This method achieves end-to-end blind image quality assessment with global dependency modeling, local distortion capture, and adaptive quality prediction, improving the accuracy of image quality assessment and reducing the deviation between the assessment results and human subjective perception. Attached Figure Description
[0015] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0016] Figure 1 A flowchart of an image quality assessment method provided in an embodiment of the present invention; Figure 2 The overall flowchart of the adaptive visual Mamba network for objective image quality assessment provided in the embodiments of the present invention; Figure 3 This is a schematic diagram of an improved bidirectional window scanning strategy provided in an embodiment of the present invention; Figure 4 This is a schematic diagram of the internal structure of the Mamba block provided in an embodiment of the present invention; Figure 5 A schematic diagram of the core structure for adaptive quality perception rule generation and quality score prediction provided in an embodiment of the present invention; Figure 6 A structural block diagram of an image quality assessment device provided in an embodiment of the present invention; Figure 7 This is a structural block diagram of an electronic device provided in an embodiment of the present invention.
[0017] Icons: 410 - Acquisition module; 420 - Extraction module; 430 - Generation module; 440 - Evaluation module; 101 - Memory; 102 - Processor; 103 - Communication interface. Detailed Implementation
[0018] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. The components of the embodiments of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.
[0019] Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.
[0020] It should be noted that, in this document, the use of relational terms such as "first" and "second" is merely to distinguish one entity or operation from another, and does not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0021] In the description of this application, it should be noted that the terms "upper", "lower", "inner", "outer", etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings, or the orientation or positional relationship that the product of this application is usually placed in. They are only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this application. Example
[0022] The following detailed description of some embodiments of this application is provided in conjunction with the accompanying drawings. Unless otherwise specified, the various embodiments and features described below can be combined with each other.
[0023] Please refer to Figure 1 and Figure 2 , Figure 1 A flowchart of an image quality assessment method provided in an embodiment of the present invention; Figure 2This is an overall flowchart of an adaptive visual Mamba network for objective image quality assessment provided in an embodiment of the present invention. This embodiment provides an image quality assessment method, including the following steps: Step S210: Obtain the image to be evaluated; In this embodiment, the image to be evaluated can be a distorted image, such as a real distorted image captured by a mobile device or a synthetic distorted image. The image format can be RGB three-channel, and the resolution can be adaptive (such as 224×224, 384×384, etc.), without the need for an original reference image.
[0024] Step S220: Use the Mamba model to extract semantic features from the image to be evaluated to obtain multi-scale content features, which include local distortion detail information and global content information; In this embodiment, the Mamba model can generate highly discriminative semantic features rich in contextual information for the image to be evaluated. The 2D image can be converted into a 1D sequence, then linearly projected. A Selective State-Space Model (SSM) is used to dynamically determine which information to remember or forget based on the content of the current image patch, ultimately capturing multi-scale semantic information to obtain multi-scale content features.
[0025] Introducing the Mamba model into image semantic feature extraction, through its selective, content-aware mechanism, the Mamba model can generate highly discriminative, context-rich semantic features for the image to be evaluated, which is helpful for image quality assessment.
[0026] In some embodiments, the step of extracting semantic features from the image to be evaluated using the Mamba model to obtain multi-scale content features includes: First, the image to be evaluated is scanned based on a bidirectional window scanning strategy to obtain a pixel block sequence; In this embodiment, global horizontal and vertical scanning can be performed to achieve feature extraction with no gaps between locally adjacent markers and no loss of global information. Please refer to... Figure 3 , Figure 3 This is a schematic diagram of an improved bidirectional window scanning strategy provided in an embodiment of the present invention. The first row of the image above is scanned using the most basic scanning method, sequentially scanning each row from left to right. You can see that the first pixel in the first row and the second pixel in the second row are actually adjacent, but after scanning row by row, they are separated by a full row. The same principle applies to vertical scanning. For example, taking a "32×32 pixel block" image as an example, there are 4 rows and 4 columns of pixel blocks, denoted as P. 11 ~P 44(Rows 1-4, Columns 1-4), the entire image (without window division) is scanned, and its global scan path and index are as follows: [P 11 (1) → P 12 (2) → P 13 (3) → P 14 (4)] ← Horizontal forward scan (first row); [P 14 (5) ← P 13 (6) ← P 12 (7) ← P 11 (8)] ← Horizontal reverse scan (first row); [P 21 (9) → P 22 (10) → P 23 (11) → P 24 (12)] ← Horizontal forward scan (second row); [P 24 (13) ← P 23 (14) ← P 22 (15) ← P 21 (16)] ← Horizontal reverse scan (second row); [P 31 (17) → P 32 (18) → P 33 (19) → P 34 (20)] ← Horizontal forward scan (third row); [P 34 (21) ← P 33 (22) ← P 32 (23) ← P 31 (24)] ← Horizontal reverse scan (third row); [P 41 (25) → P 42 (26) → P 43 (27) → P 44 (28)] ← Horizontal forward scan (fourth row); [P 44 (29) ← P 43 (30) ← P 42 (31) ← P 41 (32)] ← Horizontal reverse scan (fourth row).
[0027] There is no additional aggregation in the vertical direction; the sequences are directly concatenated into a 1D sequence according to the row scan order, for example: P 11(Index 8, end of the first row reverse scan) and P 21 (Index 9, beginning of the second row forward scan); Spatial relationship: vertically adjacent (belonging to locally adjacent pixel blocks, due to VMamba repeating reverse traversal during cross-row scan), if the window size is an 8×8 pixel block, P 18 (The 8th pixel in the first row, index 8) and P 21 The interval of the first pixel block in the second row (index 17) is 17-8=8 (equal to the window size), resulting in the loss of local dependencies. Using a bidirectional window scanning strategy, scanning can be done in a windowed manner, reducing the interval between pixels. Unlike the scanning method above, the interval is not as large, allowing the sequence index interval between adjacent pixel blocks to be 1, completely eliminating local marker intervals. This meets the requirement of distortion detection in image quality assessment that depends on the granularity of local details.
[0028] In some embodiments, scanning the image to be evaluated based on a bidirectional window scanning strategy to obtain a pixel block sequence includes: The first step is to divide the non-overlapping windows of the image to be evaluated into multiple sub-window images; In this embodiment, the image to be evaluated is divided into non-overlapping windows, such as a window size of 8×8 pixels, which can be adaptively adjusted according to the image resolution.
[0029] The second step is to perform a bidirectional horizontal scan on each sub-window image to obtain the horizontal scan results for each sub-window. In this embodiment, a bidirectional horizontal scan is performed on each sub-window. Specifically, this can involve first traversing the pixel patches within the window from left to right (forward), and then traversing the pixel patches within the same window from right to left (reverse). During the scan, the position information of each pixel patch can be recorded as a 1D sequence index, which avoids the problem of large local marker intervals caused by cross-window scanning in VMamba. Bidirectional horizontal scanning can capture local distortion details (such as tiny noise and local blur) within the sub-window, providing fine-grained features for subsequent distortion detection.
[0030] The third step is to perform feature aggregation on the horizontal scanning results of adjacent sub-windows with the goal of having no spatial gap between adjacent windows, and obtain the aggregated window sequence. The horizontal scanning results of each sub-window include the forward sequence and the reverse sequence of each sub-window; correspondingly, the step of performing feature aggregation on the horizontal scanning results of adjacent sub-windows with the goal of having no spatial gap between adjacent windows to obtain an aggregated window sequence includes: performing feature concatenation between the forward sequence of the current sub-window and the reverse sequence of the next sub-window to obtain an aggregated window sequence.
[0031] In this embodiment, feature stitching can be performed on the horizontal scan results of adjacent sub-windows. For example, the forward scan sequence of sub-window 1 and the reverse scan sequence of sub-window 2 can be directly stitched together. This ensures that the local markers of adjacent windows have no spatial gaps, such as... Figure 2 As shown in the figure, arrows connecting adjacent windows indicate the aggregation direction.
[0032] The fourth step is to perform a bidirectional vertical scan on the aggregated window sequence to obtain a pixel block sequence.
[0033] In this embodiment, a bidirectional vertical scan is performed on the aggregated window sequence. Specifically, this can be done by first traversing the aggregated sequence of all windows from top to bottom (forward), and then traversing from bottom to top (backward), thereby completely converting the spatial information of the 2D image into an ordered 1D sequence. Bidirectional vertical scanning can capture global dependencies in the image (such as the correlation between "character outline integrity" and "facial local noise") while preserving local details, achieving synergy between local and global features.
[0034] Bidirectional horizontal scanning can capture local distortion details within a sub-window, providing fine-grained features for subsequent distortion detection. By aggregating features from the horizontal scanning results of adjacent sub-windows and then performing bidirectional vertical scanning, global dependencies in the image can be captured while preserving local details, achieving synergy between local and global features.
[0035] Then, the Mamba model is used to extract features from the pixel block sequence to obtain multi-scale content features.
[0036] In this embodiment, the Mamba model can first perform image segmentation and embedding on the pixel block sequence, then add location information, and finally extract features from multiple Mamba blocks to output a feature sequence, thereby obtaining multi-scale content features.
[0037] In some embodiments, the Mamba model includes Mamba blocks, each Mamba block comprising a feature preprocessing submodule, an SSM feature extraction submodule, and a feature aggregation submodule; correspondingly, the step of using the Mamba model to extract features from the pixel block sequence to obtain multi-scale content features includes: The first step is to preprocess the pixel block sequence by the feature preprocessing submodule to obtain an initial pixel block sequence; In this embodiment, the internal structure diagram of the MambaBlock is as follows: Figure 4 As shown, Figure 4This is a schematic diagram of the internal structure of the Mamba block provided in an embodiment of the present invention. The pixel block sequence is an ordered 1D pixel block sequence, with each pixel block having dimensions of C×H×W, where C is the number of channels, such as 32 channels; H and W are the pixel block sizes, such as 4×4. The above preprocessing process includes: ① performing "Layer Normalization (LN)" on each pixel block to avoid data distribution offset affecting subsequent SSM processing; ② compressing the pixel block dimension to D dimensions (such as D=128) through a "Linear Transformation Layer" to adapt to the input requirements of the SSM feature extraction submodule; finally, outputting the normalized 1D feature sequence (with dimensions of N×D, where N is the sequence length, i.e., the total number of pixel blocks), to obtain the initial pixel block sequence.
[0038] The second step involves the SSM feature extraction submodule extracting features from the initial pixel block sequence to obtain a bidirectional global feature sequence. In this embodiment, the SSM feature extraction submodule can be a bidirectional state-space model (Bi-SSM), comprising forward SSM units and reverse SSM units (both with the same structure but opposite processing directions). The SSM feature extraction submodule processes the initial pixel block sequence based on the discretized state-space equation, the discretization formula being as follows:
[0039] in: The discretization step size can be determined through experimental optimization; for example, it can be set to 0.1. , , This is the initial parameter matrix of the SSM; , The parameter matrix is the discretized form. The hidden state at time t (stores historical sequence information); The input features at time t; The output feature at time t; It is the identity matrix; This is for matrix exponentiation operations.
[0040] The processing steps include: 1. The forward SSM unit traverses the 1D sequence from left to right and calculates the forward hidden state; 2. The reverse SSM unit traverses the 1D sequence from right to left and calculates the reverse hidden state; 3. The forward and reverse hidden states are concatenated by channel (dimension N×2D) to obtain bidirectional global features. The output bidirectional global feature sequence has a dimension of N×2D.
[0041] The third step involves the feature aggregation submodule performing convolution operations at different scales on the bidirectional global feature sequence to obtain feature maps at multiple scales. The fifth step involves the feature aggregation submodule concatenating the feature vectors from multiple dimensions to obtain multi-scale content features.
[0042] In this embodiment, the feature aggregation submodule performs convolution operations at different scales on the global feature sequence. Specifically, it uses three parallel 1×1 convolution kernels (with 64, 128, and 256 output channels respectively) to generate feature maps at three scales. Then, it performs global average pooling (GAP) on each scale feature map to obtain three 1D feature vectors (with dimensions of 64, 128, and 256 respectively). Finally, it concatenates the three feature vectors to obtain a multi-scale global-local fusion semantic feature (with a dimension of 448), which is the multi-scale content feature. The output multi-scale fusion semantic feature (with a dimension of 448) contains both local distortion details (captured by small-scale convolution) and global content information (captured by large-scale convolution), and is used as input to the subsequent supernetwork and target network.
[0043] By employing an improved scanning strategy and a bidirectional state space model (Bi-SSM), feature extraction can be achieved that includes both local distortion details (small-scale convolutional capture) and global content information (large-scale convolutional capture), making feature extraction more accurate.
[0044] Step S230: Based on the multi-scale content features, use a hypernetwork to generate the weight parameters and bias parameters of the target network; In this embodiment, the hypernetwork can be an adaptive hypernetwork, which can generate the function of a quality-aware rule adaptive hypernetwork. That is, it can dynamically generate the weights and biases of the target network according to the semantic features of the input image, avoiding the poor generalization of traditional models with fixed weights.
[0045] Please refer to Figure 5 , Figure 5 This is a schematic diagram of the core structure of adaptive quality-aware rule generation and quality score prediction provided in an embodiment of the present invention. In some embodiments, the supernetwork generating target network includes a preliminary processing submodule, a weight generation branch submodule, and a bias generation branch submodule; The step of generating the weight parameters and bias parameters of the target network using a hypernetwork based on the multi-scale content features includes: First, the preliminary processing submodule compresses the multi-scale content features to obtain a compressed feature vector; In this embodiment, the preliminary processing submodule compresses the multi-scale content features (448 dimensions) through a 1×1 convolutional layer (256 output channels); then it performs global average pooling (GAP) to obtain a 1×256 feature vector; finally, it outputs a 1×256 compressed feature vector, which is the compressed feature vector.
[0046] Then, the weight generation branch submodule performs a convolution operation on the compressed feature vector and converts the output of the convolution operation into a weight vector to obtain the weight parameters of the target network. In this embodiment, the weight dimension needs to match the target network layer dimension. In the example above, the weight parameters of the four fully connected layers (FC1~FC4) of the target network can be generated. For example, the weight dimension of FC1 is 448×256, the weight dimension of FC2 is 256×128, the weight dimension of FC3 is 128×64, and the weight dimension of FC4 is 64×1. The specific processing steps include: ① Performing three parallel 3×3 convolution operations on the compressed feature vector (output channel numbers are 448×256, 256×128, 128×64, and 64×1, corresponding to the weight dimensions of FC1~FC4); ② Performing feature reshaping on each convolution output, converting the 2D feature map into a 1D weight vector, such as reshaping the convolution output of FC1 into a 448×256 weight matrix; finally, outputting the dynamic weight matrices of FC1~FC4 (448×256, 256×128, 128×64, and 64×1, respectively), to obtain the weight parameters of the target network.
[0047] Finally, the bias generation branch submodule performs global average pooling and full connection on the compressed feature vector to obtain the bias parameters of the target network.
[0048] In this embodiment, the bias dimension matches the output dimension of the fully connected layer. In the example above, the bias generation branch submodule can generate the bias parameters of the target network FC1~FC4, such as FC1 bias dimension of 256, FC2 of 128, FC3 of 64, and FC4 of 1. Specific processing steps: ① Perform global average pooling (GAP) on the compressed feature vector to obtain 1×1 feature values; ② Directly generate the bias vectors of FC1~FC4 through four parallel fully connected layers (output dimensions of 256, 128, 64, and 1 respectively). It should be noted that the number of bias parameters is much smaller than the number of weights, so a simplified generation method of "pooling + fully connected layers" can be used to reduce the computational load. Finally, the dynamic bias vectors of FC1~FC4 (256, 128, 64, and 1 respectively) are output, which are the bias parameters of the target network.
[0049] The supernetwork generating target network may further include a rule output submodule. After obtaining the bias and weight parameters of the target network, the weight matrix and bias vector can be packaged according to the layer correspondence. For example, the weights of FC1 correspond to the biases of FC1, the weights of FC2 correspond to the biases of FC2, and so on. Finally, multiple sets of weight-bias pairs are output as quality-aware rules, which can be directly passed to the target network.
[0050] By setting up a preliminary processing submodule, a weight generation branch submodule, and a bias generation branch submodule in the supernetwork to generate the target network, the weight parameters and bias parameters of the target network can be dynamically generated in a differentiated manner. This can solve the problem of poor generalization of existing models across datasets and help the target network to be evaluated better.
[0051] Step S240: Based on the multi-scale content features, the weight parameters and bias parameters of the target network, the image quality result is obtained by using a preset image quality prediction target network.
[0052] In this embodiment, the weight parameters and bias parameters of the target network are used as the parameters of the fully connected layer of the image quality prediction target network. Multi-scale content features are input into the image quality prediction target network, and finally the image quality result is obtained.
[0053] In some embodiments, the preset image quality prediction target network includes a dynamic weighted fully connected layer submodule and a score normalization submodule; correspondingly, the step of evaluating the image quality result using the preset image quality prediction target network based on the multi-scale content features, the weight parameters and bias parameters of the target network includes: First, the multi-scale content features, the weight parameters and bias parameters of the target network are input into the dynamic weight fully connected layer submodule to obtain the original image evaluation score; In this embodiment, the dynamic weighted fully connected layer submodule includes multiple fully connected layers. For example, in the above submodule, there are four fully connected layers (FC1~FC4). The dynamic weighted fully connected layer submodule performs the following steps according to the weight parameters and bias parameters of the input target network: ①FC1: Input a 448-dimensional feature vector, load the FC1 weights (448×256) and bias (256) generated by the supernetwork, perform a linear transformation, and then pass through the sigmoid activation function to output a 256-dimensional feature vector; ②FC2: Input a 256-dimensional feature vector, load the FC2 weights (256×128) and bias (128), and output a 128-dimensional feature vector after sigmoid activation; ③FC3: Input a 128-dimensional feature vector, load the FC3 weights (128×64) and bias (64), and output a 64-dimensional feature vector after sigmoid activation; ④FC4: Input a 64-dimensional feature vector, load the FC4 weights (64×1) and bias (1), and output a 1 after sigmoid activation. Dimensional score (range 0-1); Output: 0-1 evaluation score of the original image.
[0054] Then, the original image evaluation score is normalized by the score normalization submodule to obtain the image quality result.
[0055] In this embodiment, the normalization process includes: 1. Performing linear scaling to map the original image evaluation score to a range of 0-100. Specifically, it can be obtained by using the following formula: final score = original image evaluation score × 100. Finally, the image quality score of 0-100 (such as 74.65) is output, which is the image quality result. The higher the score, the better the image quality.
[0056] The dynamic weighted fully connected layer submodule can obtain the original image evaluation score based on multi-scale content features, and then the score normalization submodule can be used to normalize the score, so that the image quality result is more in line with the evaluation requirements.
[0057] In the above implementation process, the image to be evaluated is acquired; semantic features are extracted from the image using the Mamba model to obtain multi-scale content features, which include local distortion details and global content information; based on the multi-scale content features, a hypernetwork is used to generate the weight parameters and bias parameters of the target network; based on the multi-scale content features, the weight parameters and bias parameters of the target network, a preset image quality prediction target network is used to evaluate and obtain the image quality result. Feature extraction using the Mamba model can effectively capture global dependencies in the image, and combined with an adaptive hypernetwork, quality perception rules can be dynamically determined. Finally, image quality evaluation is performed using the image quality prediction target network, which can achieve end-to-end blind image quality evaluation with global dependency modeling, local distortion capture, and adaptive quality prediction, improving the accuracy of image quality evaluation and reducing the deviation between the evaluation results and human subjective perception.
[0058] The Mamba model can extract global semantic features, capture local distortion features by combining window scanning, and then aggregate multi-scale features (image content features of different proportions) and input them into the target network. This clarifies the fusion path of local fine-grained details and global overall information, which can solve the problem of the inability of global and local information to coordinate due to the local bias of convolutional neural networks and the high complexity of Transformer.
[0059] The scheme was tested on the TID2013 dataset and achieved Spearman's Rank Correlation Coefficient (SRCC) of 96.3% and Pearson Linear Correlation Coefficient (PLCC) of 96.0%, which are significantly better than traditional methods and can achieve a balance between efficient computation and accurate evaluation.
[0060] Please refer to Figure 6 , Figure 6 This is a structural block diagram of an image quality assessment device provided in an embodiment of the present invention. Based on the same inventive concept, this embodiment provides an image quality assessment device, including an acquisition module 410, an extraction module 420, a generation module 430, and an assessment module 440, wherein: Acquisition module 410 is used to acquire the image to be evaluated; The extraction module 420 is used to extract semantic features from the image to be evaluated using the Mamba model to obtain multi-scale content features, which include local distortion detail information and global content information. The generation module 430 is used to generate the weight parameters and bias parameters of the target network based on the multi-scale content features using a hypernetwork. Evaluation module 440 is used to evaluate the image quality result using a preset image quality prediction target network based on the multi-scale content features, the weight parameters and bias parameters of the target network.
[0061] The image to be evaluated is acquired by the acquisition module 410; the extraction module 420 uses the Mamba model to extract semantic features from the image to obtain multi-scale content features, which include local distortion details and global content information; the generation module 430 uses a hypernetwork to generate the weight parameters and bias parameters of the target network based on the multi-scale content features; and the evaluation module 440 uses a preset image quality prediction target network to evaluate the image quality based on the multi-scale content features, the weight parameters and bias parameters of the target network, and obtains the image quality result. Feature extraction using the Mamba model can effectively capture global image dependencies, and combining it with an adaptive hypernetwork can dynamically determine quality perception rules. Finally, image quality evaluation is performed using the image quality prediction target network, achieving end-to-end blind image quality evaluation with global dependency modeling, local distortion capture, and adaptive quality prediction. This improves the accuracy of image quality evaluation and reduces the deviation between the evaluation results and human subjective perception.
[0062] Please see Figure 7 , Figure 7 This is a schematic structural block diagram of an electronic device provided in an embodiment of this application. The electronic device includes a memory 101, a processor 102, and a communication interface 103. The memory 101, processor 102, and communication interface 103 are electrically connected to each other directly or indirectly to realize data transmission or interaction. For example, these components can be electrically connected to each other through one or more communication buses or signal lines. The memory 101 can be used to store software programs and modules, such as the program instructions / modules corresponding to the image quality assessment device provided in the embodiment of this application. The processor 102 executes various functional applications and data processing by executing the software programs and modules stored in the memory 101. The communication interface 103 can be used to communicate with other node devices for signaling or data.
[0063] The memory 101 may be, but is not limited to, random access memory (RAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), etc.
[0064] The processor 102 can be an integrated circuit chip with signal processing capabilities. The processor 102 can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0065] Understandable. Figure 7 The structure shown is for illustrative purposes only; the electronic device may also include components that are more advanced than those shown. Figure 7 The more or fewer components shown, or having the same Figure 7 The different configurations shown. Figure 7 The components shown can be implemented using hardware, software, or a combination thereof.
[0066] In the embodiments provided in this application, it should be understood that the disclosed apparatus and methods can also be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of apparatus, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram and / or flowchart, and combinations of blocks in block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0067] In addition, the functional modules in the various embodiments of this application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.
[0068] If the aforementioned functions are implemented as software functional modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0069] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
[0070] It will be apparent to those skilled in the art that this application is not limited to the details of the exemplary embodiments described above, and that this application can be implemented in other specific forms without departing from the spirit or essential characteristics of this application. Therefore, the embodiments should be considered illustrative and non-limiting in all respects, and the scope of this application is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be included within this application. No reference numerals in the claims should be construed as limiting the scope of the claims.
Claims
1. An image quality assessment method, characterized in that, Includes the following steps: Acquire the image to be evaluated; The Mamba model is used to extract semantic features from the image to be evaluated, resulting in multi-scale content features, which include local distortion detail information and global content information. Based on the multi-scale content features, a hypernetwork is used to generate the weight parameters and bias parameters of the target network; Based on the multi-scale content features, the weight parameters and bias parameters of the target network, the image quality results are obtained by using a preset image quality prediction target network.
2. The image quality assessment method according to claim 1, characterized in that, The Mamba model is used to extract semantic features from the image to be evaluated, resulting in multi-scale content features, including: The image to be evaluated is scanned using a bidirectional window scanning strategy to obtain a pixel block sequence; The Mamba model is used to extract features from the pixel block sequence to obtain multi-scale content features.
3. The image quality assessment method according to claim 2, characterized in that, The bidirectional window scanning strategy scans the image to be evaluated to obtain a pixel block sequence, including: The non-overlapping window of the image to be evaluated is divided to obtain multiple sub-window images; Perform bidirectional horizontal scanning on each sub-window image to obtain the horizontal scanning results for each sub-window; With the goal of eliminating spatial gaps between adjacent windows, feature aggregation is performed on the horizontal scanning results of adjacent sub-windows to obtain an aggregated window sequence. The aggregated window sequence is subjected to bidirectional vertical scanning to obtain a pixel block sequence.
4. The image quality assessment method according to claim 2, characterized in that, The Mamba model includes a Mamba block, which includes a feature preprocessing submodule, an SSM feature extraction submodule, and a feature aggregation submodule. The Mamba model is used to extract features from the pixel block sequence to obtain multi-scale content features, including: The feature preprocessing submodule preprocesses the pixel block sequence to obtain an initial pixel block sequence; The SSM feature extraction submodule extracts features from the initial pixel block sequence to obtain a bidirectional global feature sequence; The feature aggregation submodule performs convolution operations at different scales on the bidirectional global feature sequence to obtain feature maps at multiple scales. The feature aggregation submodule performs global average pooling on the feature maps at each scale to obtain feature vectors with multiple dimensions. The feature aggregation submodule concatenates the feature vectors of the multiple dimensions to obtain multi-scale content features.
5. The image quality assessment method according to claim 3, characterized in that, The horizontal scan results of each sub-window include the forward sequence and the reverse sequence of each sub-window; The step of performing feature aggregation on the horizontal scanning results of adjacent sub-windows with the goal of eliminating spatial gaps between adjacent windows, to obtain an aggregated window sequence, includes: The forward sequence of the current sub-window is concatenated with the reverse sequence of the next sub-window to obtain the aggregated window sequence.
6. The image quality assessment method according to claim 1, characterized in that, The supernetwork generates the target network, which includes a preliminary processing submodule, a weight generation branch submodule, and a bias generation branch submodule. The step of generating the weight parameters and bias parameters of the target network using a hypernetwork based on the multi-scale content features includes: The preliminary processing submodule compresses the multi-scale content features to obtain a compressed feature vector; The weight generation branch submodule performs a convolution operation on the compressed feature vector and converts the output of the convolution operation into a weight vector to obtain the weight parameters of the target network. The bias generation branch submodule performs global average pooling and full connection on the compressed feature vector to obtain the bias parameters of the target network.
7. The image quality assessment method according to claim 1, characterized in that, The preset image quality prediction target network includes a dynamic weighted fully connected layer submodule and a score normalization submodule; The process of evaluating image quality results using a pre-set image quality prediction target network based on the multi-scale content features, the weight parameters and bias parameters of the target network, includes: The multi-scale content features, the weight parameters and bias parameters of the target network are input into the dynamic weight fully connected layer submodule to obtain the original image evaluation score; The original image evaluation score is normalized by the score normalization submodule to obtain the image quality result.
8. An image quality assessment device, characterized in that, include: The acquisition module is used to acquire the image to be evaluated; The extraction module is used to extract semantic features from the image to be evaluated using the Mamba model to obtain multi-scale content features, which include local distortion detail information and global content information. The generation module is used to generate the weight parameters and bias parameters of the target network based on the multi-scale content features using a hypernetwork. The evaluation module is used to evaluate the image quality results using a preset image quality prediction target network based on the multi-scale content features, the weight parameters and bias parameters of the target network.