Method and system for evaluating screen content images modulated by a distortion level guided type of modulation
Patent Information
- Application Number
- CN202611010249.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-08
- Publication Date
- 2026-09-25
AI Technical Summary
然而,这类方法往往将失真类型、失真等级与图像质量视为隐含的监督信号,未能显式建模三者之间的内在关联,导致模型对不同内容和不同失真条件的适应性不足
[0015]由上述本发明的实施例提供的技术方案可以看出,本发明针对屏幕内容图像中失真类型与失真等级相互依赖而传统评价方法难以有效建模的问题,提出了一种失真等级引导类型调制的屏幕内容图像评价方法和系统,属于图像质量评价技术领域。该方法首先通过提示词组驱动的文本图像语义对齐模型提取失真类型特征,并利用基于排序学习的网络获取不同失真程度的等级特征。同时,对视觉主干网络进行微调以提取屏幕内容图像的语义视觉特征。随后,利用等级特征对类型特征进行线性调制,使模型能够根据失真严重程度自适应调整各类失真信息的影响。最后通过互注意力机制融合视觉特征与调制后的类型特征,以增强模型对文本、图形等敏感区域的关注度。该方法能够有效提升屏幕内容图像在多类失真条件下的质量评价准确性和泛化能力,具有较高的鲁棒性和应用价值。
Smart Images

Figure CN122820631A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image quality evaluation technology, and in particular to a method and system for evaluating screen content images using distortion level-guided type modulation. Background Technology
[0002] With the rapid development of digital media applications such as online education and remote collaboration, screen content images are playing an increasingly important role in current visual communication. Unlike natural images, screen content images are typically computer-generated, containing a large number of text, graphics, and interface elements. Their content structure is regular and their regional features vary significantly, thus exhibiting unique sensitivities in terms of distortion type, distortion level, and content semantics. How to conduct accurate and stable no-reference quality assessment of these images has become an important research direction in the field of image quality assessment.
[0003] Traditional image quality assessment methods primarily rely on manually designed features, analyzing statistical properties to detect distortions such as blur and noise. While these methods are effective under specific distortion conditions, their reliance on prior assumptions makes them ill-suited to diverse screen content and complex distortion patterns, resulting in weak generalization ability. In recent years, deep learning technology has driven the development of image quality assessment methods. Some studies model quality assessment as a multi-task problem that simultaneously predicts distortion type, distortion level, and image quality, aiming to obtain a more comprehensive feature representation. However, these methods often treat distortion type, distortion level, and image quality as implicit supervisory signals, failing to explicitly model the intrinsic relationship between them, leading to insufficient model adaptability to different content and distortion conditions. Furthermore, some research based on vision-language models attempts to use cue words to describe image quality, but most rely on global descriptions and remain insufficient in handling the interaction between distortion level and type.
[0004] Existing research indicates that the subjective quality of screen content images significantly decreases with increasing distortion level, and the impact of distortion type on quality often depends on its severity and exhibits clear content relevance. However, existing methods fail to adequately model this dependency between distortion level and distortion type, making it difficult to align with human subjective perception. Summary of the Invention
[0005] The embodiments of the present invention provide a method and system for evaluating screen content images based on distortion level guided type modulation, which is used to solve the problems existing in the prior art.
[0006] To achieve the above objectives, the present invention adopts the following technical solution.
[0007] A method for evaluating the image quality of screen content based on distortion level-guided type modulation includes: S1. Perform semantic similarity matching between the original image and semantic anchors to obtain a distortion type feature vector with semantic discrimination capability; the semantic anchors are obtained by constructing multiple descriptive cue words through a text-visual semantic alignment model driven by cue words. S2. Input the original images with the same content and the same distortion type feature vector but different distortion levels into the network model. Through the ranking constraint, the network model learns to obtain the distortion continuous level features that reflect the severity of distortion. S3. Perform domain-adaptive training on the visual backbone network so that it can adapt to the structural regions of the screen content image and output semantic visual features. S4. Perform linear modulation processing on the distortion continuous level features to obtain scaling factor and offset factor, and use the scaling factor and offset factor to obtain the type feature representation that dynamically changes with the distortion level. S5. Utilize the mutual attention mechanism, use semantic visual features as queries, and use modulated distortion type features as keys and values to enable the network model to focus on the regions most sensitive to quality. S6. The semantic visual features are fused with the modulated distortion type features, and the resulting fused features are input into the fully connected layer to obtain the subjective quality score prediction value of the screen content image. The predicted subjective quality score of the screen content image is used to adjust the processing method of the screen content image.
[0008] Preferably, the network model includes: The system comprises a distortion type feature extraction module, a distortion level feature extraction module, and a visual feature extraction module, all configured in parallel. The distortion type feature extraction module is used to perform step S1, the distortion level feature extraction module is used to perform step S2, and a ranking learning network is employed. The visual feature extraction module is used to perform step S3. The modulation module is used to perform step S4; the attention mechanism fusion module is used to perform step S5; and the fully connected layer is used to perform step S6.
[0009] Preferably, in step S1, the process of constructing multiple descriptive cue words to obtain semantic anchors through a text-visual semantic alignment model driven by cue words includes: The prompts used to describe different distortion phenomena are mapped into high-dimensional semantic vectors; Step S1 specifically includes: S11. Encode the input image to extract visual semantic features and high-dimensional semantic vectors corresponding to candidate distortion types; S12. Compare the visual semantic features with the high-dimensional semantic vector to calculate the similarity. S13. Normalize the obtained similarity values to obtain the probability distribution of the input image belonging to each candidate distortion type; S14. Select the candidate distortion type with the highest probability value from the probability distribution of the input image belonging to each candidate distortion type as the distortion type of the current input image, and take the high-dimensional semantic vector corresponding to the distortion type of the current input image as the distortion type feature vector with semantic discrimination ability.
[0010] Preferably, step S2 specifically includes: S21. Extract degradation pattern features from two input images using the convolutional structure of a sorting learning network, and use them as quality level codes. During the output quality level coding process, the ranking loss constraint of the ranking learning network makes the score of the high-distortion original image output by the fully connected layer lower than the score of the low-distortion original image, and enables the ranking learning network to establish a monotonic relationship between the distortion degree of the input image and the quality level coding. S22. Execute sub-step S21 multiple times so that the ranking learning network can establish a monotonic relationship between the distortion level of the input image and the quality level coding. S23. The ranking learning network obtained by executing sub-step S22 obtains distortion successive level features that reflect the severity of distortion.
[0011] Preferably, step S3 specifically includes: The input image is divided into multiple image blocks, each image block is mapped to a feature vector, and positional encoding is added to obtain the image feature sequence; The image feature sequence is input into a multi-layer Transformer encoder of the visual backbone network, and visual features are extracted through a multi-head self-attention mechanism and a feedforward network to obtain a visual feature sequence. This visual feature sequence contains rich local structural textures, which are used to subsequently identify phenomena such as blurred text, structural breaks, and noise interference caused by distortion.
[0012] Preferably, step S4 specifically includes: The distortion successive level features are input into the modulation parameter generator to obtain a set of scaling factors. and offset factor ; Set the distortion type feature vector as Through the formula
[0013] scaling factor and offset factor Together they act on the distortion-type feature vector Obtain the modulated distortion type feature vector In the formula, ⊙ represents element-wise multiplication or channel-wise multiplication; scaling factor Used to enhance or suppress different types of feature dimensions, when the scaling factor corresponding to a certain dimension... When the value exceeds a preset threshold, the modulated distortion type feature vector It is enhanced, and conversely suppressed; the shift factor. Used to adjust the overall position of type features. w This indicates the weights generated by distorted continuous hierarchical features.
[0014] Secondly, the present invention provides a screen content image evaluation system based on distortion level-guided type modulation, comprising: The distortion level feature extraction module is used to: perform semantic similarity matching between the original image and semantic anchor points to obtain a distortion type feature vector with semantic distinguishability; The distortion type feature extraction module is used to: input original images with the same content and the same distortion type feature vector but different distortion levels into the network model, and use ranking constraints to enable the network model to learn to obtain distortion continuous level features that reflect the severity of distortion. The visual feature extraction module is used to: perform domain-adaptive training on the visual backbone network, enabling the visual backbone network to adapt to the structural regions of the screen content image and output semantic visual features. The modulation module is used to: perform linear modulation processing on the distortion continuous level features to obtain scaling factors and offset factors, and use the scaling factors and offset factors to obtain a type feature representation that dynamically changes with the distortion level; The attention mechanism fusion module is used to: utilize the mutual attention mechanism, with semantic visual features as queries, and modulated distortion type features as keys and values, so that the network model focuses on the regions most sensitive to quality. The fully connected layer is used to perform regression operations on the fused features obtained by fusing semantic visual features with modulated distortion type features to obtain the predicted subjective quality score of the screen content image. The output module is used to output the predicted subjective quality score of the screen content image.
[0015] As can be seen from the technical solutions provided by the embodiments of the present invention above, the present invention addresses the problem that traditional evaluation methods struggle to effectively model the interdependence between distortion type and distortion level in screen content images. It proposes a screen content image evaluation method and system based on distortion level-guided type modulation, belonging to the field of image quality evaluation technology. This method first extracts distortion type features using a cue phrase-driven text-image semantic alignment model and then utilizes a ranking-based learning network to obtain level features for different distortion degrees. Simultaneously, the visual backbone network is fine-tuned to extract semantic visual features of the screen content image. Subsequently, the level features are linearly modulated to the type features, enabling the model to adaptively adjust the influence of various distortion information according to the severity of distortion. Finally, a mutual attention mechanism is used to fuse visual features and modulated type features to enhance the model's attention to sensitive areas such as text and graphics. This method effectively improves the accuracy and generalization ability of screen content image quality evaluation under multiple distortion conditions, exhibiting high robustness and application value.
[0016] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and will become apparent from the description or may be learned by practice of the invention. Attached Figure Description
[0017] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 A flowchart illustrating the processing of the distortion level-guided type modulation screen content image evaluation method provided by the present invention. Figure 2 A schematic diagram of the overall framework of a preferred embodiment of the screen content image evaluation method based on distortion level guided type modulation provided by the present invention; Figure 3 A schematic diagram of the distortion type feature extraction module based on the cue group mechanism in the distortion level guided type modulation screen content image evaluation method provided by the present invention; Figure 4 A schematic diagram of the distortion level feature extraction module based on ranking learning in the distortion level guided type modulation screen content image evaluation method provided by the present invention. Figure 5 A schematic diagram of the modulation module of the distortion level-guided type modulation screen content image evaluation method provided by the present invention; Figure 6The logical block diagram of the screen content image evaluation system with distortion level guided type modulation provided by the present invention. Detailed Implementation
[0019] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.
[0020] Those skilled in the art will understand that, unless specifically stated otherwise, the singular forms “a,” “an,” “the,” and “the” used herein may also include the plural forms. It should be further understood that the term “comprising” as used in this specification means the presence of the stated features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. It should be understood that when we say an element is “connected” or “coupled” to another element, it can be directly connected or coupled to the other element, or there may be intermediate elements. Furthermore, “connected” or “coupled” as used herein can include wireless connections or couplings. The term “and / or” as used herein includes any and all combinations of one or more of the associated listed items.
[0021] It will be understood by those skilled in the art that, unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. It should also be understood that terms such as those defined in general dictionaries should be understood to have the same meaning as in the context of the prior art, and should not be interpreted in an idealized or overly formal sense unless defined as herein.
[0022] To facilitate understanding of the embodiments of the present invention, the following will provide further explanation and description with reference to the accompanying drawings and several specific embodiments. These embodiments do not constitute a limitation on the embodiments of the present invention.
[0023] This invention provides a screen content image evaluation method based on distortion level-guided type modulation, addressing the existing technical problem of needing a quality evaluation method capable of characterizing the interaction between distortion level, distortion type, and content semantics, in order to improve the accuracy and generalization ability of screen content image quality prediction in complex scenes. Figure 1 As shown, the screen content image evaluation method provided by the present invention includes the following steps: Step S1, Distortion type feature extraction By utilizing a text-visual semantic alignment model driven by cue phrases, multiple descriptive cue words are constructed to generate a set of semantic anchors. The input image (original image) is then semantically matched with the cue phrases to obtain a distortion-type feature vector with semantic discriminative capabilities.
[0024] Step S2, Distortion Level Feature Extraction
[0025] A ranking representation network based on ranking learning is constructed. Images with the same content and the same type of distortion but different levels of distortion are input into the network, and ranking constraints enable the network to learn continuous ranking features that reflect the severity of distortion.
[0026] Step S3, Visual Feature Extraction
[0027] The visual backbone network is trained in a domain-adaptive manner to adapt to structural regions such as text and graphics in the screen content image, and output semantic visual features.
[0028] Step S4 Distortion Level Modulation Distortion Type
[0029] The extracted distortion level features are linearly modulated to generate scaling and offset factors. These factors are then used to perform a weighted transformation on the distortion type features, resulting in a type feature representation that dynamically changes with the distortion level. Specifically, a higher distortion level strengthens the type features, while a lower distortion level suppresses type differences. This aligns with the human mechanism of first perceiving overall degradation and then perceiving specific distortions.
[0030] Step S5, Mutual Attention Fusion Features
[0031] By utilizing the mutual attention mechanism, visual features are used as queries, and modulated type features are used as keys and values, enabling the model to focus on the regions most sensitive to quality, such as text, edges, and contours, thereby improving the regional accuracy of quality assessment.
[0032] Step S6, Quality Regression
[0033] The fused features are input into a fully connected layer to obtain the predicted subjective quality score of the screen content image.
[0034] This subjective quality score prediction is used to quantitatively evaluate the perceived quality of screen content images and can serve as a quality feedback signal during image generation, transmission, compression, enhancement, and reconstruction. This score can determine the current image quality and guide parameter adjustments and result optimization in subsequent processing, guiding the image processing towards a direction that better aligns with human subjective perception of quality and improving the performance of the image processing system.
[0035] The screen content image quality assessment method provided by this invention employs a phased training approach. First, pre-training is performed to extract distortion types, levels, and visual features. Then, modulation-based training is conducted, with the screen content image as input and the corresponding subjective quality prediction value as output. This method explicitly models the distortion level-guided distortion type dependency mechanism, making the prediction results more consistent with human subjective perception of screen content images.
[0036] In a preferred embodiment provided by the present invention, the network model (evaluation model) has the following architecture: Figure 2 As shown, it comprises four main parts along the data flow direction: The original image is input into three parallel modules: a distortion level feature extraction module, a distortion type feature extraction module, and a visual feature extraction module, which are used to execute the first three steps mentioned above simultaneously.
[0037] The distortion type feature vector output by the distortion type feature extraction module, together with the distortion continuity level feature output by the distortion level feature extraction module, is input into the modulation module to perform the fourth step of the above process.
[0038] The output of the modulation module and the semantic visual features output by the visual feature extraction module are fed into the attention mechanism fusion module. This module uses the mutual attention mechanism, with the semantic visual features as the query and the modulated distortion type features as the key and value, so that the network model focuses on the region most sensitive to quality.
[0039] The fused feature, which combines semantic visual features with modulated distortion type features, is input into a fully connected layer to obtain the predicted subjective quality score of the screen content image.
[0040] In the process of acquiring distortion type features, to obtain distortion type features with semantic discriminative capabilities, this invention constructs multiple cue phrases describing different distortion phenomena as semantic anchors, such as "image blurry," "obvious compression artifacts," "reduced contrast," and "increased noise particles." The text encoder maps these descriptions into high-dimensional semantic vectors. The input image undergoes visual encoding to extract its visual semantic features, and similarity is calculated with the semantic vectors of all cue words. Through comparative training, the visual features gradually approach the semantic anchors corresponding to their true distortion types, thereby constructing a stable and interpretable distortion type feature space. After training, the visual features serve as distortion type features, used to represent the semantic tendency of an image to belong to various distortion types.
[0041] Specifically, this process can be performed through the distortion type feature extraction module as follows: Figure 3As shown, prompt phrases such as "motion blur," "Gaussian noise," and "contrast loss" are constructed and input into the built-in text encoder of this module to obtain the first distortion type representation (distortion type representation 1), the second distortion type representation (distortion type representation 2), and the third distortion type representation (distortion type representation 3), respectively. Simultaneously, the original image is input into the visual encoder for processing to obtain visual semantic features. Similarity is calculated between these features and the three distortion type representations mentioned above, obtaining similarity scores 1, 2, and 3. Based on the three similarities, the candidate distortion type with the highest probability distribution is selected as the distortion type of the current input image, and the high-dimensional semantic vector corresponding to this distortion type is used as the aforementioned distortion type feature vector with semantic discriminative capabilities.
[0042] In some preferred embodiments, the following sub-steps may also be used: S11. Encode the input image to extract visual semantic features and high-dimensional semantic vectors corresponding to candidate distortion types; S12. Compare the visual semantic features with the high-dimensional semantic vector to calculate the similarity. S13. Normalize the obtained similarity values to obtain the probability distribution of the input image belonging to each candidate distortion type; S14. Select the candidate distortion type with the highest probability value from the probability distribution of each candidate distortion type in the input image as the distortion type of the current input image, and use the high-dimensional semantic vector corresponding to the distortion type of the current input image as the distortion type feature vector with semantic distinguishing ability mentioned above.
[0043] In step S2, to obtain continuous level features representing the severity of distortion, this invention employs a ranking learning network. During training, image pairs with the same content and distortion type but different distortion levels are input into the network. The network first extracts degradation pattern features from the two images through a convolutional structure, and then outputs level codes through fully connected layers. By constraining the ranking loss to "the level code of a high-distortion image should be lower than that of a low-distortion image," the network gradually learns to construct a continuous distortion level axis consistent with the actual subjective quality. The final output level features can accurately reflect the comprehensive degradation magnitude, including blur degree, compression intensity, and noise amplitude, providing central information for the subsequent modulation process.
[0044] This process can be performed by the distortion level feature extraction module, as follows: Figure 4 As shown, this module employs a ranking learning network with two parallel branches. First, two images with the same content and distortion type but different distortion levels are selected as input. The degradation pattern features of the two input images are then extracted sequentially through the grade encoder, average pooling layer, and fully connected layer of each branch, outputting a quality grade code. These two parallel branches share weights.
[0045] In some embodiments, the ranking learning network employs a Siamese shared weight structure. This network includes two structurally identical feature extraction branches with shared parameters. Each branch receives an image and outputs a scalar score representing the severity of distortion. During training, two images with identical content and the same distortion type but different distortion levels are input into the two branches respectively. A ranking loss is constructed by comparing the output scores of the two branches, enabling the network to learn the relative order between distortion levels.
[0046] Each feature extraction branch employs a VGG-like convolutional network structure, comprising 5 convolutional stages and a total of 10 convolutional layers. Each convolutional stage sequentially performs feature extraction and downsampling, followed by a batch normalization layer and a LeakyReLU activation function. Taking an input image size of 256×256 and a base number of channels of 64 as an example, the number of network feature channels changes sequentially to 64, 128, 256, 512, and 256, with the spatial dimensions downsampled sequentially to 128×128, 64×64, 32×32, 16×16, and 8×8. A 256-dimensional feature vector is then obtained through global average pooling, followed by two fully connected layers (256→100→1), outputting a single distortion severity prediction score. The prediction scores from the two branches are further input into a ranking loss function to constrain the relative order between samples of different distortion levels.
[0047] The following sub-steps can be used during the execution process: S21. Extract degradation pattern features from two input images using the convolutional structure of the sorting learning network, and output quality level codes. During the output quality level coding process, the ranking loss constraint of the ranking learning network makes the score of the high-distortion original image output by the fully connected layer lower than the score of the low-distortion original image, and enables the ranking learning network to establish a monotonic relationship between the distortion degree of the input image and the quality level coding. S22. Execute sub-step S21 multiple times so that the ranking learning network can establish a monotonic relationship between the distortion level of the input image and the quality level coding. S23. The ranking learning network obtained by executing sub-step S22 obtains distortion successive level features that reflect the severity of distortion.
[0048] In the embodiment of step S3, the screen content image (input image) contains a large amount of structural information such as text, graphics, and lines with clear boundaries. To better capture these areas closely related to perceptual quality, the method of this invention performs domain-adaptive training on the visual backbone network, enabling it to highlight regional features such as text edges, stroke structures, and interface icon outlines. The visual feature sequence output by the network in the intermediate layers contains rich local structural textures, which helps in subsequent identification of phenomena such as blurred text, structural breaks, and noise interference caused by distortion.
[0049] For example, in some preferred embodiments, the input image is first divided into multiple image blocks, each image block is mapped to a feature vector, and position encoding is added to obtain an image feature sequence; The image feature sequence is then input into a multi-layer Transformer encoder of the visual backbone network. Visual features are extracted through a multi-head self-attention mechanism and a feedforward network to obtain a visual feature sequence.
[0050] In the embodiment executing step S4, since the distortion level is the dominant factor in subjective perception, and the distortion type manifests differently at different levels, the distortion level features obtained in step S2 are input into the modulation parameter generator of the modulation module, for example... Figure 5 As shown, this modulation parameter generator has parallel scaling and bias networks, which can respectively obtain a set of scaling factors and bias factors. The scaling factors are used to enhance or suppress the feature dimensions of different distortion types, while the bias factors are used to adjust the overall position of the distortion type features. These factors are along... Figure 5 The data flow shown collectively affects the distortion type features obtained in step S1, significantly enhancing the type dimension associated with significant image degradation when the distortion level is high, while weakening type differences when the distortion level is low. This aligns with the perceptual pattern that makes it difficult to distinguish types under mild distortion. The modulated type features better reflect the differences in human perception of distortion at different degrees of severity.
[0051] In other preferred embodiments, distortion level features can be input into several linear mapping layers. The overall architecture of these layers is similar to that of the modulation parameter generator, and they also have parallel scaling and bias networks that can generate scaling factors, bias factors, and weighting coefficients, respectively. Then, the distortion type features are linearly modulated using the scaling factors and bias factors, that is, the distortion type features are multiplied by the scaling factors by channel or element and then the bias factors are added. Finally, the weighting coefficients generated from the distortion continuous level features are used to perform a weighted transformation on the modulated type features to obtain a type feature representation that dynamically changes with the distortion level.
[0052] This process can be represented as follows: the distortion continuous-level features first pass through a linear layer to obtain the scaling factor γ and the offset factor β, and then the distortion type features z are processed...t Perform the following processing:
[0053] Simultaneously, weights w are generated from the distortion continuous level features to further weight the modulated features:
[0054] in, This refers to the type feature representation that dynamically changes with the distortion level. Through the above processing, the linear modulation module can make the distortion type feature controlled by the distortion level, enabling the model to generate different feature representations for different degrees of distortion.
[0055] In the above process, the scaling factor γ and the distortion type feature vector z t Multiplication is used to control the response intensity of different feature dimensions. When the corresponding γ for a certain dimension is large, the feature of that dimension is enhanced; when the corresponding γ is small, the feature of that dimension is suppressed. The offset factor β is added to the scaled feature to adjust the feature distribution by translation, so that the modulated type feature can adapt to different distortion levels.
[0056] Therefore, the combined effect of scaling and offset factors can be understood as follows: first, the various dimensions of the distortion type features are amplified or weakened; then, the overall feature distribution is shifted, ultimately resulting in a distortion type feature representation that dynamically changes with the distortion level. When the distortion level is high, the feature dimensions related to obvious degradation phenomena can be given a stronger response, thus making subsequent quality assessments focus more on distortion information that seriously affects subjective quality.
[0057] In step S5, to achieve deep interaction between the modulated type features and visual features, this invention employs a mutual attention mechanism. Visual features serve as the query, and the modulated type features serve as the key and value, enabling the network to automatically focus on the most quality-sensitive regions, such as text and graphic edges.
[0058] The fused contextual features contain both semantic degradation information and spatial structure information. Subsequently, the fused contextual features are input into a fully connected module. This module first flattens the fused features to obtain an image-level fused representation; then, it performs a nonlinear mapping on this image-level fused representation through at least one fully connected layer built into the module to enhance the feature responses related to perceived quality; finally, it uses the output layer to regress and obtain a quality prediction value consistent with the subjective rating.
[0059] The present invention also provides an embodiment for exemplarily demonstrating the effects of implementing the present invention.
[0060] (1) Training and testing process
[0061] The experiments were conducted on an NVIDIA RTX 3090 GPU. Training and testing of the model were performed using a deep learning framework, and adaptive moment estimation (IME) was employed to update the model parameters. During training, the learning rate was set to 0.0001, the batch size to 16, and the training epochs to 200. Cosine annealing was used to adjust the learning rate to ensure stability and convergence during training. To improve the reliability of the experimental results, all experiments were repeated ten times under different random seeds, and the median result was taken as the final result.
[0062] To comprehensively evaluate the performance of the distortion level-guided distortion type modulation screen content image quality assessment method proposed in this invention, the experiment adopted two commonly used metrics in the field of image quality assessment: Pearson Linear Correlation Coefficient (PLCC): which measures the linear consistency between predicted scores and subjective ratings; and Spearman Rank Order Correlation Coefficient (SRCC): which measures the consistency between predicted results and subjective rankings.
[0063] The experiments were conducted on two mainstream screen content image quality assessment datasets, SIQAD and SCID. Table 1 shows the comparison results between the method of this invention and existing representative algorithms on the above two datasets.
[0064] As clearly shown in Table 1, the present invention outperforms traditional handcrafted feature methods, deep learning methods, multi-task learning methods, and visual-language model methods in both PLCC and SRCC on the two datasets, achieving state-of-the-art performance. This fully demonstrates that the present invention, through explicit modeling of the dependency between distortion level and distortion type, can more accurately simulate the human quality perception mechanism, thereby significantly improving the accuracy and generalization ability of screen content image quality prediction.
[0065] Table 1. Quality evaluation results of different algorithms on the SIQAD and SCID test sets.
[0066] (2) Cross-dataset evaluation
[0067] To further evaluate the generalization ability of the method of this invention, the experiment adopted a cross-dataset testing approach, that is, training the model on the SIQAD dataset and testing it on the SCID dataset, and training it on the SCID dataset and testing it on SIQAD. Table 2 shows the quality prediction results across the datasets.
[0068] As shown in Table 2, the method proposed in this invention achieved excellent performance in transfer tests in both directions, significantly outperforming the comparative methods. This indicates that by modeling a semantic modulation mechanism dominated by distortion levels, this invention can effectively reduce content bias between different datasets and improve the model's cross-domain generalization ability.
[0069] Table 2 Results of Cross-Dataset Tests
[0070] Secondly, the present invention provides a screen content image evaluation system for performing distortion level-guided type modulation of the above-described method, such as... Figure 6 As shown, it includes: The distortion level feature extraction module 601 is used to: perform semantic similarity matching between the original image and semantic anchor points to obtain a distortion type feature vector with semantic distinguishability; The distortion type feature extraction module 602 is used to: input original images with the same content and the same distortion type feature vector but different distortion levels into the network model, and enable the network model to learn distortion continuous level features that reflect the severity of distortion through ranking constraints; The visual feature extraction module 603 is used to: perform domain-adaptive training on the visual backbone network so that the visual backbone network can adapt to the structural regions of the screen content image and output semantic visual features. The modulation module 604 is used to: perform linear modulation processing on the distortion continuous level features to obtain scaling factor and offset factor, and use the scaling factor and offset factor to obtain a type feature representation that dynamically changes with the distortion level; Attention mechanism fusion module 605 is used to: utilize the mutual attention mechanism, with semantic visual features as queries, and modulated distortion type features as keys and values, so that the network model focuses on the region most sensitive to quality. Fully connected layer 606 is used to perform regression operations on the fused features obtained by fusing semantic visual features and modulated distortion type features to obtain the subjective quality score prediction value of the screen content image. Output module 607 is used to output the predicted subjective quality score of the screen content image.
[0071] In summary, this invention addresses the problem that traditional evaluation methods struggle to effectively model the interdependence between distortion type and distortion level in screen content images. It proposes a distortion level-guided type modulation method and system for evaluating screen content images, belonging to the field of image quality assessment technology. This method first extracts distortion type features using a cue phrase-driven text-image semantic alignment model and then utilizes a ranking-based learning network to obtain level features for different distortion degrees. Simultaneously, the visual backbone network is fine-tuned to extract semantic visual features of the screen content image. Subsequently, the level features are linearly modulated to the type features, enabling the model to adaptively adjust the influence of various distortion information according to the severity of distortion. Finally, a mutual attention mechanism is used to fuse visual features and modulated type features to enhance the model's attention to sensitive areas such as text and graphics. This method effectively improves the accuracy and generalization ability of screen content image quality assessment under multiple distortion conditions, exhibiting high robustness and application value.
[0072] Those skilled in the art will understand that the accompanying drawings are merely schematic diagrams of one embodiment, and the modules or processes shown in the drawings are not necessarily essential for implementing the present invention.
[0073] As can be seen from the above description of the embodiments, those skilled in the art can clearly understand that the present invention can be implemented by means of software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of the present invention.
[0074] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, for apparatus or system embodiments, since they are basically similar to method embodiments, the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments. The apparatus and system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without creative effort.
[0075] The above description is merely a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A method for evaluating screen content images based on distortion level-guided type modulation, characterized in that, include: S1. Perform semantic similarity matching between the original image and semantic anchor points to obtain a distortion type feature vector with semantic distinguishability; Semantic anchors are obtained by constructing multiple descriptive cue words through a text-visual semantic alignment model driven by cue phrases; S2. Input the original images with the same content and the same distortion type feature vector but different distortion levels into the network model. Through the ranking constraint, the network model learns to obtain the distortion continuous level features that reflect the severity of distortion. S3. Perform domain-adaptive training on the visual backbone network so that it can adapt to the structural regions of the screen content image and output semantic visual features. S4. Perform linear modulation processing on the distortion continuous level features to obtain scaling factor and offset factor, and use the scaling factor and offset factor to obtain the type feature representation that dynamically changes with the distortion level. S5. Utilize the mutual attention mechanism, use semantic visual features as queries, and use modulated distortion type features as keys and values to enable the network model to focus on the regions most sensitive to quality. S6. The semantic visual features are fused with the modulated distortion type features, and the resulting fused features are input into the fully connected layer to obtain the subjective quality score prediction value of the screen content image. The predicted subjective quality score of the screen content image is used to adjust the processing method of the screen content image.
2. The screen content image evaluation method according to claim 1, characterized in that, The network model includes: The system comprises a distortion type feature extraction module, a distortion level feature extraction module, and a visual feature extraction module configured in parallel. The distortion type feature extraction module is used to perform step S1, the distortion level feature extraction module is used to perform step S2 and employs a ranking learning network, and the visual feature extraction module is used to perform step S3. The modulation module is used to perform step S4; the attention mechanism fusion module is used to perform step S5; and the fully connected layer is used to perform step S6.
3. The screen content image evaluation method according to claim 1 or 2, characterized in that, In step S1, the process of constructing multiple descriptive cue words to obtain semantic anchors through a text-visual semantic alignment model driven by cue words includes: The prompts used to describe different distortion phenomena are mapped into high-dimensional semantic vectors; Step S1 specifically includes: S11. Encode the input image to extract visual semantic features and high-dimensional semantic vectors corresponding to candidate distortion types; S12. Compare the visual semantic features with the high-dimensional semantic vector to calculate the similarity. S13. Normalize the obtained similarity values to obtain the probability distribution of the input image belonging to each candidate distortion type; S14. Select the candidate distortion type with the highest probability value from the probability distribution of the input image belonging to each candidate distortion type as the distortion type of the current input image, and use the high-dimensional semantic vector corresponding to the distortion type of the current input image as the distortion type feature vector with semantic distinguishing ability.
4. The screen content image evaluation method according to claim 3, characterized in that, Step S2 specifically includes: S21. Extract degradation pattern features from two input images using the convolutional structure of a sorting learning network, and use them as quality level codes. During the output quality level coding process, the ranking loss constraint of the ranking learning network makes the score of the high-distortion original image output by the fully connected layer lower than the score of the low-distortion original image, and enables the ranking learning network to establish a monotonic relationship between the distortion degree of the input image and the quality level coding. S22. Execute sub-step S21 multiple times so that the ranking learning network can establish a monotonic relationship between the distortion level of the input image and the quality level coding. S23. The ranking learning network obtained by executing sub-step S22 obtains distortion successive level features that reflect the severity of distortion.
5. The screen content image evaluation method according to claim 4, characterized in that, Step S3 specifically includes: The input image is divided into multiple image blocks, each image block is mapped to a feature vector, and positional encoding is added to obtain the image feature sequence; The image feature sequence is input into a multi-layer Transformer encoder of the visual backbone network, and visual features are extracted through a multi-head self-attention mechanism and a feedforward network to obtain a visual feature sequence. This visual feature sequence contains rich local structural textures, which are used to subsequently identify phenomena such as blurred text, structural breaks, and noise interference caused by distortion.
6. The screen content image evaluation method according to claim 5, characterized in that, Step S4 specifically includes: The distortion successive level features are input into the modulation parameter generator to obtain a set of scaling factors. and offset factor ; Set the distortion type feature vector as Through the formula scaling factor and offset factor Together they act on the distortion-type feature vector Obtain the modulated distortion type feature vector In the formula, ⊙ represents element-wise multiplication or channel-wise multiplication; scaling factor Used to enhance or suppress different types of feature dimensions, when the scaling factor corresponding to a certain dimension... When the value exceeds a preset threshold, the modulated distortion type feature vector It is enhanced, and conversely suppressed; the shift factor. Used to adjust the overall position of type features. w This indicates the weights generated by distorted continuous hierarchical features.
7. A screen content image evaluation system based on distortion level-guided type modulation, characterized in that, include: The distortion level feature extraction module is used to: perform semantic similarity matching between the original image and semantic anchor points to obtain a distortion type feature vector with semantic distinguishability; The distortion type feature extraction module is used to: input original images with the same content and the same distortion type feature vector but different distortion levels into the network model, and use ranking constraints to enable the network model to learn to obtain distortion continuous level features that reflect the severity of distortion. The visual feature extraction module is used to: perform domain-adaptive training on the visual backbone network, enabling the visual backbone network to adapt to the structural regions of the screen content image and output semantic visual features. The modulation module is used to: perform linear modulation processing on the distortion continuous level features to obtain scaling factors and offset factors, and use the scaling factors and offset factors to obtain a type feature representation that dynamically changes with the distortion level; The attention mechanism fusion module is used to: utilize the mutual attention mechanism, with semantic visual features as queries, and modulated distortion type features as keys and values, so that the network model focuses on the regions most sensitive to quality. The fully connected layer is used to perform regression operations on the fused features obtained by fusing semantic visual features with modulated distortion type features to obtain the predicted subjective quality score of the screen content image. The output module is used to output the predicted subjective quality score of the screen content image.