Natural scene image color quality evaluation method

By extracting the distortion, semantic and color features of natural scene images, combining nonlinear relationship modeling and loss functions, the quality evaluation model is trained, and the subjectivity and adaptability of image color quality evaluation in the prior art is solved, and more accurate and consistent automated evaluation is achieved.

CN120278960APending Publication Date: 2025-07-08QINGHAI NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510341023.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The existing image color quality evaluation methods have problems such as high subjectivity, inconsistent results, lack of analysis and insufficient adaptability of image semantic information, resulting in poor reliability and repeatability of evaluation results, making it difficult to adapt to the needs of different application scenarios.

Method used

By extracting the distortion characteristics, semantic characteristics and color characteristics of natural scene images, a quality evaluation model is used to model nonlinear relationships, and combining the color distribution loss and consistency loss function of natural scenes, the quality evaluation model is trained and verified to achieve automated evaluation.

Benefits of technology

It improves the objectivity and reliability of image color quality evaluation, and can dynamically adjust the evaluation standards according to different application scenarios, ensure the accuracy and consistency of evaluation results, and reduce the impact of artificial deviations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120278960A_ABST
    Figure CN120278960A_ABST
Patent Text Reader

Abstract

The invention relates to a natural scene image color quality evaluation method, which comprises the steps of obtaining sample data of a natural scene image, performing standardization processing on the sample data, and taking the standardized sample data as training data; extracting distortion features, semantic features and color features of the training data; carrying out nonlinear relation modeling on the distortion features, the semantic features, the color features and the perceptual quality scores, and training a quality evaluation model; after the training of the quality evaluation model is completed, verifying the quality evaluation model; after the verification of the quality evaluation model is passed, releasing the quality evaluation model; wherein the quality evaluation model takes natural scene color distribution loss and natural scene color consistency loss as loss functions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of image color quality evaluation, and particularly to a method for evaluating the color quality of natural scene images. Background Art

[0002] In the field of digital image processing, color quality assessment is an important research direction, especially in color transfer and generation of natural scene images. With the development of computer vision and deep learning, the application scenarios have been continuously expanded, including photographic color matching, video color correction, and color restoration, etc. Although great progress has been made in image quality evaluation technology, there are still some challenges and difficulties in specific color quality assessment: First, the evaluation methods often rely on the personal preferences and experiences of observers, resulting in a high degree of subjectivity in the results. This not only affects the consistency of the evaluation but also makes the results between different evaluators vary significantly, reducing the reliability and repeatability. Second, the existing methods usually ignore the internal relationship between color and semantic scenes, resulting in the evaluation results being unable to fully reflect the actual quality of the images. The lack of extraction and analysis of image semantic information makes the evaluation results out of touch with the scene requirements. In addition, the existing methods often adopt fixed evaluation criteria and are difficult to adapt to the different requirements of different application scenarios for color quality. And the lack of dynamic adjustment of color features leads to insufficient generality and practicality of the evaluation results in cross-scene applications. Summary of the Invention

[0003] To overcome at least to some extent the problem of inaccurate evaluation of image color quality in related technologies, this application provides a method for evaluating the color quality of natural scene images.

[0004] The solution of this application is as follows:

[0005] A method for evaluating the color quality of natural scene images, comprising:

[0006] Obtaining sample data of a natural scene image, performing normalization processing on the sample data, and using the normalized sample data as training data;

[0007] Extracting distortion features, semantic features, and color features of the training data;

[0008] Performing non-linear relationship modeling on the distortion features, semantic features, and color features with the perceived quality score to train a quality evaluation model;

[0009] After the quality evaluation model is trained, verifying the quality evaluation model;

[0010] After the quality evaluation model passes the verification, releasing the quality evaluation model;

[0011] Among them, the quality evaluation model uses the natural scene color distribution loss and the natural scene color consistency loss as loss functions.

[0012] Preferably, the method further includes:

[0013] Input the natural scene image to be evaluated into the quality evaluation model to obtain the quality evaluation score output by the quality evaluation model.

[0014] Preferably, the standardization process for the sample data includes:

[0015] Convert the sample data from the RGB color space to the LAB color space;

[0016] Normalize the sample data according to the mean and standard deviation of the sample data, and adjust the pixel values of the sample data to a preset range;

[0017] Crop or pad the sample data to a unified size.

[0018] Preferably, extracting the distortion features, semantic features, and color features of the training data includes:

[0019] Divide the feature maps of each scale into non-overlapping windows, and perform multi-head self-attention calculations within the windows to capture the local distortion, semantics, and color features of the natural scene images in the training data;

[0020] Pass information between different windows through the shifted window mechanism to capture the global distortion, semantics, and color features of the natural scene images in the training data;

[0021] After the self-attention calculation of each window, perform a non-linear transformation on the captured features through a multi-layer perceptron;

[0022] Fuse multi-scale features through 1x1 convolution and global average pooling operations to generate a multi-scale feature vector.

[0023] Preferably, modeling the non-linear relationship between the distortion features, semantic features, and color features and the perceptual quality score includes:

[0024] Model the non-linear relationship between the distortion features, semantic features, and color features and the perceptual quality score through a Kolmogorov - Arnold network;

[0025] Let \(x\in R\) 3 represent the input feature vector; the input feature vector includes: distortion features, semantic features, and color features;

[0026] The non-linear transformation through the Kolmogorov-Arnold network maps x to the mass fraction y ∈ R, and the Kolmogorov-Arnold network is expressed as:

[0027]

[0028] Among them, ψ ij (·) represents a continuous unary function for transforming the input features; x j , φ i (·) represents a continuous unary function for combining the transformed features; K represents the number of basis functions used in the network; n represents the dimension of the input features.

[0029] Preferably, the loss function of the natural scene color distribution loss includes:

[0030] Γ distribution = KL(P image ||P natural );

[0031] Among them, P image represents the color histogram of the input image; P natural represents the natural scene color distribution reference map; KL(·) represents the KL divergence, which is used to quantify the difference between the color distribution of the input image and the natural scene color distribution;

[0032] The loss function of the natural scene color consistency loss includes:

[0033]

[0034] Among them, represents the color gradient of the input image in the LAB color space; ||·||1 represents the L1 norm, which is used to minimize the gradient magnitude.

[0035] Preferably, the method further includes:

[0036] Assigning weights to the natural scene color distribution loss function and the natural scene color consistency loss function to obtain a comprehensive loss function:

[0037] Γ total = λ1Γ distribution + λ2Γ consistency ;

[0038] Among them, λ1 and λ2 respectively represent the weight parameters of the natural scene color distribution loss function and the natural scene color consistency loss function.

[0039] Preferably, validating the quality evaluation model includes:

[0040] Obtain natural scene images as the validation dataset;

[0041] Perform different types of degradation processing on the validation dataset to obtain validation subsets corresponding to each type of degradation processing;

[0042] Input the validation dataset and each validation subset into the quality evaluation model in sequence to obtain the quality evaluation scores of the validation dataset and each validation subset output by the quality evaluation model;

[0043] Verify the quality evaluation model through the quality evaluation scores of the validation dataset and each validation subset.

[0044] The technical solution provided by this application may include the following beneficial effects:

[0045] By extracting the distortion features, semantic features, and color features of the image, this application evaluates the quality of the image from multiple perspectives. Through the comprehensive extraction of these three types of features, it can ensure that the quality evaluation is more comprehensive and accurate, avoiding the limitations brought by single-feature analysis. By using non-linear relationship modeling, it is possible to capture more complex interaction effects between different features, thus obtaining more accurate quality scores. Since there may be different preferences or requirements for color quality in different application scenarios (such as photography, video color correction, etc.), the non-linear modeling method can flexibly adjust the weights of each feature according to the actual application requirements, avoiding the limitations of using a fixed evaluation standard. Based on the trained quality evaluation model, automated quality assessment can be achieved, avoiding the influence of human bias on the evaluation results and enhancing the objectivity and reliability of the evaluation process.

[0046] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit this application. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with this application, and are used together with the specification to explain the principles of this application.

[0048] Figure 1 is a flowchart showing a method for evaluating the color quality of natural scene images provided by an embodiment of this application;

[0049] Figure 2 is a comparison graph of the quality evaluation scores of a validation dataset and each validation subset provided by an embodiment of this application;

[0050] Figure 3 is a distribution graph of the quality evaluation scores of a validation dataset and each validation subset divided by group provided by an embodiment of this application;

[0051] Figure 4It is a distribution diagram of the quality evaluation scores of a verification data set divided by groups and each verification subset provided by an embodiment of the present application. Detailed implementation manners

[0052] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0053] Embodiment 1

[0054] Figure 1 It is a schematic flowchart of a method for evaluating the color quality of a natural scene image provided by an embodiment of the present application. Refer to Figure 1 , a method for evaluating the color quality of a natural scene image, includes:

[0055] S11: Obtain sample data of a natural scene image, perform standardization processing on the sample data, and use the standardized sample data as training data;

[0056] In this technical solution, a sample data set specifically for evaluating the color quality of natural scene images is constructed. This sample data set uses the natural scene colors adjusted by artists as reference pictures to ensure that the color quality of the images meets the high standards of human visual perception.

[0057] It should be noted that performing standardization processing on the sample data includes:

[0058] Convert the sample data from the RGB color space to the LAB color space;

[0059] Perform normalization processing on the sample data according to the mean and standard deviation of the sample data, and adjust the pixel values of the sample data to a preset range;

[0060] Crop or pad the sample data to a unified size.

[0061] First, convert the image from the RGB color space to the LAB color space.

[0062] The RGB color space (Red, Green, Blue) is a color model based on the additive principle and is usually used in display devices (such as monitors, cameras, etc.). The colors in the RGB space are represented by different intensity values of the three components of red, green, and blue, and have strong device dependence.

[0063] The LAB color space is a device-independent color space designed to be closer to the perceptual characteristics of the human visual system. The LAB space consists of three components: L* (luminance), a* (green - red contrast), and b* (blue - yellow contrast). The LAB color space is designed to be more in line with human visual perception, so it can better handle problems such as color distortion and brightness changes.

[0064] Although the RGB color space is intuitive, there is a strong correlation between its color channels, which is not conducive to the independent analysis of color features. In contrast, the LAB color space separates color information into luminance (L) and two chromaticity channels (A and B), and can better simulate the human visual system's perception of color. This transformation not only helps to capture color distribution and distortion information, but also significantly improves the accuracy of color quality assessment, especially when dealing with color contrast and color deviation problems.

[0065] Secondly, to ensure the stability and convergence of model training, the sample data is normalized according to the mean and standard deviation of the sample data, and the pixel values of the sample data are adjusted to a specific preset range. This processing can effectively reduce the deviation of data distribution and avoid the performance degradation caused by inconsistent input data scales during model training.

[0066] Perform a standardization operation on each pixel value of the sample data. Usually, by subtracting the mean of the data set and dividing by the standard deviation, the Z - score of the data is calculated. This process helps to eliminate the scale differences between different features and ensure that all input features have a similar numerical range.

[0067] For example, for the three channels L*, a*, and b* in the LAB color space, first calculate the mean and standard deviation of each channel, and then normalize each pixel:

[0068]

[0069] where x represents the pixel value, μ represents the mean, σ represents the standard deviation, and x normalized represents the value after normalization.

[0070] Finally, during the pre - processing process, the sample data is cropped or padded to a unified size to ensure the consistency of the input data. Since the sizes and aspect ratios of different images may vary greatly, directly inputting the model will lead to low computational efficiency and unstable performance. Through cropping and padding operations, the images can be adjusted to a fixed size, which not only retains the main features of the images but also avoids problems caused by inconsistent sizes.

[0071] Cropping: When the image size is larger than the preset target size, the central area or other important areas of the image can be selected for cropping to retain the information of the key parts in the image.

[0072] Padding: When the image size is smaller than the target size, the edges of the image are usually padded with zeros (or other padding colors) to expand it to the preset size.

[0073] S12: Extract the distortion features, semantic features, and color features of the training data;

[0074] In specific practice, Swin Transformer is adopted as the core model for image feature extraction. The advantage of Swin Transformer lies in its efficient multi-level feature extraction ability, which can gradually capture the detailed information and overall structure of the image color from local to global. At the same time, it can better model the global dependencies of the image through global context awareness, which is particularly important for tasks that need to comprehensively consider color distribution and semantic information in color quality assessment.

[0075] Based on this, extract the distortion features, semantic features, and color features of the training data, including:

[0076] Divide the feature maps of each scale into non-overlapping windows, and perform multi-head self-attention calculations within the windows to capture the local distortion, semantics, and color features of the natural scene images in the training data;

[0077] Pass information between different windows through the shifted window mechanism to capture the global distortion, semantics, and color features of the natural scene images in the training data;

[0078] After the self-attention calculation of each window, perform a non-linear transformation on the captured features through a multi-layer perceptron;

[0079] Fuse multi-scale features through 1x1 convolution and global average pooling operations to generate multi-scale feature vectors.

[0080] It should be noted that in the distortion feature extraction part, multi-scale feature maps are extracted through Swin Transformer, and the global and local information of the image is captured from different levels (such as Stage 2, Stage 3, Stage 4). The feature maps of each scale are divided into non-overlapping windows, and multi-head self-attention calculations are performed within the windows to capture local distortions (such as noise, blur, artifacts, etc.). By introducing the shifted window mechanism, the model can pass information between different windows, thereby capturing a larger range of distortion situations. After the self-attention calculation of each window, a multi-layer perceptron (MLP) is used to perform a non-linear transformation on the features, and multi-scale features are fused through 1x1 convolution and global average pooling operations to generate multi-scale distortion feature vectors. These feature vectors contain both local distortion information and retain the context information of global distortion, and finally are used as the input of the target network for distortion quality assessment.

[0081] For color features, the Swin Transformer is also used for extraction, and multi-scale features are combined to capture the global and local color distributions. Through the local color perception module, multi-scale feature maps are extracted from different stages of the Swin Transformer. The feature maps of each scale are divided into non-overlapping windows, and the relationship between pixels is calculated using the multi-head self-attention mechanism within the windows to capture color distortions in the local area (such as color contrast, saturation, exposure, etc.). Through the shifted window mechanism, the model can transfer information between different windows, thereby capturing a wider range of color distributions and color distortion situations.

[0082] For semantic features, the Swin Transformer is also used for extraction, which will not be elaborated here.

[0083] By fusing the distortion features, semantic features, and color features together, it is possible to simultaneously capture the local and global distortion information of the image, as well as the correlation between the color distribution and the semantic scene. It can dynamically adjust the color quality evaluation criteria according to different semantic scenes, making the evaluation results more in line with human visual perception. The hierarchical window mechanism and self-attention calculation of the Swin Transformer further enhance the global and local perception capabilities of the model, ensuring that the model can effectively capture the distortion information and color distribution of the image, thereby comprehensively improving the performance of image quality evaluation.

[0084] S13: Model the non-linear relationship between the distortion features, semantic features, and color features and the perceptual quality score to train the quality evaluation model;

[0085] In specific practice, the Kolmogorov - Arnold network (KAN) is introduced to model the complex non-linear relationship between the features (distortion features, semantic features, and color features) extracted from the Swin Transformer and the perceptual quality score. The Kolmogorov - Arnold representation theorem provides

[0086] a theoretical basis for the Kolmogorov - Arnold network. This theorem states that any multivariate continuous function can be represented as a superposition of continuous univariate functions and addition operations. This enables the Kolmogorov - Arnold network to efficiently approximate complex non-linear mappings. The structure of the Kolmogorov - Arnold network consists of multiple layers of non-linear transformations. The non-linear relationship between the distortion features, semantic features, and color features and the perceptual quality score is modeled through the Kolmogorov - Arnold network;

[0087] Let \(x\in\mathbb{R}\) 3 represent the input feature vector; the input feature vector includes: distortion features, semantic features, and color features;

[0088] The nonlinear transformation through the Kolmogorov - Arnold network maps x to the mass fraction y ∈ R, and the Kolmogorov - Arnold network is expressed as:

[0089]

[0090] where ψ ij (·) represents a continuous unary function used to transform the input features; x j , φ i (·) represents a continuous unary function used to combine the transformed features; K represents the number of basis functions used in the network; n represents the dimension of the input features.

[0091] The functions ψ ij (·) and φ i (·) are usually implemented by learnable neural networks. This enables the Kolmogorov - Arnold network to adaptively learn the optimal transformation of the input features during the training process. The nonlinear characteristics of these functions enable the network to capture the complex relationship between the input features and the mass fraction. In the framework proposed in this technical solution, the Kolmogorov - Arnold network is integrated with the feature extraction module based on Swin Transformer. The Swin Transformer extracts hierarchical features from the input image, and then inputs these features into the Kolmogorov - Arnold network for mass regression. This combination takes advantage of the strengths of both architectures: the Swin Transformer can capture global and local dependencies in the image, while the Kolmogorov - Arnold network can model complex nonlinear mappings.

[0092] S14: After the quality evaluation model is trained, verify the quality evaluation model;

[0093] S15: After the quality evaluation model passes the verification, release the quality evaluation model;

[0094] Among them, the quality evaluation model uses the natural scene color distribution loss and the natural scene color consistency loss as the loss functions.

[0095] It should be noted that in order to enhance the model's ability to evaluate the color quality of natural scenes, a comprehensive loss function is designed in this technical solution, which combines the Natural Scene Color Distribution Loss and the Color Consistency Loss. This loss function aims to guide the model to evaluate high-quality images that conform to the color characteristics of natural scenes.

[0096] The color distribution of natural scenes usually conforms to the Gaussian distribution statistical law. In order to ensure that the color distribution of the generated image is consistent with the color distribution of natural scenes, the Natural Scene Color Distribution Loss is introduced.

[0097] The loss function of the Natural Scene Color Distribution Loss includes:

[0098] Γ distribution = KL(P image ||P natural );

[0099] where P image represents the color histogram of the input image; P natural represents the reference map of the natural scene color distribution; KL(·) represents the KL divergence, which is used to quantify the difference between the color distribution of the input image and the color distribution of the natural scene;

[0100] The color transition in natural scenes is usually relatively smooth, and the color consistency in local areas is relatively high. In order to evaluate the smoothness of the color transition in the image, the Color Consistency Loss is introduced.

[0101] The loss function of the Color Consistency Loss includes:

[0102]

[0103] where represents the color gradient of the input image in the LAB color space; ||·||1 represents the L1 norm, which is used to minimize the gradient magnitude. By minimizing the Color Consistency Loss, the model can evaluate images with smooth color transitions.

[0104] Furthermore, weights are assigned to the Natural Scene Color Distribution Loss function and the Color Consistency Loss function to obtain the comprehensive loss function:

[0105] Γ total = λ1Γ distribution + λ2Γ consistency ;

[0106] where λ1 and λ2 respectively represent the weight parameters of the Natural Scene Color Distribution Loss function and the Color Consistency Loss function.

[0107] By adjusting the weight parameters, the trade-off between the color distribution and color consistency of the model can be flexibly controlled.

[0108] It should be noted that the method further includes:

[0109] Input the natural scene image to be evaluated into the quality evaluation model to obtain the quality evaluation score output by the quality evaluation model.

[0110] After the quality evaluation model is released and put into use, the natural scene image to be evaluated can be input into the quality evaluation model to obtain the quality evaluation score of the natural scene image to be evaluated through the quality evaluation model.

[0111] Embodiment 2

[0112] It should be noted that validating the quality evaluation model includes:

[0113] Obtain natural scene images as the validation dataset;

[0114] Perform different types of degradation processing on the validation dataset to obtain validation subsets corresponding to each type of degradation processing;

[0115] Input the validation dataset and each validation subset into the quality evaluation model in sequence to obtain the quality evaluation scores of the validation dataset and each validation subset output by the quality evaluation model;

[0116] Validate the quality evaluation model through the quality evaluation scores of the validation dataset and each validation subset.

[0117] In specific practice, obtain natural scene images as the validation dataset, perform different types of degradation processing on the validation dataset to obtain validation subsets corresponding to each type of degradation processing. For example, obtain 3625 benchmark pictures as the validation dataset, perform 5 different types of color degradation processing (such as contrast reduction, saturation change, color cast, etc.) on the benchmark pictures, generate 18125 degraded images and divide them into 5 validation subsets according to the degradation processing categories.

[0118] Such as Figure 2 As shown, input the validation dataset and each validation subset into the quality evaluation model in sequence to obtain the quality evaluation scores of the validation dataset and each validation subset output by the quality evaluation model. Figure 2 It includes the original image (Original) and 5 degraded images (Type1 to Type5), and the vertical axis represents the quality score. From Figure 2It can be seen that the quality evaluation model performs excellently on images of different degradation types and can effectively evaluate the color quality of images. Specifically, the quality score of the original image is 0.7137, which serves as a baseline reference; the quality scores of Type1 and Type2 images are 0.6841 and 0.6836 respectively, slightly lower than the original image, indicating a slight decline in their color quality; the quality scores of Type3 and Type4 images are 0.6707 and 0.6539 respectively, showing a further decline, indicating poor color quality; while the quality score of Type5 image is 0.7047, close to the original image, indicating good color quality. The experimental results show that this technical solution can effectively distinguish images with different color qualities and is consistent with the visual perception results, verifying the effectiveness and practicality of this technical solution.

[0119] As Figure 3 shown, the original image and 5 types of degraded images are each divided into three groups (Ns_42, Ns_87, Ns_91). The horizontal axis represents the image color degradation type, including the original image (Original) and 5 types of degraded images (Type1 to Type5), and the vertical axis represents the quality score, ranging from 0.5 to 1.0. It can be seen from the figure that the original image has the highest quality score, while the quality scores of the adjusted image types (Type1 to Type5) are all lower than that of the original image. This trend indicates that the quality of the degraded images is generally lower than that of the original image. The experimental results verify that this technical solution can effectively distinguish images with different color qualities and is consistent with the visual perception results.

[0120] As Figure 4 shown, the original image and 5 types of degraded images are each divided into three groups (Group 42, Group 87, Group 91). Figure 4 shows the comparison of the quality scores of the original image and 5 types of degraded images under different groups (Group 42, Group 87, Group 97), further presenting the quality score results of the images divided by group and type. The horizontal axis represents the images of different groups (Group 42, Group 87, Group 97), and the vertical axis represents the quality score. The black broken line represents the original image, and the blue, orange, green, red, and purple broken lines represent the five adjusted image types respectively. From Figure 4 it can be seen that in any group, the quality score of the original image is higher than that of the adjusted image types. In addition, the quality score trends among different groups are basically the same, indicating that this technical solution has stable performance in different groups. This result further verifies the generalization ability and reliability of the method.

[0121] It can be understood that the same or similar parts in the above embodiments can be referred to each other, and for the content not detailed in some embodiments, reference can be made to the same or similar content in other embodiments.

[0122] It should be noted that in the description of the present application, the terms "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance. In addition, in the description of the present application, unless otherwise specified, the meaning of "a plurality" refers to at least two.

[0123] Any process or method description in the flowchart or described in other ways herein can be understood to represent a module, segment, or part of code including one or more executable instructions for implementing a specific logical function or process. The scope of the preferred embodiments of the present application includes additional implementations, where the functions can be executed in a substantially simultaneous manner or in the reverse order according to the involved functions, rather than in the order shown or discussed, which should be understood by those skilled in the technical field of the embodiments of the present application.

[0124] It should be understood that each part of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following well-known technologies in the art can be used: discrete logic circuits with logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits with appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0125] Those of ordinary skill in the technical field of the present application can understand that all or part of the steps carried by the methods of the above embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.

[0126] In addition, each functional unit in various embodiments of the present application can be integrated into a processing module, or each unit can exist physically alone, or two or more units can be integrated into one module. The above integrated module can be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.

[0127] The above-mentioned storage medium can be a read-only memory, a magnetic disk, an optical disk, etc.

[0128] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc., mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of this application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in a suitable manner in any one or more embodiments or examples.

[0129] Although the embodiments of this application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limitations on this application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of this application.

Claims

1. A method for evaluating the color quality of natural scene images, characterized in that, Including: Obtain sample data of natural scene images, perform normalization processing on the sample data, and use the normalized sample data as training data; Extract the distortion features, semantic features, and color features of the training data; Perform non-linear relationship modeling on the distortion features, semantic features, and color features with the perceived quality score to train a quality evaluation model; After the quality evaluation model is trained, verify the quality evaluation model; After the quality evaluation model passes the verification, release the quality evaluation model; Among them, the quality evaluation model uses the natural scene color distribution loss and the natural scene color consistency loss as loss functions.

2. The method according to claim 1, characterized in that The method further includes: Input the natural scene image to be evaluated into the quality evaluation model to obtain the quality evaluation score output by the quality evaluation model.

3. The method according to claim 1, wherein Performing normalization processing on the sample data includes: Convert the sample data from the RGB color space to the LAB color space; Normalize the sample data according to the mean and standard deviation of the sample data, and adjust the pixel values of the sample data to a preset range; Crop or pad the sample data to a unified size.

4. The method according to claim 1, wherein Extracting the distortion features, semantic features, and color features of the training data includes: Divide the feature map of each scale into non-overlapping windows, and perform multi-head self-attention calculation within the windows to capture the local distortion, semantics, and color features of the natural scene images in the training data; Use the shifted window mechanism to transfer information between different windows to capture the global distortion, semantics, and color features of the natural scene images in the training data; After the self-attention calculation of each window, perform non-linear transformation on the captured features through a multi-layer perceptron; Fuse multi-scale features through 1x1 convolution and global average pooling operations to generate a multi-scale feature vector.

5. The method according to claim 1, wherein Performing non-linear relationship modeling on the distortion features, semantic features, and color features with the perceived quality score includes: Perform non-linear relationship modeling on the distortion features, semantic features, and color features with the perceived quality score through a Kolmogorov-Arnold network; Let \(x\in R\) 3 represents an input feature vector; the input feature vector includes: a distortion feature, a semantic feature, and a color feature; If the Kolmogorov-Arnold network performs non-linear transformation to map x to the quality score y∈R, the Kolmogorov-Arnold network is expressed as: Among them, ψ ij (·) represents a continuous unary function for transforming input features; x j , φ i (·) represents a continuous unary function for combining the transformed features; K represents the number of basis functions used in the network; n represents the dimension of the input features.

6. The method according to claim 1, characterized in that, The loss function of the natural scene color distribution loss includes: Γ distribution = KL(P image ||P natural ); Among them, P image represents the color histogram of the input image; P natural represents the reference map of the natural scene color distribution; KL(·) represents the KL divergence, which is used to quantify the difference between the color distribution of the input image and the natural scene color distribution; The loss function of the natural scene color consistency loss includes: Among them, represents the color gradient of the input image in the LAB color space; ||·||1 represents the L1 norm, which is used to minimize the gradient magnitude.

7. The method according to claim 6, wherein The method further includes: Assign weights to the natural scene color distribution loss function and the natural scene color consistency loss function to obtain a comprehensive loss function: Γ total = λ1Γ distribution + λ2Γ consistency ; Among them, λ1 and λ2 respectively represent the weight parameters of the natural scene color distribution loss function and the natural scene color consistency loss function.

8. The method according to claim 1, wherein Verifying the quality evaluation model includes: Obtain natural scene images as a validation dataset; Perform different types of degradation processing on the validation dataset to obtain validation subsets corresponding to each degradation processing type; Input the validation dataset and each validation subset into the quality evaluation model in turn to obtain the quality evaluation scores of the validation dataset and each validation subset output by the quality evaluation model; Verify the quality evaluation model by means of the quality evaluation scores of the verification dataset and each verification subset.