An image quality determination method, apparatus, device and medium

By constructing a dual-layer recognition system of subject semantics and visual aesthetics, and analyzing the visual focus and feature quantification information of images, this solves the problem that existing image quality determination methods cannot determine the effectiveness of content delivery, and ensures the effectiveness of image quality and visual appeal.

CN121616606BActive Publication Date: 2026-05-29E-JOINED INTERNET & TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
E-JOINED INTERNET & TECH CO LTD
Filing Date
2026-02-03
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing methods for determining image quality are based solely on easily quantifiable basic parameters, which cannot effectively assess the effectiveness of image data in conveying content and its visual appeal in real-world scenarios, thus failing to guarantee the effectiveness of information dissemination.

Method used

A two-layer progressive recognition system based on subject semantics and visual aesthetics is constructed. By acquiring target image and scene information, analyzing visual focus pixel regions, intent information and feature quantification information, and combining with a deep convolutional neural network model, image quality is determined.

Benefits of technology

It enables the parsing of the expressive content of image data, ensuring the effectiveness of content delivery and visual appeal with a certain image quality, and ensuring the efficient dissemination of images in different application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121616606B_ABST
    Figure CN121616606B_ABST
Patent Text Reader

Abstract

The application discloses an image quality determination method, device, equipment and medium, the method comprises the following steps: acquiring a target image and image scene information, performing subject object analysis on the target image to obtain a visual focus pixel area; performing semantic analysis on the target image to obtain intention information, and determining a target proportion threshold based on the intention information and the image scene information; in the case that the ratio of the number of pixels in the visual focus pixel area to the total number of pixels of the target image exceeds the target proportion threshold, performing quantitative analysis processing on the target image to obtain feature quantitative information; based on the feature quantitative information and the preset quality evaluation condition, the image quality of the target image is determined. The technical scheme constructs a double-layer progressive identification system of subject semantics and visual aesthetics, realizes expression content analysis of image data, and guarantees the content transmission effectiveness and visual attraction of the identified image quality determination.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the general field of image data processing or generation technology, and specifically relates to an image quality determination method, apparatus, device and medium. Background Technology

[0002] With the rapid development of network communication technology and the multimedia information industry, images, as an intuitive and efficient information carrier, have become a core means of information dissemination in scenarios such as product introductions, knowledge popularization, and social sharing. High-quality image data is image data that can significantly improve attention and information transmission efficiency. In various information dissemination scenarios, pre-determining image quality and selecting high-quality image data while avoiding the misuse of low-quality image data are key prerequisites for ensuring content attractiveness and achieving efficient information transmission.

[0003] Existing methods for determining image quality often focus on easily quantifiable basic parameters such as resolution, file size, aspect ratio, basic noise level, and format compatibility. However, simple identification based on these basic parameters can only quickly and mechanically filter the usability of image data, excluding severely blurry, undersized, or corrupted images. It cannot determine the effectiveness of image data in conveying content in real-world scenarios, and therefore cannot guarantee that it achieves the expected visual appeal and information dissemination effects. Summary of the Invention

[0004] This application provides an image quality determination method, apparatus, device, and medium, aiming to achieve the analysis of the expressive content of image data by constructing a two-layer progressive recognition system of subject semantics and visual aesthetics, thereby ensuring the effectiveness of the content delivery and visual appeal of the identified image quality determination.

[0005] In a first aspect, this application provides an image quality determination method, the method comprising:

[0006] Acquire target image and image scene information, and perform subject object analysis on the target image to obtain the visual focus pixel region;

[0007] Semantic analysis is performed on the target image to obtain intent information, and a target proportion threshold is determined based on the intent information and the image scene information;

[0008] If the ratio of the number of pixels in the visual focus pixel region to the total number of pixels in the target image exceeds the target proportion threshold, the target image is subjected to quantitative analysis to obtain feature quantization information; wherein, the feature quantization information includes color quantization information, expression quantization information, and style quantization information;

[0009] Based on the feature quantification information and preset quality assessment conditions, the image quality of the target image is determined.

[0010] Optionally, the step of performing semantic analysis on the target image to obtain intent information includes:

[0011] The target image is semantically segmented to obtain semantic labels and the pixel regions corresponding to the semantic labels;

[0012] The positions of the visual focus pixel region and the pixel region corresponding to the semantic label are matched, and the semantic label with the highest proportion in the visual focus pixel region is determined as the main semantic label;

[0013] Intent information is determined based on the spatial location and relative scale information of the visual focal pixel region and the subject semantic label.

[0014] Optionally, determining the intent information based on the spatial location information and relative scale information of the visual focal pixel region and the subject semantic label includes:

[0015] Candidate visual expression templates are determined based on the subject semantic tags and the association between the pre-constructed subject semantic tags and the preset visual expression templates; wherein, the visual expression templates include visual focus spatial location features and visual focus relative scale features;

[0016] Based on the spatial location information and relative scale information of the visual focal pixel region, the similarity between the visual focal pixel region and the candidate visual expression template is determined;

[0017] The candidate visual expression template with the highest similarity is identified as the intent information.

[0018] Optionally, determining the target proportion threshold based on the intent information and the image scene information includes:

[0019] Determine the basic proportion threshold based on the image scene information;

[0020] The proportion correction coefficient is determined based on the intent information, and the confidence level of the proportion correction coefficient is determined based on the similarity corresponding to the intent information.

[0021] The target percentage threshold is determined based on the basic percentage threshold, the percentage correction coefficient, and the confidence level.

[0022] Optionally, the step of performing quantization analysis on the target image to obtain feature quantization information includes:

[0023] The average saturation, the difference in the number of warm and cool pixels, and the maximum brightness difference of the target image are calculated, and the average saturation, the difference in the number of warm and cool pixels, and the maximum brightness difference are weighted and summed to obtain color quantization information.

[0024] Calculate the average color difference and average edge intensity between the visual focal pixel region and other pixel regions in the target image, and perform a weighted summation of the average color difference and average edge intensity to obtain the quantitative information of expression.

[0025] The target image is input into a pre-trained deep convolutional neural network model to obtain a target style feature vector. The target style feature vector is then matched with a large data vector library to obtain style quantification information.

[0026] Optionally, matching the target style feature vector with a large data vector library to obtain style quantification information includes:

[0027] Calculate the average vector distance between the target style feature vector and all style feature vectors in the big data vector library;

[0028] The average vector distance between each style feature vector in the big data vector library and all other style feature vectors is calculated as the baseline average vector distance.

[0029] The average vector distance is sorted with the benchmark average vector distance, and the position distribution percentage is determined as style quantification information based on the sorting position of the average vector distance in the sorting results.

[0030] Optionally, the step of performing subject object analysis on the target image to obtain the visual focus pixel region includes:

[0031] Generate a saliency heatmap corresponding to the target image, and select candidate focal regions in the saliency heatmap according to a preset heatmap threshold;

[0032] If the overlapping area of ​​the bounding rectangles of two candidate focal regions exceeds a preset area threshold, the bounding rectangles of the two candidate focal regions are weighted and merged according to the total regional thermal value of the two candidate focal regions to obtain a merged bounding rectangle.

[0033] The overlapping area between the two candidate focal regions and the fused bounding rectangle is determined as the new candidate focal region, and the original two candidate focal regions are deleted.

[0034] If the overlapping area of ​​the bounding rectangles of any two candidate focal regions does not exceed a preset area threshold, the candidate focal region is determined as the visual focal pixel region.

[0035] Secondly, this application provides an image quality determination apparatus, the apparatus comprising:

[0036] The visual focus determination module is used to acquire target image and image scene information, and to perform subject object analysis on the target image to obtain the visual focus pixel region;

[0037] The target proportion determination module is used to perform semantic analysis on the target image to obtain intent information, and determine the target proportion threshold based on the intent information and the image scene information;

[0038] The quantization feature determination module is used to perform quantization analysis on the target image to obtain feature quantization information when the ratio of the number of pixels in the visual focus pixel region to the total number of pixels in the target image exceeds the target proportion threshold; wherein, the feature quantization information includes color quantization information, expression quantization information, and style quantization information;

[0039] The image quality determination module is used to determine the image quality of the target image based on the feature quantization information and preset quality assessment conditions.

[0040] Optionally, the target proportion determination module is specifically used for:

[0041] The target image is semantically segmented to obtain semantic labels and the pixel regions corresponding to the semantic labels;

[0042] The positions of the visual focus pixel region and the pixel region corresponding to the semantic label are matched, and the semantic label with the highest proportion in the visual focus pixel region is determined as the main semantic label;

[0043] Intent information is determined based on the spatial location and relative scale information of the visual focal pixel region and the subject semantic label.

[0044] Optionally, the target proportion determination module is specifically used for:

[0045] Candidate visual expression templates are determined based on the subject semantic tags and the association between the pre-constructed subject semantic tags and the preset visual expression templates; wherein, the visual expression templates include visual focus spatial location features and visual focus relative scale features;

[0046] Based on the spatial location information and relative scale information of the visual focal pixel region, the similarity between the visual focal pixel region and the candidate visual expression template is determined;

[0047] The candidate visual expression template with the highest similarity is identified as the intent information.

[0048] Optionally, the target proportion determination module is specifically used for:

[0049] Determine the basic proportion threshold based on the image scene information;

[0050] The proportion correction coefficient is determined based on the intent information, and the confidence level of the proportion correction coefficient is determined based on the similarity corresponding to the intent information.

[0051] The target percentage threshold is determined based on the basic percentage threshold, the percentage correction coefficient, and the confidence level.

[0052] Optionally, the quantization feature determination module is specifically used for:

[0053] The average saturation, the difference in the number of warm and cool pixels, and the maximum brightness difference of the target image are calculated, and the average saturation, the difference in the number of warm and cool pixels, and the maximum brightness difference are weighted and summed to obtain color quantization information.

[0054] Calculate the average color difference and average edge intensity between the visual focal pixel region and other pixel regions in the target image, and perform a weighted summation of the average color difference and average edge intensity to obtain the quantitative information of expression.

[0055] The target image is input into a pre-trained deep convolutional neural network model to obtain a target style feature vector. The target style feature vector is then matched with a large data vector library to obtain style quantification information.

[0056] Optionally, the quantization feature determination module is specifically used for:

[0057] Calculate the average vector distance between the target style feature vector and all style feature vectors in the big data vector library;

[0058] The average vector distance between each style feature vector in the big data vector library and all other style feature vectors is calculated as the baseline average vector distance.

[0059] The average vector distance is sorted with the benchmark average vector distance, and the position distribution percentage is determined as style quantification information based on the sorting position of the average vector distance in the sorting results.

[0060] Optionally, the visual focus determination module is specifically used for:

[0061] Generate a saliency heatmap corresponding to the target image, and select candidate focal regions in the saliency heatmap according to a preset heatmap threshold;

[0062] If the overlapping area of ​​the bounding rectangles of two candidate focal regions exceeds a preset area threshold, the bounding rectangles of the two candidate focal regions are weighted and merged according to the total regional thermal value of the two candidate focal regions to obtain a merged bounding rectangle.

[0063] The overlapping area between the two candidate focal regions and the fused bounding rectangle is determined as the new candidate focal region, and the original two candidate focal regions are deleted.

[0064] If the overlapping area of ​​the bounding rectangles of any two candidate focal regions does not exceed a preset area threshold, the candidate focal region is determined as the visual focal pixel region.

[0065] Thirdly, this application provides an electronic device including a processor, a memory, and a program or instructions stored in the memory and executable on the processor, wherein the program or instructions, when executed by the processor, implement the method described in the first aspect.

[0066] Fourthly, this application provides a readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the method described in the first aspect.

[0067] In this application, a target image and image scene information are acquired. Subject object analysis is performed on the target image to obtain a visual focus pixel region. Semantic analysis is performed on the target image to obtain intent information, and a target proportion threshold is determined based on the intent information and the image scene information. If the ratio of the number of pixels in the visual focus pixel region to the total number of pixels in the target image exceeds the target proportion threshold, quantitative analysis is performed on the target image to obtain feature quantization information. The feature quantization information includes color quantization information, expression quantization information, and style quantization information. Based on the feature quantization information and preset quality assessment conditions, the image quality of the target image is determined. This image quality determination method, by constructing a two-layer progressive recognition system of subject semantics and visual aesthetics, achieves the parsing of the expressive content of image data, ensuring the effectiveness of the content delivery and visual appeal of the identified image quality determination. Attached Figure Description

[0068] Figure 1 This is a flowchart illustrating an image quality determination method provided in an embodiment of this application;

[0069] Figure 2 This is a flowchart illustrating another image quality determination method provided in an embodiment of this application;

[0070] Figure 3This is an example diagram showing the position matching results of semantic tags and visual focus pixel regions provided in the embodiments of this application;

[0071] Figure 4 This is a flowchart illustrating another image quality determination method provided in an embodiment of this application;

[0072] Figure 5 This is a schematic diagram of the structure of an image quality determination device provided in an embodiment of this application;

[0073] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0074] To make the objectives, technical solutions, and advantages of this application clearer, specific embodiments of this application will be described in further detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are merely for explaining this application and not for limiting it. It should also be noted that, for ease of description, only the parts relevant to this application are shown in the drawings, not all of them. Before discussing exemplary embodiments in more detail, it should be mentioned that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts describe operations (or steps) as sequential processes, many of these operations can be performed in parallel, concurrently, or simultaneously. Furthermore, the order of the operations can be rearranged. The process can be terminated when its operation is completed, but may also have additional steps not included in the drawings. The process can correspond to a method, function, procedure, subroutine, subprogram, etc.

[0075] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.

[0076] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0077] The image quality determination method, apparatus, device, and medium provided in this application will be described in detail below with reference to the accompanying drawings and through specific embodiments and application scenarios.

[0078] First, this application applies to scenarios that have clear requirements for the core expressive effect and visual quality of images, such as compliance testing of e-commerce product images, quality rating of photographic works, and review of advertising image materials. It is understood that the subject of this application can be a terminal device with image processing capabilities.

[0079] Figure 1 This is a flowchart illustrating an image quality determination method provided in an embodiment of this application. Figure 1 As shown, the specific steps include the following:

[0080] S101, acquire the target image and image scene information, and perform subject object analysis on the target image to obtain the visual focus pixel region.

[0081] The target image can be image data whose quality needs to be determined. Correspondingly, the image scene information can be the specific application scenario of the target image. Specifically, the image scene information varies depending on the scenario in which the image quality is determined. For example, if the image quality determination scenario is for the review of advertising image materials, the corresponding image scene information could be the product type, which could include skincare, digital products, food, clothing, and bags, etc.

[0082] In one embodiment, the target image and image scene information can be obtained by having the user upload the target image and image scene information.

[0083] Among them, the visual focus pixel region can be the core pixel region in the target image that can first attract human visual attention.

[0084] In one embodiment, the method of obtaining the visual focus pixel region by performing subject object analysis on the target image can be to locate the subject object in the target image through edge detection, contour extraction and region growing, and then determine its corresponding pixel region as the visual focus pixel region.

[0085] In one embodiment, performing subject object analysis on the target image to obtain the visual focus pixel region includes: generating a saliency heatmap corresponding to the target image, and selecting candidate focus regions in the saliency heatmap according to a preset heatmap threshold; if the overlap area of ​​the bounding rectangles of two candidate focus regions exceeds a preset area threshold, weighted fusion of the bounding rectangles of the two candidate focus regions according to the total heatmap value of the two candidate focus regions to obtain a fused bounding rectangle; determining the overlap area between the two candidate focus regions and the fused bounding rectangle as a new candidate focus region and deleting the original two candidate focus regions; if the overlap area of ​​the bounding rectangles of any two candidate focus regions does not exceed a preset area threshold, determining the candidate focus region as the visual focus pixel region.

[0086] Among them, the saliency heatmap can be an image that represents the visual saliency of each pixel in the target image in the form of a two-dimensional matrix. The value of each pixel in the matrix (heat value) corresponds to the degree to which the pixel attracts human visual attention. The higher the value, the more likely the pixel is to become the visual focus.

[0087] In one embodiment, the method for generating a saliency heatmap corresponding to a target image can be to input the target image into a pre-trained deep learning-based saliency detection model (e.g., U-2-Net, DUTS-TR pre-trained model). The saliency detection model extracts multi-scale semantic features of the target image (e.g., texture features, edge features, color contrast features, etc.) and performs weighted fusion calculation on the multi-scale semantic features to obtain the saliency heatmap.

[0088] The preset thermal threshold can be a critical value used to filter pixels with sufficient significance. For example, when the thermal value range is 0-255, the preset thermal threshold can be set to 120.

[0089] The candidate focal region can be a set of connected pixels in the saliency heatmap where all pixel values ​​are higher than a preset heatmap threshold.

[0090] In one embodiment, the method of selecting candidate focal regions in a salient heatmap based on a preset thermal threshold can be achieved by marking pixels with thermal values ​​higher than the preset thermal threshold as foreground pixels (value 1) and pixels with thermal values ​​lower than or equal to the preset thermal threshold as background pixels (value 0) to obtain a binarized image. Then, a connected component analysis algorithm is used to identify all connected foreground regions in the binarized image to obtain candidate focal regions.

[0091] The bounding rectangle of the candidate focal region can be the smallest rectangle that can completely enclose a single candidate focal region; the overlapping area of ​​the bounding rectangles can be the total number of pixels of the overlapping part of the two bounding rectangles in the image coordinate system; the preset area threshold can be a critical value used to determine whether there is significant overlap between the two candidate focal regions, and can be set based on the area ratio of the bounding rectangles, for example, set to 30% of the area of ​​the smaller bounding rectangle of the two bounding rectangles.

[0092] The total thermal value of a candidate focal region can be the sum of the thermal values ​​of all pixels within a single candidate focal region.

[0093] The fused circumscribed rectangle can be a new rectangle obtained by weighting and adjusting the circumscribed rectangles of two candidate focal regions that have significant overlap.

[0094] In one embodiment, the method of weightedly fusing the bounding rectangles of the two candidate focal regions based on the total regional thermal values ​​of the two candidate focal regions to obtain the fused bounding rectangle can be achieved by assigning weight coefficients to the bounding rectangles of the two candidate focal regions with a sum of 1 based on the total regional thermal values ​​of the two candidate focal regions, and then calculating the coordinates of the corresponding angles of the fused bounding rectangle by weighted summation of the coordinates of the corresponding angles of the bounding rectangles of the two candidate focal regions.

[0095] In one embodiment, the method of determining the overlapping area of ​​two candidate focal regions and the fused outer rectangle as the new candidate focal regions and deleting the original two candidate focal regions can be achieved by calculating the intersection area of ​​the two candidate focal regions and the fused outer rectangle respectively, determining the union of the two intersection areas as the new candidate focal regions, and deleting the pixel markers corresponding to the original two candidate focal regions to complete the update of the candidate focal regions.

[0096] In one embodiment, if the overlapping area of ​​the bounding rectangles of any two candidate focal regions does not exceed a preset area threshold, the candidate focal region can be determined as the visual focal pixel region by traversing all candidate focal regions and calculating the overlapping area of ​​their bounding rectangles pairwise. If all overlapping areas do not exceed the preset area threshold, it indicates that each candidate focal region is independent and has no significant interference, and it can be determined as the final visual focal pixel region.

[0097] The advantage of this approach is that it ensures that the selection of visual focus conforms to the laws of human visual attention. By using weighted fusion of overlapping regions, it solves the problem of blurred focus determination caused by the overlap of multiple candidate regions, and avoids the core focus being split or missed.

[0098] S102, semantic analysis is performed on the target image to obtain intent information, and the target proportion threshold is determined based on the intent information and the image scene information.

[0099] Intent information can be the core expressive intent carried by the target image. Specifically, the corresponding intent information will vary depending on the scenario in which the image quality is determined. For example, if the scenario in which the image quality is determined is the review of advertising image materials, the corresponding intent information can be the intent to showcase product selling points, which may include highlighting the product itself, emphasizing the atmosphere of the scene, or highlighting both the product and the scene.

[0100] In one embodiment, the method of obtaining intent information by semantic analysis of the target image can be to input the target image into a pre-trained image semantic understanding model (such as CLIP or ViT-B / 32 model), and have the image semantic understanding model directly output the intent information.

[0101] Among them, the target proportion threshold can be a pixel proportion critical value determined based on intent information and image scene information, used to determine whether the visual focus pixel region meets the core expression requirements.

[0102] In one embodiment, determining the target percentage threshold based on intent information and image scene information can be achieved by pre-constructing a correlation between intent information, image scene information, and percentage threshold. The stored data of this correlation is then queried using the current intent information and image scene information as query conditions. The resulting query includes the target percentage threshold. For example, in the case of image quality determination in the context of advertising image material review, if the intent information is to highlight the product and the image scene information is digital, the target percentage threshold is 85%; if the intent information is to highlight both the product and the scene and the image scene information is food, the target percentage threshold is 45%; and if the intent information is to highlight the scene atmosphere and the image scene information is skincare, the target percentage threshold is 35%.

[0103] S103, if the ratio of the number of pixels in the visual focus pixel region to the total number of pixels in the target image exceeds the target proportion threshold, the target image is subjected to quantitative analysis processing to obtain feature quantization information; wherein, the feature quantization information includes color quantization information, expression quantization information and style quantization information.

[0104] If the ratio of the number of pixels in the visual focus pixel region to the total number of pixels in the target image exceeds the target proportion threshold, it indicates that the core subject region of the target image is sufficiently prominent and can effectively convey the core expressive intent of the target image, thus meeting the basic conditions for image quality assessment. Therefore, further quantitative analysis of the deep features of the target image can be performed. If the ratio of the number of pixels in the visual focus pixel region to the total number of pixels in the target image does not exceed the target proportion threshold, the image quality of the target image is directly determined to be low quality.

[0105] Feature quantization information can be a set of parameters that represent the visual expressive features of a target image in a numerical manner, and may include color quantization information, expression quantization information, and style quantization information. Specifically, color quantization information can be quantization parameters used to describe the color expression effect of the target image; expression quantization information can be quantization parameters used to describe the expression effect of the core subject of the target image; and style quantization information can be quantization parameters used to describe the style prominence effect of the target image.

[0106] In one embodiment, the method for obtaining feature quantization information by performing quantization analysis on the target image can be as follows: converting the target image to the HSV color space, calculating the standard deviation of pixel values ​​in each channel, and normalizing the standard deviation to obtain color balance as color quantization information; using the Laplacian operator to perform edge detection on the visual focus pixel region, calculating the mean of edge gradient values ​​and standardizing them to obtain subject sharpness as expression quantization information; performing semantic analysis on the target image to obtain target style labels, and determining the target style prominence evaluation score as style quantization information based on the target style labels and the pre-constructed correlation between style labels and style prominence evaluation scores.

[0107] S104, Based on the feature quantization information and preset quality assessment conditions, determine the image quality of the target image.

[0108] The preset quality assessment conditions can be a set of thresholds for feature quantization information used to determine image quality. These can include single-dimensional thresholds for color quantization information, expression quantization information, and style quantization information, as well as a comprehensive threshold for color quantization information, expression quantization information, and style quantization information.

[0109] The image quality of the target image can be an evaluation result reflecting whether the target image meets the application scenario requirements in terms of core expression and visual effect. Specifically, it can be an image quality level divided based on the matching result of feature quantification information and quality assessment conditions, including but not limited to two levels: high quality and low quality.

[0110] In one embodiment, the image quality of a target image is determined based on feature quantization information and preset quality assessment conditions. This can be achieved by determining that the image quality of the target image is high when the color quantization information exceeds the corresponding single-dimensional threshold, the expression quantization information exceeds the corresponding single-dimensional threshold, the style quantization information exceeds the corresponding single-dimensional threshold, and the comprehensive evaluation score obtained by weighted summation of the color quantization information, expression quantization information, and style quantization information exceeds the comprehensive threshold. Otherwise, the image quality of the target image is determined to be low.

[0111] In this embodiment, a target image and image scene information are acquired. Subject object analysis is performed on the target image to obtain a visual focus pixel region. Semantic analysis is performed on the target image to obtain intent information, and a target proportion threshold is determined based on the intent information and the image scene information. If the ratio of the number of pixels in the visual focus pixel region to the total number of pixels in the target image exceeds the target proportion threshold, quantitative analysis is performed on the target image to obtain feature quantization information. The feature quantization information includes color quantization information, expression quantization information, and style quantization information. Based on the feature quantization information and preset quality assessment conditions, the image quality of the target image is determined. This image quality determination method, by constructing a two-layer progressive recognition system of subject semantics and visual aesthetics, achieves the parsing of the expressive content of image data, ensuring the effectiveness of the content delivery and visual appeal of the identified image quality determination.

[0112] Figure 2 This is a flowchart illustrating another image quality determination method provided in an embodiment of this application. Figure 2 As shown, the specific steps include the following:

[0113] S201, acquire the target image and image scene information, and perform subject object analysis on the target image to obtain the visual focus pixel region.

[0114] S202, perform semantic segmentation on the target image to obtain semantic labels and the pixel regions corresponding to the semantic labels.

[0115] The semantic label can be a set of labels used to characterize the category attribute of a pixel region in the target image, including but not limited to subject category labels (people, animals, plants, food, digital products, clothing, furniture, etc.), background category labels (sky, earth, buildings, roads, water bodies, solid color backgrounds, etc.), and functional attribute labels (promotional labels, brand trademarks, text information, etc.). The pixel region corresponding to the semantic label can be a set of all pixels in the target image that belong to the category attribute of that semantic label.

[0116] Figure 3 This is an example diagram showing the positional matching results of semantic tags and visual focus pixel regions provided in the embodiments of this application. For example... Figure 3 As shown, the light gray area represents the pixel area corresponding to the background (semantic label), the medium gray area represents the pixel area corresponding to the person (semantic label), and the dark gray area represents the pixel area corresponding to the food (semantic label).

[0117] In one embodiment, the method of semantic segmentation of the target image to obtain semantic labels and the corresponding pixel regions can be achieved by inputting the target image into a deep learning-based semantic segmentation model (such as Mask R-CNN, U-Net++, or DeepLabv3+). The semantic segmentation model extracts multi-scale image features and generates a semantic segmentation mask map with the same size as the target image. The value of each pixel in the semantic segmentation mask map corresponds to the category index of the semantic label.

[0118] S203, perform position matching between the visual focus pixel region and the pixel region corresponding to the semantic label, and determine the semantic label with the highest proportion in the visual focus pixel region as the main semantic label.

[0119] In one embodiment, the method of matching the position of the visual focus pixel region with the pixel region corresponding to the semantic label can be achieved by calculating the intersection of the set of pixel coordinates of the visual focus pixel region and the set of coordinates of the pixel region corresponding to each semantic label.

[0120] Among them, the main semantic label can be a semantic label that can represent the core main category of the target image, and it is the category label with the highest proportion in the visual focus pixel area.

[0121] In one embodiment, the method of determining the semantic label with the highest proportion in the visual focus pixel region as the main semantic label can be achieved by counting the number of pixels at the intersection of the set of pixel coordinates of the visual focus pixel region and the set of coordinates of the pixel region corresponding to each semantic label, and then determining the semantic label corresponding to the intersection with the highest number of pixels as the main semantic label.

[0122] like Figure 3 As shown, Figure 3 The black-framed area in the image is the visual focus pixel area, according to Figure 3 The positional matching results between semantic tags and the visual focus pixel region can reveal... Figure 3 The corresponding semantic tag is "person".

[0123] Optionally, if the image scene information of the target image does not correspond to the determined subject semantic label, the image quality of the target image can be directly determined as low quality.

[0124] S204, determine the intent information based on the spatial location information and relative scale information of the visual focus pixel region and the subject semantic label.

[0125] Among them, the spatial location information can be the coordinates of the center position of the visual focus pixel region in the image coordinate system of the target image; the relative scale information can be the size ratio of the visual focus pixel region relative to the target image and the size ratio between visual focus pixel regions.

[0126] In one embodiment, determining intent information based on the spatial location and relative scale information of the visual focal pixel region and the subject semantic label can be achieved by pre-constructing the association between spatial location information, relative scale information, subject semantic label, and intent information, and then querying the stored data of the association using the current spatial location information, relative scale information, and subject semantic label as query conditions to obtain the intent information. For example, if the spatial location information is (400, 300) (assumed to be at the center of the target image), the relative scale information is 50%, and the subject semantic label is food, then the corresponding intent information is to highlight the product subject.

[0127] In one embodiment, determining intent information based on the spatial location and relative scale information of the visual focus pixel region and the subject semantic tag includes: determining candidate visual expression templates based on the subject semantic tag and the association between the pre-constructed subject semantic tag and the preset visual expression template; wherein the visual expression template includes visual focus spatial location features and visual focus relative scale features; determining the similarity between the visual focus pixel region and the candidate visual expression template based on the spatial location and relative scale information of the visual focus pixel region; and determining the candidate visual expression template with the highest similarity as the intent information.

[0128] The preset visual expression template can be a standardized template of visual features corresponding to a specific expressive intent, summarized from a large number of high-quality sample images. It can include spatial location features of the visual focus and relative scale features of the visual focus. Specifically, the spatial location features of the visual focus can be a range of spatial location parameters, including the coordinate interval of the center position, quadrant affiliation, and edge distance threshold; the relative scale features of the visual focus can be a range of size ratio parameters.

[0129] Correspondingly, the relationship between the main semantic tag and the preset visual expression template can be a one-to-many mapping table, that is, one main semantic tag can correspond to multiple preset visual expression templates with different expressive intentions.

[0130] Among them, the candidate visual expression template can be a set of all preset visual expression templates corresponding to the subject semantic tag after querying the mapping relationship table through the subject semantic tag.

[0131] In one embodiment, the method for determining candidate visual expression templates based on the subject semantic tags and the association between the pre-built subject semantic tags and preset visual expression templates can be to query the mapping relationship table using the current subject semantic tags as query conditions, and all preset visual expression templates included in the query results are candidate visual expression templates.

[0132] The similarity between the visual focus pixel region and the candidate visual expression template can be a quantitative value representing the degree of fit between the spatial location information and relative scale information of the visual focus pixel region and the visual focus spatial location features and visual focus relative scale features of the candidate visual expression template.

[0133] In one embodiment, the method for determining the similarity between the visual focus pixel region and the candidate visual expression template based on the spatial location information and relative scale information of the visual focus pixel region can be achieved by converting the spatial location information and relative scale information into a vector format to obtain a first vector, and converting the spatial location features and relative scale features of the visual focus of the candidate visual expression template into a vector format to obtain a second vector corresponding to the candidate visual expression template, and calculating the cosine similarity between the first vector and the second vector as the similarity between the visual focus pixel region and the candidate visual expression template.

[0134] The advantage of this approach is that it transforms abstract intent information into quantifiable visual feature parameters by using preset visual expression templates, avoiding the subjectivity of traditional association rules and making intent inference more objective and consistent.

[0135] S205, determine the target proportion threshold based on the intent information and the image scene information.

[0136] In one embodiment, determining a target proportion threshold based on the intent information and the image scene information includes: determining a basic proportion threshold based on the image scene information; determining a proportion correction coefficient based on the intent information, and determining the confidence level of the proportion correction coefficient based on the similarity corresponding to the intent information; and determining a target proportion threshold based on the basic proportion threshold, the proportion correction coefficient, and the confidence level.

[0137] Among them, the basic proportion threshold can be a benchmark value preset for a specific image scene, reflecting the minimum proportion requirement of the main body area in that scene.

[0138] In one embodiment, determining the basic percentage threshold based on image scene information can be achieved by pre-constructing a correlation between image scene information and the basic percentage threshold. The stored data of this correlation is then queried using the current image scene information as the query condition, and the query result includes the basic percentage threshold. For example, in the case of image quality determination in the context of advertising image material review, if the image scene information is food, the basic percentage threshold is 50%; if the image scene information is skincare, the basic percentage threshold is 35%.

[0139] Among them, the proportion correction coefficient can be a quantitative coefficient that dynamically adjusts the basic proportion threshold based on intent information.

[0140] In one embodiment, determining the proportion correction coefficient based on intent information can be achieved by pre-constructing a correlation between intent information and the proportion correction coefficient, and then querying the stored data of the correlation using the current intent information as the query condition. The resulting query result includes the proportion correction coefficient. For example, if the intent information is a first preset visual expression template (a subject prominence template under the current image scene information), the proportion correction coefficient is 1.2; if the intent information is a second preset visual expression template (a scene atmosphere enhancement template under the current image scene information), the proportion correction coefficient is 0.4.

[0141] Among them, the confidence level of the proportion correction coefficient can be a quantitative indicator that characterizes the adaptability of the proportion correction coefficient.

[0142] In one embodiment, the confidence level of the proportion correction coefficient can be determined based on the similarity corresponding to the intent information by directly determining the similarity corresponding to the intent information as the confidence level of the corresponding proportion correction coefficient.

[0143] In one embodiment, the target percentage threshold can be determined based on the base percentage threshold, the percentage correction coefficient, and the confidence level by calculating the difference between 1 and the percentage correction coefficient, multiplying the difference by the confidence level, and subtracting the result of the multiplication from 1 to obtain the final percentage correction coefficient. The target percentage threshold is then obtained by multiplying the base percentage threshold by the final percentage correction coefficient.

[0144] The advantage of setting the target percentage threshold in this way is that it ensures that the threshold not only fits the needs of the scenario and intent, but also has a high degree of stability and rationality.

[0145] S206, if the ratio of the number of pixels in the visual focus pixel region to the total number of pixels in the target image exceeds the target proportion threshold, the target image is subjected to quantitative analysis processing to obtain feature quantization information; wherein, the feature quantization information includes color quantization information, expression quantization information and style quantization information.

[0146] S207, Based on the feature quantization information and preset quality assessment conditions, determine the image quality of the target image.

[0147] The advantage of this approach is that it achieves pixel-level classification through semantic segmentation, and combined with position matching of the visual focus pixel region, it avoids misjudgment of the subject category caused by relying solely on saliency detection. It ensures that the subject semantic label is consistent with the core object of the image, and integrates subject semantic label, spatial location and relative scale information to overcome the limitations of single feature inference, making the intent information more consistent with the actual expression purpose of the image.

[0148] Figure 4This is a flowchart illustrating another image quality determination method provided in an embodiment of this application. Figure 4 As shown, the specific steps include the following:

[0149] S401, acquire the target image and image scene information, and perform subject object analysis on the target image to obtain the visual focus pixel region.

[0150] S402, perform semantic analysis on the target image to obtain intent information, and determine the target proportion threshold based on the intent information and the image scene information.

[0151] S403, if the ratio of the number of pixels in the visual focus pixel region to the total number of pixels in the target image exceeds the target proportion threshold, calculate the average saturation, the difference between the number of warm and cool pixels, and the maximum brightness difference of the target image, and perform a weighted summation calculation on the average saturation, the difference between the number of warm and cool pixels, and the maximum brightness difference to obtain color quantization information.

[0152] Among them, the average saturation can be the arithmetic mean of the saturation values ​​of all pixels in the target image; the difference between the number of warm and cool pixels can be the absolute difference between the total number of warm-toned pixels and the total number of cool-toned pixels in the target image; and the maximum brightness difference can be the difference between the maximum and minimum brightness values ​​of all pixels in the target image.

[0153] In one embodiment, the method for calculating the average saturation, the difference between the number of warm and cool pixels, and the maximum brightness difference of the target image can be as follows: convert the target image from the RGB color space to the HSV color space, obtain the three components of saturation (S), hue (H), and brightness (V) for each pixel, iterate through all pixels and accumulate the saturation value of each pixel to calculate the average saturation; count the number of pixels with hue in the warm color range and the number of pixels with hue in the cool color range, calculate the difference between the two pixel counts to obtain the difference between the number of warm and cool pixels, iterate through the brightness values ​​of all pixels to filter out the maximum and minimum values, and calculate the difference between the maximum and minimum values ​​as the maximum brightness difference.

[0154] In one embodiment, the method of obtaining color quantization information by weighted summation of average saturation, the difference in the number of warm and cool pixels, and the difference in maximum brightness can be achieved by normalizing the average saturation, the difference in the number of warm and cool pixels, and the difference in maximum brightness, and then weighting and summing the average saturation, the difference in the number of warm and cool pixels, and the difference in maximum brightness according to preset weight coefficients (e.g., 0.4, 0.3, and 0.3 respectively) to obtain color quantization information.

[0155] S404, calculate the average color difference and average edge intensity between the visual focus pixel region and other pixel regions in the target image, and perform a weighted summation of the average color difference and the average edge intensity to obtain the expression quantization information.

[0156] The average color difference can be the mean of the Euclidean distance between the average color of the visual focal pixel region and the average color of other pixel regions; the average edge intensity can be the mean of the gradient values ​​of all edge pixels at the boundary between the visual focal pixel region and other pixel regions.

[0157] In one embodiment, the method for calculating the average color difference and average edge intensity between the visual focal pixel region and other pixel regions in the target image can be as follows: calculate the average color value of the visual focal pixel region and the average color value of other pixel regions respectively, and substitute them into the CIE Lab color difference formula to calculate the average color difference; use the Sobel operator to perform edge detection on the target image, extract the set of boundary pixels between the visual focal pixel region and other pixel regions, calculate the gradient magnitude of each boundary pixel, and traverse all boundary pixels to calculate the average value of the gradient magnitude as the average edge intensity.

[0158] In one embodiment, the method of obtaining quantitative information by weighted summation of average color difference and average edge intensity can be achieved by normalizing the average color difference and average edge intensity, and then weighting and summing the average color difference and average edge intensity according to preset weight coefficients (e.g., 0.4 and 0.6 respectively) to obtain quantitative information.

[0159] S405, the target image is input into a pre-trained deep convolutional neural network model to obtain a target style feature vector, and the target style feature vector is matched with a large data vector library to obtain style quantification information.

[0160] Here, the style feature vector can be a high-dimensional numerical vector representing the artistic style attributes of an image. Correspondingly, the target style feature vector can be a specific style feature vector extracted from the target image by a deep convolutional neural network model.

[0161] Among them, the deep convolutional neural network model can be a convolutional neural network model pre-trained based on the style transfer task (such as VGG-19, ResNet-50, MobileNetV3), or a lightweight model designed specifically for image style classification (such as EfficientNet-B0).

[0162] In one embodiment, the method of inputting the target image into a pre-trained deep convolutional neural network model to obtain the target style feature vector can be achieved by preprocessing the target image, such as size normalization and pixel value standardization, and then inputting the preprocessed target image into the pre-trained deep convolutional neural network model, which outputs the target style feature vector.

[0163] Among them, the big data vector library can be a structured database containing style feature vectors of massive images of different styles.

[0164] In one embodiment, the method of matching the target style feature vector with a large data vector library to obtain style quantification information can be achieved by calculating the average cosine similarity between the target style feature vector and each style feature vector in the large data vector library as the style quantification information.

[0165] In one embodiment, matching the target style feature vector with a large data vector library to obtain style quantification information includes: calculating the average vector distance between the target style feature vector and all style feature vectors in the large data vector library; calculating the average vector distance between each style feature vector in the large data vector library and all other style feature vectors as a baseline average vector distance; sorting the average vector distance and the baseline average vector distance, and determining the position distribution percentage as style quantification information based on the sorting position of the average vector distance in the sorting result.

[0166] The average vector distance can be the arithmetic mean of the cosine distances between the target style feature vector and all style feature vectors in the big data vector library.

[0167] In one embodiment, the average vector distance between the target style feature vector and all style feature vectors in the big data vector library can be calculated by calculating the cosine similarity between the target style feature vector and each style feature vector in the big data vector library, subtracting the cosine similarity from 1 to obtain the cosine distance, and then calculating the arithmetic mean of all cosine distances to obtain the average vector distance.

[0168] The baseline average vector distance can be the set of average vector distances between each style feature vector in the big data vector library and all other style feature vectors in the library.

[0169] In one embodiment, the average vector distance between each style feature vector in the big data vector library and all other style feature vectors can be used as a benchmark average vector distance. This can be achieved by traversing each style feature vector in the big data vector library and calculating its average vector distance with all other style feature vectors in the library.

[0170] In one embodiment, the average vector distance and the baseline average vector distance can be sorted in ascending order to obtain the sorting result.

[0171] The position distribution percentage can be the proportion of the average vector distance of the target vector in the sorting results.

[0172] In one embodiment, the method of determining the position distribution percentage as style quantification information based on the ranking position of the average vector distance in the ranking results can be achieved by counting the total number of distances in the ranking results, determining the ranking position of the average vector distance in the ranking results, dividing the ranking position by the total number of distances, and then dividing the result of the division into percentages to obtain the position distribution percentage as style quantification information.

[0173] The advantage of this approach is that by comparing the average vector distance of the target image with the baseline average vector distance of all styles in the vector library, and using the style distribution characteristics of the large data vector library itself as a reference, style prominence quantification becomes more referential and conforms to the objective laws of style distribution.

[0174] S406, Based on the feature quantization information and preset quality assessment conditions, determine the image quality of the target image.

[0175] The advantages of this approach are that it addresses color quantization from three core dimensions: saturation, warm / cool color balance, and brightness / contrast, comprehensively covering key attributes such as color vibrancy, emotional tone, and light / dark levels, thus avoiding the limitations of a single color parameter. Expression quantification focuses on the color difference and edge intensity between the subject and background, directly reflecting the prominence and visual distinctiveness of the image subject, providing a quantifiable basis for expression clarity. Style quantification relies on pre-trained deep convolutional neural networks to extract high-dimensional style feature vectors, combined with cosine similarity matching from a large data vector library, to achieve objective quantification of style attributes, eliminating subjective bias.

[0176] Figure 5 This is a schematic diagram of the structure of an image quality determination device provided in an embodiment of this application. Figure 5 As shown, the device includes:

[0177] The visual focus determination module 510 is used to acquire target image and image scene information, and to perform subject object analysis on the target image to obtain the visual focus pixel region;

[0178] The target proportion determination module 520 is used to perform semantic analysis on the target image to obtain intent information, and determine the target proportion threshold based on the intent information and the image scene information;

[0179] The quantization feature determination module 530 is used to perform quantization analysis on the target image to obtain feature quantization information when the ratio of the number of pixels in the visual focus pixel region to the total number of pixels in the target image exceeds the target proportion threshold; wherein, the feature quantization information includes color quantization information, expression quantization information and style quantization information;

[0180] The image quality determination module 540 is used to determine the image quality of the target image based on the feature quantization information and preset quality evaluation conditions.

[0181] Optionally, the target proportion determination module 520 is specifically used for:

[0182] The target image is semantically segmented to obtain semantic labels and the pixel regions corresponding to the semantic labels;

[0183] The positions of the visual focus pixel region and the pixel region corresponding to the semantic label are matched, and the semantic label with the highest proportion in the visual focus pixel region is determined as the main semantic label;

[0184] Intent information is determined based on the spatial location and relative scale information of the visual focal pixel region and the subject semantic label.

[0185] Optionally, the target proportion determination module 520 is specifically used for:

[0186] Candidate visual expression templates are determined based on the subject semantic tags and the association between the pre-constructed subject semantic tags and the preset visual expression templates; wherein, the visual expression templates include visual focus spatial location features and visual focus relative scale features;

[0187] Based on the spatial location information and relative scale information of the visual focal pixel region, the similarity between the visual focal pixel region and the candidate visual expression template is determined;

[0188] The candidate visual expression template with the highest similarity is identified as the intent information.

[0189] Optionally, the target proportion determination module 520 is specifically used for:

[0190] Determine the basic proportion threshold based on the image scene information;

[0191] The proportion correction coefficient is determined based on the intent information, and the confidence level of the proportion correction coefficient is determined based on the similarity corresponding to the intent information.

[0192] The target percentage threshold is determined based on the basic percentage threshold, the percentage correction coefficient, and the confidence level.

[0193] Optionally, the quantization feature determination module 530 is specifically used for:

[0194] The average saturation, the difference in the number of warm and cool pixels, and the maximum brightness difference of the target image are calculated, and the average saturation, the difference in the number of warm and cool pixels, and the maximum brightness difference are weighted and summed to obtain color quantization information.

[0195] Calculate the average color difference and average edge intensity between the visual focal pixel region and other pixel regions in the target image, and perform a weighted summation of the average color difference and average edge intensity to obtain the quantitative information of expression.

[0196] The target image is input into a pre-trained deep convolutional neural network model to obtain a target style feature vector. The target style feature vector is then matched with a large data vector library to obtain style quantification information.

[0197] Optionally, the quantization feature determination module 530 is specifically used for:

[0198] Calculate the average vector distance between the target style feature vector and all style feature vectors in the big data vector library;

[0199] The average vector distance between each style feature vector in the big data vector library and all other style feature vectors is calculated as the baseline average vector distance.

[0200] The average vector distance is sorted with the benchmark average vector distance, and the position distribution percentage is determined as style quantification information based on the sorting position of the average vector distance in the sorting results.

[0201] Optionally, the visual focus determination module 510 is specifically used for:

[0202] Generate a saliency heatmap corresponding to the target image, and select candidate focal regions in the saliency heatmap according to a preset heatmap threshold;

[0203] If the overlapping area of ​​the bounding rectangles of two candidate focal regions exceeds a preset area threshold, the bounding rectangles of the two candidate focal regions are weighted and merged according to the total regional thermal value of the two candidate focal regions to obtain a merged bounding rectangle.

[0204] The overlapping area between the two candidate focal regions and the fused bounding rectangle is determined as the new candidate focal region, and the original two candidate focal regions are deleted.

[0205] If the overlapping area of ​​the bounding rectangles of any two candidate focal regions does not exceed a preset area threshold, the candidate focal region is determined as the visual focal pixel region.

[0206] In this embodiment, a visual focus determination module is used to acquire a target image and image scene information, and perform subject object analysis on the target image to obtain a visual focus pixel region; a target proportion determination module is used to perform semantic analysis on the target image to obtain intent information, and determine a target proportion threshold based on the intent information and the image scene information; a quantization feature determination module is used to perform quantization analysis on the target image to obtain feature quantization information when the ratio of the number of pixels in the visual focus pixel region to the total number of pixels in the target image exceeds the target proportion threshold; wherein, the feature quantization information includes color quantization information, expression quantization information, and style quantization information; and an image quality determination module is used to determine the image quality of the target image based on the feature quantization information and preset quality assessment conditions. The above-mentioned image quality determination device, by constructing a two-layer progressive recognition system of subject semantics and visual aesthetics, realizes the parsing of the expressive content of image data, ensuring the effectiveness of the content delivery and visual appeal of the identified image quality determination.

[0207] The image quality determination device in this application embodiment can be a device, or a component, integrated circuit, or chip in a terminal. The device can be a mobile electronic device or a non-mobile electronic device. For example, mobile electronic devices can be mobile phones, tablets, laptops, PDAs, in-vehicle electronic devices, wearable devices, ultra-mobile personal computers (UMPCs), netbooks, or personal digital assistants (PDAs), etc., while non-mobile electronic devices can be servers, network attached storage (NAS), personal computers (PCs), televisions (TVs), ATMs, or self-service machines, etc. This application embodiment does not impose specific limitations.

[0208] The image quality determination device in this application embodiment can be a device with an operating system. This operating system can be Android, iOS, or other possible operating systems; this application embodiment does not specifically limit it.

[0209] The image quality determination device provided in this application embodiment can realize the various processes implemented in the above embodiments, and will not be described again here to avoid repetition.

[0210] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. For example... Figure 6 As shown, this application embodiment also provides an electronic device 600, including a processor 601, a memory 602, and a program or instructions stored in the memory 602 and executable on the processor 601. When the program or instructions are executed by the processor 601, they implement the various processes of the above-described image quality determination method embodiment and achieve the same technical effect. To avoid repetition, they will not be described again here.

[0211] It should be noted that the electronic devices in the embodiments of this application include the mobile electronic devices and non-mobile electronic devices described above.

[0212] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described image quality determination method embodiments and achieve the same technical effect. To avoid repetition, they will not be described again here.

[0213] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0214] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0215] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0216] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

[0217] The above description is merely a preferred embodiment and the technical principles employed in this application. This application is not limited to the specific embodiments described herein, and various obvious changes, readjustments, and substitutions that can be made by those skilled in the art will not depart from the scope of protection of this application. Therefore, although this application has been described in detail through the above embodiments, this application is not limited to the above embodiments, and may include more other equivalent embodiments without departing from the concept of this application, the scope of which is determined by the scope of the claims.

Claims

1. A method for determining image quality, characterized in that, The method includes: Acquire target image and image scene information, and perform subject object analysis on the target image to obtain the visual focus pixel region; Semantic analysis is performed on the target image to obtain intent information, and a target proportion threshold is determined based on the intent information and the image scene information; If the ratio of the number of pixels in the visual focus pixel region to the total number of pixels in the target image exceeds the target proportion threshold, the target image is subjected to quantitative analysis to obtain feature quantization information; wherein, the feature quantization information includes color quantization information, expression quantization information, and style quantization information; Based on the feature quantification information and preset quality assessment conditions, the image quality of the target image is determined.

2. The image quality determination method according to claim 1, characterized in that, The semantic analysis of the target image to obtain intent information includes: The target image is semantically segmented to obtain semantic labels and the pixel regions corresponding to the semantic labels; The positions of the visual focus pixel region and the pixel region corresponding to the semantic label are matched, and the semantic label with the highest proportion in the visual focus pixel region is determined as the main semantic label; Intent information is determined based on the spatial location information and relative scale information of the visual focus pixel region and the subject semantic label; wherein, the relative scale information is the size ratio of the visual focus pixel region relative to the target image and the size ratio between visual focus pixel regions.

3. The image quality determination method according to claim 2, characterized in that, The determination of intent information based on the spatial location and relative scale information of the visual focal pixel region and the subject semantic label includes: Candidate visual expression templates are determined based on the subject semantic tags and the association between the pre-constructed subject semantic tags and the preset visual expression templates; wherein, the visual expression templates include visual focus spatial location features and visual focus relative scale features; Based on the spatial location information and relative scale information of the visual focal pixel region, the similarity between the visual focal pixel region and the candidate visual expression template is determined; The candidate visual expression template with the highest similarity is identified as the intent information.

4. The image quality determination method according to claim 3, characterized in that, Determining the target proportion threshold based on the intent information and the image scene information includes: Determine the basic proportion threshold based on the image scene information; The proportion correction coefficient is determined based on the intent information, and the confidence level of the proportion correction coefficient is determined based on the similarity corresponding to the intent information. The target percentage threshold is determined based on the basic percentage threshold, the percentage correction coefficient, and the confidence level.

5. The image quality determination method according to claim 1, characterized in that, The step of performing quantization analysis on the target image to obtain feature quantization information includes: The average saturation, the difference in the number of warm and cool pixels, and the maximum brightness difference of the target image are calculated, and the average saturation, the difference in the number of warm and cool pixels, and the maximum brightness difference are weighted and summed to obtain color quantization information. Calculate the average color difference and average edge intensity between the visual focal pixel region and other pixel regions in the target image, and perform a weighted summation of the average color difference and average edge intensity to obtain the quantitative information of expression. The target image is input into a pre-trained deep convolutional neural network model to obtain a target style feature vector. The target style feature vector is then matched with a large data vector library to obtain style quantification information.

6. The image quality determination method according to claim 5, characterized in that, The step of matching the target style feature vector with a large data vector library to obtain style quantification information includes: Calculate the average vector distance between the target style feature vector and all style feature vectors in the big data vector library; The average vector distance between each style feature vector in the big data vector library and all other style feature vectors is calculated as the baseline average vector distance. The average vector distance is sorted with the benchmark average vector distance, and the position distribution percentage is determined as style quantification information based on the sorting position of the average vector distance in the sorting results.

7. The image quality determination method according to claim 1, characterized in that, The step of performing subject object analysis on the target image to obtain the visual focus pixel region includes: Generate a saliency heatmap corresponding to the target image, and select candidate focal regions in the saliency heatmap according to a preset heatmap threshold; If the overlapping area of ​​the bounding rectangles of two candidate focal regions exceeds a preset area threshold, the bounding rectangles of the two candidate focal regions are weighted and merged according to the total regional thermal value of the two candidate focal regions to obtain a merged bounding rectangle. The overlapping area between the two candidate focal regions and the fused bounding rectangle is determined as the new candidate focal region, and the original two candidate focal regions are deleted. If the overlapping area of ​​the bounding rectangles of any two candidate focal regions does not exceed a preset area threshold, the candidate focal region is determined as the visual focal pixel region.

8. An image quality determination device, characterized in that, The device includes: The visual focus determination module is used to acquire target image and image scene information, and to perform subject object analysis on the target image to obtain the visual focus pixel region; The target proportion determination module is used to perform semantic analysis on the target image to obtain intent information, and determine the target proportion threshold based on the intent information and the image scene information; The quantization feature determination module is used to perform quantization analysis on the target image to obtain feature quantization information when the ratio of the number of pixels in the visual focus pixel region to the total number of pixels in the target image exceeds the target proportion threshold; wherein, the feature quantization information includes color quantization information, expression quantization information, and style quantization information; The image quality determination module is used to determine the image quality of the target image based on the feature quantization information and preset quality assessment conditions.

9. An electronic device, characterized in that, It includes a processor, a memory, and a program or instructions stored in the memory and executable on the processor, wherein the program or instructions, when executed by the processor, implement the image quality determination method as described in any one of claims 1-7.

10. A readable storage medium, characterized in that, The readable storage medium stores a program or instructions that, when executed by a processor, implement the image quality determination method as described in any one of claims 1-7.