Graph selection method, device and equipment based on real person model generated graph evaluation and medium

By using the comprehensive evaluation model of ResNet50 and CLIP models in the evaluation of live model generation diagrams, the one-sided problem of evaluation results in the prior art is solved, and a comprehensive evaluation of multiple factors of the image is achieved, and the screening efficiency and standardization of results is improved.

CN120070346APending Publication Date: 2025-05-30ZIXUN TECHNOLOGY (FUJIAN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510108217.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

When evaluating the aesthetic value of live model generated pictures, the prior art ignores the impact of multiple factors on the overall aesthetics of the image, resulting in one-sided evaluation results, which cannot fully reflect the actual aesthetic value of the image, and is inefficient.

Method used

The evaluation model based on ResNet50 and CLIP models is adopted, and a comprehensive score is generated through color evaluation, composition evaluation and model aesthetic evaluation, combined with the adaptive pooling layer and concat function to achieve multi-dimensional evaluation of the image.

Benefits of technology

Through this evaluation model, pictures that conform to the public's aesthetic view can be quickly selected, work efficiency, screening differences, and shop competitiveness can be enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070346A_ABST
    Figure CN120070346A_ABST
Patent Text Reader

Abstract

The invention provides an image selection method, device and equipment based on real person model generated image evaluation and a medium, and the method comprises the steps: carrying out the color evaluation, composition evaluation and model beauty evaluation of an input image, and obtaining a marked image; an evaluation model is built, the evaluation model comprises a ResNet50 model and a CLIP model, a full connection layer of output parts of the ResNet50 model and the CLIP model is deleted, and model output of the same dimension is obtained by using an adaptive pooling layer; the output dimensions comprise six dimensions which are respectively corresponding to good color, poor color, good composition, poor composition, beautiful model and probability value of poor model; training the evaluation model through the marked image to obtain a trained evaluation model; inputting pictures needing to be screened into the trained evaluation model, and carrying out image screening according to scores; screening is fast, and the working efficiency is greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image selection, and in particular to an image selection method, device, equipment and medium based on the evaluation of images generated by real models. Background Art

[0002] In the product display scene of e-commerce platforms, the aesthetic evaluation of real-life model images is crucial to attracting consumers and improving purchase conversion rates. For online shopping platforms, product images are the most important display window for products, which directly affects the click-through rate of products. Excellent product images usually bring higher click-through rates, thereby increasing sales.

[0003] In the prior art, AI functions are used to generate real-life model images for product display in online stores. Therefore, it is necessary to select images that meet the needs from the generated images. However, most current aesthetic evaluation methods only focus on a single dimension (such as image clarity or model appearance), ignoring the impact of multiple factors of the image on the overall aesthetics. For example, factors such as the coordination of color matching of the model image, the rationality of the composition, and the beauty of the model, all of which will significantly affect consumers' perception of the product and purchase decisions. The existing single-dimensional evaluation method often leads to one-sided evaluation results and cannot fully reflect the actual aesthetic value of the image; therefore, it is impossible to score and select images through existing methods, and the selection can only be done manually, resulting in low efficiency; and due to each person's subjective factors, there is also a large gap between the images selected by different people. Summary of the invention

[0004] The technical problem to be solved by the present invention is to provide a method, device, equipment and medium for selecting pictures based on the evaluation of pictures generated by real models, so that the screened pictures are more standardized and in line with the public's aesthetic views. Through the evaluation model of the present invention, pictures can be screened quickly, greatly improving work efficiency.

[0005] In a first aspect, the present invention provides a method for selecting images based on evaluation of images generated by real models, comprising the following steps:

[0006] Step 1, evaluate the color, composition and model beauty of the input image, and use 0 or 1 to mark good color, bad color, good composition, bad composition, good model and bad model to obtain a marked image;

[0007] Step 2: Build an evaluation model. The evaluation model includes a ResNet50 model and a CLIP model. Delete the fully connected layers in the output parts of the ResNet50 model and the CLIP model, and use an adaptive pooling layer to obtain model outputs with the same dimension. The outputs of the evaluation model are concatenated into the same output using the concat function, and a fully connected layer with an output dimension of 6 is used to obtain the final result. The output dimension includes 6 dimensions, corresponding to the probabilities of good color, bad color, good composition, bad composition, good-looking model, and bad-looking model respectively. Train the evaluation model with the labeled images to obtain the trained evaluation model.

[0008] Step 3: Input the images to be screened into the trained evaluation model, and screen the images according to the scores.

[0009] In a second aspect, the present invention provides a picture selection device for evaluating pictures generated based on real models, including:

[0010] An evaluation module that performs color evaluation, composition evaluation, and model beauty evaluation on the input images, and marks good color, bad color, good composition, bad composition, good-looking model, and bad-looking model with 0 or 1 to obtain labeled images;

[0011] A model acquisition module that builds an evaluation model. The evaluation model includes a ResNet50 model and a CLIP model. Delete the fully connected layers in the output parts of the ResNet50 model and the CLIP model, and use an adaptive pooling layer to obtain model outputs with the same dimension. The outputs of the evaluation model are concatenated into the same output using the concat function, and a fully connected layer with an output dimension of 6 is used to obtain the final result. The output dimension includes 6 dimensions, corresponding to the probabilities of good color, bad color, good composition, bad composition, good-looking model, and bad-looking model respectively. Train the evaluation model with the labeled images to obtain the trained evaluation model.

[0012] A screening module that inputs the images to be screened into the trained evaluation model and screens the images according to the scores.

[0013] In a third aspect, the present invention provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the method described in the first aspect is implemented.

[0014] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the method described in the first aspect is implemented.

[0015] One or more technical solutions provided by the present invention have at least the following technical effects or advantages:

[0016] Through the technical solution of the present invention, the selected pictures can be made more standardized and conform to the aesthetic concept of the public. Through the evaluation model of the present invention, pictures can be screened quickly, greatly reducing the screening differences and improving the screening efficiency. Merchants can quickly select the pictures as the display pictures of their online stores, greatly enhancing the competitiveness of the stores.

[0017] The above description is only an overview of the technical solution of the present invention. In order to understand the technical means of the present invention more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present invention more obvious and understandable, the specific embodiments of the present invention are specifically exemplified below. Brief Description of the Drawings

[0018] The present invention will be further described below with reference to the accompanying drawings in conjunction with the embodiments.

[0019] Figure 1 It is a flowchart of the method in Embodiment 1 of the present invention;

[0020] Figure 2 It is a structural schematic diagram of the device in Embodiment 2 of the present invention. Detailed Embodiments

[0021] The overall idea of the technical solution in the embodiments of the present application is as follows:

[0022] In order to obtain the integrated learning model for picture evaluation, the training data is processed first. The training data will be processed into six dimensions, corresponding to six classification categories, where the categories are: good color, bad color, good composition, bad composition, good-looking model, bad-looking model.

[0023] For the annotation acquisition of the data labels of good or bad color: First, use cv2.cvtColor to convert the input image to the HSV color space, and extract the three channels of hue, saturation, and value respectively. Use cv2.calcHist to calculate the histogram distribution features for each channel, and obtain the statistical values of key quantiles such as 25%, 50%, and 75% through np.percentile. Calculate the mean H_mean and standard deviation H_std of the hue, the mean S_mean and standard deviation S_std of the saturation, and the mean V_mean and standard deviation V_std of the value. When S_mean is in the range of [0.3, 0.7] and V_mean is in the range of [0.4, 0.8], and at the same time H_std is moderate, it is marked as good color. At this time, good color is recorded as 1 and bad color is recorded as 0; when S_mean < 0.3 or > 0.7, or V_mean < 0.4 or > 0.8, or H_std is too large, it is marked as bad color; good color is recorded as 0 and bad color is recorded as 1;

[0024] For the annotation and acquisition of data labels for good or bad composition: Use cv2.Laplacian to calculate the sharpness score sharpness_score of the image, and use the cv2.Sobel operator to extract the edge features in the horizontal and vertical directions to calculate edge_strength. Calculate the centroid position of the image based on cv2.moments to obtain symmetry_score, and divide the image into a 3x3 grid to calculate the feature response intensity golden_ratio_score at the golden section point position. When sharpness_score > 0.7 and edge_strength is moderate, and at the same time symmetry_score > 0.6 and golden_ratio_score > 0.5, it is marked as a good composition. At this time, a good composition is recorded as 1, and a bad composition is recorded as 0; when sharpness_score < 0.7 or edge_strength is too low, or symmetry_score < 0.6 or golden_ratio_score < 0.5, it is marked as a bad composition; a good composition is recorded as 0, and a bad composition is recorded as 1.

[0025] For the annotation and acquisition of data labels for whether the model is beautiful or not: Use cv2.Laplacian to calculate the sharpness score model_sharpness_score of the model, use the QwenVL2 vision model to score the beauty of the model image, and use the prompt to obtain qwen_beauty_score: "Please evaluate the overall beauty of the model in this picture. Please score from the following aspects: 1. Whether the pose of the model is natural and elegant; 2. Whether the facial expression is natural and appropriate; 3. Whether the clothing matching is reasonable and beautiful; 4. Whether the overall styling is harmonious. Please give a comprehensive score between 0 and 1, with 1 point indicating very beautiful and 0 point indicating not beautiful." When model_sharpness_score > 0.7 and qwen_beauty_score > 0.7, it is marked as the model being good-looking. At this time, the model being good-looking is recorded as 1, and the model being not good-looking is recorded as 0; in other cases, it is marked as the model being not good-looking; at this time, the model being good-looking is recorded as 0, and the model being not good-looking is recorded as 1.

[0026] After obtaining the data of the three types of indicators, model training is carried out. The model architecture uses two types of model bases for splicing and combination to enhance the learning ability of image features such as color, composition, and the beauty of models. First, use torchvision.models.resnet50(pretrained=True) to load the pre-trained ResNet50 as the base of the small aesthetic evaluation model, and use torch.load('clip-vit-base-patch32') to load the CLIP model as the base of the large aesthetic evaluation model. Delete the fully connected layers in the model output parts of the two models, use the adaptive pooling layer to obtain model outputs of the same dimension, and finally use the concat function to splice the outputs into the same output, and use a fully connected layer with an output dimension of 6 to obtain the final result. The output dimension is a total of six dimensions, corresponding to six category probabilities, where the categories are: good color, bad color, good composition, bad composition, good-looking model, and bad-looking model.

[0027] Next, use 100,000 real model images of different styles to fine-tune the above combined model architecture. Use the torch.optim.Adam optimizer and adopt the MSE loss function to calculate the error between the predicted score and the true score (the true score is set as before; for example: good color is recorded as 1, bad color is recorded as 0, good composition is recorded as 1, bad composition is recorded as 0, good-looking model is recorded as 1, bad-looking model is recorded as 0; calculate the error by using the MSE loss function for the predicted score and the corresponding marked value; for example, calculate the probability value of good color and the marked value 1 of good color). After training 100 rounds, a model that can appropriately evaluate color features, composition features, and the beauty of models is obtained. Among them, the model will obtain six-dimensional probabilities after inference, which are the probability values of good color, bad color, good composition, bad composition, good-looking model, and bad-looking model respectively.

[0028] For the probability outputs of these six dimensions, first use np.normalize to standardize the probabilities of each dimension. Considering the strong correlations among the three groups of dimensions: good color vs. bad color, good composition vs. bad composition, and good-looking model vs. not good-looking model, appropriately reduce the weights of these related dimensions to avoid duplicate calculations. Specifically, set the weights of the probability of good color and the probability of bad color to 0.8, the weights of the probability of good composition and the probability of bad composition to 1.0, and the weights of the probability of good-looking model and the probability of not good-looking model to 1.2. Use np.average to achieve weighted average to obtain the final evaluation score (evaluation score = (probability value of good color × 0.8 + probability value of good composition × 1.0 + probability value of good-looking model × 1.2) ÷ 3), and limit the score to the range of 0 to 1 through np.clip. Set 0.7 as the screening threshold for high-quality samples, and only retain the samples with scores greater than the threshold as the final high-quality images.

[0029] Example 1

[0030] As Figure 1 shown, this embodiment provides a method for selecting images based on the evaluation of images generated from real models, including the following steps:

[0031] Step 1: Perform color evaluation, composition evaluation, and model aesthetics evaluation on the input image, and mark good color, bad color, good composition, bad composition, good-looking model, and not good-looking model with 0 or 1 to obtain a marked image;

[0032] Step 2: Build an evaluation model. The evaluation model includes a ResNet50 model and a CLIP model. Delete the fully connected layers in the output parts of the ResNet50 model and the CLIP model, and use an adaptive pooling layer to obtain model outputs of the same dimension; Concatenate the outputs of the evaluation model using the concat function into the same output, and use a fully connected layer with an output dimension of 6 to obtain the final result. The output dimension includes 6 dimensions, corresponding to the probability values of good color, bad color, good composition, bad composition, good-looking model, and not good-looking model respectively; Train the evaluation model with the marked image to obtain a trained evaluation model;

[0033] Step 3: Input the images to be screened into the trained evaluation model, and screen the images according to the scores.

[0034] Preferably, in this embodiment, step 1 is specifically: Perform color evaluation, composition evaluation, and model aesthetics evaluation on the input image, and mark good color, bad color, good composition, bad composition, good-looking model, and not good-looking model with 0 or 1 to obtain a marked image;

[0035] The color evaluation is specifically as follows: color space conversion and channel extraction; use cv2.cvtColor to convert the input image from the RGB color space to the HSV color space, and extract the three channels of hue, saturation, and value respectively;

[0036] Histogram distribution feature calculation: For the three channels of hue, saturation, and value, use cv2.calcHist to calculate the corresponding histogram distribution features respectively; obtain the statistical values of the set quantiles for each channel through np.percentile;

[0037] Calculate the mean H_mean and standard deviation H_std of hue, the mean S_mean and standard deviation S_std of saturation, and the mean V_mean and standard deviation V_std of value; when S_mean is in the range of [0.3, 0.7] and V_mean is in the range of [0.4, 0.8], and at the same time H_std is less than the set first threshold, then record the color as good = 1 and the color as bad = 0; when S_mean < 0.3 or > 0.7, or V_mean < 0.4 or > 0.8, or H_std is greater than or equal to the first set threshold, then record the color as good = 0 and the color as bad = 1;

[0038] The composition evaluation is specifically as follows: clarity and edge feature calculation: use cv2.Laplacian to calculate the clarity score sharpness_score of the image;

[0039] Use the cv2.Sobel operator to extract the edge features in the horizontal and vertical directions of the image, and calculate to obtain edge_strength;

[0040] Symmetry and golden ratio point feature calculation: Based on cv2.moments, calculate the centroid position of the image to obtain symmetry_score;

[0041] Divide the image into a 3x3 grid, and calculate the feature response intensity golden_ratio_score at the position of the golden ratio point;

[0042] When sharpness_score > 0.7 and edge_strength is greater than the second set threshold, and at the same time symmetry_score > 0.6 and golden_ratio_score > 0.5, then record the composition as good = 1 and the composition as bad = 0; when sharpness_score ≤ 0.7 or edge_strength is less than or equal to the second set threshold, or symmetry_score ≤ 0.6 or golden_ratio_score ≤ 0.5, then record the composition as good = 0 and the composition as bad = 1;

[0043] The specific evaluation of the model's aesthetics is as follows: Calculation of the model's clarity and aesthetics scores: Use cv2.Laplacian to calculate the model's clarity score model_sharpness_score;

[0044] Use the QwenVL2 vision model to evaluate the aesthetics of the model image, set the prompt words, and the prompt words include: Please evaluate the overall aesthetics of the model in this picture, and score from the following aspects: 1. Whether the pose of the model is natural and elegant; 2. Whether the facial expression is natural and appropriate; 3. Whether the clothing matching is reasonable and beautiful; 4. Whether the overall styling is harmonious; Please give a comprehensive score between 0 and 1, where 1 point means very beautiful and 0 point means not beautiful; Obtain the qwen_beauty_score.

[0045] When model_sharpness_score > 0.7 and qwen_beauty_score > 0.7, mark the model as good-looking as 1 and not good-looking as 0; in other cases, mark the model as good-looking as 0 and not good-looking as 1.

[0046] In this embodiment, preferably, step 2 is specifically: Building the evaluation model architecture, and the evaluation model includes the aesthetic evaluation small model base and the aesthetic evaluation large model base:

[0047] Use torchvision.models.resnet50 to load the pre-trained ResNet50 model as the aesthetic evaluation small model base;

[0048] Use torch.load to load the CLIP model as the aesthetic evaluation large model base;

[0049] Delete the fully connected layers in the output parts of the ResNet50 model and the CLIP model, and use the adaptive pooling layer to obtain model outputs with the same dimension;

[0050] Finally, the output is concatenated into the same output using the concat function, and the final result is obtained using a fully connected layer with an output dimension of 6. The output dimension includes 6 dimensions, corresponding to the probabilities of good color, bad color, good composition, bad composition, good-looking model, and not good-looking model respectively;

[0051] Model training: Input the labeled images into the evaluation model in sequence, and perform fine-tuning training on the evaluation model architecture: Use the torch.optim.Adam optimizer and calculate the error between the predicted score and the set value using the MSE loss function; When the error is less than the set threshold, stop training to obtain the trained evaluation model.

[0052] In this embodiment, preferably, step 3 is specifically:

[0053] Evaluation score calculation and high-quality sample screening:

[0054] Input the pictures to be screened into the trained evaluation model to obtain probability values in six dimensions; Probability normalization and weighted average

[0055] For the probability values in six dimensions obtained after the inference of the trained evaluation model,

[0056] First, use np.normalize for normalization: set the weights of the probability of good color and the probability of bad color to 0.8, the weights of the probability of good composition and the probability of bad composition to 1.0, and the weights of the probability of good-looking model and the probability of bad-looking model to 1.2;

[0057] Use np.average to achieve weighted average to obtain the final evaluation score, and limit the score within the range of 0 to 1 through np.clip;

[0058] High-quality sample screening: Set 0.7 as the screening threshold for high-quality samples. If the score is greater than the screening threshold, it is a high-quality image and is retained; otherwise, it is a non-high-quality image and is deleted.

[0059] Based on the same inventive concept, the present application also provides an apparatus corresponding to the method in Embodiment 1. For details, see Embodiment 2.

[0060] Embodiment 2

[0061] As Figure 2 shown, in this embodiment, a picture selection apparatus for evaluating pictures generated based on real models is provided, including:

[0062] An evaluation module that performs color evaluation, composition evaluation, and model beauty evaluation on the input image, and marks good color, bad color, good composition, bad composition, good-looking model, and bad-looking model with 0 or 1 to obtain a marked image;

[0063] A model acquisition module for building an evaluation model. The evaluation model includes a ResNet50 model and a CLIP model. Delete the fully connected layers in the output parts of the ResNet50 model and the CLIP model, and use an adaptive pooling layer to obtain model outputs with the same dimension; The outputs of the evaluation model are concatenated into the same output using the concat function, and a fully connected layer with an output dimension of 6 is used to obtain the final result. The output dimension includes 6 dimensions, corresponding to the probability values of good color, bad color, good composition, bad composition, good-looking model, and bad-looking model respectively; Train the evaluation model with the marked image to obtain the trained evaluation model;

[0064] A screening module that inputs the pictures to be screened into the trained evaluation model and screens the images according to the scores.

[0065] In this embodiment, preferably, the evaluation module is specifically configured to: perform color evaluation, composition evaluation, and model beauty evaluation on the input image, mark good color, bad color, good composition, bad composition, good-looking model, and bad-looking model with 0 or 1 to obtain a marked image;

[0066] The color evaluation is specifically: color space conversion and channel extraction; use cv2.cvtColor to convert the input image from the RGB color space to the HSV color space, and extract the three channels of hue, saturation, and value respectively;

[0067] Histogram distribution feature calculation: for the three channels of hue, saturation, and value, use cv2.calcHist to calculate the corresponding histogram distribution features respectively; obtain the statistical values of the set quantiles of each channel through np.percentile;

[0068] Calculate the mean H_mean and standard deviation H_std of the hue, the mean S_mean and standard deviation S_std of the saturation, and the mean V_mean and standard deviation V_std of the value; when S_mean is within the range of [0.3, 0.7] and V_mean is within the range of [0.4, 0.8], and at the same time H_std is less than the set first threshold, then mark good color as 1 and bad color as 0; when S_mean < 0.3 or > 0.7, or V_mean < 0.4 or > 0.8, or H_std is greater than or equal to the first set threshold, then mark good color as 0 and bad color as 1;

[0069] The composition evaluation is specifically: clarity and edge feature calculation: use cv2.Laplacian to calculate the clarity score sharpness_score of the image;

[0070] Use the cv2.Sobel operator to extract the edge features in the horizontal and vertical directions of the image, and calculate to obtain edge_strength;

[0071] Symmetry and golden ratio point feature calculation: based on cv2.moments, calculate the centroid position of the image to obtain symmetry_score;

[0072] Divide the image into a 3x3 grid, and calculate the feature response intensity golden_ratio_score at the position of the golden ratio point;

[0073] When sharpness_score > 0.7 and edge_strength is greater than the second set threshold, and at the same time symmetry_score > 0.6 and golden_ratio_score > 0.5, then mark the composition as good (1) and the composition as bad (0); when sharpness_score ≤ 0.7 or edge_strength is less than or equal to the second set threshold, or symmetry_score ≤ 0.6 or golden_ratio_score ≤ 0.5, then mark the composition as good (0) and the composition as bad (1).

[0074] The evaluation of the model's beauty is specifically as follows: Calculation of the model's clarity and beauty scores: Use cv2.Laplacian to calculate the model's clarity score model_sharpness_score;

[0075] Use the QwenVL2 vision model to evaluate the beauty of the model image, and set the prompt words, which include: Please evaluate the overall beauty of the model in this picture, and score from the following aspects: 1. Whether the model's pose is natural and elegant; 2. Whether the facial expression is natural and appropriate; 3. Whether the clothing matching is reasonable and beautiful; 4. Whether the overall styling is harmonious; Please give a comprehensive score between 0 and 1, where 1 point means very beautiful and 0 point means not beautiful; Obtain qwen_beauty_score.

[0076] When model_sharpness_score > 0.7 and qwen_beauty_score > 0.7, mark the model as good-looking (1) and the model as not good-looking (0); in other cases, mark the model as good-looking (0) and the model as not good-looking (1).

[0077] In this embodiment, preferably, the model acquisition module is specifically: Building the evaluation model architecture, and the evaluation model includes the aesthetic evaluation small model base and the aesthetic evaluation large model base:

[0078] Use torchvision.models.resnet50 to load the pre-trained ResNet50 model as the aesthetic evaluation small model base;

[0079] Use torch.load to load the CLIP model as the aesthetic evaluation large model base;

[0080] Delete the fully connected layers in the output parts of the ResNet50 model and the CLIP model, and use the adaptive pooling layer to obtain model outputs with the same dimension;

[0081] The final output is concatenated into the same output using the concat function, and the final result is obtained using a fully connected layer with an output dimension of 6. The output dimension includes 6 dimensions, corresponding to the probabilities of good color, bad color, good composition, bad composition, good-looking model, and not good-looking model respectively;

[0082] Model training: Input the labeled images into the evaluation model in sequence, and fine-tune the evaluation model architecture: Use the torch.optim.Adam optimizer, and use the MSE loss function to calculate the error between the predicted score and the set value; When the error is less than the set threshold, stop training to obtain the trained evaluation model.

[0083] In this embodiment, preferably, the screening module is specifically:

[0084] Evaluation score calculation and high-quality sample screening:

[0085] Input the images to be screened into the trained evaluation model to obtain the probability values of six dimensions; Probability normalization and weighted average

[0086] For the probability values of the six dimensions obtained after the inference of the trained evaluation model,

[0087] First, use np.normalize for normalization: Set the weights of the probability of good color and the probability of bad color to 0.8, the weights of the probability of good composition and the probability of bad composition to 1.0, and the weights of the probability of good-looking model and the probability of not good-looking model to 1.2;

[0088] Use np.average to achieve weighted average to obtain the final evaluation score, and limit the score within the range of 0 to 1 through np.clip;

[0089] High-quality sample screening: Set 0.7 as the screening threshold for high-quality samples. If the score is greater than the screening threshold, it is a high-quality image and is retained; otherwise, it is a non-high-quality image and is deleted.

[0090] Since the device introduced in the second embodiment of the present invention is the device adopted for implementing the method of the first embodiment of the present invention, based on the method introduced in the first embodiment of the present invention, those skilled in the art can understand the specific structure and deformation of the device, so it will not be elaborated here. Any device adopted for the method of the first embodiment of the present invention belongs to the scope of protection of the present invention.

[0091] Based on the same inventive concept, this application provides an electronic device embodiment corresponding to the first embodiment, as shown in Embodiment Three.

[0092] Embodiment Three

[0093] This embodiment provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, any implementation manner in Embodiment 1 can be implemented.

[0094] Since the electronic device introduced in this embodiment is the device used to implement the method in Embodiment 1 of this application, based on the method introduced in Embodiment 1 of this application, those skilled in the art can understand the specific implementation manners and various forms of variation of the electronic device in this embodiment. Therefore, the specific implementation of how this electronic device implements the method in the embodiments of this application will not be described in detail here. As long as the device used by those skilled in the art to implement the method in the embodiments of this application belongs to the scope protected by this application.

[0095] Based on the same inventive concept, this application provides a storage medium corresponding to Embodiment 1, details of which can be found in Embodiment 4.

[0096] Embodiment 4

[0097] This embodiment provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, any implementation manner in Embodiment 1 can be implemented.

[0098] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.

[0099] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0100] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to work in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the functions specified in one or more processes and / or blocks Figure 1 in one or more processes and / or blocks Figure 1 specified in the function.

[0101] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are executed on the computer or other programmable apparatus to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one or more processes and / or blocks Figure 1 in one or more processes and / or blocks Figure 1 specified in the function.

[0102] Although the specific embodiments of the present invention have been described above, those skilled in the art of this technology should understand that the specific embodiments we described are illustrative rather than limiting the scope of the present invention. Equivalent modifications and variations made by those skilled in the art in accordance with the spirit of the present invention should be covered by the scope of the claims of the present invention.

Claims

1. A method for selecting images based on evaluation of images generated by real models, characterized in that: The steps include: Step 1, evaluate the color, composition and model beauty of the input image, and use 0 or 1 to mark good color, bad color, good composition, bad composition, good model and bad model to obtain a marked image; Step 2: Building an evaluation model, wherein the evaluation model includes a ResNet50 model and a CLIP model. The fully connected layers of the output parts of the ResNet50 model and the CLIP model are deleted, and an adaptive pooling layer is used to obtain model outputs of the same dimension. The evaluation model outputs are concatenated into the same output using a concat function, and a fully connected layer with an output dimension of 6 is used to obtain the final result. The output dimension includes 6 dimensions, which correspond to the probability values ​​of good color, bad color, good composition, bad composition, good-looking model, and bad-looking model respectively. Training the evaluation model using the labeled images to obtain a trained evaluation model; Step 3: Input the images to be screened into the trained evaluation model and perform image screening based on the scores.

2. The image selection method based on evaluation of images generated by real models according to claim 1, characterized in that: The step 1 specifically includes: performing color evaluation, composition evaluation and model beauty evaluation on the input image, marking good color, bad color, good composition, bad composition, good model and bad model with 0 or 1 to obtain a marked image; Color evaluation is specifically as follows: color space conversion and channel extraction; using cv2.cvtColor to convert the input image from RGB color space to HSV color space, and extracting the three channels of hue, saturation and brightness respectively; Histogram distribution feature calculation: For the three channels of hue, saturation and brightness, use cv2.calcHist to calculate the corresponding histogram distribution features respectively; use np.percentile to obtain the statistical value of the set quantile of each channel; Calculate the mean H_mean and standard deviation H_std of hue, the mean S_mean and standard deviation S_std of saturation, and the mean V_mean and standard deviation V_std of brightness; when S_mean is in the range of [0.3, 0.7] and V_mean is in the range of [0.4, 0.8], and H_std is less than the set first threshold, then the good color is recorded as 1, and the bad color is recorded as 0; when S_mean<0.3 or>0.7, or V_mean<0.4 or>0.8, or H_std is greater than or equal to the first set threshold, then the good color is recorded as 0, and the bad color is recorded as 1; The composition evaluation is as follows: cv2.Laplacian is used to calculate the sharpness score of the image; Use cv2.Sobel operator to extract the edge features in the horizontal and vertical directions of the image, and calculate edge_strength; Calculate the image centroid position based on cv2.moments and get the symmetry_score; Divide the image into a 3x3 grid and calculate the feature response intensity golden_ratio_score at the golden section point; When sharpness_score>0.7 and edge_strength is greater than the second set threshold, and symmetry_score>0.6 and golden_ratio_score>0.5, good composition is recorded as 1, and bad composition is recorded as 0; when sharpness_score≤0.7 or edge_strength is less than or equal to the second set threshold, or symmetry_score≤0.6 or golden_ratio_score≤0.5, good composition is recorded as 0, and bad composition is recorded as 1; The model beauty evaluation is as follows: use cv2.Laplacian to calculate the model's sharpness score model_sharpness_score; Use the QwenVL2 visual model to score the beauty of the model image, set the prompt word, and obtain qwen_beauty_score; When model_sharpness_score>0.7 and qwen_beauty_score>0.7, the model is good-looking as 1 and the model is not good-looking as 0; in other cases, the model is good-looking as 0 and the model is not good-looking as 1.

3. The image selection method based on evaluation of images generated by real models according to claim 1, characterized in that: The step 2 is specifically as follows: constructing an evaluation model framework, wherein the evaluation model includes a small aesthetic evaluation model base and a large aesthetic evaluation model base: Use torchvision.models.resnet50 to load the pre-trained ResNet50 model as the base of the aesthetic evaluation model; Use torch.load to load the CLIP model as the basis of the aesthetic evaluation model; Delete the fully connected layers in the output part of the ResNet50 model and the CLIP model, and use the adaptive pooling layer to obtain the model output of the same dimension; The final output is concatenated into the same output using the concat function, and the final result is obtained using a fully connected layer with an output dimension of 6. The output dimension includes 6 dimensions, corresponding to the probability values ​​of good color, bad color, good composition, bad composition, good model, and bad model; Model training: Input the labeled images into the evaluation model in sequence, and fine-tune the evaluation model architecture: Use the torch.optim.Adam optimizer and the MSE loss function to calculate the error between the predicted score and the set value; when the error is less than the set threshold, stop training and obtain the trained evaluation model.

4. The image selection method based on evaluation of images generated by real models according to claim 1, characterized in that: The step 3 is specifically as follows: Evaluation score calculation and high-quality sample screening: Input the images to be screened into the trained evaluation model to obtain the probability values ​​of the six dimensions; Probability Normalization and Weighted Average The probability values ​​of the six dimensions obtained after inference of the training evaluation model, First, use np.normalize to perform normalization: set the weights of the probability of good color and bad color to 0.8, the weights of the probability of good composition and bad composition to 1.0, and the weights of the probability of good model and bad model to 1.2; Use np.average to achieve weighted averaging to get the final evaluation score, and use np.clip to limit the score to the range of 0 to 1; High-quality sample screening: 0.7 is set as the screening threshold for high-quality samples. If the score is greater than the screening threshold, it is a high-quality image and is retained; If not, it is a low-quality image and will be deleted.

5. A device for selecting images based on evaluation of images generated by real models, characterized in that: include: An evaluation module evaluates the color, composition and model beauty of the input image, and marks good color, bad color, good composition, bad composition, good model and bad model with 0 or 1 to obtain a marked image; Model acquisition module, evaluation model construction, the evaluation model includes ResNet50 model and CLIP model, delete the fully connected layer of the output part of ResNet50 model and CLIP model, use adaptive pooling layer to obtain model output of the same dimension; the evaluation model output is spliced ​​into the same output using concat function, and the final result is obtained using a fully connected layer with an output dimension of 6. The output dimension includes 6 dimensions, which correspond to the probability values ​​of good color, bad color, good composition, bad composition, good model and bad model; Training the evaluation model using the labeled images to obtain a trained evaluation model; The screening module inputs the images to be screened into the trained evaluation model and screens the images based on the scores.

6. The image selection device based on the evaluation of the image generated by the real model according to claim 5, characterized in that: The evaluation module specifically performs color evaluation, composition evaluation and model beauty evaluation on the input image, and marks good color, bad color, good composition, bad composition, good model and bad model with 0 or 1 to obtain a marked image; Color evaluation is specifically as follows: color space conversion and channel extraction; using cv2.cvtColor to convert the input image from RGB color space to HSV color space, and extracting the three channels of hue, saturation and brightness respectively; Histogram distribution feature calculation: For the three channels of hue, saturation and brightness, use cv2.calcHist to calculate the corresponding histogram distribution features respectively; use np.percentile to obtain the statistical value of the set quantile of each channel; Calculate the mean H_mean and standard deviation H_std of hue, the mean S_mean and standard deviation S_std of saturation, and the mean V_mean and standard deviation V_std of brightness; when S_mean is in the range of [0.3, 0.7] and V_mean is in the range of [0.4, 0.8], and H_std is less than the set first threshold, then the good color is recorded as 1, and the bad color is recorded as 0; when S_mean<0.3 or>0.7, or V_mean<0.4 or>0.8, or H_std is greater than or equal to the first set threshold, then the good color is recorded as 0, and the bad color is recorded as 1; The composition evaluation is as follows: cv2.Laplacian is used to calculate the sharpness score of the image; Use cv2.Sobel operator to extract the edge features in the horizontal and vertical directions of the image, and calculate edge_strength; Calculate the image centroid position based on cv2.moments and get the symmetry_score; Divide the image into a 3x3 grid and calculate the feature response intensity golden_ratio_score at the golden section point; When sharpness_score>0.7 and edge_strength is greater than the second set threshold, and symmetry_score>0.6 and golden_ratio_score>0.5, good composition is recorded as 1, and bad composition is recorded as 0; when sharpness_score≤0.7 or edge_strength is less than or equal to the second set threshold, or symmetry_score≤0.6 or golden_ratio_score≤0.5, good composition is recorded as 0, and bad composition is recorded as 1; Model beauty evaluation is as follows: Model clarity and beauty score calculation: Use cv2.Laplacian to calculate the model's clarity score model_sharpness_score; Use the QwenVL2 visual model to score the beauty of the model image, set the prompt word, and obtain qwen_beauty_score; When model_sharpness_score>0.7 and qwen_beauty_score>0.7, the model is good-looking as 1 and the model is not good-looking as 0; in other cases, the model is good-looking as 0 and the model is not good-looking as 1.

7. The image selection device based on evaluation of images generated by real models according to claim 5, characterized in that: The model acquisition module specifically includes: building an evaluation model framework, where the evaluation model includes an aesthetic evaluation small model base and an aesthetic evaluation large model base: Use torchvision.models.resnet50 to load the pre-trained ResNet50 model as the base of the aesthetic evaluation model; Use torch.load to load the CLIP model as the basis of the aesthetic evaluation model; Delete the fully connected layers in the output part of the ResNet50 model and the CLIP model, and use the adaptive pooling layer to obtain the model output of the same dimension; The final output is concatenated into the same output using the concat function, and the final result is obtained using a fully connected layer with an output dimension of 6. The output dimension includes 6 dimensions, corresponding to the probability values ​​of good color, bad color, good composition, bad composition, good model, and bad model; Model training: Input the labeled images into the evaluation model in sequence, and fine-tune the evaluation model architecture: Use the torch.optim.Adam optimizer and the MSE loss function to calculate the error between the predicted score and the set value; when the error is less than the set threshold, stop training and obtain the trained evaluation model.

8. The image selection device based on evaluation of images generated by real models according to claim 5, characterized in that: The screening module is specifically: Evaluation score calculation and high-quality sample screening: Input the images to be screened into the trained evaluation model to obtain the probability values ​​of the six dimensions; Probability Normalization and Weighted Average The probability values ​​of the six dimensions obtained after inference of the training evaluation model, First, use np.normalize to perform normalization: set the weights of the probability of good color and bad color to 0.8, the weights of the probability of good composition and bad composition to 1.0, and the weights of the probability of good model and bad model to 1.2; Use np.average to achieve weighted averaging to get the final evaluation score, and use np.clip to limit the score to the range of 0 to 1; High-quality sample screening: 0.7 is set as the screening threshold for high-quality samples. If the score is greater than the screening threshold, it is a high-quality image and is retained; If not, it is a low-quality image and will be deleted.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the method according to any one of claims 1 to 4 is implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 4 is implemented.