Model training method and device, image evaluation method and device and electronic equipment
By determining a small number of labeled sample images from the user images and extending the second sample image for model retraining, the problem of low accuracy of the open source data set is solved and efficient image aesthetic evaluation is achieved.
Patent Information
- Application Number
- CN202510175673.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-07-04
AI Technical Summary
The image aesthetic evaluation model using open source aesthetic evaluation datasets in the prior art is low in accuracy and availability, and re-customizing the dataset requires a lot of labor and time costs.
By determining a small number of labeled first sample images from the images submitted by the user, using pre-trained image features to extract and evaluate the model, expanding a large number of second sample images for model retraining, and transferring labels using label search, feature search or teacher model methods to train an accurate image evaluation model.
It reduces the labor and time cost of training data, improves the accuracy of model evaluation images, and meets the project's aesthetic evaluation needs.
Smart Images

Figure CN120259809A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of data processing, and in particular, to a model training method, an image evaluation method, an apparatus, and an electronic device. Background Art
[0002] Image aesthetics evaluation is to quantify the aesthetic quality of a given image. Usually, it is evaluated according to aesthetic-related criteria such as the color, composition, lines, and light of the image. In related technologies, deep learning methods are usually used for aesthetics evaluation. For example, an aesthetics evaluation model is obtained by training a deep model using an aesthetics evaluation data set. However, when using an open-source aesthetics evaluation data set, the characteristics of its image data and the criteria for aesthetic evaluation may be quite different from the actual evaluation requirements, with low accuracy and usability and being difficult to directly apply; if a new aesthetics evaluation data set is customized according to the actual evaluation requirements, it will require a large amount of human and time costs. Summary of the Invention
[0003] In view of this, the purpose of the present disclosure is to provide a model training method, an image evaluation method, an apparatus, and an electronic device. A small number of labeled first sample images are determined from the images submitted by the user, and then a large number of second sample images are expanded from the images submitted by the user according to the first sample images. The pre-trained model is retrained with the second sample images to reduce the human and time costs of determining training data, so that the finally obtained model can learn accurate evaluation criteria and improve the accuracy of the model in evaluating images.
[0004] In a first aspect, an embodiment of the present disclosure provides a model training method, which includes: determining first sample images from the images submitted by the user according to a pre-trained image feature extraction model and a pre-trained image evaluation model; wherein, the first sample images carry image annotations, and the image annotations are used to indicate the image evaluation results; determining target images similar to the first sample images from the images submitted by the user through a preset method, copying the image annotations carried by the first sample images to the target images, and determining the target images as second sample images; training the pre-trained image feature extraction model with the second sample images to obtain a trained image feature extraction model; training the pre-trained image evaluation model with the second sample images to obtain a trained image evaluation model.
[0005] Second aspect, embodiments of the present disclosure provide an image evaluation method, the method comprising: obtaining a specified image from an image submitted by a user, inputting the specified image into a trained image feature extraction model obtained by any of the model training methods in the first aspect, to obtain an image feature vector corresponding to the specified image; inputting the image feature vector into a trained image evaluation model obtained by any of the model training methods in the first aspect, to obtain an image evaluation result of the specified image.
[0006] The above-mentioned trained image evaluation model includes at least one image evaluation sub-model; the step of inputting the image feature vector into a trained image evaluation model obtained by any of the model training methods in the first aspect to obtain an image evaluation result of the specified image includes: for each image evaluation sub-model, inputting the image feature vector into the image evaluation sub-model to obtain an initial evaluation result corresponding to the image evaluation sub-model; performing data processing on the obtained at least one initial evaluation result to obtain an image evaluation result of the specified image.
[0007] Third aspect, embodiments of the present disclosure provide a model training apparatus, the apparatus comprising: a first sample image determination module, configured to determine a first sample image from an image submitted by a user according to a pre-trained image feature extraction model and a pre-trained image evaluation model; wherein the first sample image carries an image annotation, and the image annotation is used to indicate an image evaluation result; a second sample image determination module, configured to determine a target image similar to the first sample image from an image submitted by a user by a preset method, copy the image annotation carried by the first sample image to the target image, and determine the target image as the second sample image; a model training module, configured to train the pre-trained image feature extraction model with the second sample image to obtain a trained image feature extraction model; and train the pre-trained image evaluation model with the second sample image to obtain a trained image evaluation model.
[0008] Fourth aspect, embodiments of the present disclosure provide an image evaluation apparatus, the apparatus comprising: an image feature vector extraction module, configured to obtain a specified image from an image submitted by a user, input the specified image into a trained image feature extraction model obtained by any of the model training methods in the first aspect, to obtain an image feature vector corresponding to the specified image; an image evaluation result determination module, configured to input the image feature vector into a trained image evaluation model obtained by any of the model training methods in the first aspect, to obtain an image evaluation result of the specified image.
[0009] Fifth aspect, embodiments of the present disclosure provide an electronic device, comprising a processor and a memory, the memory storing computer executable instructions that can be executed by the processor, and the processor executing the computer executable instructions to implement any of the model training methods in the first aspect, or implement any of the image evaluation methods in the second aspect.
[0010] In a sixth aspect, embodiments of the present disclosure provide a computer-readable storage medium storing computer-executable instructions, which, when called and executed by a processor, cause the processor to implement the model training method according to any one of the first aspect, or implement the image evaluation method according to any one of the second aspect.
[0011] The embodiments of the present disclosure bring the following beneficial effects:
[0012] The present disclosure provides a model training method, an image evaluation method, an apparatus, and an electronic device. According to a pre-trained image feature extraction model and a pre-trained image evaluation model, a first sample image is determined from the images submitted by a user; wherein, the first sample image carries an image annotation for indicating an image evaluation result; through a preset method, a target image similar to the first sample image is determined from the images submitted by the user, the image annotation carried by the first sample image is copied onto the target image, and the target image is determined as a second sample image; the pre-trained image feature extraction model is trained by the second sample image to obtain a trained image feature extraction model; the pre-trained image evaluation model is trained by the second sample image to obtain a trained image evaluation model. In this way, a small number of labeled first sample images are determined from the images submitted by the user, and then a large number of second sample images are expanded from the images submitted by the user according to the first sample image. The pre-trained models are trained again by the second sample images, reducing the human and time costs for determining training data, enabling the finally obtained model to learn accurate evaluation criteria, and improving the accuracy of the model for evaluating images.
[0013] Other features and advantages of the present disclosure will be described in the following specification, and part of them will be obvious from the specification, or will be understood by implementing the present disclosure. The objectives and other advantages of the present disclosure are achieved and obtained by the structures specifically pointed out in the specification, the claims, and the drawings.
[0014] To make the above objectives, features, and advantages of the present disclosure more obvious and understandable, the following specific preferred embodiments are given, and in conjunction with the accompanying drawings, the detailed description is as follows. Description of the Drawings
[0015] To more clearly illustrate the specific embodiments of the present disclosure or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the following drawings are some embodiments of the present disclosure. For those skilled in the art, without creative efforts, other drawings can be obtained according to these drawings.
[0016] Figure 1 Flow chart of a model training method provided by an embodiment of the present disclosure;
[0017] Figure 2 Flow chart of an image evaluation method provided by an embodiment of the present disclosure;
[0018] Figure 3 Structural schematic diagram of a model training device provided by an embodiment of the present disclosure;
[0019] Figure 4 Structural schematic diagram of an image evaluation device provided by an embodiment of the present disclosure;
[0020] Figure 5 Structural schematic diagram of an electronic device provided by an embodiment of the present disclosure. Detailed implementation manners
[0021] To make the objectives, technical solutions and advantages of the embodiments of the present disclosure clearer, the technical solutions of the present disclosure will be clearly and completely described below with reference to the accompanying drawings. Apparently, the described embodiments are some but not all of the embodiments of the present disclosure. Based on the embodiments in the present disclosure, all other embodiments obtained by those skilled in the art without creative efforts fall within the scope of protection of the present disclosure.
[0022] Image aesthetics evaluation is to quantify the aesthetic quality of a given image. Usually, it is evaluated according to aesthetic-related criteria such as the color, composition, lines, and light of the image. Since different people's feelings about the same picture are subjective and diverse, in the problem of image aesthetics evaluation, what is studied is the common aesthetic feelings and evaluations of people about pictures, so as to distinguish the quality level between different pictures.
[0023] In some related technologies, a certain number of testers are used to conduct aesthetic evaluations on the images to be tested. By ensuring that the number of testers is higher than a certain number, it is ensured that the obtained results can represent general aesthetic evaluation conclusions. When using this method, before the manual evaluation, the evaluation rules and scales that the testers need to follow can be agreed in advance, and then the aesthetic evaluations performed by the testers are all based on this, so as to obtain results that more meet the project requirements. However, to ensure the usability of the evaluation results, a certain number of evaluators need to be arranged, and the required labor cost is relatively high. In addition, when using the manual method for aesthetic evaluation, the processing speed is generally slow, and it is difficult to meet the real-time and throughput requirements of online services.
[0024] In some related technologies, deep learning methods are usually used for aesthetic evaluation. For example, a deep model is trained using an aesthetic evaluation dataset to obtain an aesthetic evaluation model. However, when using an open-source aesthetic evaluation dataset, the characteristics of the image data and the criteria for aesthetic evaluation may be quite different from the actual evaluation requirements, resulting in low accuracy and usability and being difficult to apply directly. If a new aesthetic evaluation dataset is customized according to the actual evaluation requirements, it will require a large amount of human and time costs. Based on this, an embodiment of the present disclosure provides a model training method, an image evaluation method, a device, and an electronic device. This technology can be applied to devices such as mobile phones, laptops, computers, and servers, especially to devices configured with image evaluation functions.
[0025] To facilitate the understanding of this embodiment, a model training method disclosed in an embodiment of the present disclosure will be introduced in detail first. This model training method generally refers to a model training method for image aesthetic evaluation. As Figure 1 shown, the method includes the following steps:
[0026] Step S102, determine a first sample image from the images submitted by the user according to a pre-trained image feature extraction model and a pre-trained image evaluation model; wherein, the first sample image carries an image annotation, and the image annotation is used to indicate the image evaluation result;
[0027] The above-mentioned pre-trained image feature extraction model can be obtained by training a deep learning model in advance through an open-source dataset, or an open-source pre-trained deep learning model can be directly used as the above-mentioned pre-trained image feature extraction model. The above-mentioned pre-trained image evaluation model can be obtained by training a deep learning model in advance through a large amount of labeled data, and the large amount of labeled data can be an open-source dataset or a part of the data selected from the open-source dataset.
[0028] Optionally, the preliminary evaluation results of some of the images submitted by the user can be determined through the pre-trained image feature extraction model and the pre-trained image evaluation model, and the part of the images submitted by the user and the preliminary evaluation results can be directly determined as the first sample image. It is also possible to manually verify the preliminary evaluation results, re-evaluate the images that do not conform to the actual evaluation results to obtain the manual evaluation results, and determine the part of the images submitted by the user and the manual evaluation results of the images as the first sample image. The above-mentioned image annotation is used to indicate the standard evaluation result of the image.
[0029] Step S104, determine a target image similar to the first sample image from the images submitted by the user through a preset method, copy the image annotation carried by the first sample image to the target image, and determine the target image as the second sample image;
[0030] Using only generalized aesthetic standards is often insufficient to meet the aesthetic needs of a project. For example, colorful pictures can arouse people's perception of beauty more effectively, while in practical applications, it is desired that well-drawn black-and-white pictures can also be recommended equally. In such cases, based on the first sample image, and according to the project requirements, the sample needs to be expanded in a specific direction. The expansion of the project dataset depends on the data provided by the project (i.e., the first sample image).
[0031] The above preset methods are any one of the following: label retrieval method, feature retrieval method, teacher model method.
[0032] Optionally, the above first sample image includes an image label, which is usually used to indicate the image type, image style, etc. of the image. The image submitted by the user also includes an image label, which is usually selected by the user or can be automatically configured by the model. Through the label retrieval method, images similar to the image label of the first sample image are determined from the images submitted by the user based on the first sample image, and the image annotation of the first sample image is directly transferred to this image to obtain the second sample image.
[0033] Optionally, through the feature retrieval method, a pre-trained image feature extraction model is used to extract image feature vectors from the first sample image and the images submitted by the user. Then, according to the similarity of the feature vectors, images with a high similarity to the first sample image are obtained from the images submitted by the user, and then the image annotation of the first sample image can be directly transferred to this image to obtain the second sample image.
[0034] Optionally, through the teacher model method, first use the first sample image to train a teacher model. Then, input the image entered by the user into the teacher model, output the evaluation result of the image, and determine this image and the evaluation result of the image as the second sample image.
[0035] Step S106, use the second sample image to train the pre-trained image feature extraction model to obtain a trained image feature extraction model; use the second sample image to train the pre-trained image evaluation model to obtain a trained image evaluation model.
[0036] Input the second sample image into the pre-trained image feature extraction model, obtain the image feature vector of the second sample image through the forward propagation algorithm, calculate the loss value according to the loss function, and then calculate the gradient through the backward propagation algorithm to update the model parameters until the termination condition is reached to obtain the trained image feature extraction model, where the termination condition can be reaching the specified number of iterations, the loss value meeting the preset loss value, etc.
[0037] Similarly, input the image feature vector of the second sample image into the image evaluation model to obtain the predicted evaluation result of the second sample image. Update the model parameters according to the predicted evaluation result and the image annotation of the second sample image until the difference between the predicted evaluation result and the image annotation of the second sample image meets the preset threshold, and obtain the trained image evaluation model.
[0038] Optionally, directly train the pre-trained image feature extraction model with the images submitted by the user to obtain the trained image feature extraction model. For example, at the initial stage of the project, an open-source pre-trained deep learning model can be directly used as the image feature extraction module. However, since there may be significant differences in the feature distributions between the training data of the open-source model and the business data of the project, which may affect the effect of downstream tasks, after the project stabilizes, a large amount of unlabeled project data (i.e., the images submitted by the user) can be used to train the pre-trained image feature extraction model, or a deep learning model can be re-trained as the project-specific image feature extraction model to enable it to more effectively extract the feature information of project images.
[0039] As the upper limit time of the model is longer, the business regularly checks the prediction results of the online aesthetic evaluation service through manual review and spot checks, and organizes the results of manual inspections into the labeled data of aesthetic evaluation (that is, new second sample images can be continuously directly obtained without re-determination), and gives regular feedback to iteratively update the image evaluation model regularly to continuously improve the model accuracy.
[0040] A model training method provided by an embodiment of the present disclosure determines a first sample image from the images submitted by the user according to a pre-trained image feature extraction model and a pre-trained image evaluation model; wherein, the first sample image carries an image annotation, and the image annotation is used to indicate the image evaluation result; by a preset method, determine a target image similar to the first sample image from the images submitted by the user, copy the image annotation carried by the first sample image to the target image, and determine the target image as the second sample image; train the pre-trained image feature extraction model with the second sample image to obtain the trained image feature extraction model; train the pre-trained image evaluation model with the second sample image to obtain the trained image evaluation model. In this way, a small number of labeled first sample images are determined from the images submitted by the user, and then a large number of second sample images are expanded from the images submitted by the user according to the first sample image. The pre-trained model is trained again with the second sample image, reducing the human and time costs of determining training data, enabling the finally obtained model to learn accurate evaluation criteria, and improving the accuracy of the model for evaluating images.
[0041] For the step of determining the first sample image from the images submitted by the user, a possible implementation is as follows: obtain a first image from the images submitted by the user; determine the predicted evaluation result of the first image according to a pre-trained image feature extraction model and a pre-trained image evaluation model; verify the predicted evaluation result to obtain the standard evaluation result of the first image; determine the first image and the standard evaluation result of the first image as the first sample image.
[0042] Generally, the predicted evaluation result is verified manually. The main function of the image feature extraction model is to compress and encode the input image into a feature vector of a fixed length. Objectively speaking, these image feature vectors should contain as much feature information of the image itself as possible to describe the main features of the input image with a limited amount of data. The richer the feature information extracted by the image feature extraction module, the more available information it has when applied to the aesthetic evaluation task, thus improving the final accuracy.
[0043] The above-mentioned image evaluation model includes at least one image evaluation sub-model; different image evaluation sub-models are used to indicate different image evaluation criteria; the image annotation carried by the second sample image includes at least one image sub-label; the image evaluation sub-model corresponds to the image sub-label one by one.
[0044] In practical applications, the aesthetic criteria of images may cover multiple aspects, such as color richness, rationality of human body structure, completion degree of the picture, whether it is an AI drawing, etc. If only a single deep learning model is used to predict the evaluation result, it may be difficult to meet the project requirements. Therefore, according to the specific requirements of the project, the aesthetic evaluation criteria should be decomposed into several sub-problems (i.e., different image evaluation criteria), and then for each sub-problem, a corresponding deep learning model (i.e., the image evaluation sub-model) is trained. In this way, the image evaluation model is classified according to the image evaluation criteria, and different image evaluation models are used to evaluate the score of a certain aspect of the image, improving the accuracy of the evaluation result.
[0045] For the step of training the image evaluation model with the second sample image to obtain the trained image evaluation model, a possible implementation is as follows: input the second sample image into the trained image feature extraction model to obtain the image feature vector of the second sample image; for each image evaluation sub-model, input the image feature vector of the second sample image into the image evaluation sub-model to obtain the initial evaluation result of the second sample image; update the image evaluation sub-model according to the initial evaluation result and the image sub-label corresponding to the image evaluation sub-model until the termination condition is reached to obtain the trained image evaluation sub-model.
[0046] Each image evaluation sub-model needs to be trained. For example, for problems related to the aesthetics of pictures, an image evaluation sub-model can be trained to evaluate the aesthetics of images. For problems related to the human body structure, an image evaluation sub-model can be trained to evaluate the human body structure in images, etc.
[0047] In the above method, by training multiple image evaluation sub-models, images can be evaluated from different aspects, which simplifies the model training process.
[0048] The above method further includes: obtaining first training data, training an image feature extraction model with the first training data to obtain a pre-trained image feature extraction model; obtaining second training data, training an image evaluation model with the second training data to obtain a pre-trained image evaluation model.
[0049] The above first training data can be an open-source dataset or other images. The above pre-training methods can be at least any one of the following: masked image modeling method, contrastive learning, image classification pre-training method, etc.
[0050] For example, using the Masked Image Modeling (MIM) method, the first training data (unlabeled images) for pre-training is divided into several blocks (split into multiple small patches), encoded as visual tokens (i.e., each small patch is transformed into a low-dimensional vector through embedding), and then the information of some visual tokens is randomly masked, and the model to be trained is required to reconstruct the masked visual tokens, thereby training the image feature extraction model.
[0051] If the contrastive learning or image classification pre-training method is used to train the image feature extraction module, the first training data is labeled images.
[0052] Optionally, the above second training data includes a first training image and a second training image. The first training image carries a first label, and the second training image carries a second label; wherein, the first label indicates that the image score is greater than a first preset threshold, and the second label indicates that the image score is less than a second preset threshold.
[0053] The above second training data usually refers to labeled images, which can be obtained from an open-source dataset. For example, by using an open-source model or an open-source dataset to obtain a batch of relatively certain high-quality picture data (i.e., the first training images) and low-quality picture data (the second training images), high-quality pictures (high-quality picture data) and low-quality pictures (low-quality picture data) with high confidence can be screened out by setting high and low score thresholds as the second training data.
[0054] Generally, for pictures with particularly high aesthetic evaluations, their image quality can also be recognized by most people and it is not easy to have misjudgments for project requirements; conversely, pictures with particularly low aesthetic evaluations are mostly low-quality pictures that need to be blocked within the project. For pictures with aesthetic evaluations in the middle part, there are relatively more disagreements between their evaluation results and project requirements. Therefore, high-quality pictures and low-quality pictures can be used as secondary training data in the initial stage of the project to improve the accuracy of the model.
[0055] An embodiment of the present disclosure provides an image evaluation method, as Figure 2 shown, the method includes the following steps:
[0056] Step S202, obtain a specified image from the images submitted by the user, input the specified image into the trained image feature extraction model obtained by the model training method in the foregoing embodiment, and obtain an image feature vector corresponding to the specified image;
[0057] The main function of the image feature extraction module is to compress and encode the input image into a feature vector of a fixed length. Objectively speaking, these feature vectors should contain as much feature information of the image itself as possible to describe the main features of the input image with a limited amount of data. The richer the feature information extracted by this module, the more available information there will be when it is applied to the aesthetic evaluation task, thereby improving the final accuracy.
[0058] Step S204, input the image feature vector into the trained image evaluation model obtained by the model training method in the foregoing embodiment, and obtain an image evaluation result of the specified image.
[0059] Optionally, the above-mentioned trained image evaluation model includes at least one image evaluation sub-model; optionally, for each image evaluation sub-model, input the image feature vector into the image evaluation sub-model to obtain an initial evaluation result corresponding to the image evaluation sub-model; perform data processing on the obtained at least one initial evaluation result to obtain an image evaluation result of the specified image.
[0060] Optionally, a weighted average calculation can be performed on the at least one initial evaluation result to obtain an image evaluation result of the specified image. Optionally, an average calculation can be performed on the at least one initial evaluation result to obtain an image evaluation result of the specified image. The specific data processing method can be set according to actual needs. In this method, the image is evaluated from different aspects by different image evaluation sub-models, and finally multiple evaluation results are processed, further improving the accuracy of image evaluation.
[0061] An embodiment of the present disclosure provides an image evaluation method, which obtains a specified image from the images submitted by a user, inputs the specified image into a trained image feature extraction model obtained by any of the model training methods in the first aspect, and obtains an image feature vector corresponding to the specified image; inputs the image feature vector into a trained image evaluation model obtained by any of the model training methods in the first aspect, and obtains an image evaluation result of the specified image. In this method, a small number of labeled first sample images are determined from the images submitted by the user, and then a large number of second sample images are expanded from the images submitted by the user according to the first sample images. The pre-trained model is trained again through the second sample images, reducing the manpower and time costs for determining training data, enabling the finally obtained model to learn accurate evaluation criteria, and improving the accuracy of the model for evaluating images.
[0062] The above method further includes: determining the specified image and the image evaluation result of the specified image as a third sample image; training the trained image evaluation model through the third sample image to update the model parameters of the trained image evaluation model.
[0063] After the model has been on the line for a period of time and the project is stable, the images submitted by the user and the evaluation results obtained by the model are regularly used as new training data to perform iterative training on the model. Specifically, the evaluation results are manually checked, and a certain amount of labeled data can be obtained from the business, providing dataset expansion for the iteration of the image evaluation model and improving the accuracy of the model.
[0064] Whether it is a method based on deep learning or a method based on manual evaluation, the aesthetic evaluation results obtained by a single method often have inevitable errors, which affect the effect of aesthetic evaluation in business use. The present disclosure uses a deep learning-based image feature extraction model and an image evaluation model to obtain aesthetic evaluation results through calculation. In practical applications, when the aesthetic evaluation results affect sensitive business operations, such as image shielding, artist signing, etc., manual recheck or spot check can be used to reduce the impact brought by the errors.
[0065] Corresponding to the embodiment of the above model training method, an embodiment of the present disclosure provides a model training device, as Figure 3 shown. The device includes:
[0066] A first sample image determination module 301, configured to determine a first sample image from the images submitted by the user according to a pre-trained image feature extraction model and a pre-trained image evaluation model; wherein, the first sample image carries an image annotation, and the image annotation is used to indicate an image evaluation result;
[0067] The second sample image determination module 302 is configured to determine, by a preset method, a target image similar to the first sample image from the images submitted by the user, copy the image annotation carried by the first sample image to the target image, and determine the target image as the second sample image;
[0068] The model training module 303 is configured to train the pre-trained image feature extraction model with the second sample image to obtain a trained image feature extraction model; and train the pre-trained image evaluation model with the second sample image to obtain a trained image evaluation model.
[0069] The embodiment of the present disclosure provides a model training device, which determines a first sample image from the images submitted by the user according to the pre-trained image feature extraction model and the pre-trained image evaluation model; wherein, the first sample image carries an image annotation, and the image annotation is used to indicate the image evaluation result; by a preset method, a target image similar to the first sample image is determined from the images submitted by the user, the image annotation carried by the first sample image is copied to the target image, and the target image is determined as the second sample image; the pre-trained image feature extraction model is trained with the second sample image to obtain a trained image feature extraction model; and the pre-trained image evaluation model is trained with the second sample image to obtain a trained image evaluation model. In this way, a small number of labeled first sample images are determined from the images submitted by the user, and then a large number of second sample images are expanded from the images submitted by the user according to the first sample image. The pre-trained model is trained again with the second sample image, reducing the manpower and time costs for determining the training data, enabling the finally obtained model to learn accurate evaluation criteria, and improving the accuracy of the model for evaluating images.
[0070] The above first sample image determination module is further configured to: obtain a first image from the images submitted by the user; determine the predicted evaluation result of the first image according to the pre-trained image feature extraction model and the pre-trained image evaluation model; verify the predicted evaluation result to obtain the standard evaluation result of the first image; and determine the first image and the standard evaluation result of the first image as the first sample image.
[0071] The above preset method is any one of the following: label retrieval method, feature retrieval method, teacher model method.
[0072] The above image evaluation model includes at least one image evaluation sub-model; different image evaluation sub-models are used to indicate different image evaluation criteria; the image annotation carried by the second sample image includes at least one image sub-label; and the image evaluation sub-model corresponds to the image sub-label one by one.
[0073] The above model training module is further configured to: input the second sample image into the trained image feature extraction model to obtain the image feature vector of the second sample image; for each image evaluation sub-model, input the image feature vector of the second sample image into the image evaluation sub-model to obtain the initial evaluation result of the second sample image; update the image evaluation sub-model according to the initial evaluation result and the image sub-label corresponding to the image evaluation sub-model until the termination condition is reached, so as to obtain the trained image evaluation sub-model.
[0074] The above device further includes a model pre-training module, configured to: obtain first training data, and train an image feature extraction model with the first training data to obtain a pre-trained image feature extraction model; obtain second training data, and train an image evaluation model with the second training data to obtain a pre-trained image evaluation model.
[0075] The above second training data includes first training images and second training images, the first training images carry first labels, and the second training images carry second labels; wherein, the first labels indicate that the image scores are greater than a first preset threshold, and the second labels indicate that the image scores are less than a second preset threshold.
[0076] Corresponding to the embodiments of the above image evaluation method, embodiments of the present disclosure provide an image evaluation device, as Figure 4 shown, the device includes:
[0077] An image feature vector extraction module 401, configured to obtain a specified image from the images submitted by the user, input the specified image into the trained image feature extraction model obtained by any of the model training methods in the first aspect, and obtain the image feature vector corresponding to the specified image;
[0078] An image evaluation result determination module 402, configured to input the image feature vector into the trained image evaluation model obtained by any of the model training methods in the first aspect, and obtain the image evaluation result of the specified image.
[0079] An embodiment of the present disclosure provides an image evaluation device, which obtains a specified image from the images submitted by a user, inputs the specified image into a trained image feature extraction model obtained by any one of the model training methods in the first aspect, and obtains an image feature vector corresponding to the specified image; inputs the image feature vector into a trained image evaluation model obtained by any one of the model training methods in the first aspect, and obtains an image evaluation result of the specified image. In this manner, a small number of labeled first sample images are determined from the images submitted by the user, and then a large number of second sample images are expanded from the images submitted by the user according to the first sample images. The pre-trained model is trained again through the second sample images, reducing the human and time costs for determining training data, enabling the finally obtained model to learn accurate evaluation criteria, and improving the accuracy of the model for evaluating images.
[0080] The above-mentioned trained image evaluation model includes at least one image evaluation sub-model; the above-mentioned image evaluation result determination module is further configured to: for each image evaluation sub-model, input the image feature vector into the image evaluation sub-model to obtain an initial evaluation result corresponding to the image evaluation sub-model; perform data processing on the obtained at least one initial evaluation result to obtain an image evaluation result of the specified image.
[0081] The above-mentioned device further includes an update module, which is configured to: determine the specified image and the image evaluation result of the specified image as a third sample image; train the trained image evaluation model through the third sample image to update the model parameters of the trained image evaluation model.
[0082] The image evaluation device provided by the embodiment of the present disclosure has the same technical features as the image evaluation method provided by the above-mentioned embodiment, so it can also solve the same technical problems and achieve the same technical effects.
[0083] This embodiment further provides an electronic device, including a processor and a memory. The memory stores machine-executable instructions that can be executed by the processor, and the processor executes the machine-executable instructions to implement the above-mentioned model training method and image evaluation method. The electronic device can be a server or a terminal device.
[0084] See Figure 5 As shown, the electronic device includes a processor 100 and a memory 101. The memory 101 stores machine-executable instructions that can be executed by the processor 100, and the processor 100 executes the machine-executable instructions to implement the above-mentioned model training method and image evaluation method.
[0085] Further, Figure 5 The electronic device shown further includes a bus 102 and a communication interface 103. The processor 100, the communication interface 103, and the memory 101 are connected through the bus 102.
[0086] Among them, the memory 101 may include a high-speed random access memory (RAM), and may also include a non-volatile memory, such as at least one disk memory. The communication connection between the system network element and at least one other network element is realized through at least one communication interface 103 (which can be wired or wireless), and the Internet, wide area network, local area network, metropolitan area network, etc. can be used. The bus 102 can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, an EISA (Extended Industry Standard Architecture) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, Figure 5 only a bidirectional arrow is used in the figure, but it does not mean that there is only one bus or one type of bus.
[0087] The processor 100 may be an integrated circuit chip with the ability to process signals. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in the processor 100 or the instructions in the form of software. The above-mentioned processor 100 may be a general-purpose processor, including a central processing unit (CPU for short), a network processor (NP for short), etc.; it may also be a digital signal processor (DSP for short), an application specific integrated circuit (ASIC for short), a field-programmable gate array (FPGA for short) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present disclosure. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present disclosure can be directly embodied as being executed and completed by a hardware decoding processor, or executed and completed by a combination of hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 101, and the processor 100 reads the information in the memory 101 and combines its hardware to complete the steps of the method in the foregoing embodiments.
[0088] The processor in the above electronic device can implement the following operations in the above model training method by executing machine-executable instructions:
[0089] According to the pre-trained image feature extraction model and the pre-trained image evaluation model, determine the first sample image from the images submitted by the user; wherein, the first sample image carries an image annotation, and the image annotation is used to indicate the image evaluation result; by a preset method, determine the target image similar to the first sample image from the images submitted by the user, copy the image annotation carried by the first sample image to the target image, and determine the target image as the second sample image; train the pre-trained image feature extraction model with the second sample image to obtain a trained image feature extraction model; train the pre-trained image evaluation model with the second sample image to obtain a trained image evaluation model.
[0090] The steps of determining the first sample image from the images submitted by the user include: obtaining a first image from the images submitted by the user; determining a predicted evaluation result of the first image according to a pre-trained image feature extraction model and a pre-trained image evaluation model; verifying the predicted evaluation result to obtain a standard evaluation result of the first image; and determining the first image and the standard evaluation result of the first image as the first sample image.
[0091] The above preset method is any one of the following: label retrieval method, feature retrieval method, teacher model method.
[0092] The above image evaluation model includes at least one image evaluation sub-model; different image evaluation sub-models are used to indicate different image evaluation criteria; the image annotation carried by the second sample image includes at least one image sub-label; and the image evaluation sub-models and the image sub-labels are in one-to-one correspondence.
[0093] The steps of training the image evaluation model with the second sample image to obtain a trained image evaluation model include: inputting the second sample image into the trained image feature extraction model to obtain an image feature vector of the second sample image; for each image evaluation sub-model, inputting the image feature vector of the second sample image into the image evaluation sub-model to obtain an initial evaluation result of the second sample image; and updating the image evaluation sub-model according to the initial evaluation result and the image sub-label corresponding to the image evaluation sub-model until a termination condition is reached to obtain a trained image evaluation sub-model.
[0094] The above method further includes: obtaining first training data, and training the image feature extraction model with the first training data to obtain a pre-trained image feature extraction model; obtaining second training data, and training the image evaluation model with the second training data to obtain a pre-trained image evaluation model.
[0095] The above second training data includes a first training image and a second training image, the first training image carries a first label, and the second training image carries a second label; wherein, the first label indicates that the image score is greater than a first preset threshold, and the second label indicates that the image score is less than a second preset threshold.
[0096] The processor in the above electronic device can implement the following operations in the above image evaluation method by executing machine-executable instructions:
[0097] Obtaining a specified image from the images submitted by the user, inputting the specified image into the trained image feature extraction model obtained by any one of the model training methods in the first aspect to obtain an image feature vector corresponding to the specified image; and inputting the image feature vector into the trained image evaluation model obtained by any one of the model training methods in the first aspect to obtain an image evaluation result of the specified image.
[0098] The above-mentioned trained image evaluation model includes at least one image evaluation sub-model; the step of inputting the image feature vector into the trained image evaluation model obtained by any one of the model training methods in the first aspect to obtain the image evaluation result of the specified image includes: for each image evaluation sub-model, inputting the image feature vector into the image evaluation sub-model to obtain the initial evaluation result corresponding to the image evaluation sub-model; performing data processing on the obtained at least one initial evaluation result to obtain the image evaluation result of the specified image.
[0099] The above method further includes: determining the specified image and the image evaluation result of the specified image as the third sample image; training the trained image evaluation model with the third sample image to update the model parameters of the trained image evaluation model.
[0100] In this way, a small number of labeled first sample images are determined from the images submitted by the user, and then a large number of second sample images are expanded from the images submitted by the user according to the first sample images. The pre-trained model is trained again with the second sample images, reducing the labor and time costs for determining the training data, enabling the finally obtained model to learn accurate evaluation criteria, and improving the accuracy of the model in evaluating images.
[0101] This embodiment further provides a machine-readable storage medium, which stores machine-executable instructions. When the machine-executable instructions are called and executed by a processor, the machine-executable instructions prompt the processor to implement the above-mentioned model training method and image evaluation method.
[0102] The processor in the above electronic device can implement the following operations in the above-mentioned model training method by executing the machine-executable instructions:
[0103] According to the pre-trained image feature extraction model and the pre-trained image evaluation model, determine the first sample image from the images submitted by the user; wherein, the first sample image carries an image annotation, and the image annotation is used to indicate the image evaluation result; by a preset method, determine the target image similar to the first sample image from the images submitted by the user, copy the image annotation carried by the first sample image to the target image, and determine the target image as the second sample image; train the pre-trained image feature extraction model with the second sample image to obtain the trained image feature extraction model; train the pre-trained image evaluation model with the second sample image to obtain the trained image evaluation model.
[0104] The steps of determining the first sample image from the images submitted by the user include: obtaining a first image from the images submitted by the user; determining a predicted evaluation result of the first image according to a pre-trained image feature extraction model and a pre-trained image evaluation model; verifying the predicted evaluation result to obtain a standard evaluation result of the first image; and determining the first image and the standard evaluation result of the first image as the first sample image.
[0105] The above-mentioned preset method is any one of the following: label retrieval method, feature retrieval method, teacher model method.
[0106] The above-mentioned image evaluation model includes at least one image evaluation sub-model; different image evaluation sub-models are used to indicate different image evaluation criteria; the image annotation carried by the second sample image includes at least one image sub-label; and the image evaluation sub-model corresponds to the image sub-label one by one.
[0107] The steps of training the image evaluation model with the second sample image to obtain a trained image evaluation model include: inputting the second sample image into the trained image feature extraction model to obtain an image feature vector of the second sample image; for each image evaluation sub-model, inputting the image feature vector of the second sample image into the image evaluation sub-model to obtain an initial evaluation result of the second sample image; and updating the image evaluation sub-model according to the initial evaluation result and the image sub-label corresponding to the image evaluation sub-model until a termination condition is reached to obtain a trained image evaluation sub-model.
[0108] The above-mentioned method further includes: obtaining first training data, and training the image feature extraction model with the first training data to obtain a pre-trained image feature extraction model; obtaining second training data, and training the image evaluation model with the second training data to obtain a pre-trained image evaluation model.
[0109] The above-mentioned second training data includes a first training image and a second training image, the first training image carries a first label, and the second training image carries a second label; wherein, the first label indicates that the image score is greater than a first preset threshold, and the second label indicates that the image score is less than a second preset threshold.
[0110] The processor in the above-mentioned electronic device can implement the following operations in the above-mentioned image evaluation method by executing machine-executable instructions:
[0111] Obtain a specified image from the images submitted by the user, input the specified image into the trained image feature extraction model obtained by any one of the model training methods in the first aspect to obtain an image feature vector corresponding to the specified image; input the image feature vector into the trained image evaluation model obtained by any one of the model training methods in the first aspect to obtain an image evaluation result of the specified image.
[0112] The above-mentioned trained image evaluation model includes at least one image evaluation sub-model; the step of inputting the image feature vector into the trained image evaluation model obtained by any of the model training methods in the first aspect to obtain the image evaluation result of the specified image includes: for each image evaluation sub-model, inputting the image feature vector into the image evaluation sub-model to obtain the initial evaluation result corresponding to the image evaluation sub-model; performing data processing on the obtained at least one initial evaluation result to obtain the image evaluation result of the specified image.
[0113] The above method further includes: determining the specified image and the image evaluation result of the specified image as the third sample image; training the trained image evaluation model with the third sample image to update the model parameters of the trained image evaluation model.
[0114] In this method, a small number of labeled first sample images are determined from the images submitted by the user, and then a large number of second sample images are expanded from the images submitted by the user according to the first sample images. The pre-trained model is trained again with the second sample images, reducing the human and time costs for determining the training data, enabling the finally obtained model to learn accurate evaluation criteria, and improving the accuracy of the model for evaluating images.
[0115] The computer program product of the model training method, image evaluation method, device and electronic device provided by the embodiments of the present disclosure includes a computer-readable storage medium storing program codes, and the instructions included in the program codes can be used to execute the methods described in the foregoing method embodiments. For the specific implementation, reference can be made to the method embodiments and will not be elaborated herein.
[0116] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described systems and devices can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0117] In addition, in the description of the embodiments of the present disclosure, unless otherwise clearly specified and limited, the terms "installation", "connection" and "connection" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two components. For those skilled in the art, the specific meanings of the above terms in the present disclosure can be understood according to specific situations.
[0118] When the above-described functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present disclosure, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present disclosure. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs.
[0119] In the description of the present disclosure, it should be noted that the orientation or positional relationship indicated by the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present disclosure and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation to the present disclosure. In addition, the terms "first", "second", "third" are only used for descriptive purposes and should not be construed as indicating or implying relative importance.
[0120] Finally, it should be noted that the above embodiments are only specific implementation manners of the present disclosure, used to illustrate the technical solutions of the present disclosure, rather than limiting them. The protection scope of the present disclosure is not limited thereto. Although the present disclosure has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: any person skilled in the art within the technical scope disclosed by the present disclosure can still modify the technical solutions recorded in the foregoing embodiments, or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes, or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present disclosure, and should all be covered by the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be subject to the protection scope of the claims.
Claims
1. A model training method, characterized in that, The method includes: Determining a first sample image from the images submitted by the user according to a pre-trained image feature extraction model and a pre-trained image evaluation model; wherein, the first sample image carries an image annotation for indicating an image evaluation result; Determining a target image similar to the first sample image from the images submitted by the user by a preset method, copying the image annotation carried by the first sample image to the target image, and determining the target image as a second sample image; Training the pre-trained image feature extraction model with the second sample image to obtain a trained image feature extraction model; training the pre-trained image evaluation model with the second sample image to obtain a trained image evaluation model.
2. The method according to claim 1, characterized in that, The step of determining a first sample image from the images submitted by the user includes: Obtaining a first image from the images submitted by the user; Determining a predicted evaluation result of the first image according to the pre-trained image feature extraction model and the pre-trained image evaluation model; Verifying the predicted evaluation result to obtain a standard evaluation result of the first image; Determining the first image and the standard evaluation result of the first image as the first sample image.
3. The method according to claim 1, characterized in that, The preset method is any one of the following: a label retrieval method, a feature retrieval method, and a teacher model method.
4. The method according to claim 1, wherein The image evaluation model includes at least one image evaluation sub-model; different image evaluation sub-models are used to indicate different image evaluation criteria; The image annotation carried by the second sample image includes at least one image sub-label; the image evaluation sub-model corresponds to the image sub-label one by one.
5. The method according to claim 4, wherein The step of training the image evaluation model with the second sample image to obtain a trained image evaluation model includes: Inputting the second sample image into the trained image feature extraction model to obtain an image feature vector of the second sample image; For each image evaluation sub-model, inputting the image feature vector of the second sample image into the image evaluation sub-model to obtain an initial evaluation result of the second sample image; Updating the image evaluation sub-model according to the initial evaluation result and the image sub-label corresponding to the image evaluation sub-model until a termination condition is reached to obtain a trained image evaluation sub-model.
6. The method according to claim 1, characterized in that The method further includes: Obtaining first training data and training the image feature extraction model with the first training data to obtain a pre-trained image feature extraction model; Obtaining second training data and training the image evaluation model with the second training data to obtain a pre-trained image evaluation model.
7. The method according to claim 6, characterized in that, The second training data includes a first training image and a second training image, the first training image carries a first label, and the second training image carries a second label; wherein, the first label indicates that the image score is greater than a first preset threshold, and the second label indicates that the image score is less than a second preset threshold.
8. An image evaluation method, characterized in that, The method includes: Obtain a specified image from the images submitted by the user, input the specified image into the trained image feature extraction model obtained by the model training method according to any one of claims 1-7, and obtain the image feature vector corresponding to the specified image; Input the image feature vector into the trained image evaluation model obtained by the model training method according to any one of claims 1-7, and obtain the image evaluation result of the specified image.
9. The method according to claim 8, characterized in that, The trained image evaluation model includes at least one image evaluation sub-model; The step of inputting the image feature vector into the trained image evaluation model obtained by the model training method according to any one of claims 1-7 to obtain the image evaluation result of the specified image includes: For each of the image evaluation sub-models, input the image feature vector into the image evaluation sub-model to obtain the initial evaluation result corresponding to the image evaluation sub-model; Perform data processing on the obtained at least one initial evaluation result to obtain the image evaluation result of the specified image.
10. The method according to claim 8, wherein The method further includes: Determine the specified image and the image evaluation result of the specified image as the third sample image; Train the trained image evaluation model with the third sample image to update the model parameters of the trained image evaluation model.
11. A model training device, characterized in that, The device includes: A first sample image determination module, configured to determine a first sample image from the images submitted by the user according to a pre-trained image feature extraction model and a pre-trained image evaluation model; wherein, the first sample image carries an image annotation, and the image annotation is used to indicate the image evaluation result; A second sample image determination module, configured to determine a target image similar to the first sample image from the images submitted by the user by a preset method, copy the image annotation carried by the first sample image to the target image, and determine the target image as the second sample image; A model training module, configured to train the pre-trained image feature extraction model with the second sample image to obtain a trained image feature extraction model; and train the pre-trained image evaluation model with the second sample image to obtain a trained image evaluation model.
12. An image evaluation device, characterized in that, The device includes: An image feature vector extraction module, configured to obtain a specified image from the images submitted by the user, input the specified image into the trained image feature extraction model obtained by the model training method according to any one of claims 1-7, and obtain the image feature vector corresponding to the specified image; An image evaluation result determination module, configured to input the image feature vector into the trained image evaluation model obtained by the model training method according to any one of claims 1-7 to obtain the image evaluation result of the specified image.
13. An electronic device, characterized in that, It includes a processor and a memory, the memory stores computer-executable instructions that can be executed by the processor, and the processor executes the computer-executable instructions to implement the model training method according to any one of claims 1-7, or implement the image evaluation method according to any one of claims 8-10.
14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are called and executed by a processor, the computer-executable instructions cause the processor to implement the model training method according to any one of claims 1-7, or to implement the image evaluation method according to any one of claims 8-10.