An image scoring model training method, an image scoring method, and related devices

By training the image scoring model cyclically, using the image relative scoring data set and sample processing module, the problem of being unable to score relative appearance on two face images in the prior art is solved, and training the relative appearance scoring model of face images is realized.

CN117671771BActive Publication Date: 2025-05-27GUANGZHOU HUYA TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202311735687.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-12-14
Publication Date
2025-05-27
Estimated Expiration
2043-12-14

AI Technical Summary

Technical Problem

The existing training methods for rating face images can only score absolute appearance on a single image, but cannot score relative appearance on two face images.

Method used

By acquiring the image relative scoring data set, the sample data is input into the image scoring model and sample processing module, the predicted preference probability and preference probability distribution are obtained, the loss information is determined, and the training is cycled until the preset conditions are met, and a mature image scoring model is obtained.

Benefits of technology

The training of the relative appearance scoring model between different face images is achieved, and the appearance preference between the two images can be more accurately evaluated.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117671771B_ABST
    Figure CN117671771B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention provides an image scoring model training method, an image scoring method and related devices, which relate to the field of machine learning. By inputting sample data into an image scoring model, obtaining the predicted preference probability of the sample data, inputting the sample data into a sample processing module, obtaining the preference probability distribution of the sample data, determining loss information according to the predicted preference probability and the probability distribution, cyclically inputting the sample data of the data set into the image scoring model and the sample processing module, obtaining the loss information of the current training cycle until the loss information corresponding to the Nth training cycle meets a first preset condition, then ending the training, and using the trained image scoring model as a mature image scoring model. Thus, a relative face value scoring model between different face images can be trained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of machine learning, and more particularly, to a method for training an image scoring model, an image scoring method, and related devices. Background Art

[0002] Existing methods for training face image scoring are all based on the same abstract objective standard, scoring and training a single image to obtain a scoring model for a single face image, and these models can only evaluate the appearance value of a single face image. Summary of the Invention

[0003] In view of this, the purpose of the embodiments of the present invention is to provide a method for training an image scoring model, an image scoring method, and related devices to at least partially improve the above problems.

[0004] To achieve the above purpose, the technical solutions adopted in the embodiments of the present invention are as follows:

[0005] In a first aspect, an embodiment of the present application provides a method for training an image scoring model. The method is applied to a model training system, and the model training system includes a sample processing module and an image scoring model. The method includes:

[0006] Obtain an image relative scoring data set; the image relative scoring data set includes at least one set of sample data, and any one of the sample data includes two images and a set of relative scores, and the relative scores represent the preference probability for the two images;

[0007] Input the sample data into the image scoring model to obtain the predicted preference probability of the sample data; the preference probability represents the probability of preferring one of the two images;

[0008] Input the sample data into the sample processing module to obtain the preference probability distribution of the sample data; the probability distribution represents the probability distribution of the crowd preferring one of the two images in the sample data;

[0009] Determine loss information according to the predicted preference probability and the probability distribution; the loss information represents the difference between the predicted preference probability and the mode of the probability distribution;

[0010] Input the sample data of the image relative scoring data set into the image scoring model and the sample processing module in a loop to obtain the loss information of the current training cycle until the loss information corresponding to the Nth training cycle meets the first preset condition, then end the training, and use the trained image scoring model as a mature image scoring model.

[0011] Optionally, the image scoring model includes: a convolutional neural network model and a feature aggregation contrast model. The step of inputting the sample data into the image scoring model to obtain the predicted preference probability of the sample data includes:

[0012] Input the sample data into the convolutional neural network model to obtain a first deep feature and a second deep feature; the first deep feature and the second deep feature respectively represent the deep features of the two images;

[0013] Input the first deep feature and the second deep feature into the feature aggregation contrast model to obtain the predicted preference probability of the sample data.

[0014] Optionally, the feature aggregation contrast model includes: a tensor unfolding layer and a fully connected layer. The step of inputting the first deep feature and the second deep feature into the feature aggregation contrast model to obtain the predicted preference probability of the sample data includes:

[0015] Use the tensor unfolding layer to unfold the first deep feature and the second deep feature into one dimension to obtain the unfolded first deep feature and second deep feature;

[0016] According to the unfolded first deep feature and the second deep feature, use the fully connected layer to predict the preference probability of the sample data to obtain the predicted preference probability of the sample data.

[0017] Optionally, after obtaining the predicted preference probability of the sample data, it further includes:

[0018] Multiply the predicted preference probability by a preset value to obtain a predicted relative score;

[0019] Add the predicted relative score to the sample data.

[0020] Optionally, the feature aggregation contrast model further includes a linear threshold activation layer, and the linear threshold activation layer is used to limit the range of the predicted preference probability to 0 to 1.

[0021] Optionally, after obtaining the image relative score dataset, it further includes:

[0022] Perform a normalization process on the image relative score dataset to obtain a standardized image relative score dataset that eliminates personal biases, and replace the image relative score dataset with the standardized image relative score dataset.

[0023] Optionally, the method further includes:

[0024] Normalize the relative image scoring dataset to obtain a normalized relative image scoring dataset, and replace the relative image scoring dataset with the normalized relative image scoring dataset.

[0025] In a second aspect, an embodiment of the present invention provides an image scoring method, which is applied to an image scoring model. The image scoring model includes: a convolutional neural network model and a feature aggregation and contrast model. The method includes:

[0026] Input two images to be scored into the convolutional neural network model to obtain a third deep feature and a fourth deep feature;

[0027] Input the third deep feature and the fourth deep feature into the feature aggregation and contrast model to obtain the predicted preference probability of the two images to be scored.

[0028] In a third aspect, an embodiment of the present invention provides an image scoring model training device, which is applied to a model training system. The model training system includes a sample processing module and an image scoring model. The device includes:

[0029] A dataset acquisition module, configured to acquire a relative image scoring dataset; the relative image scoring dataset includes at least one set of sample data, and any one of the sample data includes two images and a set of relative scores, and the relative scores represent the preference probability for the two images;

[0030] A predicted preference probability module, configured to input the sample data into the image scoring model to obtain the predicted preference probability of the sample data; the preference probability represents the probability of preferring one of the two images;

[0031] A probability distribution calculation module, configured to input the sample data into the sample processing module to obtain the preference probability distribution of the sample data; the probability distribution represents the probability distribution of the crowd preferring one of the two images in the sample data;

[0032] A loss information confirmation module, configured to determine loss information according to the predicted preference probability and the probability distribution; the loss information represents the difference between the predicted preference probability and the mode of the probability distribution;

[0033] An iteration module, configured to cyclically input the sample data of the relative image scoring dataset into the image scoring model and the sample processing module to obtain the loss information of the current training cycle, and end the training until the loss information corresponding to the Nth training cycle meets a first preset condition, and use the trained image scoring model as a mature image scoring model.

[0034] Fourthly, an embodiment of the present invention provides an image scoring device, which is applied to an image scoring model. The image scoring model includes: a convolutional neural network model and a feature aggregation and contrast model. The device includes:

[0035] A deep feature acquisition module, configured to input two images to be scored into the convolutional neural network model to obtain a third deep feature and a fourth deep feature;

[0036] A preference probability acquisition module, configured to input the third deep feature and the fourth deep feature into the feature aggregation and contrast model to obtain the predicted preference probabilities of the two images to be scored.

[0037] Fifthly, an embodiment of the present invention provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the method described in any one of the above first aspects is implemented; and / or, the method described in the above second aspect.

[0038] Fourthly, an embodiment of the present invention provides a storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method described in any one of the above first aspects is implemented; and / or, the method described in the above second aspect.

[0039] Compared with the prior art, an image scoring model training method, an image scoring method, and related devices provided by an embodiment of the present invention input sample data into an image scoring model to obtain the predicted preference probabilities of the sample data, input the sample data into a sample processing module to obtain the preference probability distribution of the sample data, determine loss information according to the predicted preference probabilities and the probability distribution, and cyclically input the sample data of the data set into the image scoring model and the sample processing module to obtain the loss information of the current training cycle until the loss information corresponding to the Nth training cycle meets a first preset condition, then the training is ended, and the trained image scoring model is used as a mature image scoring model. Thereby, a relative face value scoring model between different face images can be trained. Description of the Drawings

[0040] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the embodiments. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0041] Figure 1 It is a schematic diagram of an electronic device provided by an embodiment of the present invention;

[0042] Figure 2 Schematic diagram of a model training system provided by an embodiment of the present invention;

[0043] Figure 3 Flow schematic diagram of an image scoring model training method provided by an embodiment of the present invention;

[0044] Figure 4 One of the schematic diagrams of another model training system provided by an embodiment of the present invention;

[0045] Figure 5 One of the other flow schematic diagrams of an image scoring model training method provided by an embodiment of the present invention;

[0046] Figure 6 Two of the schematic diagrams of another model training system provided by an embodiment of the present invention;

[0047] Figure 7 Two of the other flow schematic diagrams of an image scoring model training method provided by an embodiment of the present invention;

[0048] Figure 8 Three of the schematic diagrams of another model training system provided by an embodiment of the present invention;

[0049] Figure 9 Three of the other flow schematic diagrams of an image scoring model training method provided by an embodiment of the present invention;

[0050] Figure 10 Four of the other flow schematic diagrams of an image scoring model training method provided by an embodiment of the present invention;

[0051] Figure 11 Five of the other flow schematic diagrams of an image scoring model training method provided by an embodiment of the present invention;

[0052] Figure 12 Flow schematic diagram of an image scoring method provided by an embodiment of the present invention;

[0053] Figure 13 Schematic diagram of an image scoring model training device provided by an embodiment of the present invention;

[0054] Figure 14 Schematic diagram of an image scoring device provided by an embodiment of the present invention.

[0055] Icons: 100 - Electronic device; 101 - Memory; 102 - Communication interface; 103 - Processor; 104 - Bus; 10 - Model training system; 11 - Sample processing module; 12 - Image scoring model; 121 - Convolutional neural network model; 122 - Feature aggregation and contrast model; 1221 - Tensor unfolding layer; 1222 - Fully connected layer; 1223 - Linear threshold activation layer; 300 - Image scoring model training device; 310 - Dataset acquisition module; 320 - Prediction preference probability module; 330 - Probability distribution calculation module; 340 - Loss information confirmation module; 350 - Iteration module; 400 - Image scoring device; 410 - Deep feature acquisition module; 420 - Preference probability acquisition module. Detailed implementation manners

[0056] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention. The components of the embodiments of the present invention described and illustrated herein can be arranged and designed in various different configurations.

[0057] Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.

[0058] It should be noted that: Similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. At the same time, in the description of the present invention, the terms "first", "second", etc. are only used for descriptive distinction and cannot be construed as indicating or implying relative importance.

[0059] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the presence of additional identical elements in the process, method, article or device including the said element.

[0060] In the prior art, for the training of a face image beauty score model, feature extraction is performed on a single face image, and based on the same abstract objective standard, deep learning is performed on multiple single face images to complete the training of the model.

[0061] However, these trained models can only perform absolute beauty scoring on a single face image and cannot perform relative beauty scoring on two face images.

[0062] To solve the above problems, the present invention proposes an image score model training method, an image score method and related devices. By inputting sample data into the image score model, the predicted preference probability of the sample data is obtained, and the sample data is input into the sample processing module to obtain the preference probability distribution of the sample data. According to the predicted preference probability and the probability distribution, loss information is determined, and the sample data of the data set is cyclically input into the image score model and the sample processing module to obtain the loss information of the current training cycle until the loss information corresponding to the Nth training cycle meets the first preset condition, then the training is ended, and the trained image score model is used as a mature image score model. Thus, a relative beauty score model between different face images can be trained.

[0063] To implement the process steps and functions of each example of the present invention, please refer to Figure 1 , Figure 1 FIG. is a schematic structural block diagram of an electronic device provided by an embodiment of the present invention. The electronic device 100 includes a memory 101 and a processor 103, and the memory 101 and the processor 103 are directly or indirectly electrically connected to each other to realize data transmission or interaction. For example, these components can be electrically connected to each other through one or more communication buses 104 or signal lines. The memory 101 can be used to store software programs and modules, and the processor 103 can execute various functional applications and data processing by executing the software programs and modules stored in the memory 101.

[0064] The electronic device 100 can be, but is not limited to, a personal computer (PC), a server, a distributed computer, etc. It can be understood that the electronic device 100 is not limited to a physical server, and can also be a virtual machine on a physical server, a virtual machine built on a cloud platform, etc., which can provide a computer with the same functions as the server or virtual machine. The operating system of the electronic device 100 can be, but is not limited to, the Windows system, the Linux system, etc.

[0065] Among them, the memory 101 is used to store programs. After receiving the execution instruction, the processor 103 executes the programs to implement an image scoring model training method and an image scoring method disclosed in the embodiments of the present invention.

[0066] Among them, the memory 101 can be, but is not limited to, a random access memory (RAM), a read only memory (ROM), a programmable read only memory (PROM), an erasable programmable read only memory (EPROM), an electrically erasable programmable read only memory (EEPROM), etc.

[0067] The communication connection between the electronic device 100 and an external device is realized through at least one communication interface 102 (which can be wired or wireless).

[0068] The processor 103 may be an integrated circuit chip with signal processing capabilities. In the implementation process, the steps of the embodiments of the present invention can be completed by the integrated logic circuit in the hardware of the processor 103 or the instructions in software form. The processor 103 can be a general-purpose processor, including a central processor, a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0069] It can be understood Figure 1The structure shown is only schematic, and the electronic device 100 may also include more or fewer components than those shown in Figure 1 or have a different configuration from that shown in Figure 1 . Figure 1 Each component shown in can be implemented by hardware, software, or a combination thereof.

[0070] Next, an exemplary description will be given of the image scoring model training method provided by the present invention. Specifically, Figure 2 is a schematic diagram of a model training system provided by an embodiment of the present invention. Referring to Figure 2 , this method is applied to the model training system 10, which includes a sample processing module 11 and an image scoring model 12; further, Figure 3 is a flowchart of an image scoring model training method provided by an embodiment of the present invention. Combining Figure 2 and Figure 3 , this method includes the steps:

[0071] Step S210: Obtain an image relative scoring data set;

[0072] Among them, the image relative scoring data set includes at least one set of sample data, and any one of the sample data includes two images and a set of relative scores, and the relative score represents the preference probability for the two images.

[0073] Step S220: Input the sample data into the image scoring model to obtain the predicted preference probability of the sample data;

[0074] Among them, the preference probability represents the probability of preferring one of the two images more.

[0075] Step S230: Input the sample data into the sample processing module to obtain the preference probability distribution of the sample data;

[0076] Among them, the probability distribution represents the probability distribution of the crowd preferring one of the two images in the sample data.

[0077] Step S240: Determine loss information according to the predicted preference probability and the probability distribution;

[0078] Among them, the loss information represents the difference between the predicted preference probability and the mode of the probability distribution.

[0079] Step S250: Input the sample data of the image relative scoring data set into the image scoring model and the sample processing module in a loop to obtain the loss information of the current training cycle.

[0080] Step S260: Determine whether the loss information in the current training cycle meets the first preset condition; if so, execute Step S270; if not, return to execute Step S220 and Step S230.

[0081] Step S270: End the training, and use the trained image scoring model as the mature image scoring model.

[0082] To better understand the above method, the above method will be described exemplarily as follows:

[0083] The relative image scoring dataset obtained in Step S210 includes multiple groups of sample data. Each group of sample data includes two face images A and B, and a group of relative scores. This group of relative scores includes the relative scores of multiple annotators for one of the images. For example, it is the score of annotator for A relative to B.

[0084] Taking the j-th group of sample data as an example, the j-th group of sample data contains the relative scores of m annotators for image A

[0085] Input the j-th group of sample data into the image scoring model to obtain the predicted preference probability p' for image A.

[0086] Input the j-th group of sample data into the sample processing module to obtain the preference probability distribution of all annotators for image A.

[0087] Optionally, a beta distribution can be constructed for the j-th group of sample data to obtain the preference probability distribution S of the annotators for image A j , and the specific expression is as follows:

[0088]

[0089] where B(α j ,β j ) is the beta function.

[0090] Determine the loss information according to the predicted preference probability p' and the preference probability distribution S j .

[0091] Optionally, the value of the preference probability distribution S j at the predicted preference probability p' divided by the L1 distance between the value at the mode of the preference probability distribution S j and 1 can be used as the loss information, and the specific expression is as follows:

[0092]

[0093]

[0094]

[0095] Among them, P Sj (p′; α j , β j ) is the value of the preference probability distribution S j at the predicted preference probability p', mode(S j ) is the value at the mode of the preference probability distribution S j and loss is the loss information.

[0096] The sample data of the image relative scoring dataset is cyclically input into the image scoring model and the sample processing module to obtain the loss information of the current training cycle.

[0097] When loss meets the first preset condition, the training ends, and the trained image scoring model is used as a mature image scoring model.

[0098] Optionally, the first preset condition can be that loss is less than or equal to 0.0001.

[0099] When loss does not meet the first preset condition, the next sample data is taken, and steps S220 and S230 are returned for execution.

[0100] The image scoring model training method provided in the implementation of the present invention obtains the predicted preference probability of the sample data by inputting the sample data into the image scoring model, obtains the preference probability distribution of the sample data by inputting the sample data into the sample processing module, and determines the loss information according to the predicted preference probability and the probability distribution. The relative appearance scoring loss calculation method based on the beta distribution enables the image scoring model to better capture the general preference tendency of the crowd from the sample data. This method has high implementation efficiency and strong versatility.

[0101] In a possible implementation manner, in order to obtain richer scoring data of the sample data, the scoring can also be the absolute scoring of a single image in the sample data. After obtaining the image relative scoring dataset, the following steps can also be included:

[0102] When the scoring is absolute scoring, the difference between any two sets of adjacent data in the scoring of the sample data is obtained.

[0103] And the difference is used as the relative scoring in the sample data.

[0104] When the scoring in a set of sample data is absolute scoring, it needs to be converted into an approximate relative scoring. Take any two adjacent scores in the scoring data, calculate the difference as a percentage, and obtain the approximate relative scoring as the relative scoring.

[0105] In a possible implementation, to reduce the bias caused by personal factors in the sample data, all the scores of each annotator are normalized by subtracting the mean and dividing by the standard deviation.

[0106] Referring to Figure 10 , Figure 10 FIG. 4 is another schematic flowchart of an image scoring model training method provided by an embodiment of the present invention. After step S210, the method may further include:

[0107] Step S211: Normalize the image relative scoring data set to obtain a standardized image relative scoring data set with personal bias eliminated, and replace the image relative scoring data set with the standardized image relative scoring data set.

[0108] Optionally, all the scores of each annotator are normalized by subtracting the mean and dividing by the standard deviation to obtain a standardized relative scoring data set with personal bias eliminated. The specific expression is as follows:

[0109]

[0110] wherein, is the j-th relative score or approximate relative score annotation of the i-th annotator, is the μ i and σ i are the mean and standard deviation of the relative scores of the i-th annotator, respectively.

[0111] In a possible implementation, to reduce the bias caused by the annotator giving extremely high or low scores in the sample data, the image relative scoring data set may be normalized. Referring to Figure 11 , Figure 11 FIG. 5 is another schematic flowchart of an image scoring model training method provided by an embodiment of the present invention. After step S210, the method may further include:

[0112] Step S212: Normalize the image relative scoring data set to obtain a normalized image scoring data set, and replace the image relative scoring data set with the normalized image relative scoring data set.

[0113] Optionally, calculate the relative score values at the 95% and 5% quantiles, and normalize the entire data set with these two values as the upper and lower limits to obtain a normalized standardized relative scoring data set. The specific expression is as follows:

[0114]

[0115] where Q .95 and Q .5Represent the relative scoring values at the 95% and 5% quantiles, which are the normalized standard relative scores.

[0116] In one possible implementation, referring to Figure 4 , Figure 4 which is one of the schematic diagrams of another model training system provided by an embodiment of the present invention, the image scoring model 12 includes: a convolutional neural network model 121 and a feature aggregation contrast model 122; further, Figure 5 which is one of the schematic diagrams of another process of an image scoring model training method provided by an embodiment of the present invention. Combining Figure 4 and Figure 5 , the steps of inputting sample data into the image scoring model to obtain the predicted preference probability of the sample data include:

[0117] Step S221: Input the sample data into the convolutional neural network model to obtain a first review feature and a second deep feature;

[0118] Among them, the first deep feature and the second deep feature respectively represent the deep features of two images.

[0119] Step S222: Input the first deep feature and the second deep feature into the feature aggregation contrast model to obtain the predicted preference probability of the sample data.

[0120] Optionally, the convolutional neural network model can adopt the backbone of any general feature extraction model. For example, ResNet-18 without a classification head is adopted.

[0121] In one possible implementation, referring to Figure 6 , Figure 6 which is the second of the schematic diagrams of another model training system provided by an embodiment of the present invention. The feature aggregation contrast model 122 includes: a tensor unfolding layer 1221 and a fully connected layer 1222; further, Figure 7 which is the second of the schematic diagrams of another process of an image scoring model training method provided by an embodiment of the present invention. Combining Figure 6 and Figure 7 , the steps of inputting the first deep feature and the second deep feature into the feature aggregation contrast model to obtain the predicted preference probability of the sample data include:

[0122] Step S2221: Use the tensor unfolding layer to unfold the first deep feature and the second deep feature into one dimension to obtain the unfolded first deep feature and the second deep feature.

[0123] Step S2222: According to the unfolded first deep feature and the second deep feature, use the fully connected layer to predict the preference probability of the sample data to obtain the predicted preference probability of the sample data.

[0124] In a possible implementation, it is necessary to limit the range of the predicted preference probability. Refer to Figure 8 , Figure 8 FIG. 3 is another schematic diagram of a model training system provided by an embodiment of the present invention. The feature aggregation and comparison model 122 may further include: a linear threshold activation layer 1223; preferably, the linear threshold activation layer 1223 is used to limit the range of the predicted preference probability to 0 to 1.

[0125] In a possible implementation, in order to continuously enrich the sample data and achieve a more perfect model training purpose, refer to Figure 9 , Figure 9 FIG. 3 is another flowchart of an image scoring model training method provided by an embodiment of the present invention. After obtaining the predicted preference probability of the sample data, the following steps are further included:

[0126] Step S310: Multiply the predicted preference probability by a preset value to obtain a predicted relative score.

[0127] Step S320: Add the predicted relative score to the sample data.

[0128] After obtaining the predicted preference probability of a sample data, multiply the predicted preference probability by a preset value. Preferably, the preset value may be 100 to obtain a predicted relative score, and add the predicted relative score to the sample data. The relative score in the sample data will continuously increase, and the preference probability distribution S obtained in step S230 j will be more reasonably distributed, making the model training more efficient.

[0129] After completing the above model training, a mature image scoring model is obtained and deployed to the target device through an inference engine. The inference engine may include: TensorRT, MNN, etc., to make a relative score for two face images. An embodiment of the present invention provides an image scoring method, which is applied to an image scoring model. The image scoring model includes: a convolutional neural network model and a feature aggregation and comparison model. Figure 12 FIG. 4 is a flowchart of an image scoring method provided by an embodiment of the present invention. Refer to Figure 12 , and the method includes:

[0130] Step S610: Input two images to be scored into the convolutional neural network model to obtain a third deep feature and a fourth deep feature;

[0131] Step S620: Input the third deep feature and the fourth deep feature into the feature aggregation and comparison model to obtain the predicted preference probability of the two images to be scored.

[0132] Extract the deep features of two images using a convolutional neural network model. The neural network model can adopt the backbone of any general feature extraction model. Here, ResNet-18 without a classification head is used. The feature aggregation and comparison model includes a tensor unfolding layer and a fully connected layer. The tensor unfolding layer unfolds the deep features of the two images into one dimension, and after unfolding, the fully connected layer is used to predict the preference probability. Further, the predicted preference probability can be multiplied by a preset value to obtain the relative appearance score. For example, if the two images are C and D, the preset value can be 100. If the predicted preference probability for C is 0.7, then the relative appearance score of C is 0.7 * 100 = 70, and the relative appearance score of D is 100 - 70 = 30.

[0133] Based on the above example, the following provides an image scoring model training device, which can be deployed in the model training system 10. Specifically, Figure 13 Schematic diagram of an image scoring model training device 300 provided by an embodiment of the present invention. Refer to Figure 13 This device includes: a dataset acquisition module 310, a predicted preference probability module 320, a probability distribution calculation module 330, a loss information confirmation module 340, and an iteration module 350;

[0134] The dataset acquisition module 310 is used to acquire an image relative scoring dataset; the image relative scoring dataset includes at least one set of sample data. Any set of sample data includes two images and a set of relative scores. The relative score represents the preference probability for the two images.

[0135] The predicted preference probability module 320 is used to input the sample data into the image scoring model to obtain the predicted preference probability of the sample data; the preference probability represents the probability of preferring one of the two images more.

[0136] The probability distribution calculation module 330 is used to input the sample data into the sample processing module to obtain the preference probability distribution of the sample data; the probability distribution represents the probability distribution of the crowd preferring one of the two images in the sample data.

[0137] The loss information confirmation module 340 is used to determine the loss information according to the predicted preference probability and the probability distribution; the loss information represents the difference between the predicted preference probability and the mode of the probability distribution.

[0138] The iteration module 350 is used to repeatedly input the sample data of the image relative scoring dataset into the image scoring model and the sample processing module to obtain the loss information of the current training cycle until the loss information corresponding to the Nth training cycle meets the first preset condition, then the training is terminated, and the trained image scoring model is used as a mature image scoring model.

[0139] Furthermore, the present invention also provides an image scoring device. Specifically, Figure 14 A schematic diagram of an image scoring device provided by an embodiment of the present invention is shown in Figure 14 , the image scoring device 400 is applied to an image scoring model, and the image scoring model includes: a convolutional neural network model and a feature aggregation and contrast model. The device includes: a deep feature acquisition module 410 and a preference probability acquisition module 420;

[0140] The deep feature acquisition module 410 is configured to input two images to be scored into the convolutional neural network model to obtain a third deep feature and a fourth deep feature.

[0141] The preference probability acquisition module 420 is configured to input the third deep feature and the fourth deep feature into the feature aggregation and contrast model to obtain the predicted preference probabilities of the two images to be scored.

[0142] In summary, for an image scoring model training method, an image scoring method, and related devices provided by embodiments of the present invention, by obtaining a relative scoring data set of images and performing standardization and normalization processing on the data set, personal errors and the like in the data are reduced. The sample data is input into the image scoring model to obtain the predicted preference probabilities of the sample data. The sample data is input into the sample processing module to obtain the preference probability distribution of the sample data. According to the predicted preference probabilities and the probability distribution, loss information is determined. The sample data of the data set is cyclically input into the image scoring model and the sample processing module to obtain the loss information of the current training cycle until the loss information corresponding to the Nth training cycle meets the first preset condition, then the training is ended, and the trained image scoring model is used as a mature image scoring model. Thus, a relative face value scoring model between different face images can be trained.

[0143] In the embodiments provided by the present invention, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of devices, methods, and computer program products according to multiple embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and a module, a program segment, or a part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.

[0144] In addition, each functional module in various embodiments of the present invention may be integrated together to form an independent part, or each module may exist alone, or two or more modules may be integrated to form an independent part.

[0145] If the function is implemented in the form of a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a computer-readable storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0146] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

[0147] It is obvious to those skilled in the art that the present invention is not limited to the details of the above-described exemplary embodiments, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be embraced within the present invention. Any reference signs in the claims should not be construed as limiting the claims involved.

Claims

1. A method for training an image scoring model, characterized in that, the method is applied to a model training system, the model training system includes a sample processing module and an image scoring model, and the method includes: Obtain an image relative scoring data set; the image relative scoring data set includes at least one set of sample data, and any one of the sample data includes two images and a set of relative scores, and the relative scores represent the preference probability for the two images; Input the sample data into the image scoring model to obtain the predicted preference probability of the sample data; the preference probability represents the probability of preferring one of the two images; Input the sample data into the sample processing module to obtain the preference probability distribution of the sample data; the probability distribution represents the probability distribution of the crowd preferring one of the two images in the sample data; Determine loss information according to the predicted preference probability and the probability distribution; the loss information represents the difference between the predicted preference probability and the mode of the probability distribution; Circularly input the sample data of the image relative scoring data set into the image scoring model and the sample processing module to obtain the loss information of the current training cycle until the loss information corresponding to the Nth training cycle meets the first preset condition, then end the training, and use the trained image scoring model as a mature image scoring model.

2. The method according to claim 1, characterized in that, the image scoring model includes: a convolutional neural network model and a feature aggregation contrast model, and the step of inputting the sample data into the image scoring model to obtain the predicted preference probability of the sample data includes: Input the sample data into the convolutional neural network model to obtain a first deep feature and a second deep feature; the first deep feature and the second deep feature respectively represent the deep features of the two images; Input the first deep feature and the second deep feature into the feature aggregation contrast model to obtain the predicted preference probability of the sample data.

3. The method according to claim 2, characterized in that, the feature aggregation contrast model includes: a tensor unfolding layer and a fully connected layer, and the step of inputting the first deep feature and the second deep feature into the feature aggregation contrast model to obtain the predicted preference probability of the sample data includes: Use the tensor unfolding layer to one-dimensionally unfold the first deep feature and the second deep feature to obtain the unfolded first deep feature and second deep feature; According to the unfolded first deep feature and second deep feature, use the fully connected layer to predict the preference probability of the sample data to obtain the predicted preference probability of the sample data.

4. The method according to claim 1 or 2 or 3, characterized in that, after obtaining the predicted preference probability of the sample data, it further includes: Multiply the predicted preference probability by a preset value to obtain a predicted relative score; Add the predicted relative score to the sample data.

5. The method according to claim 2 or 3, characterized in that, The feature aggregation contrast model further includes a linear threshold activation layer, which is used to limit the range of the predicted preference probability to 0 to 1.

6. The method according to claim 1, wherein, after obtaining the image relative scoring data set, it further includes: performing a normalization process on the image relative scoring data set to obtain a standardized image relative scoring data set with personal biases eliminated, and replacing the image relative scoring data set with the standardized image relative scoring data set.

7. The method according to claim 1 or 6, wherein, the method further includes: performing a normalization process on the image relative scoring data set to obtain a normalized image scoring data set, and replacing the image relative scoring data set with the normalized image relative scoring data set.

8. An image scoring method, wherein, the method is applied to an image scoring model trained by the image scoring model training method according to any one of claims 1 to 7. The image scoring model includes: a convolutional neural network model and a feature aggregation contrast model. The method includes: inputting two images to be scored into the convolutional neural network model to obtain a third deep feature and a fourth deep feature; inputting the third deep feature and the fourth deep feature into the feature aggregation contrast model to obtain the predicted preference probability of the two images to be scored.

9. An image scoring model training device, wherein, the device is applied to a model training system, and the model training system includes a sample processing module and an image scoring model. The device includes: a data set acquisition module, configured to acquire an image relative scoring data set; the image relative scoring data set includes at least one set of sample data, and any one of the sample data includes two images and a set of relative scores, and the relative scores represent the preference probability for the two images; a predicted preference probability module, configured to input the sample data into the image scoring model to obtain the predicted preference probability of the sample data; the preference probability represents the probability of preferring one of the two images; a probability distribution calculation module, configured to input the sample data into the sample processing module to obtain the preference probability distribution of the sample data; the probability distribution represents the probability distribution of the crowd preferring one of the two images in the sample data; a loss information confirmation module, configured to determine loss information according to the predicted preference probability and the probability distribution; the loss information represents the difference between the predicted preference probability and the mode of the probability distribution; an iteration module, configured to cyclically input the sample data of the image relative scoring data set into the image scoring model and the sample processing module to obtain the loss information of the current training cycle, and end the training until the loss information corresponding to the Nth training cycle meets a first preset condition, and use the trained image scoring model as a mature image scoring model.

10. An image scoring device, wherein, The device is applied to an image scoring model trained by the image scoring model training method according to any one of claims 1 to 7. The image scoring model includes: a convolutional neural network model and a feature aggregation and contrast model. The device includes: A deep feature acquisition module, configured to input two images to be scored into the convolutional neural network model to obtain a third deep feature and a fourth deep feature; A preference probability acquisition module, configured to input the third deep feature and the fourth deep feature into the feature aggregation and contrast model to obtain the predicted preference probability of the two images to be scored.

11. An electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein, when the processor executes the program, it implements the method according to any one of claims 1 to 7 and / or claim 8.

12. A storage medium, on which a computer program is stored, wherein: when the computer program is executed by the processor, it implements the method according to any one of claims 1 to 7 and / or claim 8.

Citation Information

Patent Citations

  • Image display method, evaluation model generation method, device and electronic equipment

    CN110766052A