Network model training method, image processing method, device and electronic device

By constructing and training the image aesthetic evaluation model, using multiple loss functions to train the basic network model, the shortcomings in accuracy and stability of the image aesthetic evaluation model are solved, and higher scoring accuracy and stability are achieved.

CN114402356BActive Publication Date: 2025-05-06SHENZHEN HEYTAP TECHNOLOGY CO LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN201980100428.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2019-11-13
Publication Date
2025-05-06
Estimated Expiration
2039-11-13

AI Technical Summary

Technical Problem

The existing image aesthetic evaluation methods have shortcomings in terms of accuracy and stability, especially when processing slight changes in images, the prediction results of the model are unstable.

Method used

By obtaining the image sample set with the initial scoring distribution data, building the basic network model and defining multiple loss functions, using these loss functions to train the network model until the model converges, thereby improving the accuracy of image aesthetic scores.

Benefits of technology

It achieves higher accuracy and stability of the image aesthetic scoring model, can better capture changes in image scoring distribution, and provide more detailed aesthetic quality evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114402356B_ABST
    Figure CN114402356B_ABST
Patent Text Reader

Abstract

The embodiment of the present application discloses a network model training method, an image processing method, a device and an electronic device. The network model training method includes: obtaining an image sample set; constructing a basic network model and multiple loss functions corresponding to the basic network model; training the basic network model according to the image sample set and the multiple loss functions until the basic network model converges; using the converged basic network model as a scoring model for image aesthetic scoring.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of machine learning, and in particular to a network model training method, an image processing method, a device and an electronic device. Background Art

[0002] With the rapid development of mobile Internet and the rapid popularization of smart phones, the amount of visual content data such as images and videos is increasing day by day. The perception and understanding of these visual contents has become a research direction in multiple interdisciplinary disciplines such as computer vision, computational photography and human psychology. Among them, image aesthetics assessment is a recent research hotspot in the field of computer vision perception and understanding. Image aesthetics reflects the human visual pursuit and yearning for "beautiful" things. Therefore, visual aesthetics assessment is of great significance in the fields of photography, videography, advertising design and art production.

[0003] With the rapid development of machine learning in recent years, the development of repeatable and objective image aesthetic evaluation methods has been greatly promoted. Machine learning, especially deep learning systems, can efficiently and accurately imitate human thinking and processing methods. Therefore, using machine learning or deep learning methods to evaluate image aesthetics is an important research topic. Summary of the invention

[0004] The embodiments of the present application provide a network model training method, an image processing method, a device and an electronic device, which can improve the accuracy of network model training.

[0005] In a first aspect, an embodiment of the present application provides a method for training a network model, comprising:

[0006] Acquire an image sample set, wherein the image sample set includes a plurality of to-be-rated images with initial rating distribution data;

[0007] Constructing a basic network model and a plurality of loss functions corresponding to the basic network model;

[0008] Inputting the image sample set into the basic network model for aesthetic scoring, and obtaining scoring distribution data corresponding to each image to be scored;

[0009] Training the basic network model according to the score distribution data, the initial score distribution data, and the multiple loss functions until the basic network model converges;

[0010] The converged base network model is used as the scoring model for aesthetic scoring of images.

[0011] In a second aspect, an embodiment of the present application provides an image processing method, including:

[0012] Receive aesthetic rating requests;

[0013] Acquiring a target image that needs to be aesthetically scored according to the aesthetic scoring request;

[0014] Call the pre-trained scoring model;

[0015] Performing an aesthetic scoring on the target image according to the scoring model to obtain a target scoring score corresponding to the target image;

[0016] The scoring model is trained using the network model training method provided in this embodiment.

[0017] In a third aspect, an embodiment of the present application provides a network model training device, including:

[0018] A first acquisition module, configured to acquire an image sample set, wherein the image sample set includes a plurality of images to be rated with initial rating distribution data;

[0019] A construction module, used to construct a basic network model and multiple loss functions corresponding to the basic network model;

[0020] A first scoring module is used to input the image sample set into the basic network model for aesthetic scoring, and obtain score distribution data corresponding to each image to be scored;

[0021] A training module, configured to train the basic network model according to the score distribution data, the initial score distribution data and the multiple loss functions until the basic network model converges;

[0022] A determination module is used to use the converged basic network model as a scoring model for performing aesthetic scoring on the image.

[0023] In a fourth aspect, an embodiment of the present application provides an image processing device, including:

[0024] A receiving module, used for receiving an aesthetic scoring request;

[0025] A second acquisition module is used to acquire a target image that needs to be aesthetically scored according to the aesthetic scoring request;

[0026] Call model, used to call the pre-trained scoring model;

[0027] A second scoring module is used to perform aesthetic scoring on the target image according to the scoring model to obtain a target scoring score corresponding to the target image;

[0028] The scoring model is trained using the network model training method provided in this embodiment.

[0029] In a fifth aspect, an embodiment of the present application provides a storage medium on which a computer program is stored, wherein when the computer program is executed on a computer, the computer executes the network model training method or image processing method provided in this embodiment.

[0030] In a sixth aspect, an embodiment of the present application provides an electronic device, including a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to execute, by calling the computer program stored in the memory:

[0031] Acquire an image sample set, wherein the image sample set includes a plurality of to-be-rated images with initial rating distribution data;

[0032] Constructing a basic network model and a plurality of loss functions corresponding to the basic network model;

[0033] Inputting the image sample set into the basic network model for aesthetic scoring, and obtaining scoring distribution data corresponding to each image to be scored;

[0034] Training the basic network model according to the score distribution data, the initial score distribution data, and the multiple loss functions until the basic network model converges;

[0035] The converged base network model is used as the scoring model for aesthetic scoring of images.

[0036] In a seventh aspect, an embodiment of the present application provides an electronic device, including a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to execute, by calling the computer program stored in the memory:

[0037] Receive aesthetic rating requests;

[0038] Acquiring a target image that needs to be aesthetically scored according to the aesthetic scoring request;

[0039] Call the pre-trained scoring model;

[0040] Performing an aesthetic scoring on the target image according to the scoring model to obtain a target scoring score corresponding to the target image;

[0041] The scoring model is trained using the network model training method provided in the embodiment of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] The technical solution and beneficial effects of the present application will be made apparent by describing in detail the specific implementation methods of the present application in conjunction with the accompanying drawings.

[0043] Figure 1 This is a first flow chart of the network model training method provided in the embodiment of the present application.

[0044] Figure 2 This is a second flow chart of the network model training method provided in the embodiment of the present application.

[0045] Figure 3 It is a distribution diagram of initial score distribution data and expected score distribution data provided in the embodiments of the present application.

[0046] Figure 4 It is a flowchart of the image processing method provided in the embodiment of the present application.

[0047] Figure 5 It is a structural diagram of a training device for a network model provided in an embodiment of the present application.

[0048] Figure 6 It is a structural schematic diagram of an image processing device provided in an embodiment of the present application.

[0049] Figure 7 This is a schematic diagram of the first structure of an electronic device provided in an embodiment of the present application.

[0050] Figure 8 This is a second structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0051] Please refer to the drawings, where the same component symbols represent the same components. The principle of the present application is illustrated by implementing it in an appropriate computing environment. The following description is based on the illustrated specific embodiments of the present application, which should not be considered as limiting other specific embodiments of the present application that are not described in detail herein.

[0052] The embodiment of the present application provides a training method for a network model. Among them, the execution subject of the training method of the network model can be the training device of the network model provided in the embodiment of the present application, or an electronic device integrating the training device of the network model. The training device of the network model can be implemented in hardware or software, and the electronic device can be a smart phone, a tablet computer, a PDA, a laptop computer, or a desktop computer, etc., which is equipped with a processor and has processing capabilities. For the convenience of description, the following will be illustrated by taking the execution subject of the training method of the network model as an electronic device.

[0053] See also Figure 1 , Figure 1 This is a first flow chart of the network model training method provided in the embodiment of the present application. The process of the network model training method may include:

[0054] 101. Obtain an image sample set.

[0055] Among them, the electronic device can obtain the image sample set by wired connection or wireless connection. Among them, the image sample set can adopt the existing first large-scale aesthetic quality assessment database (A Large-Scale Database for Aesthetic Visual Analysis, AVA). For the convenience of description, the first large-scale aesthetic quality assessment database is referred to as the AVA public data set below. And all sample images in the AVA public data set are used as images to be scored in this application. It should be noted that the AVA public data set is a database for aesthetic quality assessment. The AVA public data set includes approximately 256,000 sample images, and each sample image is aesthetically scored by multiple different users, wherein the score of the aesthetic score is a multiple natural number between [1,10], for example, the score can be 1, 2 or 10, etc. Further, the initial score distribution data corresponding to the sample image can be generated according to the score data of multiple different users, and the initial score distribution data includes the number of scorers corresponding to each score. It can be understood that the higher the score of the aesthetic score, the higher the aesthetic quality of the sample image.

[0056] It should be noted that the number of people who scored each sample image in the AVA public dataset is between 78 and 539. On average, 210 people participated in the aesthetic scoring of each sample image. Therefore, the large amount of aesthetic scoring data for each sample image can better reflect the public's evaluation and perception of an image. Therefore, this dataset is a recognized benchmark test set in the field of image aesthetic evaluation. Therefore, this application uses the AVA public dataset to train the original network model, and can obtain a more accurate scoring model for image aesthetic scoring.

[0057] 102. Construct a basic network model and multiple loss functions corresponding to the basic network model.

[0058] Among them, since the lightweight deep neural network model is smaller and faster, it is widely used in embedded electronic devices such as smart phones. Therefore, a lightweight deep neural network model can be constructed and used as a basic network model, so that the trained basic network model can be applied to electronic devices such as smart phones, thereby enabling smart phones to perform aesthetic scoring on images, further improving the intelligence of smart phones. Among them, the MobileNets model can be used as the basic network model. Specifically, the core of MobileNets is the depth-separable convolution composed of deep convolution (depthwise conv) and point-to-point convolution (Point conv). It is this depth-separable convolution structure that enables the MobileNets model to reduce network parameters and computation without reducing network performance.

[0059] In some implementations, the MobileNetV2 network model can be constructed as the basic network model. Further, in order to make the score distribution data output by the basic network model between [0,1], the softmax function can be used as the output function of the output layer of the basic network model, so that the score distribution data output by the basic network model includes the probability value corresponding to each score.

[0060] In some embodiments, after building the MobileNetV2 network model, it is also necessary to pre-train the MobileNetV2 network model on the ImageNet database, and use the pre-trained MobileNetV2 network model as the basic network model. It should be noted that the ImageNet database is a large-scale visualization database for visual object recognition software research. The ImageNet database contains more than 15 million images, of which 1.2 million images are divided into 1000 categories (about 1 million images contain bounding boxes and annotations). The MobileNetV2 network model is pre-trained on the ImageNet database, and a MobileNetV2 network model with good network parameters can be obtained at the end of the training. Using the pre-trained MobileNetV2 network model as the basic network model can greatly shorten the training time of the basic network model.

[0061] Among them, since the rating distribution data output by the basic network model includes the probability distribution data of each rating score, it is necessary to obtain the corresponding functions for processing the probability distribution as multiple loss functions of the corresponding basic network model. Through the joint action of multiple loss functions, the difference between the rating distribution data of each image to be rated and the initial rating distribution data can be better captured, and the rating distribution data can be better fitted. Therefore, the basic network model can be better trained through multiple loss functions to obtain a more accurate rating model.

[0062] 103. Input the image sample set into the basic network model for aesthetic scoring, and obtain the score distribution data corresponding to each image to be scored.

[0063] The images to be rated in the image sample set are input into a basic network model such as the MobileNetV2 network model for aesthetic scoring, and the scoring distribution data corresponding to each image to be rated output by the basic network model is obtained. The scoring distribution data includes the probability value corresponding to each scoring score. For example, the output scoring distribution data is where p s1 is the probability value when the score is 1, is the probability value when the score is 2, is the probability value when the score is N-1, is the probability value when the scoring score is N.

[0064] 104. Train the basic network model according to the rating distribution data, the initial rating distribution data and multiple loss functions until the basic network model converges.

[0065] Among them, in order to make the score distribution data output by the basic network model consistent with the initial score distribution data of the user's real score, the basic model can be trained through multiple loss functions. The specific training process can be as follows: after the aesthetic scoring of a batch of images to be scored is completed, the score distribution data and initial score distribution data of a batch of images to be scored can be input into multiple loss functions for calculation to obtain the corresponding target loss value. According to the target loss value, it is determined whether the basic network model meets the convergence condition. When the target loss value gradually approaches a certain value, or fluctuates around a certain value, and the loss change is less than a very small positive number, it can be confirmed that the basic network model converges. If the target loss value does not meet the above conditions, it means that the basic network model does not meet the convergence condition. At this time, the target loss value is returned to the basic network model, and the network parameters of the basic network model, namely the weights and biases, are adjusted according to the back propagation algorithm. And obtain the images to be scored that have not been aesthetically scored by the basic network model, so as to continue to train the adjusted basic network model according to the images to be scored that have not been aesthetically scored by the basic network model and multiple loss functions until the basic network model converges.

[0066] 105. The converged base network model is used as a scoring model for aesthetic scoring of images.

[0067] The converged basic network model is used as a scoring model for aesthetically scoring images. The scoring model can be applied to an electronic device to aesthetically score multiple images stored in the electronic device according to the scoring model, sort the multiple images according to the scoring scores, and display the images according to the sorting results, so that the user can preferentially browse images with high aesthetic scores, i.e., images with higher aesthetic quality.

[0068] As can be seen from the above, the training method of the network model provided in the embodiment of the present application obtains an image sample set, which includes multiple images to be scored with initial score distribution data; constructs a basic network model and multiple loss functions corresponding to the basic network model; inputs the image sample set into the basic network model for aesthetic scoring, and obtains the score distribution data corresponding to each image to be scored; trains the basic network model according to the score distribution data, the initial score distribution data, and multiple loss functions until the basic network model converges; and uses the converged basic network model as a scoring model for aesthetic scoring of images. In this way, the basic network model can be trained through multiple loss functions to obtain a scoring model with higher accuracy, thereby improving the accuracy of the scoring model.

[0069] See also Figure 2 , Figure 21 is a second flow chart of the network model training method provided in the embodiment of the present application. The flow of the network model training method may include:

[0070] 201. Obtain an initial image sample set.

[0071] Among them, the initial image sample set includes multiple sample images with initial rating distribution data. The initial image sample set can adopt the existing AVA public data set. The AVA public data set is a database for aesthetic quality assessment. The AVA public data set includes approximately 256,000 sample images, and each sample image is aesthetically rated by multiple different users, wherein the rating score of the aesthetic rating is a multiple natural number between [1,10]. For example, the rating score can be 10 rating scores. Furthermore, the initial rating distribution data corresponding to the sample image can be generated based on the rating data of multiple different users, and the initial rating distribution data includes the number of raters corresponding to each rating score. It can be understood that the higher the score of the aesthetic rating, the higher the aesthetic quality of the sample image.

[0072] It should be noted that the number of people who scored each sample image in the AVA public dataset is between 78 and 539. On average, 210 people participated in the aesthetic scoring of each sample image. Therefore, the large amount of aesthetic scoring data for each sample image can better reflect the public's evaluation and perception of an image. Therefore, this dataset is a recognized benchmark test set in the field of image aesthetic evaluation. Therefore, this application uses the AVA public dataset to train the original network model, and can obtain a more accurate scoring model for image aesthetic scoring.

[0073] 202. Perform image preprocessing on each sample image to obtain a plurality of first sample images with initial score distribution data.

[0074] Among them, the image preprocessing method includes conventional multi-scale scaling, random flipping, random translation and random cropping, etc., so as to obtain multiple first sample images by image preprocessing of the sample image, and the first sample image is also input into the basic model as the image to be scored for aesthetic scoring, so as to enhance the data volume of the input image, i.e., the image to be scored. However, considering that the composition and color information of the image need to be used as the basis for aesthetic scoring, the present application does not perform image preprocessing with large-scale random cropping and random color changes to ensure that the first sample image and the corresponding sample image are similar in composition and color information. In addition, it can be understood that random scaling, random flipping and random translation of the sample image will not change the composition and color information of the sample image too much. In other words, the first sample image obtained is similar to the corresponding sample image in composition and color information, so the initial score distribution data of the corresponding sample image can be used as the initial score distribution data of the first sample image.

[0075] 203. Add random noise to the first sample image to obtain a first target sample image.

[0076] Among them, since the first sample image is obtained by randomly changing the sample image, such as multi-scale scaling, random flipping, random translation, and random cropping, there must be a slight difference between the first sample data and the sample data. Although this slight difference does not produce a difference in visual recognition, it will directly affect the score distribution data output by the basic network model, so that the score distribution data corresponding to the first sample image is different from the score distribution data of the corresponding sample image. Therefore, after image preprocessing is performed on each sample image, the slight difference operation caused by the image preprocessing will affect the aesthetic score of the first sample image obtained, making the score distribution data predicted by the basic network model unstable.

[0077] Based on this problem, a tiny random noise such as Gaussian noise can be added to each pixel of the first sample image to obtain the first target sample image. It is understandable that for general images, a slight change in image pixels will not produce a difference in visual recognition, that is, a slight change in image pixel values ​​is difficult for the user's eyes to distinguish. In other words, the first target sample image and the first sample image are the same in the user's vision, so the user's rating of the first target sample image and the first sample image will not change. Therefore, under ideal conditions, the score distribution data of the basic network model corresponding to the first sample image and the first target sample image should not be very different. Therefore, it is necessary to train the basic network model through the first target sample image, so that the basic network model can adapt to slight changes in the input sample image, making the basic network model more robust to slight changes in pixel values, thereby making the predicted score distribution data more stable.

[0078] 204. Use the sample image and the first target sample image as images to be scored to obtain an image sample set.

[0079] The sample image and the first target sample image are both used as images to be rated to obtain an image sample set, so that the data volume of the images to be rated in the final image sample set is greater than that of the initial image sample set. The image sample set with a larger data volume is more suitable for training the basic network model, so that the prediction result of the scoring model after training is more accurate. In addition, the sample image and the first target sample image are used as images to be rated to train the basic network model.

[0080] 205. Perform indexation processing on each score data in the initial score distribution data corresponding to each image to be rated to obtain corresponding expected score distribution data.

[0081] Among them, although the initial rating distribution data of each image to be rated can better reflect the public's evaluation and perception of an image, the standard deviation of the initial rating distribution data is large, and the rating scores are mostly concentrated in the middle part, that is, between the rating scores [3, 7], making the initial rating distribution data relatively flat. Therefore, when the basic network model is trained and learned through the initial rating distribution data, the standard deviation of the rating distribution data output by the trained rating model will also be large. As a result, the aesthetic qualities of images with different rating scores are similar, and it is difficult to make detailed differences in aesthetic quality. For example, the aesthetic qualities of images corresponding to the rating scores from 3 to 7 output by the trained rating model are similar, and it is impossible to better reflect the aesthetic quality of the image based on the rating scores.

[0082] Therefore, it is necessary to perform exponential processing on each score data in the initial score distribution data to obtain expected score distribution data with a smaller standard deviation. For example, each score data in the initial score distribution data is exponentially calculated to obtain the expected score distribution data.

[0083] See also Figure 3 , Figure 3 A schematic diagram of the distribution of initial rating distribution data and expected rating distribution data provided in the embodiment of the present application. As shown in the figure, curve A is the initial rating distribution data corresponding to a certain image to be rated provided in the embodiment of the present application, Figure 3 The horizontal axis is the score, and the vertical axis is the probability value corresponding to the score. When the exponent is 2, each score data in the initial score distribution data is squared to obtain the expected score distribution data with a smaller standard deviation, that is, Figure 3 The curve B shown. Figure 3It can be seen that the probability distribution of the score in curve B is more concentrated, which means that the standard deviation of curve B is smaller than that of curve A. Therefore, each score data in the initial score distribution data can be indexed to obtain the corresponding expected score distribution data with a smaller standard deviation. When the basic network model is trained and learned through the expected score distribution data with a smaller standard deviation, a scoring model with a smaller standard deviation of the output result can be obtained.

[0084] In some embodiments, please refer to Figure 3 , we can also obtain a Gaussian function related to the score, and multiply the Gaussian function with the initial score distribution data to obtain the expected score distribution data, that is, Figure 3 Curve C shown. The Gaussian function is a Gaussian function with a small standard deviation, that is, the probability distribution of the Gaussian function is relatively concentrated. Therefore, the Gaussian function with a concentrated probability distribution can be multiplied with the initial score distribution data to reduce the standard deviation of the initial score distribution data and obtain the expected score distribution data. However, the probability distribution of the Gaussian function is relatively absolute, that is, there will only be a single peak close to the middle value. However, considering that there may be a multi-peak situation in the probability distribution of image rating scores, that is, some people think it is very good and some people think it is very bad. For example, the probability value is large when the score is 3, and the probability value is also large when the score is 8. At this time, two peaks will be formed in the initial score distribution data at the score of 3 and the score of 8. In this case, when the Gaussian function is used to reduce the standard deviation of the initial score distribution data, the score with a large probability value will be close to the middle value such as the score of 5 points, making the expected score distribution data obtained inaccurate. Therefore, when adjusting the initial score distribution data according to the Gaussian function, it is necessary to select according to the actual situation.

[0085] 206. Construct a basic network model and a first loss function and a second loss function corresponding to the basic network model.

[0086] Among them, since the lightweight deep neural network model is smaller and faster, it is widely used in embedded electronic devices such as smart phones. Therefore, a lightweight deep neural network model can be constructed and used as a basic network model, so that the trained basic network model can be applied to electronic devices such as smart phones, thereby enabling smart phones to perform aesthetic scoring on images, further improving the intelligence of smart phones. Among them, the MobileNets model can be used as the basic network model. Specifically, the core of MobileNets is the depth-separable convolution composed of deep convolution (depthwise conv) and point-to-point convolution (Point conv). It is this depth-separable convolution structure that enables the MobileNets model to reduce network parameters and computation without reducing network performance.

[0087] In some implementations, the MobileNetV2 network model can be constructed as the basic network model. Further, in order to make the score distribution data output by the basic network model between [0,1], the softmax function can be used as the output function of the output layer of the basic network model, so that the score distribution data output by the basic network model includes the probability value corresponding to each score.

[0088] In some embodiments, after building the MobileNetV2 network model, it is also necessary to pre-train the MobileNetV2 network model on ImageNet, and use the pre-trained MobileNetV2 network model as the basic network model. It should be noted that ImageNet is a large-scale visualization database for visual object recognition software research. The ImageNet database contains more than 15 million images, of which 1.2 million images are divided into 1000 categories (about 1 million images contain bounding boxes and annotations). The MobileNetV2 network model is pre-trained on ImageNet, and a MobileNetV2 network model with better model parameters can be obtained at the end of the training, and the pre-trained MobileNetV2 network model is used as the basic network model, which can greatly shorten the training time of the basic network model.

[0089] Among them, since the rating distribution data output by the basic network model includes the probability distribution data of each rating score, it is necessary to obtain the corresponding different functions for processing the probability distribution as the first loss function and the second loss function of the corresponding basic network model. Through the joint action of the first loss function and the second loss function, the difference between the rating distribution data of each image to be rated and the initial rating distribution data can be better captured, and the rating distribution data can be better fitted. Therefore, the basic network model can be better trained through multiple loss functions to obtain a more accurate rating model.

[0090] Specifically, a first loss function may be constructed using a function for measuring the distance between two distributions. The first loss function may be an Earth Mover's Distance (EMD) loss function. The first loss function is:

[0091]

[0092] in, is the expected rating distribution data, p is the rating distribution data, N is the number of rating scores, N = 10, function CDF p (k) and function is the cumulative probability distribution function, k is the score, CDF p (k) represents the cumulative probability value when the score is k in the score distribution data, represents the cumulative probability value when the score in the expected score distribution data is k, and l is an exponential value. In some embodiments, l=2.

[0093] in, k is the score, Represents the probability value of score i in the score distribution data; k is the score, Represents the probability value of score i in the expected score distribution data.

[0094] Specifically, the second loss function may be a CJS loss function, and the second loss function is:

[0095]

[0096] in, is the expected rating distribution data, p is the rating distribution data, N is the number of rating scores, N = 10, function CDF p (k) and function is the cumulative probability distribution function, k is the score, CDF p (k) represents the cumulative probability value when the score is k in the score distribution data, represents the cumulative probability value when the score in the expected score distribution data is k, and l is an exponential value. In some embodiments, l=2.

[0097] in, k is the score, Represents the probability value of score i in the score distribution data; k is the score, Represents the probability value of score i in the expected score distribution data.

[0098] 207. Input the image sample set into the basic network model for aesthetic scoring, and obtain the score distribution data corresponding to each image to be scored.

[0099] The images to be rated in the image sample set are input into a basic network model such as the MobileNetV2 network model for aesthetic scoring, and the scoring distribution data corresponding to each image to be rated output by the basic network model is obtained. The scoring distribution data includes the probability value corresponding to each scoring score. For example, the output scoring distribution data is where p s1 is the probability value when the score is 1, is the probability value when the score is 2, is the probability value when the score is N-1, is the probability value when the scoring score is N.

[0100] 208. Input the score distribution data and the expected score distribution data into a first loss function to obtain a first loss value.

[0101] The score distribution data and expected score distribution data corresponding to each image to be rated are input into the first loss function to obtain a first loss value corresponding to each image to be rated.

[0102] 209. Input the score distribution data and the expected score distribution data into a second loss function to obtain a second loss value.

[0103] The rating distribution data and expected rating distribution data corresponding to each image to be rated are input into the second loss function to obtain a second loss value corresponding to each image to be rated.

[0104] 210. Determine a target loss value according to the first loss value and the second loss value.

[0105] Among them, after obtaining the first loss value and the second loss value, the electronic device can determine the target loss value according to the first loss value and the second loss value. Specifically, the first loss value can be multiplied by the first weight value to obtain the third loss value. And the second loss value can be multiplied by the second weight value to obtain the fourth loss value. Finally, the third loss value and the fourth loss value are added to obtain the target loss value. Among them, the first weight value and the second weight value can be set according to actual conditions, and the first weight value and the second weight value can be two values ​​with equal values, or two values ​​with different values. For example, the first weight value can be 0.6, and the second weight value can be 0.4. Or the first weight value and the second weight value can both be 0.5.

[0106] 211. Adjust the parameters of the basic network model according to the target loss value until the basic network model converges.

[0107] Among them, the target loss value is used to determine whether the basic network model meets the convergence conditions. When the target loss value gradually approaches a certain value, or fluctuates around a certain value, and the loss change is less than a very small positive number, the basic network model can be confirmed to have converged. If the target loss value does not meet the above conditions, it means that the basic network model does not meet the convergence conditions. At this time, the target loss value is transmitted back to the basic network model, and the network parameters of the basic network model, namely the weights and biases, are adjusted according to the back-propagation algorithm. The adjusted basic network model is trained continuously until the basic network model converges.

[0108] In some embodiments, after each batch of images to be scored is input to train the basic network model, a network model with adjusted parameters can be obtained, and the electronic device can obtain a batch of verification images from the verification set and input them into the network model with adjusted parameters to verify the accuracy of the network model with adjusted parameters. When the accuracy obtained this time is greater than the accuracy obtained last time, the electronic device can save the parameters of the network model with adjusted parameters. When the accuracy obtained this time is less than the accuracy obtained last time, the electronic device may not save the network model with adjusted parameters. When the accuracy of the network model with adjusted parameters obtained multiple times does not increase, for example, when the accuracy of the network model with adjusted parameters obtained multiple times is 87%, 86.9%, 86.7%, and 86.8% respectively, the electronic device can confirm that the training of the basic network model is completed, that is, the basic network model converges.

[0109] 212. The converged base network model is used as a scoring model for aesthetic scoring of images.

[0110] The converged basic network model is obtained, and the converged basic network model is used as a scoring model for aesthetically scoring images. The scoring model can be applied to electronic devices to aesthetically score images stored in the electronic devices according to the scoring model, sort the images according to the scoring scores, and display the images according to the sorting results, so that users can preferentially browse images with high aesthetic scores, i.e., images with higher aesthetic quality.

[0111] As can be seen from the above, the training method of the network model provided in the embodiment of the present application obtains an image sample set, which includes multiple images to be scored with initial score distribution data; constructs a basic network model and multiple loss functions corresponding to the basic network model; inputs the image sample set into the basic network model for aesthetic scoring, and obtains the score distribution data corresponding to each image to be scored; trains the basic network model according to the score distribution data, the initial score distribution data, and multiple loss functions until the basic network model converges; and uses the converged basic network model as a scoring model for aesthetic scoring of images. In this way, the basic network model can be trained through multiple loss functions to obtain a scoring model with higher accuracy, thereby improving the accuracy of the scoring model.

[0112] See also Figure 4 , Figure 4 : is a schematic diagram of a process flow of an image processing method provided in an embodiment of the present application. The process flow of the image processing method may include:

[0113] 301. Receive an aesthetic scoring request.

[0114] Among them, when the electronic device receives a touch operation of the target component, a preset voice operation, or a start instruction of a preset target application, a request for generating an aesthetic score is triggered. In addition, the electronic device can also automatically trigger the generation of an aesthetic score request at intervals of a preset duration or based on certain triggering rules. For example, when the electronic device detects that the current display interface includes multiple images, such as when it detects that the electronic device starts a browser application to browse an article page containing images, it can automatically trigger the generation of an aesthetic score request, and perform aesthetic scoring on multiple images in the current page according to the scoring model. This allows the electronic device to sort multiple images according to different scoring scores, and give priority to displaying images with high scoring scores, that is, images with good aesthetic quality.

[0115] 302. Acquire a target image that needs to be aesthetically scored according to the aesthetic scoring request.

[0116] The target image may be an image stored in an electronic device. In this case, the aesthetic scoring request includes path information indicating the location where the target image is stored. The electronic device can obtain the target image that needs to be aesthetically scored through the path information. Of course, when the target image is not an image stored in the electronic device, the electronic device can obtain the target image that needs to be aesthetically scored through a wired connection or a wireless connection according to the aesthetic scoring request.

[0117] 303. Call the pre-trained scoring model.

[0118] The scoring model is obtained by training using the network model training method provided in this embodiment. The specific network model training process can refer to the relevant description of the above embodiment, which will not be repeated here.

[0119] 304. Perform aesthetic scoring on the target image according to the scoring model to obtain a target scoring score corresponding to the target image.

[0120] The target image is input into the scoring model for aesthetic scoring to obtain a target scoring score corresponding to the target image. The target scoring score can represent the aesthetic quality of the target image. The higher the target scoring score, the higher the aesthetic quality of the target image, that is, the target image is more in line with the public's aesthetics.

[0121] In some embodiments, the step of performing aesthetic scoring on the target image according to the scoring model to obtain a target scoring score corresponding to the target image includes: performing aesthetic scoring on the target image according to the scoring model to obtain target scoring distribution data corresponding to the target image, the target scoring distribution data being probability distribution data of each scoring score; and taking the scoring score with the largest probability value in the target scoring distribution data as the target scoring score.

[0122] In some embodiments, when the electronic device scores multiple target images in an album or image library stored in the electronic device, after obtaining the target score corresponding to each target image, it can also detect whether the target score is greater than a preset score value, and delete the target image whose target score is less than or equal to the preset score value. It is understandable that when the target score is less than or equal to the preset score value, it indicates that the aesthetic quality of the target image is not high, that is, the target image may be an image with unclear images or an image with incomplete composition. It is understandable that such unclear images and images with incomplete composition are highly likely to be invalid images obtained by the user accidentally pressing the shooting key, so such images are not images that the user needs to save, and such images will also occupy the storage space of the electronic device. Therefore, the electronic device of the present application can regularly trigger the corresponding aesthetic scoring request to screen multiple images stored in the electronic device through the scoring model, and intelligently delete the target images whose aesthetic scores are less than or equal to the preset score value, which can more intelligently help users manage images in the electronic device and save the memory space of the electronic device.

[0123] As can be seen from the above, the image processing method provided in the embodiment of the present application receives an aesthetic scoring request and obtains a target image that needs to be aesthetically scored according to the aesthetic scoring request; calls a pre-trained scoring model; performs aesthetic scoring on the target image according to the scoring model to obtain a target scoring score corresponding to the target image; and thereby obtains an aesthetic score for the target image through preset training.

[0124] See also Figure 5 , Figure 5 A schematic diagram of the structure of a network model training device provided in an embodiment of the present application. The network model training device may include: a first acquisition module 41, a construction module 42, a first scoring module 43, a training module 44 and a determination module 45.

[0125] The first acquisition module 41 is used to acquire an image sample set, where the image sample set includes a plurality of images to be rated with initial rating distribution data.

[0126] The construction module 42 is used to construct a basic network model and a plurality of loss functions corresponding to the basic network model.

[0127] The first scoring module 43 is used to input the image sample set into the basic network model for aesthetic scoring, and obtain scoring distribution data corresponding to each image to be scored.

[0128] The training module 44 is used to train the basic network model according to the score distribution data, the initial score distribution data and the multiple loss functions until the basic network model converges.

[0129] A determination module 45 is used to use the converged basic network model as a scoring model for performing aesthetic scoring on the image.

[0130] In some embodiments, the first acquisition module 41 is specifically used to acquire an initial image sample set, wherein the initial image sample set includes multiple sample images with initial rating distribution data; perform image preprocessing on each of the sample images to obtain multiple first sample images with initial rating distribution data; add random noise to the first sample images to obtain a first target sample image; obtain an image sample set based on the sample images and the first target sample images, and use the sample images and the first target sample images as images to be rated.

[0131] In some embodiments, before the step of constructing the basic network model and multiple loss functions corresponding to the basic network model, the construction module 42 is also used to: adjust the initial score distribution data corresponding to each image to be scored to obtain corresponding expected score distribution data, wherein the standard deviation of the expected score distribution data corresponding to each image to be scored is smaller than the standard deviation of the initial score distribution data.

[0132] In some embodiments, when the construction module 42 adjusts the initial score distribution data corresponding to each image to be rated to obtain the corresponding expected score distribution data, it is specifically used to index each score data in the initial score distribution data corresponding to each image to be rated to obtain the corresponding expected score distribution data.

[0133] In some implementations, the training module 44 is specifically configured to train the basic network model according to the rating distribution data, the expected rating distribution data, and the multiple loss functions.

[0134] In some embodiments, multiple loss functions include a first loss function and a second loss function, and the training module 44 is specifically used to input the score distribution data and the expected score distribution data into the first loss function to obtain a first loss value; input the score distribution data and the expected score distribution data into the second loss function to obtain a second loss value; determine a target loss value based on the first loss value and the second loss value; and adjust the parameters of the basic network model based on the target loss value.

[0135] In some embodiments, the training module 44 multiplies the first loss value by the first weight value to obtain a third loss value; multiplies the second loss value by the second weight value to obtain a fourth loss value; and adds the third loss value and the fourth loss value to obtain a target loss value.

[0136] As can be seen from the above, the network model training device provided in the embodiment of the present application obtains an image sample set through the first acquisition module 41, and the image sample set includes multiple images to be scored with initial score distribution data; the construction module 42 constructs a basic network model and multiple loss functions corresponding to the basic network model; the first scoring module 43 inputs the image sample set into the basic network model for aesthetic scoring, and obtains the score distribution data corresponding to each image to be scored; the training module 44 trains the basic network model according to the score distribution data, the initial score distribution data and multiple loss functions until the basic network model converges; the determination module 45 uses the converged basic network model as a scoring model for aesthetic scoring of images. In this way, the basic network model can be trained through multiple loss functions to obtain a scoring model with higher accuracy, thereby improving the accuracy of the scoring model.

[0137] It should be noted that the network model training device provided in the embodiment of the present application and the network model training method in the above embodiment belong to the same concept. Any method provided in the network model training method embodiment can be run on the network model training device. The specific implementation process is detailed in the network model training method embodiment, which will not be repeated here.

[0138] See also Figure 6 , Figure 6 The schematic diagram of the structure of the image processing device 500 provided in the embodiment of the present application is as follows. The image processing device may include: a receiving module 51, a second acquisition module 52, a calling model 53, and a second scoring module 54.

[0139] A receiving module 51 is used to receive an aesthetic scoring request;

[0140] A second acquisition module 52 is used to acquire a target image that needs to be aesthetically scored according to the aesthetic scoring request;

[0141] Calling model 53, used to call the pre-trained scoring model;

[0142] A second scoring module 54 is used to perform aesthetic scoring on the target image according to the scoring model to obtain a target scoring score corresponding to the target image;

[0143] The scoring model is trained using the network model training method provided in the embodiment of the present application.

[0144] In some embodiments, the second scoring module 54 is specifically configured to perform aesthetic scoring on the target image according to the scoring model to obtain target scoring distribution data corresponding to the target image; and use the scoring score with the largest probability value in the target scoring distribution data as the target scoring score.

[0145] As can be seen from the above, the image processing device provided in the embodiment of the present application receives an aesthetic scoring request through the receiving module 51; the second acquisition module 52 obtains the target image that needs to be aesthetically scored according to the aesthetic scoring request; the calling model 53 calls the pre-trained scoring model; the second scoring module 54 aesthetically scores the target image according to the scoring model to obtain a target scoring score corresponding to the target image; in this way, the aesthetic scoring of the target image is obtained through preset training.

[0146] It should be noted that the image processing device provided in the embodiment of the present application and the image processing method in the above embodiment belong to the same concept. Any method provided in the image processing method embodiment can be run on the image processing device. The specific implementation process is detailed in the image processing method embodiment, which will not be repeated here.

[0147] An embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon. When the computer program stored thereon is executed on a computer, the computer executes a network model training method or an image processing method as provided in an embodiment of the present application.

[0148] The storage medium may be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).

[0149] An embodiment of the present application also provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor calls the computer program stored in the memory to execute a network model training method or an image processing method as provided in an embodiment of the present application.

[0150] For example, the electronic device may be a mobile terminal such as a tablet computer or a smart phone. Figure 7 , Figure 7 A first structural schematic diagram of an electronic device provided in an embodiment of the present application.

[0151] The electronic device 600 may include components such as a memory 601 and a processor 602. Those skilled in the art will appreciate that Figure 7 The electronic device structure shown in the figure does not constitute a limitation of the electronic device, and may include more or less components than shown in the figure, or combine certain components, or arrange the components differently.

[0152] The memory 601 can be used to store software programs and modules, and the processor 602 executes various functional applications and data processing by running the computer programs and modules stored in the memory 601. The memory 601 can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, a computer program required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area can store data created according to the use of the electronic device, etc.

[0153] The processor 602 is the control center of the electronic device. It uses various interfaces and lines to connect various parts of the entire electronic device. It executes various functions of the electronic device and processes data by running or executing applications stored in the memory 601 and calling data stored in the memory 601, thereby monitoring the electronic device as a whole.

[0154] In addition, the memory 601 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage devices. Accordingly, the memory 601 may also include a memory controller to provide the processor 602 with access to the memory 601.

[0155] In this embodiment, the processor 602 in the electronic device will load the executable code corresponding to the process of one or more application programs into the memory 601 according to the following instructions, and the processor 602 will run the application program stored in the memory 601, thereby implementing the process:

[0156] Acquire an image sample set, wherein the image sample set includes a plurality of to-be-rated images with initial rating distribution data;

[0157] Constructing a basic network model and a plurality of loss functions corresponding to the basic network model;

[0158] Inputting the image sample set into the basic network model for aesthetic scoring, and obtaining scoring distribution data corresponding to each image to be scored;

[0159] Training the basic network model according to the score distribution data, the initial score distribution data, and the multiple loss functions until the basic network model converges;

[0160] The converged base network model is used as the scoring model for aesthetic scoring of images.

[0161] In some implementations, before the processor 602 executes the construction of a basic network model and a plurality of loss functions corresponding to the basic network model, the processor 602 may execute:

[0162] The initial rating distribution data corresponding to each image to be rated is adjusted to obtain corresponding expected rating distribution data, wherein the standard deviation of the expected rating distribution data corresponding to each image to be rated is smaller than the standard deviation of the initial rating distribution data.

[0163] In some implementations, when the processor 602 performs training on the basic network model according to the score distribution data, the initial score distribution data, and the multiple loss functions, the processor 602 may perform:

[0164] The basic network model is trained according to the score distribution data, the expected score distribution data and the multiple loss functions.

[0165] In some implementations, the multiple loss functions include a first loss function and a second loss function. When the processor 602 executes the training of the basic network model according to the score distribution data, the expected score distribution data, and the multiple loss functions, the following steps may be performed:

[0166] Inputting the score distribution data and the expected score distribution data into a first loss function to obtain a first loss value;

[0167] Inputting the score distribution data and the expected score distribution data into a second loss function to obtain a second loss value;

[0168] Determine a target loss value according to the first loss value and the second loss value;

[0169] The parameters of the basic network model are adjusted according to the target loss value.

[0170] In some implementations, when the processor 602 determines the target loss value according to the first loss value and the second loss value, the processor 602 may execute:

[0171] Multiplying the first loss value by the first weight value to obtain a third loss value;

[0172] Multiplying the second loss value by the second weight value to obtain a fourth loss value;

[0173] The third loss value and the fourth loss value are added to obtain a target loss value.

[0174] In some implementations, when the processor 602 adjusts the initial rating distribution data corresponding to each image to be rated to obtain the corresponding expected rating distribution data, the following steps may be performed:

[0175] Each score data in the initial score distribution data corresponding to each image to be rated is subjected to indexation processing to obtain corresponding expected score distribution data.

[0176] In some implementations, when the processor 602 executes acquiring an image sample set, the processor 602 may execute:

[0177] Acquire an initial image sample set, wherein the initial image sample set includes a plurality of sample images with initial score distribution data;

[0178] Performing image preprocessing on each of the sample images to obtain a plurality of first sample images with initial score distribution data;

[0179] Adding random noise to the first sample image to obtain a first target sample image;

[0180] An image sample set is obtained according to the sample image and the first target sample image, and the sample image and the first target sample image are used as images to be scored.

[0181] In this embodiment, the processor 602 in the electronic device will load the executable code corresponding to the process of one or more application programs into the memory 601 according to the following instructions, and the processor 602 will run the application program stored in the memory 601, thereby implementing the process:

[0182] Receive aesthetic rating requests;

[0183] Acquiring a target image that needs to be aesthetically scored according to the aesthetic scoring request;

[0184] Call the pre-trained scoring model;

[0185] Performing an aesthetic scoring on the target image according to the scoring model to obtain a target scoring score corresponding to the target image;

[0186] The scoring model is trained using the network model training method provided in the embodiment of the present application.

[0187] In some implementations, when the processor 602 performs aesthetic scoring on the target image according to the scoring model to obtain a target scoring score corresponding to the target image, the following steps may be performed:

[0188] Performing aesthetic scoring on the target image according to the scoring model to obtain target score distribution data corresponding to the target image;

[0189] The score with the largest probability value in the target score distribution data is taken as the target score.

[0190] Please refer to Figure 8 , Figure 8 A second structural diagram of an electronic device provided in an embodiment of the present application, and Figure 7The difference between the electronic device shown is that the electronic device further includes: a camera component 603, a radio frequency circuit 604, an audio circuit 605 and a power supply 606. The display 603, the radio frequency circuit 604, the audio circuit 605 and the power supply 606 are electrically connected to the processor 602 respectively.

[0191] The display 603 may be used to display information input by a user or information provided to a user and various graphical user interfaces, which may be composed of graphics, text, icons, videos, and any combination thereof. The display 603 may include a display panel, and in some embodiments, the display panel may be configured in the form of a liquid crystal display (LCD) or an organic light-emitting diode (OLED).

[0192] The radio frequency circuit 604 can be used to send and receive radio frequency signals to establish wireless communication with network devices or other electronic devices through wireless communication, and to send and receive signals between network devices or other electronic devices.

[0193] The audio circuit 605 may be used to provide an audio interface between a user and the electronic device through a speaker and a microphone.

[0194] The power supply 606 can be used to supply power to various components of the electronic device 600. In some embodiments, the power supply 606 can be logically connected to the processor 602 through a power management system, so that the power management system can manage charging, discharging, power consumption and other functions.

[0195] although Figure 8 Not shown, the electronic device 600 may also include a camera component, a Bluetooth module, etc. The camera component may include an image processing circuit, which may be implemented using hardware and / or software components and may include various processing units that define an image signal processing (Image Signal Processing) pipeline. The image processing circuit may include at least: multiple cameras, an image signal processor (Image Signal Processor, ISP processor), a control logic, an image memory, and a display, etc. Each camera may include at least one or more lenses and an image sensor. The image sensor may include a color filter array (such as a Bayer filter). The image sensor may obtain light intensity and wavelength information captured by each imaging pixel of the image sensor, and provide a set of raw image data that can be processed by an image signal processor.

[0196] In the above embodiments, the description of each embodiment has its own focus. For the parts that are not described in detail in a certain embodiment, please refer to the detailed description of the training method / image processing method of the network model above, which will not be repeated here.

[0197] The network model training method / image processing method device provided in the embodiment of the present application belongs to the same concept as the network model training method / image processing method in the above embodiment. Any method provided in the network model training method / image processing method embodiment can be run on the network model training method / image processing method device. The specific implementation process is detailed in the network model training method / image processing method embodiment, which will not be repeated here.

[0198] It should be noted that, for the training method of the network model / image processing method described in the embodiment of the present application, a person of ordinary skill in the art can understand that all or part of the process of implementing the training method of the network model / image processing method described in the embodiment of the present application can be completed by controlling the relevant hardware through a computer program, and the computer program can be stored in a computer-readable storage medium, such as stored in a memory, and executed by at least one processor, and during the execution process, it may include the process of the embodiment of the training method of the network model / image processing method. Among them, the storage medium can be a disk, an optical disk, a read-only memory (ROM), a random access memory (RAM), etc.

[0199] For the training method / image processing method device of the network model in the embodiment of the present application, each functional module can be integrated into a processing chip, or each module can exist physically separately, or two or more modules can be integrated into one module. The above-mentioned integrated module can be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium, such as a read-only memory, a disk or an optical disk.

[0200] The above is a detailed introduction to a network model training method, image processing method, device, storage medium and electronic device provided in the embodiments of the present application. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea; at the same time, for technical personnel in this field, according to the ideas of the present application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.

Claims

1. A method for training a network model, wherein: include: Acquire an image sample set, wherein the image sample set includes a plurality of to-be-rated images with initial rating distribution data; Exponentially process each score data in the initial score distribution data corresponding to each image to be rated to obtain corresponding expected score distribution data, wherein the standard deviation of the expected score distribution data corresponding to each image to be rated is less than the standard deviation of the initial score distribution data; Constructing a basic network model and a plurality of loss functions corresponding to the basic network model; Inputting the image sample set into the basic network model for aesthetic scoring, and obtaining scoring distribution data corresponding to each image to be scored; Training the basic network model according to the score distribution data, the expected score distribution data, and the multiple loss functions until the basic network model converges; The converged base network model is used as the scoring model for aesthetic scoring of images.

2. The method according to claim 1, wherein: The multiple loss functions include a first loss function and a second loss function, and the step of training the basic network model according to the score distribution data, the expected score distribution data and the multiple loss functions includes: Inputting the score distribution data and the expected score distribution data into a first loss function to obtain a first loss value; Inputting the score distribution data and the expected score distribution data into a second loss function to obtain a second loss value; Determine a target loss value according to the first loss value and the second loss value; The parameters of the basic network model are adjusted according to the target loss value.

3. The method according to claim 2, wherein: The step of determining a target loss value according to the first loss value and the second loss value comprises: Multiplying the first loss value by the first weight value to obtain a third loss value; Multiplying the second loss value by the second weight value to obtain a fourth loss value; The third loss value and the fourth loss value are added to obtain a target loss value.

4. The method according to any one of claims 2 to 3, wherein: The first loss function is: in, is the expected rating distribution data, p is the rating distribution data, N is the number of rating scores, and the function CDF p (k) and function is the cumulative probability distribution function, k is the score, CDF p (k) represents the cumulative probability value when the score is k in the score distribution data, It represents the cumulative probability value when the score is k in the expected score distribution data, and l is the exponential value; The second loss function is: in, is the expected rating distribution data, p is the rating distribution data, N is the number of rating scores, and the function CDF p (k) and function is the cumulative probability distribution function, k is the score, CDF p (k) represents the cumulative probability value when the score is k in the score distribution data, Represents the cumulative probability value when the score is k in the expected score distribution data.

5. The method according to any one of claims 1 to 3, wherein: The step of obtaining an image sample set comprises: Acquire an initial image sample set, wherein the initial image sample set includes a plurality of sample images with initial score distribution data; Performing image preprocessing on each of the sample images to obtain a plurality of first sample images with initial score distribution data; Adding random noise to the first sample image to obtain a first target sample image; The sample image and the first target sample image are both used as images to be scored to obtain a sample image set.

6. A method for processing an image, wherein: include: Receive aesthetic rating requests; Acquiring a target image that needs to be aesthetically scored according to the aesthetic scoring request; Call the pre-trained scoring model; Performing an aesthetic scoring on the target image according to the scoring model to obtain a target scoring score corresponding to the target image; Wherein, the scoring model is trained using the training method of the network model described in any one of claims 1 to 5.

7. The method according to claim 6, wherein: The step of performing aesthetic scoring on the target image according to the scoring model to obtain a target scoring score corresponding to the target image comprises: Performing aesthetic scoring on the target image according to the scoring model to obtain target score distribution data corresponding to the target image; The score with the largest probability value in the target score distribution data is taken as the target score.

8. A network model training device, wherein: include: A first acquisition module is used to acquire an image sample set, wherein the image sample set includes a plurality of images to be rated with initial rating distribution data, and to perform exponential processing on each rating data in the initial rating distribution data corresponding to each image to be rated to obtain corresponding expected rating distribution data, wherein a standard deviation of the expected rating distribution data corresponding to each image to be rated is smaller than a standard deviation of the initial rating distribution data; A construction module, used to construct a basic network model and multiple loss functions corresponding to the basic network model; A first scoring module is used to input the image sample set into the basic network model for aesthetic scoring, and obtain score distribution data corresponding to each image to be scored; A training module, configured to train the basic network model according to the score distribution data, the expected score distribution data and the multiple loss functions until the basic network model converges; A determination module is used to use the converged basic network model as a scoring model for performing aesthetic scoring on the image.

9. An image processing device, wherein: include: A receiving module, used for receiving an aesthetic scoring request; A second acquisition module is used to acquire a target image that needs to be aesthetically scored according to the aesthetic scoring request; Call model, used to call the pre-trained scoring model; A second scoring module is used to perform aesthetic scoring on the target image according to the scoring model to obtain a target scoring score corresponding to the target image; Wherein, the scoring model is trained using the training method of the network model described in any one of claims 1 to 5.

10. A storage medium, wherein: The storage medium stores a computer program, and when the computer program runs on a computer, the computer executes the network model training method described in any one of claims 1 to 5 or the image processing method described in any one of claims 6 to 7.

11. An electronic device, wherein: The electronic device includes a processor and a memory, wherein a computer program is stored in the memory, and the processor is configured to execute, by calling the computer program stored in the memory: Acquire an image sample set, wherein the image sample set includes a plurality of to-be-rated images with initial rating distribution data; Exponentially process each score data in the initial score distribution data corresponding to each image to be rated to obtain corresponding expected score distribution data, wherein the standard deviation of the expected score distribution data corresponding to each image to be rated is less than the standard deviation of the initial score distribution data; Constructing a basic network model and a plurality of loss functions corresponding to the basic network model; Inputting the image sample set into the basic network model for aesthetic scoring, and obtaining scoring distribution data corresponding to each image to be scored; Training the basic network model according to the score distribution data, the expected score distribution data, and the multiple loss functions until the basic network model converges; The converged base network model is used as the scoring model for aesthetic scoring of images.

12. The electronic device according to claim 11, wherein: The processor is configured to execute: Inputting the score distribution data and the expected score distribution data into a first loss function to obtain a first loss value; Inputting the score distribution data and the expected score distribution data into a second loss function to obtain a second loss value; Determine a target loss value according to the first loss value and the second loss value; The parameters of the basic network model are adjusted according to the target loss value.

13. The electronic device according to claim 12, wherein: The processor is configured to execute: Multiplying the first loss value by the first weight value to obtain a third loss value; Multiplying the second loss value by the second weight value to obtain a fourth loss value; The third loss value and the fourth loss value are added to obtain a target loss value.

14. The electronic device according to any one of claims 11 to 13, wherein: The processor is configured to execute: Acquire an initial image sample set, wherein the initial image sample set includes a plurality of sample images with initial score distribution data; Performing image preprocessing on each of the sample images to obtain a plurality of first sample images with initial score distribution data; Adding random noise to the first sample image to obtain a first target sample image; The sample image and the first target sample image are used as images to be scored to obtain an image sample set.

15. An electronic device, wherein: The electronic device includes a processor and a memory, wherein a computer program is stored in the memory, and the processor is configured to execute, by calling the computer program stored in the memory: Receive aesthetic rating requests; Acquiring a target image that needs to be aesthetically scored according to the aesthetic scoring request; Call the pre-trained scoring model; Performing an aesthetic scoring on the target image according to the scoring model to obtain a target scoring score corresponding to the target image; Wherein, the scoring model is trained using the training method of the network model described in any one of claims 1 to 5.

16. The electronic device according to claim 15, wherein: The processor is configured to execute: Performing aesthetic scoring on the target image according to the scoring model to obtain target score distribution data corresponding to the target image; The score with the largest probability value in the target score distribution data is taken as the target score.

Citation Information

Patent Citations

  • Image evaluation method and device and computer readable storage medium

    CN110223292A