Image processing method, device, computer equipment and storage medium
By constructing similar training samples of high-quality picture collections and using neural networks to generate a style evaluation model, the objectivity problem of picture style judgment is solved, the accuracy of style quality score is improved, and the effective acquisition of picture features and the judgment of community tone is achieved.
Patent Information
- Application Number
- CN202011219812.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-11-04
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2040-11-04
AI Technical Summary
In the prior art, there is a lack of objective standards for judging picture styles, resulting in insufficient accuracy of picture style quality scores.
By obtaining high-quality and different styles of picture collections, computer vision technology extracts similar pictures from low-quality sets, constructs training samples, and uses neural networks in machine learning to train a classifier to generate a style evaluation model to score the quality of picture styles.
It improves the accuracy of picture style quality scores, can effectively obtain effective features of the picture, and helps to judge and establish the tone of the picture community.
Smart Images

Figure CN113392865B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to an image processing method, apparatus, computer equipment, and storage medium. Background Art
[0002] The style of a picture and the classification of a picture are two different technical issues. For the problem of picture classification, most processing methods are to first determine the category label, then manually mark the category to which the picture belongs, and then train the classifier on these picture data with category labels to obtain the classification results of the picture type. However, there is no category for judging the style of a picture. Therefore, judging the style of a picture is a very subjective issue. How to improve the accuracy of the style quality rating of a picture is a hot issue today. Summary of the Invention
[0003] The embodiments of the present application provide an image processing method, apparatus, computer device, and storage medium, which can improve the accuracy of the style quality rating of an image.
[0004] On one hand, an embodiment of the present application discloses a method for processing an image, the method comprising:
[0005] Acquire a first picture set and a second picture set, where the style of pictures included in the first picture set is different from the style of pictures included in the second picture set, and the quality of pictures included in the first picture set is higher than the quality of pictures included in the second picture set;
[0006] Obtaining, from the second picture set, k pictures that are similar to each picture in the first picture set, and determining k groups of training samples based on the first picture set and the k pictures similar to each picture, each group of training samples including the first picture set and one of the k pictures similar to each picture, where k is an integer greater than or equal to 1;
[0007] K classifiers are trained using the k groups of training samples, and a style evaluation model is generated based on the k classifiers obtained through training. The style evaluation model is used to score the style quality of the picture.
[0008] In one aspect, an embodiment of the present application discloses an image processing device, which includes:
[0009] an acquiring unit, configured to acquire a first picture set and a second picture set, wherein the style of pictures included in the first picture set is different from the style of pictures included in the second picture set, and the quality of pictures included in the first picture set is higher than the quality of pictures included in the second picture set;
[0010] The acquisition unit is further configured to acquire, from the second picture set, k pictures that are similar to each picture in the first picture set;
[0011] a determining unit, configured to determine k groups of training samples based on the first picture set and k pictures similar to each of the pictures, each group of training samples including the first picture set and one of the k pictures similar to each of the pictures, where k is an integer greater than or equal to 1;
[0012] A processing unit, configured to train k classifiers using the k groups of training samples;
[0013] The determining unit is further configured to generate a style evaluation model based on the k classifiers obtained through training, and the style evaluation model is configured to score the style quality of the image.
[0014] On one hand, an embodiment of the present application discloses a computer device, which includes a memory and a processor: the memory is used to store a computer program; the processor runs the computer program to implement the above-mentioned image processing method.
[0015] On one hand, an embodiment of the present application discloses a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the above-mentioned image processing method is executed.
[0016] In one aspect, embodiments of the present application disclose a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the above-described image processing method.
[0017] In an embodiment of the present application, a computer device obtains a first picture set and a second picture set, wherein the style of the pictures included in the first picture set is different from the style of the pictures included in the second picture set, and the quality of the pictures included in the first picture set is higher than the quality of the pictures included in the second picture set; the computer device obtains k pictures similar to each picture in the first picture set from the second picture set, and determines k groups of training samples based on the first picture set and the k pictures similar to each picture, each group of training samples includes the first picture set and one of the k pictures similar to each picture, k is an integer greater than or equal to 1, and the training samples are determined by this method to achieve the marginalization of the main semantic information of the picture; k classifiers are trained using the k groups of training samples, and a style evaluation model is generated based on the k classifiers obtained by training, and the style evaluation model is used to score the style quality of the picture. Through the above embodiment, the accuracy of the style quality scoring of the picture can be improved, and further based on the style quality score of the picture, the effective features of the picture can be effectively obtained, which is also helpful for judging and establishing the tonality of the picture community. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0019] Figure 1a This is a schematic diagram of the architecture of an image processing system disclosed in an embodiment of the present application;
[0020] Figure 1b This is a schematic diagram of a form of image processing product disclosed in an embodiment of the present application;
[0021] Figure 1c This is a schematic diagram of the structure of a classifier disclosed in an embodiment of the present application;
[0022] Figure 1d This is a schematic diagram of the structure of a residual learning network disclosed in an embodiment of the present application;
[0023] Figure 1e Schematic diagram of the structure of a global average pooling layer disclosed in an embodiment of the present application;
[0024] Figure 2 This is a flowchart of an image processing method disclosed in an embodiment of the present application;
[0025] Figure 3 This is a flowchart of another image processing method disclosed in an embodiment of the present application;
[0026] Figure 4 This is a schematic diagram of a framework of an image processing method disclosed in an embodiment of the present application;
[0027] Figure 5 This is a schematic diagram of the structure of an image processing device disclosed in an embodiment of the present application;
[0028] Figure 6 It is a structural diagram of a computer device disclosed in an embodiment of the present application. DETAILED DESCRIPTION
[0029] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0030] Artificial Intelligence (AI) refers to the theories, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. AI also studies the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making.
[0031] Artificial intelligence (AI) technology is a comprehensive discipline encompassing a wide range of fields, encompassing both hardware and software technologies. Foundational AI technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, speech processing, natural language processing, and machine learning / deep learning.
[0032] This application relates to computer vision and machine learning, both of which fall under the umbrella of artificial intelligence. Computer vision (CV) is the science of making machines "see." Specifically, it refers to machine vision techniques that use cameras and computers to replace the human eye to identify, track, and measure targets, and further image processing to transform computer-generated images into images more suitable for human observation or transmission to instrumentation. As a scientific discipline, computer vision studies related theories and technologies, attempting to establish artificial intelligence systems capable of extracting information from images or multidimensional data. Computer vision technologies typically include image processing, image recognition, image semantic understanding, image retrieval, optical character recognition (OCR), video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, and also common biometric recognition technologies such as face recognition and fingerprint recognition. Machine learning (ML) is a multidisciplinary interdisciplinary subject that encompasses probability theory, statistics, approximation theory, convex analysis, and algorithmic complexity theory. Machine learning is the study of how computers can simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is at the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications span all areas of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and self-learning.
[0033] With the research and advancement of artificial intelligence technology, artificial intelligence technology has been studied and applied in many fields, such as common smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, unmanned driving, autonomous driving, drones, robots, smart medical care, smart customer service, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.
[0034] The solutions provided in the embodiments of this application involve artificial intelligence computer vision technology and machine learning technology, which are specifically illustrated by the following embodiments:
[0035] A computer device obtains a first picture set and a second picture set, wherein the style of the pictures included in the first picture set is different from the style of the pictures included in the second picture set, and the quality of the pictures included in the first picture set is higher than the quality of the pictures included in the second picture set; the computer device uses computer vision technology to obtain k pictures similar to each picture in the first picture set from the second picture set, and determines k groups of training samples based on the first picture set and the k pictures similar to each picture, each group of training samples includes the first picture set and one of the k pictures similar to each picture, k is an integer greater than or equal to 1, and determining the training samples by this method can achieve the marginalization of the main semantic information of the picture; using a neural network in machine learning to train the k groups of training samples to obtain k classifiers, and generating a style evaluation model based on the trained k classifiers, the style evaluation model is used to score the style quality of the picture. Through the above embodiment, the accuracy of the style quality scoring of the picture can be improved, and further based on the style quality score of the picture, the effective features of the picture can be effectively obtained, which is also helpful for judging and establishing the tonality of the picture community.
[0036] See Figure 1a , Figure 1a This is a schematic diagram of the architecture of an image processing system disclosed in an embodiment of the present application. Figure 1a As shown, the image processing system architecture 100 may include a client 101, a computer device 102, and a storage platform 103. The client 101, computer device 102, and storage platform 103 may be communicatively connected. The client 101 is primarily used to send images to be predicted to the computer device 102 and to send various images to the storage platform 103. The computer device 102 is primarily used to train the classifier using training samples and to assign style quality scores to the images to be predicted. The storage platform 103 is primarily used to store images from various data sources.
[0037] In one possible implementation, the computer device 102 obtains a first picture set and a second picture set from the storage platform 103, wherein the style of the pictures included in the first picture set is different from the style of the pictures included in the second picture set, and the quality of the pictures included in the first picture set is higher than the quality of the pictures included in the second picture set; the computer device 102 obtains k pictures similar to each picture in the first picture set from the second picture set, and determines k groups of training samples based on the first picture set and the k pictures similar to each picture, each group of training samples includes the first picture set and one of the k pictures similar to each picture, k is an integer greater than or equal to 1; the computer device 102 uses the k groups of training samples to train k classifiers, and generates a style evaluation model based on the trained k classifiers, and the style evaluation model is used to score the style quality of the picture.
[0038] In a possible implementation, the computer device 102 receives the image to be detected sent by the client 101, and the computer device 102 pre-processes the image to be predicted to obtain the pre-processed image to be predicted, and inputs the pre-processed image to be predicted into the style evaluation model to obtain the style quality score of the image to be predicted, and determines the style classification of the image to be predicted according to the style quality score, and further stores the image to be predicted in the storage platform 103. Furthermore, in a possible implementation, the product expression corresponding to the image processing method can be as follows: Figure 1b As shown, the user uploads the image to be detected to the corresponding terminal interface through the client interface. After the computer device detects the image to be predicted, it performs the above processing on the image to be predicted and obtains a style quality score, such as Figure 1b , the style quality score of the picture on the left is 0.03087026, and the style quality score of the picture on the right is 0.9428519. At the same time, the computer device can also judge the image category based on the style quality score and display the result on Figure 1b In the interface shown ( Figure 1b (not shown in the figure). In some possible scenarios, the community tone can be established based on the style quality score. Among them, the higher the style quality score, the higher the possibility that the image belongs to a high-quality community image. Using this method, a quantitative scoring of the style quality of the image can be achieved, which can help humans to more quickly screen out possible high-quality images. At the same time, in some possible scenarios, the above-mentioned image processing methods can be integrated into an interface to provide effective image features for services such as content recommendation.
[0039] Among them, the structure of the above classifier can be as follows Figure 1c As shown in the figure, the classifier includes a deep residual network, a global average pooling layer, a fully connected layer, and a Sigmoid activation function (the Sigmoid activation function is also in the fully connected layer). After the image is processed by the deep residual network, the global average pooling layer, the fully connected layer, and the Sigmoid activation function, a high-quality probability value and a low-quality probability value are obtained for the image. The structure of the classifier is explained in detail as follows:
[0040] The Deep Residual Network (ResNet) converts ordinary deep convolutional neural networks into corresponding residual versions by inserting short-circuit connections. Deep convolutional neural networks have achieved outstanding results in image classification. However, deep convolutional neural networks face the degradation problem, that is, the deeper the network, the saturation of accuracy or even decreases. Since identity mapping can ensure that the error of deep networks is not higher than that of shallow networks, residual learning can be used to solve the degradation problem. Figure 1d As shown in the figure, instead of letting the stacked layers learn the focused mapping H(x), it is better to let it learn the residual mapping F(x) = H(x) - x, so that the original focused mapping can be expressed as F(x) + x. Among them, F(x) + x can be implemented by a feedforward neural network with a "shortcut connection". Figure 1d The identity mapping in is a short-circuit connection.
[0041] The global average pooling layer (GAP) averages each feature map. In this embodiment, the pooling layer is used to average multiple image features of each image to obtain a 2048-dimensional vector, such as Figure 1e As shown, Figure 1e It intuitively shows how GAP is implemented.
[0042] The fully connected layer (FC layer) maps the input x to z = activation(Wx + b), where W and b are parameters and activation is the activation function. In this embodiment, the fully connected layer performs linear dimensionality reduction on the 2048-dimensional image vector, converting the 2048-dimensional vector into a 1024-dimensional vector.
[0043] The Sigmoid activation function maps any real number x to a real number between 0 and 1.
[0044]
[0045] Correspondingly, the loss function is Binary Cross Entropy Loss. The output of the Sigmoid activation function is between 0 and 1. Generally, in binary classification tasks, the output of the sigmoid function is the event probability, that is, when the output meets a certain probability condition, it is determined to be a positive class.
[0046] Explain the client 101. The "client" used herein includes but is not limited to user equipment, handheld devices with wireless communication functions, vehicle-mounted devices, wearable devices or computing devices. For example, the user terminal can be a mobile phone, a tablet computer or a computer with wireless transceiver functions. The client can also be a virtual reality (VR) terminal device, an augmented reality (AR) terminal device, a wireless terminal device in industrial control, a wireless terminal device in unmanned driving, a wireless terminal device in telemedicine, a wireless terminal device in a smart grid, a wireless terminal device in a smart city, a wireless terminal device in a smart home, etc. In the embodiment of the present application, the device for realizing the function of the client can be a terminal; it can also be a device that can support the terminal device to realize the function, such as a chip system, which can be installed in the terminal device. In the technical solution provided in the embodiment of the present application, the technical solution provided in the embodiment of the present application is described by taking the device for realizing the function of the client as an example.
[0047] The computer device 102 is explained. The computer device 102 can specifically be a server. The server here can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. The embodiments of this application are not limited here. In the technical solutions provided in the embodiments of this application, the computer device is used as an example to describe the technical solutions provided in the embodiments of this application.
[0048] See Figure 2 , Figure 2 This is a flowchart of an image processing method disclosed in an embodiment of the present application, which may mainly include the following steps:
[0049] S201. A computer device obtains a first picture set and a second picture set. The style of pictures included in the first picture set is different from the style of pictures included in the second picture set, and the quality of pictures included in the first picture set is higher than the quality of pictures included in the second picture set.
[0050] In one possible implementation, a computer device obtains a first picture set and a second picture set from different data sources, respectively. The data source may be a high / low quality picture content community, and the low quality picture content community is relative to the high quality picture content community. In an embodiment of the present application, it is exemplified that the quality of the pictures contained in the first picture set is higher than the quality of the pictures contained in the second picture set, and at the same time, the style of the pictures contained in the first picture set is different from the style of the pictures contained in the second picture set. Different styles can specifically refer to different pictures of the same theme. For example, if the target theme of the picture taken is railroad tracks, different styles refer to different exposure, color, saturation, etc. of pictures corresponding to different pictures of railroad tracks.
[0051] It should be noted that the relationship between the first picture set and the second picture set is only exemplarily indicated. In some possible implementations, the quality of the pictures included in the second picture set may be higher than the quality of the pictures included in the first picture set.
[0052] S202. The computer device obtains k pictures similar to each picture in the first picture set from the second picture set, and determines k groups of training samples based on the first picture set and the k pictures similar to each picture, each group of training samples includes the first picture set and one of the k pictures similar to each picture, where k is an integer greater than or equal to 1.
[0053] In one possible implementation, each image in the first image set is set to be of relatively high quality. Based on this, the computer device obtains k images similar to each image in the first image set from the second image set. Based on the first image set and the k images similar to each image, k groups of training samples are determined. Each group of training samples includes the first image set and one of the k images similar to each image. In each two groups of training samples, one of the k images similar to each image is not the same image.
[0054] For example, first determine a picture a in the first picture set, and then determine k pictures similar to picture a from the second picture set, where similarity means that the subject in picture a is the same as or similar to the subject in the k pictures.
[0055] Furthermore, the computer device obtains k pictures from the second picture set that are similar to each picture in the first picture set. Specifically, the computer device obtains the similarity between each picture in the second picture set and picture a based on the subject of each picture in the first picture set, such as picture a. The subject of the picture refers to the subject of the picture. The computer device can extract picture features from picture a and each picture in the second picture set based on the graph matching system, and then determine the k pictures with the highest similarity to picture a from the second picture set based on the picture features. The k pictures with the highest similarity are used as the k pictures similar to picture a. For each picture in the first picture set, the computer device selects one picture from the k pictures similar to each picture, and uses the first picture set and the one picture selected for each picture as a set of training samples. Therefore, the number of selection operations corresponds to the number of pictures in the first picture set. For example, if there are n pictures in the first picture set, we need to select n pictures from the second picture set accordingly to form a set of training samples. Since each picture in the first picture set has k corresponding similar pictures, we can determine k groups of training samples, and each group of training samples includes 2*n pictures.
[0056] For example, if the first image set includes images a, b, and c, then correspondingly, based on similarity, k images (a1, a2, ..., ak) similar to image a, k images (b1, b2, ..., bk) similar to image b, and k images (c1, c2, ..., ck) similar to image c are obtained from the second image set. Then, based on the first image set and the k images similar to each image, k groups of training samples are determined. The corresponding first group of training samples is (a, a1, b, b1, c, c1), the second group of training samples is (a, a2, b, b2, c, c2), and the kth group of training samples is (a, ak, b, bk, c, ck). In each training sample, a1, a2, and ak are different images, b1, b2, and bk are different images, and c1, c2, and ck are also different images. That is, different training samples include different images selected from the k similar images.
[0057] It should be noted that k is an integer greater than or equal to 1. For each picture in the first picture set, the pictures obtained from the second picture set based on the similarity may be greater than k. This application only selects the k pictures with the highest similarity for example.
[0058] S203 , the computer device uses the k groups of training samples to train k classifiers, and generates a style evaluation model based on the trained k classifiers. The style evaluation model is used to score the style quality of the picture.
[0059] In one possible implementation, before the computer device uses k groups of training samples to train k classifiers, since the sizes of the images obtained from different data sources are different, the computer device also needs to preprocess each group of training samples in the k groups of training, and the preprocessing here includes at least one of size conversion and normalization. In an embodiment of the present application, the computer device performs size conversion on the images in each group of training samples in the k groups of training samples, specifically converting the size of each image to 224*224, and then normalizing the pixel values of each image to ensure the consistency of the input data and avoid introducing large errors in the training results. After the computer device preprocesses the k groups of training samples, the computer device further uses each group of preprocessed training samples to train the initialized classifier to obtain k classifiers. Among them, the initialized classifier can refer to that the parameters corresponding to the classifier are empty, or it can refer to that the parameters corresponding to the classifier are fixed.
[0060] In one possible implementation, the initialized classifier may include a residual learning network, a pooling layer, a fully connected layer, and a Sigmoid activation function. The computer device then trains the initialized classifier using each set of preprocessed training samples to obtain k classifiers, which may specifically include:
[0061] The computer device uses a residual learning network to extract the image features of each image in any group of training samples after preprocessing. The image features of each image may include multiple features, and then the pooling layer is used to average the multiple image features of each image. In an embodiment of the present application, each image is specifically processed using a pooling layer to obtain a 2048-dimensional vector, and then a fully connected layer is used to perform linear dimensionality reduction processing on the image of the 2048-dimensional vector, and the 2048-dimensional vector is changed to a 1024-dimensional vector. Finally, the Sigmoid activation function is used to obtain the probability of each image being predicted as a high-quality image. Correspondingly, the probability of each image being predicted as a low-quality image is obtained by subtracting the probability of being predicted as a high-quality image from 1. The computer device then adjusts the parameters of the initialized classifier according to the probability of each image being predicted as a high-quality image to obtain the corresponding trained classifier. For example, after training a classifier on a particular image, if the probability value obtained is not within the set threshold range, or if the training does not reach the iterative stopping condition, the computer device will further use the images in the training sample to adjust the parameters of the initialized classifier. When the training stopping condition is met or the probability value obtained is within the set threshold range, the corresponding trained classifier is obtained. Each set of training samples is trained in this order. Therefore, if there are k sets of training samples, k trained classifiers can be obtained.
[0062] In one possible implementation, a computer device uses k sets of training samples to train k classifiers, and then integrates the k classifiers to generate a style evaluation model. In this process, the computer device first obtains the weight coefficient corresponding to each of the k classifiers. The weight coefficient can be an equal weight coefficient or an unequal weight coefficient learned during the classifier training process. After obtaining the weight coefficient corresponding to each classifier, the computer device integrates the k classifiers with equal weight coefficients through a "majority voting" mechanism to obtain the style evaluation model, or the computer device averages the weights of the k classifiers with unequal weight coefficients, and then integrates the k classifiers according to the voting mechanism to obtain the style evaluation model. This style evaluation model is mainly used to score the style quality of an image.
[0063] In the implementation of this application, a computer device obtains a first picture set and a second picture set, wherein the style of the pictures included in the first picture set is different from the style of the pictures included in the second picture set, and the quality of the pictures included in the first picture set is higher than the quality of the pictures included in the second picture set; the computer device obtains k pictures similar to each picture in the first picture set from the second picture set, and determines k groups of training samples based on the first picture set and the k pictures similar to each picture, each group of training samples includes the first picture set and one of the k pictures similar to each picture, k is an integer greater than or equal to 1; k classifiers are trained using the k groups of training samples, and a style evaluation model is generated based on the k classifiers obtained by training, the style evaluation model is used to score the style quality of the picture. Through the above embodiment, the training samples are determined by this method, the marginalization of the main semantic information of the picture can be achieved, and the accuracy of the style quality score of the picture can be improved.
[0064] See Figure 3 , Figure 3 This is a flowchart of another image processing method disclosed in an embodiment of the present application. The flowchart mainly includes two parts: one part is the generation process of the style evaluation model (i.e., the training process), and the other part is the use process of the style evaluation model (i.e., the prediction process). The flowchart may include the following steps:
[0065] S301. A computer device obtains a first picture set and a second picture set. The style of pictures included in the first picture set is different from the style of pictures included in the second picture set, and the quality of pictures included in the first picture set is higher than the quality of pictures included in the second picture set.
[0066] S302. The computer device obtains k pictures similar to each picture in the first picture set from the second picture set, and determines k groups of training samples based on the first picture set and the k pictures similar to each picture, each group of training samples includes the first picture set and one of the k pictures similar to each picture, where k is an integer greater than or equal to 1.
[0067] S303: The computer device preprocesses each of the k groups of training samples, where the preprocessing includes at least one of a size conversion process and a normalization process.
[0068] S304: The computer device trains the initialized classifier using each group of preprocessed training samples to obtain k classifiers.
[0069] S305 , the computer device determines a weight coefficient of each of the k classifiers obtained through training, and integrates the k classifiers obtained through training according to the weight coefficient of each classifier to obtain a style evaluation model.
[0070] Among them, steps S301 to S305 have been Figure 2 The corresponding steps S201 to S203 have been described in detail, and will not be described one by one here.
[0071] S306: The computer device obtains the picture to be predicted and preprocesses the picture to be predicted to obtain a preprocessed picture to be predicted, where the preprocessing includes at least one of a size conversion process and a normalization process.
[0072] In a possible implementation, the computer device obtains a picture to be predicted, which may be a picture to be predicted used by R&D personnel for testing, or a picture to be predicted in an actual application scenario (such as Figure 1b In the application scenario shown in FIG. 1 , a user-uploaded image to be predicted is obtained. After obtaining the image to be predicted, the computer device preprocesses the image to be predicted to obtain a preprocessed image to be predicted. The preprocessing herein includes at least one of resizing and normalization. As can be seen from the foregoing, the image to be predicted is resized to 224*224, and then the pixel values of the image to be predicted are normalized to obtain the preprocessed image to be predicted.
[0073] S307 , the computer device inputs the pre-processed image to be predicted into the style evaluation model to obtain a style quality score of the image to be predicted, and determines the style classification of the image to be predicted according to the style quality score.
[0074] In one possible implementation, the pre-processed picture to be predicted is input into the style evaluation model. After the computer device receives the pre-processed picture to be predicted, it predicts the pre-processed picture to be predicted according to the same steps of the aforementioned training process to obtain the style quality score of the picture to be predicted. The style quality score here includes a low-quality style quality score and a high-quality style quality score. Finally, the computer device determines the style classification of the picture to be predicted based on one of the style quality scores, such as making a classification judgment based on the high-quality style quality score. Style classification refers to whether it is a high-quality type or a low-quality type. After the style classification is determined, the picture to be predicted is sent to the corresponding database for storage (it can be sent to Figure 1a The storage platform 103 in the image is used for storage) to facilitate subsequent calls to the image by various applications, for example, using the image to determine and establish the community tone of the image, or to implement content recommendation services, etc.
[0075] The computer device determines the style classification of the image to be predicted based on one of the style quality scores, and can determine the style quality score based on a specified threshold. If the style quality score is greater than or equal to the specified threshold, the image to be predicted is determined to be a high-quality image; if the style quality score is less than the specified threshold, the image to be predicted is determined to be a high-quality image. Figure 3 The entire process can be used Figure 4 To elaborate, Figure 4 It can include four parts, among which the first part is to determine k groups of training samples based on the first picture set and the second picture set, the second part is to train k classifiers, the third part is to integrate the k classifiers to obtain a style evaluation model, and the fourth part is to use the style evaluation model to predict the score of the predicted picture. Figure 4 The fourth part shows the style quality scores of each image obtained using the style evaluation model, including high-quality and low-quality scores. It can be clearly seen that according to the order from top to bottom in the fourth part, the first image obtained a high-quality score of 0.92 and a low-quality score of 0.08, so the image can be determined as a high-quality style image. The corresponding high-quality scores of the second, third, and fourth images are 0.03, 0.05, and 0.09, respectively, and the low-quality scores are 0.97, 0.95, and 0.91, respectively, so the second, third, and fourth images can be determined as low-quality style images.
[0076] In addition to being able to realize Figure 2The functions described can also highlight the style quality information of the image to achieve the classification purpose, and eliminate the impact of category bias on the image style classifier as much as possible; at the same time, the style evaluation model is used to score the images to be tested, and the scored images are applied to actual scenarios, which helps to judge and establish the tone of the image community and provide effective image features for other services such as content recommendation.
[0077] See Figure 5 , Figure 5 This is a schematic diagram of the structure of an image processing device disclosed in an embodiment of the present application. The image processing device 50 may include: an acquisition unit 501, a determination unit 502, and a processing unit 503, which are mainly used to:
[0078] An acquiring unit 501 is configured to acquire a first picture set and a second picture set, wherein the style of pictures included in the first picture set is different from the style of pictures included in the second picture set, and the quality of pictures included in the first picture set is higher than the quality of pictures included in the second picture set;
[0079] The obtaining unit 501 is further configured to obtain, from the second picture set, k pictures that are similar to each picture in the first picture set;
[0080] a determining unit 502, configured to determine k groups of training samples based on the first picture set and k pictures similar to each of the pictures, each group of training samples including the first picture set and one of the k pictures similar to each of the pictures, where k is an integer greater than or equal to 1;
[0081] A processing unit 503 is configured to train k classifiers using the k groups of training samples;
[0082] The determining unit 502 is further configured to generate a style evaluation model based on the k classifiers obtained through training, and the style evaluation model is used to score the style quality of the picture.
[0083] In a possible implementation, the processing unit 503 uses the k groups of training samples to train k classifiers, which are used to:
[0084] performing preprocessing on each of the k groups of training samples, wherein the preprocessing includes at least one of a size conversion process and a normalization process;
[0085] Each set of preprocessed training samples is used to train the initialized classifier to obtain k classifiers.
[0086] In one possible implementation, the initialized classifier includes a residual learning network, a pooling layer, and a fully connected layer. The processing unit 503 trains the initialized classifier using each set of preprocessed training samples to obtain k classifiers for:
[0087] For each image in any set of preprocessed training samples, extracting image features of each image using the residual learning network;
[0088] Processing the image features using the pooling layer and the fully connected layer to obtain a probability that each image is predicted to be a high-quality image;
[0089] The parameters of the initialized classifier are adjusted according to the probability that each picture is predicted to be a high-quality picture to obtain a corresponding trained classifier.
[0090] In a possible implementation, the determining unit 502 generates a style evaluation model based on the k classifiers obtained through training, for:
[0091] Determine a weight coefficient for each of the k classifiers obtained through training;
[0092] The k classifiers obtained through training are integrated according to the weight coefficient of each classifier to obtain a style evaluation model.
[0093] In a possible implementation, the acquiring unit 501 acquires, from the second picture set, k pictures that are similar to each picture in the first picture set, for:
[0094] For each picture in the first picture set, obtaining, based on the photographed object, a similarity between each picture included in the second picture set and the picture;
[0095] The corresponding k pictures with the highest similarity are obtained from the second picture set, and the k pictures with the highest similarity are used as the k pictures similar to each of the pictures.
[0096] In a possible implementation, the determining unit 502 determines k groups of training samples according to the first picture set and k pictures similar to each picture, to:
[0097] For each picture in the first picture set, select a picture from k pictures similar to the picture;
[0098] The first picture set and one picture selected for each picture are used as a group of training samples.
[0099] In a possible implementation, the acquisition unit 501 is further configured to acquire a picture to be predicted;
[0100] The processing unit 503 is further configured to:
[0101] Preprocessing the image to be predicted to obtain a preprocessed image to be predicted, wherein the preprocessing includes at least one of a resizing process and a normalization process;
[0102] Inputting the preprocessed image to be predicted into the style evaluation model to obtain a style quality score of the image to be predicted;
[0103] The determining unit 502 is further configured to determine the style classification of the to-be-predicted picture according to the style quality score.
[0104] In an embodiment of the present application, the acquisition unit 501 acquires a first picture set and a second picture set, wherein the style of the pictures included in the first picture set is different from the style of the pictures included in the second picture set, and the quality of the pictures included in the first picture set is higher than the quality of the pictures included in the second picture set, and k pictures similar to each picture in the first picture set are obtained from the second picture set, and the determination unit 502 determines k groups of training samples based on the first picture set and the k pictures similar to each picture, each group of training samples includes the first picture set and one of the k pictures similar to each picture, and k is an integer greater than or equal to 1; the processing unit 503 uses the k groups of training samples to train k classifiers, and generates a style evaluation model based on the trained k classifiers, the style evaluation model is used to score the style quality of the picture, and the training samples are determined by the above method, so as to achieve the marginalization of the main semantic information of the picture and improve the accuracy of the style quality score of the picture.
[0105] See Figure 6 , Figure 6 : is a structural diagram of a computer device disclosed in an embodiment of the present application, wherein the computer device 60 includes at least a processor 601, a memory 602 and a communication device 603. The processor 601, the memory 602 and the communication device 603 may be connected via a bus or other means. The communication device 603 is used to send and receive data. The memory 602 may include a computer-readable storage medium, the memory 602 is used to store a computer program, the computer program includes computer instructions, and the processor 601 is used to execute the computer instructions stored in the memory 602. The processor 601 (or CPU (Central Processing Unit)) is the computing core and control core of the computer device 60, which is suitable for implementing one or more computer instructions, and is specifically suitable for loading and executing one or more computer instructions to implement the corresponding method flow or corresponding function.
[0106] The embodiment of the present application also discloses a computer-readable storage medium (Memory), which is a memory device in the computer device 60 for storing programs and data. It can be understood that the memory 602 here can include both the built-in storage medium in the computer device 60 and, of course, the extended storage medium supported by the computer device 60. The computer-readable storage medium provides a storage space, which stores the operating system of the computer device 60. In addition, one or more computer instructions suitable for being loaded and executed by the processor 601 are also stored in the storage space. These computer instructions can be one or more computer programs (including program codes). It should be noted that the memory 602 here can be a high-speed RAM memory, or a non-volatile memory (Non-Volatile Memory), such as at least one disk memory; optionally, it can also be at least one computer-readable storage medium located away from the aforementioned processor 601.
[0107] In one implementation, the computer device 60 may be Figure 1a The computer device 102 in the picture processing system shown in FIG. 1 ; the memory 602 stores a first computer instruction; the processor 601 loads and executes the first computer instruction stored in the memory 602 to implement Figure 2 、 Figure 3 Corresponding steps in the method embodiment shown; in a specific implementation, the first computer instruction in the memory 602 is loaded by the processor 601 and executes the following steps:
[0108] Acquire a first picture set and a second picture set, where the style of pictures included in the first picture set is different from the style of pictures included in the second picture set, and the quality of pictures included in the first picture set is higher than the quality of pictures included in the second picture set;
[0109] Obtaining, from the second picture set, k pictures that are similar to each picture in the first picture set, and determining k groups of training samples based on the first picture set and the k pictures similar to each picture, each group of training samples including the first picture set and one of the k pictures similar to each picture, where k is an integer greater than or equal to 1;
[0110] K classifiers are trained using the k groups of training samples, and a style evaluation model is generated based on the k classifiers obtained through training. The style evaluation model is used to score the style quality of the picture.
[0111] In a possible implementation, the processor 601 uses the k groups of training samples to train k classifiers, which are used to:
[0112] performing preprocessing on each of the k groups of training samples, wherein the preprocessing includes at least one of a size conversion process and a normalization process;
[0113] Each set of preprocessed training samples is used to train the initialized classifier to obtain k classifiers.
[0114] In one possible implementation, the initialized classifier includes a residual learning network, a pooling layer, and a fully connected layer. The processor 601 trains the initialized classifier using each set of preprocessed training samples to obtain k classifiers for:
[0115] For each image in any set of preprocessed training samples, extracting image features of each image using the residual learning network;
[0116] Processing the image features using the pooling layer and the fully connected layer to obtain a probability that each image is predicted to be a high-quality image;
[0117] The parameters of the initialized classifier are adjusted according to the probability that each picture is predicted to be a high-quality picture to obtain a corresponding trained classifier.
[0118] In a possible implementation, the processor 601 generates a style evaluation model based on the k classifiers obtained through training, for:
[0119] Determine a weight coefficient for each of the k classifiers obtained through training;
[0120] The k classifiers obtained through training are integrated according to the weight coefficient of each classifier to obtain a style evaluation model.
[0121] In a possible implementation, the processor 601 obtains, from the second picture set, k pictures that are similar to each picture in the first picture set, for:
[0122] For each picture in the first picture set, obtaining, based on the photographed object, a similarity between each picture included in the second picture set and the picture;
[0123] The corresponding k pictures with the highest similarity are obtained from the second picture set, and the k pictures with the highest similarity are used as the k pictures similar to each of the pictures.
[0124] In a possible implementation, the processor 601 determines k groups of training samples according to the first picture set and k pictures similar to each of the pictures, and is configured to:
[0125] For each picture in the first picture set, select a picture from k pictures similar to the picture;
[0126] The first picture set and one picture selected for each picture are used as a group of training samples.
[0127] In a possible implementation, the processor 601 is further configured to:
[0128] Acquire a picture to be predicted, and preprocess the picture to be predicted to obtain a preprocessed picture to be predicted, wherein the preprocessing includes at least one of a resizing process and a normalization process;
[0129] Inputting the preprocessed image to be predicted into the style evaluation model to obtain a style quality score of the image to be predicted;
[0130] The style classification of the picture to be predicted is determined according to the style quality score.
[0131] In the implementation of this application, the processor 601 of the computer device obtains a first picture set and a second picture set, wherein the style of the pictures included in the first picture set is different from the style of the pictures included in the second picture set, and the quality of the pictures included in the first picture set is higher than the quality of the pictures included in the second picture set; k pictures similar to each picture in the first picture set are obtained from the second picture set, and k groups of training samples are determined based on the first picture set and the k pictures similar to each picture, each group of training samples includes the first picture set and one of the k pictures similar to each picture, k is an integer greater than or equal to 1; k classifiers are trained using the k groups of training samples, and a style evaluation model is generated based on the k classifiers obtained by training, the style evaluation model is used to score the style quality of the picture. Through the above embodiment, the training samples are determined by this method, the main semantic information of the picture can be marginalized, and the accuracy of the style quality score of the picture can be improved.
[0132] According to one aspect of the present application, a computer program product or computer program is also disclosed, the computer program product or computer program including computer instructions, the computer instructions being stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device can perform the aforementioned Figure 2 、 Figure 3 The method in the embodiment corresponding to the flowchart will not be described again here.
[0133] It should be noted that for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps may be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily required by this application.
[0134] In the several embodiments provided in this application, it should be understood that the disclosed devices can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules described above is merely a logical functional division. In actual implementation, other division methods may be used. For example, multiple modules or components may be combined or integrated into another system, or some features may be ignored or not implemented.
[0135] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
Claims
1. A method for image processing, characterized in that: The method comprises: Acquire a first picture set and a second picture set, where the style of pictures included in the first picture set is different from the style of pictures included in the second picture set, and the quality of pictures included in the first picture set is higher than the quality of pictures included in the second picture set; Obtaining, from the second picture set, k pictures that are similar to each picture in the first picture set, and determining k groups of training samples based on the first picture set and the k pictures similar to each picture, each group of training samples including the first picture set and one of the k pictures similar to each picture, where k is an integer greater than or equal to 1; Using the k groups of training samples to train k classifiers, and generating a style evaluation model based on the k classifiers obtained through training, wherein the style evaluation model is used to score the style quality of the picture; The step of training k classifiers using the k groups of training samples includes: Preprocessing each of the k groups of training samples, wherein the preprocessing includes at least one of a resizing process and a normalization process; wherein the initialized classifier includes a residual learning network, a pooling layer, and a fully connected layer; For each image in any set of preprocessed training samples, extracting image features of each image using the residual learning network; Processing the image features using the pooling layer and the fully connected layer to obtain a probability that each image is predicted to be a high-quality image; The parameters of the initialized classifier are adjusted according to the probability that each picture is predicted to be a high-quality picture to obtain a corresponding trained classifier.
2. The method according to claim 1, characterized in that Generating a style evaluation model based on the k classifiers obtained through training includes: Determine a weight coefficient for each of the k classifiers obtained through training; The k classifiers obtained through training are integrated according to the weight coefficient of each classifier to obtain a style evaluation model.
3. The method according to claim 1, characterized in that The acquiring, from the second picture set, k pictures that are similar to each picture in the first picture set includes: For each picture in the first picture set, obtaining, based on the photographed object, a similarity between each picture included in the second picture set and the picture; The corresponding k pictures with the highest similarity are obtained from the second picture set, and the k pictures with the highest similarity are used as the k pictures similar to each of the pictures.
4. The method according to claim 1, wherein The determining k groups of training samples according to the first picture set and k pictures similar to each of the pictures includes: For each picture in the first picture set, select a picture from k pictures similar to the picture; The first picture set and one picture selected for each picture are used as a group of training samples.
5. The method according to claim 1, wherein The method further comprises: Acquire a picture to be predicted, and preprocess the picture to be predicted to obtain a preprocessed picture to be predicted, wherein the preprocessing includes at least one of a resizing process and a normalization process; Inputting the preprocessed image to be predicted into the style evaluation model to obtain a style quality score of the image to be predicted; The style classification of the picture to be predicted is determined according to the style quality score.
6. A picture processing device, characterized in that: The device comprises: an acquiring unit, configured to acquire a first picture set and a second picture set, wherein the style of pictures included in the first picture set is different from the style of pictures included in the second picture set, and the quality of pictures included in the first picture set is higher than the quality of pictures included in the second picture set; The acquisition unit is further configured to acquire, from the second picture set, k pictures that are similar to each picture in the first picture set; a determining unit, configured to determine k groups of training samples based on the first picture set and k pictures similar to each of the pictures, each group of training samples including the first picture set and one of the k pictures similar to each of the pictures, where k is an integer greater than or equal to 1; A processing unit, configured to train k classifiers using the k groups of training samples; The determining unit is further configured to generate a style evaluation model based on the k classifiers obtained through training, wherein the style evaluation model is used to score the style quality of the picture; The processing unit is specifically configured to: Preprocessing each of the k groups of training samples, wherein the preprocessing includes at least one of a resizing process and a normalization process; wherein the initialized classifier includes a residual learning network, a pooling layer, and a fully connected layer; For each image in any set of preprocessed training samples, extracting image features of each image using the residual learning network; Processing the image features using the pooling layer and the fully connected layer to obtain a probability that each image is predicted to be a high-quality image; The parameters of the initialized classifier are adjusted according to the probability that each picture is predicted to be a high-quality picture to obtain a corresponding trained classifier.
7. A computer device, characterized in that: The computer device comprises: memory for storing computer programs; A processor, configured to run the computer program and implement the image processing method according to any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein the computer program includes program instructions, and when the program instructions are executed by a processor, the processor executes the image processing method according to any one of claims 1 to 5.
9. A computer program product, characterized in that The computer program product includes a computer program or computer instructions, and when the computer program or computer instructions are executed by a processor, they are used to implement the image processing method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Image quality evaluation method and device, electronic equipment and computer readable storage medium
CN110858394A