Image classification model acquisition method, device, electronic device and readable storage medium

By constructing the main training task and auxiliary training task for the image classification model, and combining the probability of identifying sensitive image categories to which the image belongs, the problem of poor accuracy in identifying sensitive images in the prior art is solved, and the recognition accuracy and stability of the model are improved.

CN112487226BActive Publication Date: 2025-05-23BEIJING SANKUAI ONLINE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011232795.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-11-06
Publication Date
2025-05-23
Estimated Expiration
2040-11-06

AI Technical Summary

Technical Problem

In the prior art, the image classification model used to identify sensitive pictures has misjudgment problems when classifying, and the recognition accuracy is poor.

Method used

By constructing the main training task and auxiliary training task for the image classification model to be trained, the main training task combines the auxiliary information obtained in the auxiliary training task to identify the probability that the image belongs to each sensitive image category; the auxiliary training task includes identifying the category to which the image belongs in the network platform.

Benefits of technology

It improves the identification accuracy of the image classification model, reduces the probability of misjudgment, ensures that the model takes into account the presentation method between sensitive pictures under different categories during classification, and classifies them in combination with the categories to which the pictures belong.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112487226B_ABST
    Figure CN112487226B_ABST
Patent Text Reader

Abstract

A method, device, electronic device and readable storage medium for obtaining a picture classification model. The method can construct a main training task and at least one auxiliary training task for a picture classification model to be trained; wherein the main training task is to identify the probability that a picture belongs to each sensitive picture category in combination with the auxiliary information obtained in the auxiliary training task; the auxiliary training task includes identifying the category to which the picture belongs on the network platform, and training the picture classification model to be trained according to the main training task and the auxiliary training task to obtain a target picture classification model. In this way, the picture classification model to be trained can learn how to identify the category to which the picture belongs, and the ability to identify the probability that the picture belongs to each sensitive picture category in combination with the category to which the picture belongs during the training process. Furthermore, to a certain extent, the influence of the presentation method of sensitive pictures under different categories on the classification during model classification can be eliminated, the probability of misjudgment can be reduced, and the accuracy can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of computer technology, and in particular, relates to a method, device, electronic device and readable storage medium for obtaining an image classification model. Background Art

[0002] With the continuous development of network technology, various network platforms have emerged on the Internet. In order to facilitate users to make intuitive choices on the network platform, the platform will display merchants or items in the form of pictures. In order to create a clean and good network platform environment, when displaying pictures, it is often necessary to identify the pictures to avoid showing sensitive pictures to users. And a network platform often contains multiple categories. For example, the network platform can include the categories of "takeout", "hotel accommodation", "leisure and entertainment", etc., and a category often contains a large number of pictures of network items. Therefore, how to efficiently identify sensitive pictures has become a problem of widespread concern.

[0003] In the prior art, sample images labeled as sensitive images are often used directly to train image classification models that can determine whether an image is a sensitive image or a non-sensitive image, so as to identify sensitive images through the image classification model. However, when the image classification model trained by this training method is used for classification, sometimes misjudgment occurs, and the recognition accuracy is poor. Summary of the invention

[0004] The present application provides a method, device, electronic device and readable storage medium for obtaining an image classification model, so as to improve the recognition accuracy of the model.

[0005] According to a first aspect of the present application, a method for obtaining an image classification model is provided, the method comprising:

[0006] Constructing a main training task and at least one auxiliary training task for the image classification model to be trained; wherein the main training task is to identify the probability that an image belongs to each sensitive image category in combination with the auxiliary information obtained in the auxiliary training task; the auxiliary training task includes identifying the category to which the image belongs on the network platform;

[0007] According to the main training task and the auxiliary training task, the image classification model to be trained is trained to obtain a target image classification model.

[0008] According to a second aspect of the present application, a device for acquiring a picture classification model is provided, the device comprising:

[0009] A construction module, used to construct a main training task and at least one auxiliary training task for the image classification model to be trained; wherein the main training task is to identify the probability that an image belongs to each sensitive image category in combination with the auxiliary information obtained in the auxiliary training task; and the auxiliary training task includes identifying the category to which the image belongs on the network platform;

[0010] The training module is used to train the image classification model to be trained according to the main training task and the auxiliary training task to obtain a target image classification model.

[0011] According to a third aspect of the present application, an electronic device is provided, including:

[0012] A processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the processor implements the aforementioned image classification model acquisition method when executing the program.

[0013] According to a fourth aspect of the present application, a readable storage medium is provided. When the instructions in the storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the aforementioned image classification model acquisition method.

[0014] The present application provides a method, device, electronic device and readable storage medium for obtaining a picture classification model, the method comprising: constructing a main training task and at least one auxiliary training task for the picture classification model to be trained; wherein the main training task is to identify the probability that the picture belongs to each sensitive picture category in combination with the auxiliary information obtained in the auxiliary training task; the auxiliary training task includes identifying the category to which the picture belongs in the network platform, and training the picture classification model to be trained according to the main training task and the auxiliary training task to obtain a target picture classification model. In this way, the picture classification model to be trained can learn how to identify the category to which the picture belongs, and the ability to identify the probability of the picture belonging to each sensitive picture category in combination with the category to which the picture belongs during the training process. In addition, it can be ensured that when the trained target picture classification model is used for classification in the future, the target picture classification model can take into account the presentation methods of sensitive pictures under different categories during the classification process, and classify the pictures in combination with the categories to which the pictures belong, thereby eliminating the influence of the presentation methods of sensitive pictures under different categories on the classification to a certain extent, reducing the probability of misjudgment and improving accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0016] Figure 1 This is a flowchart of a method for obtaining an image classification model provided in an embodiment of the present application;

[0017] Figure 2 is a schematic diagram of an auxiliary training task provided in an embodiment of the present application;

[0018] Figure 3 It is a schematic diagram of a model provided in an embodiment of the present application;

[0019] Figure 4 It is a flowchart of the steps of a sensitive image recognition method provided by this application;

[0020] Figure 5 A structural diagram of a device for acquiring a picture classification model in one embodiment of the present application is shown;

[0021] Figure 6 A structural diagram of an electronic device provided by an embodiment of the present application is shown. DETAILED DESCRIPTION

[0022] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0023] First, an application scenario involved in the present application is described. In order to facilitate users to make intuitive choices on the network platform, the platform will display merchants or items in the form of pictures. In order to create a clean and good network platform environment, the platform often needs to use an image classification model to identify sensitive images to avoid displaying sensitive images to users. In an existing implementation method, sample images marked as sensitive images can be used to train an image classification model that performs binary classification based on the images. The trained image classification model can determine whether the image is a sensitive image or a non-sensitive image. When the image classification model trained by this training method is classified, sometimes misjudgment problems will occur, and the recognition accuracy is poor.

[0024] To this end, an embodiment of the present application provides a method for obtaining a picture classification model, and the method for obtaining a picture classification model is described in detail below.

[0025] Figure 1 is a flowchart of a method for obtaining an image classification model provided in an embodiment of the present application. Figure 1 As shown, the method may include:

[0026] Step 101: construct a main training task and at least one auxiliary training task for the image classification model to be trained; wherein the main training task is to identify the probability that an image belongs to each sensitive image category in combination with the auxiliary information obtained in the auxiliary training task; and the auxiliary training task includes identifying the category to which the image belongs on the network platform.

[0027] In an embodiment of the present application, the main training task and at least one auxiliary training task may be a learning task of a given picture classification model to be trained. By giving the learning task of the picture classification model to be trained, the picture classification model to be trained can perform the learning task during the training process and learn the ability to accurately perform the learning task. The picture classification model to be trained may be selected according to actual needs. For example, the picture classification model to be trained may be a neural network model. Furthermore, the auxiliary information obtained in the auxiliary training task may be determined according to the specific content of the auxiliary training task. For example, in the case where the auxiliary training task is to identify the category to which the picture belongs on the network platform, the auxiliary information obtained in the auxiliary training task may be: the category to which the picture belongs on the network platform.

[0028] Further, the sensitive picture category can be used to characterize the type of picture to which the sensitive picture belongs, that is, to characterize which type of sensitive picture the sensitive picture is. The sensitive picture category can be pre-divided according to the actual situation. In actual application scenarios, the presentation methods of sensitive pictures under different categories are often quite different. Therefore, in the embodiment of the present application, an auxiliary training task of "identifying the category to which the picture belongs in the network platform" can be set, so that the picture classification model to be trained can learn the ability to identify the category to which the picture belongs in the network platform. Further, in the scenario of using the picture classification model for sensitive picture identification, by specifically identifying the probability that the picture belongs to each sensitive picture category, the probability can be used to more accurately characterize whether the picture is a sensitive picture and which category of sensitive picture it specifically belongs to. Accordingly, "combining the auxiliary information obtained in the auxiliary training task to identify the probability that the picture belongs to each sensitive picture category" is used as the main training task, so that the picture classification model to be trained can learn how to combine the category information of the picture to more accurately identify the probability that the picture belongs to each sensitive picture category in the subsequent training process. In this way, when the trained target image classification model is used for classification in the future, the target image classification model can take into account the presentation methods of sensitive images under different categories during the classification process, and classify the images according to the categories to which they belong, thereby eliminating the impact of the presentation methods of sensitive images under different categories on the classification to a certain extent, reducing the probability of misjudgment and improving the accuracy of classification.

[0029] At the same time, compared with the method of roughly and broadly dividing the images into two categories, namely, sensitive images and non-sensitive images, and training the model to identify whether the images belong to sensitive images or non-sensitive images, in the embodiment of the present application, the sensitive images are divided into sensitive image categories in a more fine-grained manner, and the training model is trained to identify the probability that the images specifically belong to the sensitive image category. In this way, to a certain extent, the model can fully learn the features of images belonging to different sensitive image categories during the training process, improve the recognition accuracy of the model, and further improve the accuracy of the subsequent use of the model for sensitive image recognition.

[0030] Step 102: training the to-be-trained picture classification model according to the main training task and the auxiliary training task to obtain a target picture classification model.

[0031] In the embodiment of the present application, the image classification model to be trained is trained by combining the main training task and the auxiliary training task. During the training process, the image classification model to be trained can learn how to identify the category to which the image belongs, and determine the probability that the image belongs to each sensitive image category in combination with the category to which the image belongs. The ability of the learning model is continuously improved. In this way, it can be ensured that the target image classification model finally obtained can more accurately identify the category to which the image belongs, and determine the probability that the image belongs to each sensitive image category in combination with the category to which the image belongs. Accordingly, after the training is completed, the trained image classification model to be trained can be determined as the target image classification model.

[0032] In summary, the method for obtaining a picture classification model provided in the embodiment of the present application is to construct a main training task and at least one auxiliary training task for the picture classification model to be trained; wherein the main training task is to identify the probability that the picture belongs to each sensitive picture category in combination with the auxiliary information obtained in the auxiliary training task; the auxiliary training task includes identifying the category to which the picture belongs in the network platform, and training the picture classification model to be trained according to the main training task and the auxiliary training task to obtain the target picture classification model. In this way, the picture classification model to be trained can learn how to identify the category to which the picture belongs, and the ability to identify the probability of the picture belonging to each sensitive picture category in combination with the category to which the picture belongs during the training process. In addition, it can be ensured that when the trained target picture classification model is used for classification in the future, the target picture classification model can take into account the presentation methods of sensitive pictures under different categories during the classification process, and classify them in combination with the categories to which the pictures belong, thereby eliminating the influence of the presentation methods of sensitive pictures under different categories on the classification to a certain extent, reducing the probability of misjudgment and improving accuracy.

[0033] Optionally, the auxiliary training tasks in the embodiments of the present application may also include at least one of the following: identifying the rotation angle of the image, perceiving the color information contained in the image, determining the aesthetic score of the image, and determining the probability of containing a sensitive human body in the image. In this way, by additionally setting these auxiliary training tasks, the image classification model to be trained can learn how to identify the rotation angle of the image, how to perceive the color information contained in the image, how to determine the aesthetic score of the image, how to determine the probability of containing a sensitive human body in the image, and how to further combine the identified rotation angle, color information, aesthetic score, and the probability of containing a sensitive human body in the image to identify the probability that the image belongs to each sensitive image category. In this way, when the trained target image classification model classifies and identifies the image, it can make more full use of the information contained in the image itself, such as category, rotation angle, color information, aesthetic score, and the probability of containing a sensitive human body in the image, thereby improving the accuracy of recognition to a certain extent.

[0034] Optionally, in the embodiment of the present application, the image classification model to be trained may include a shared layer and a network layer corresponding to each training task. Accordingly, according to the main training task and the auxiliary training task, training the image classification model to be trained may include the following steps:

[0035] Step 1021: According to the first training sample corresponding to the main training task and the second training sample corresponding to the auxiliary training task, obtain a first loss value of the to-be-trained image classification model corresponding to the main training task and a second loss value corresponding to the auxiliary training task.

[0036] In an embodiment of the present application, the sample input in the first training sample can be input into the image classification model to be trained, and then the first loss value is determined according to the output given by the image classification model to be trained and the sample output in the first training sample. The sample input in the second training sample is input into the image classification model to be trained, and then the second loss value is determined according to the output given by the image classification model to be trained and the sample output in the second training sample. Specifically, the output given and the sample output of the image classification model to be trained can be used as the input of a preset loss function, and the output of the loss function can be used as the loss value. Among them, the first loss value and the second loss value can respectively reflect the accuracy of the image classification model to be trained in currently executing the main training task and the accuracy of executing the auxiliary training task. The loss function can include a loss function corresponding to each training task. In this way, by setting a corresponding loss function for each training task, the loss value of the image classification model to be trained corresponding to each training task can be more accurately determined according to the corresponding loss function.

[0037] Step 1022: Adjust the model parameters in the shared layer according to the first loss value and the second loss value.

[0038] In the embodiment of the present application, the shared layer can also be referred to as a shallow layer in the neural network model. The shared layer can be a layer that is used when executing each training task, for example, a feature extraction layer, an output layer, and so on. Specifically, the stochastic gradient descent method can be used to adjust the model parameters in the shared layer according to the first loss value and the second loss value. In the embodiment of the present application, the model parameters in the shared layer are adjusted by combining the first loss value of the main training task corresponding to the image classification model to be trained and the second loss value corresponding to the auxiliary training task. The shared layer is optimized with multiple different training tasks, which can promote each other among multiple training tasks, thereby improving the adjustment effect of the model parameters to a certain extent and improving the model optimization effect.

[0039] Step 1023: According to the first loss value and the second loss value, adjust the model parameters in the network layer corresponding to the main training task and the model parameters in the network layer corresponding to the auxiliary training task respectively.

[0040] In an embodiment of the present application, the network layer corresponding to the main training task may be a layer used only when executing the main training task, and the network layer corresponding to the auxiliary training task may be a layer used only when executing the auxiliary training task. For example, the network layer corresponding to the main training task may be an image classification layer, and the network layer corresponding to the auxiliary training task: identifying the category to which the image belongs may be a category identification layer. By adjusting the network layers corresponding to each training task separately and in a targeted manner, the optimization effect of the network layer can be improved to a certain extent. Among them, the network layer corresponding to each training task can be regarded as an independent network for a training task, specifically an independent fully connected layer.

[0041] Optionally, the auxiliary training task in the embodiment of the present application can be implemented based on self-supervised learning technology. Among them, self-supervised learning can mine the supervisory information (high-level semantic information) of the data itself from large-scale unlabeled data. Self-supervised learning can use these semantic information in downstream processing tasks by mining the high-level semantic information contained in a large amount of unlabeled data. Through this constructed supervisory information (self-supervised label), the network is trained, and the model can learn valuable representations for downstream tasks. In one implementation, the auxiliary training task can be combined with the downstream task (main training task) in the form of a pre-training-fine-tuning model. Among them, the pre-training-fine-tuning model method refers to pre-building a network model, using the auxiliary training task for training, and saving the parameters in the network model at the end of the training, so that the trained network model can obtain better results when performing similar tasks next time. Then use the saved parameters as the initialization parameters of the main training task, for example, as the initialization parameters of the image classification task, and then make some modifications according to the results during the training process to obtain the final model. Since this pre-training-fine-tuning model method separates the auxiliary training task from the image classification task, the effect is often poor. In the embodiment of the present application, a multi-task learning method is used to simultaneously learn the image classification task and multiple self-supervised auxiliary training tasks, thereby improving the training effect and the classification performance of the trained model. Among them, multiple learning tasks are given in multi-task learning, all or part of which are related but not exactly the same. Multi-task learning can use the knowledge contained in these learning tasks to help improve the performance of the model in performing each learning task.

[0042] Optionally, in the embodiment of the present application, the first training sample corresponding to the main training task can be obtained by the following steps:

[0043] Step A: Obtain first sample images belonging to different sensitive image categories and label values ​​of the sensitive image categories to which they belong.

[0044] In an embodiment of the present application, the first sample image may be a sensitive image obtained from a network platform, and the label value of the sensitive image category to which it belongs may characterize which sensitive image category the first sample image specifically belongs to. The first sample image may be used as a sample input in a first training sample, and the label value of the sensitive image category to which it belongs may be used as a sample output in the first training sample to form a first training sample. Furthermore, the first sample image that does not belong to a sensitive image and its corresponding label value may be added to the first training sample, so that the model can learn the features of non-sensitive images, thereby improving the resolution ability of the model to a certain extent. By constructing the first training sample with first sample images belonging to different sensitive image categories, the embodiment of the present application can enable the model to fully learn the features of images belonging to different sensitive image categories, thereby improving the recognition accuracy of the trained model.

[0045] Optionally, the sensitive image category in the embodiment of the present application can be determined based on the historical sensitive images under each category contained in the network platform. Among them, the historical sensitive images can be sensitive images previously identified from the images contained in each category. The specific number of sensitive image categories can be determined by the types involved in the historical sensitive images. For example, it can be refined into seven subcategories. In the embodiment of the present application, by combining the historical sensitive images under each category to divide the sensitive image categories more finely, the characteristics of sensitive images under different categories can be learned during the model training process, and then the interference caused by the presentation differences of sensitive images under different categories in the subsequent recognition process can be reduced, and the recognition accuracy can be improved. At the same time, since the range of categories involved in the recognition method through binary classification is too large, it is often easy to cause insufficient model accuracy. Therefore, in the embodiment of the present application, by subdividing according to the category, the original classification task is converted from binary classification, for example, binary classification of sensitive images and non-sensitive images, to a fine-grained multi-classification task, and then the classification accuracy of the model can be improved to a certain extent.

[0046] Furthermore, the second training sample corresponding to the auxiliary training task can be obtained by the following steps:

[0047] Step B: If the auxiliary training task is to identify the category to which the image belongs on the network platform, obtain a second sample image belonging to a different category and the category to which it belongs.

[0048] For the auxiliary training task of identifying the category to which the image belongs in the network platform, the auxiliary training task can specifically be a category prediction auxiliary task. For example, an auxiliary task can be created first to automatically obtain a second sample image belonging to a different category from the network platform, and the category to which the second sample image belongs is used as a label. Accordingly, the second sample image can be used as the sample input in the second training sample, and the category to which the second sample image belongs can be used as the sample output in the first training sample to form a second training sample. Since the diversification of categories will cause the diversification of sensitive images, if the category to which the image belongs can be predicted, it will greatly improve the classification task. In the embodiment of the present application, the ability of the model to identify the category to which the image belongs can be further trained through the category prediction auxiliary task, thereby improving the accuracy of using the model for recognition to a certain extent.

[0049] Step C: If the auxiliary training task is to identify the rotation angle of the image, obtain third sample images belonging to different specified angles and the specified angles to which they belong.

[0050] For the auxiliary training task of identifying the rotation angle of the image, the auxiliary training task can be specifically an angle classification auxiliary task. Through the angle classification auxiliary task, the model can learn how to identify the rotation angle of the image, so as to avoid the recognition interference caused by the change of the image angle, and help the model to make correct judgments after the angle of the image is changed, thereby improving the accuracy of the model for recognition to a certain extent. For example, an auxiliary task can be created to mine the angle (rotation) information (self-supervisory information) of the image (unsupervised data). For example, by automatically rotating the same image, a third sample image with a different specified angle can be obtained. Among them, the specified angle can be set according to actual needs, for example, it can be 0 degrees, 90 degrees, 180 degrees, and 270 degrees. Among them, the four angle categories of 0 degrees, 90 degrees, 180 degrees, and 270 degrees are the labels to be predicted by the model in the angle classification auxiliary task, that is, the angle classification auxiliary task can be regarded as a four-category classification, and the model needs to identify which angle category the image belongs to. Specifically, the third sample image may be input as a sample in the second training sample, and the specified angle may be output as a sample in the first training sample to form the second training sample.

[0051] Step D: If the auxiliary training task is to perceive the color information contained in the picture, obtain a fourth sample picture with color information and its corresponding grayscale image.

[0052] For the auxiliary training task of perceiving the color information contained in the picture, the auxiliary training task can specifically be a color completion auxiliary task. For example, an auxiliary task can be created to convert a color image with color information, that is, the fourth sample image, into a grayscale image to obtain self-supervised data, and then the model can be trained through the color completion auxiliary task to restore the grayscale image back to a color image with color information. In the process of restoration, the model can learn the specific colors corresponding to different objects contained in the picture, and then establish the perception ability of the model to help the model better perceive the color information contained in the picture. Accordingly, by improving the model's perception ability of color information, the model can be able to perceive richer color information when in use, thereby improving the accuracy of using the model for recognition to a certain extent.

[0053] Furthermore, the fourth sample image with color information may be output as a sample in the second training sample, and the grayscale image may be input as a sample in the first training sample to form the second training sample.

[0054] Step E: If the auxiliary training task is to determine the aesthetic score of the image, a fifth sample image is obtained, and the aesthetic score corresponding to the fifth sample image is obtained through a preset image aesthetic scoring model.

[0055] For the auxiliary training task of determining the aesthetic score of the image, the auxiliary training task can specifically be an aesthetic scoring auxiliary task. For example, an auxiliary task can be created first to give the aesthetic score of each fifth sample image in advance through a preset image aesthetic scoring model as a label of the fifth sample image, and then the model can learn the ability to aesthetically score the image through the aesthetic scoring auxiliary task, helping the model to better understand the aesthetic meaning of the image. Among them, the preset image aesthetic scoring model can be a neural network model, and the aesthetic scoring auxiliary task can be a constructed auxiliary regression task. When the image is aesthetically scored, each object in the image can be scored, or the image as a whole can be scored. Since sensitive images often have weak aesthetic meanings and poor aesthetic quality, the ability to aesthetically score images through the training model can improve the model's perception of the aesthetic meaning of the image to a certain extent, and improve the accuracy of using the model for recognition. Furthermore, the fifth sample image can be used as a sample input in the second training sample, and the aesthetic score corresponding to the fifth sample image can be used as a sample output in the first training sample to form a second training sample.

[0056] Step F: If the auxiliary training task is to determine the probability of containing a sensitive human body in the picture, obtain a sixth sample picture, and obtain a human classification label corresponding to the sixth sample picture through a preset human recognition model; the human classification label represents the probability of containing a sensitive human body in the sixth sample picture.

[0057] For the auxiliary training task of determining the probability of containing sensitive human bodies in pictures, the auxiliary training task may be a human body recognition auxiliary task. For example, an auxiliary task may be created to give the probability of containing sensitive human bodies in each sixth sample picture in advance through a preset human body recognition model as the classification label of the sixth sample picture. The ability to determine the probability of containing sensitive human bodies in pictures through training models can improve the accuracy of using models for recognition to a certain extent.

[0058] It should be noted that the above-mentioned first sample picture, second sample picture, third sample picture, fourth sample picture, fifth sample picture and sixth sample picture may be the same picture or different pictures, and the embodiment of the present application is not limited to this.

[0059] Figure 2 is a schematic diagram of an auxiliary training task provided in an embodiment of the present application, such as Figure 2 As shown, auxiliary training tasks may include: angle classification auxiliary tasks, color completion auxiliary tasks, aesthetic scoring auxiliary tasks, human body recognition auxiliary tasks and category prediction auxiliary tasks. Figure 2 The neural network in can be the image classification model to be trained in the embodiment of the present application.

[0060] Further, Figure 3 This is a schematic diagram of a model provided in an embodiment of the present application, such as Figure 3As shown, in the embodiment of the present application, a category prediction auxiliary task can be constructed according to the category label of the picture according to the pictures belonging to different categories, an angle classification auxiliary task can be constructed through pictures of different angles, an aesthetic scoring auxiliary task can be constructed through pictures with different aesthetic scores, a human body recognition auxiliary task can be constructed through pictures with different probabilities of containing sensitive human bodies (not shown in the figure), and a color completion auxiliary task can be constructed through grayscale images of different pictures (not shown in the figure), and a fine-grained picture classification task, that is, the main training task, can be constructed through vulgar pictures of different sensitive picture categories. These auxiliary tasks are combined with the fine-grained picture classification task for multi-task learning, thereby effectively improving the accuracy and recall rate of vulgar picture recognition. Specifically, in the embodiment of the present application, when the fine-grained picture classification task and the constructed auxiliary task are jointly multi-task learned, they can be on a shallow neural network, that is, a neural network common layer, and promote each other by sharing the network, and then each individual task is passed through the corresponding network layer (the vulgar picture classification layer, the self-supervised learning angle recognition layer, and the self-supervised learning category recognition layer in the figure) and the corresponding loss function to obtain the final result. Among them, the corresponding network layer can be an independent fully connected layer. Then, the network is updated through gradient back propagation. For example, the final result can be used to jointly adjust the model parameters in the common layer of the neural network, and to adjust each independent fully connected layer separately, thereby completing the training of the model.

[0061] Optionally, after obtaining the target image classification model, the embodiment of the present application may also perform sensitive image recognition through the following steps:

[0062] Step G: Identify at least one auxiliary information of the image to be identified through the target image classification model; the auxiliary information includes the category to which the image to be identified belongs on the network platform.

[0063] Step H: According to the auxiliary information and the image to be identified, the probability that the image to be identified belongs to each sensitive image category is determined by the target image classification model.

[0064] Since the target classification model learns to identify the category to which the image belongs during the training process, and determines the probability that the image belongs to each sensitive image category in combination with the category. Therefore, in the embodiment of the present application, the image to be identified can be input into the target image classification model, and accordingly, the target image classification model can process the image to be identified to identify the category to which the image to be identified belongs in the network platform, and determine the probability that the image to be identified belongs to each sensitive image category in combination with the category to which the image to be identified belongs in the network platform.

[0065] Step I: Determine whether the image to be identified is a sensitive image based on the probability that the image to be identified belongs to each sensitive image category.

[0066] In the embodiment of the present application, the greater the probability that the image to be identified belongs to a sensitive image category, the greater the possibility that the image to be identified is a sensitive image of the sensitive image category. Therefore, in the embodiment of the present application, it can be determined first whether there is a sensitive image category with a probability greater than a preset probability threshold. If so, it can be confirmed that the image to be identified is a sensitive image, and then the sensitive image category with the highest probability among the sensitive image categories with a probability greater than the preset probability threshold can be used as the sensitive image category to which the image to be identified belongs. Correspondingly, if it does not exist, it can be considered that the image to be identified is not a sensitive image. Among them, the preset probability threshold can be set according to actual needs. For example, the preset probability threshold can be 0.5, which is not limited in the embodiment of the present application. Furthermore, after identifying that the image to be identified is a sensitive image, a rectification reminder can also be sent to the merchant to which the image to be identified belongs.

[0067] In an embodiment of the present application, multi-task joint training is used to perform recognition based on a target classification model that can combine the probability of category recognition images belonging to various sensitive image categories. The fine-grained category to which the image to be recognized belongs can be determined simply by inputting the image to be recognized into the target classification model. At the same time, the image to be recognized can be combined with category information when treating the image to be recognized, and the image to be recognized can be classified and recognized according to the fine-grained sensitive image category, thereby achieving more accurate recognition. At the same time, recognition interference caused by differences in the presentation of sensitive images under different categories can be avoided, thereby improving recognition accuracy.

[0068] Optionally, the auxiliary information also includes at least one of the following: the rotation angle of the image to be identified, the color information contained in the image to be identified, the aesthetic score of the image to be identified, and the probability that the image to be identified contains a sensitive human body. In this way, by combining the information contained in the image in multiple dimensions, the information contained in the image itself can be more fully utilized, thereby improving the accuracy of recognition to a certain extent.

[0069] In another application scenario, for example, in the homepage recommendation, the recommended merchants or items will be displayed in the form of pictures, which is convenient for users to make more intuitive choices. However, since the pictures are provided by the merchants themselves, the platform often only performs basic safety bottom-line filtering on these pictures, so there may be low-quality sensitive pictures in the displayed pictures. Due to the diversification of online platform services, the categories and picture categories contained in the online platform are often very diverse. For example, the online platform may contain both pictures of dishes and pictures of rooms and daily necessities under multiple categories such as rooms and daily necessities, which makes the identification of sensitive pictures quite difficult. Therefore, a model is needed to identify low-quality pictures for filtering.

[0070] In one implementation, a network model is used to identify whether a picture contains a human body and the posture of the human body, or the picture is directly classified into two categories, the sensitive parts and areas in the picture are identified, and the combined sensitive parts and areas are identified, or the upper and lower body of the picture are detected and identified. In these implementations, the utilization rate of the information contained in the picture itself is low, and the interference of the diversity of categories on classification and identification cannot be overcome, and the recognition accuracy is often low. In the embodiment of the present application, by combining multiple auxiliary training tasks such as identifying the rotation angle of the picture, perceiving the color information contained in the picture, determining the aesthetic score of the picture, and determining the probability of containing a sensitive human body in the picture, the target picture classification model can make full use of the information contained in the picture itself, and overcome the interference of the diversity of categories on classification and identification, thereby improving the classification accuracy to a certain extent.

[0071] Figure 4 This is a flowchart of a sensitive image recognition method provided by this application. Figure 4 As shown, the method includes:

[0072] Step 401: Determine at least one auxiliary information of the image to be identified through a trained image classification model; the auxiliary information includes the category to which the image to be identified belongs in the network platform.

[0073] Step 402: According to the auxiliary information and the picture to be identified, determine the probability that the picture to be identified belongs to each sensitive picture category through the trained picture classification model.

[0074] Step 403: Determine whether the image to be identified is a sensitive image based on the probability that the image to be identified belongs to each sensitive image category.

[0075] The trained image classification model in the embodiment of the present application can be generated by the aforementioned image classification model acquisition method. The execution subject of the sensitive image recognition method provided in the embodiment of the present application can be the same as or different from the execution subject of the aforementioned image classification model acquisition method.

[0076] In an embodiment of the present application, when the trained classification model is used to identify pictures, the category information can be combined to classify and identify the pictures according to fine-grained sensitive picture categories, thereby enabling more accurate identification. At the same time, it can avoid identification interference caused by differences in the presentation of sensitive pictures under different categories, thereby improving recognition accuracy.

[0077] Figure 5 is a structural diagram of a device for acquiring a picture classification model provided in an embodiment of the present application. The device 50 may include:

[0078] A construction module 501 is used to construct a main training task and at least one auxiliary training task for the image classification model to be trained; wherein the main training task is to identify the probability that an image belongs to each sensitive image category in combination with the auxiliary information obtained in the auxiliary training task; and the auxiliary training task includes identifying the category to which the image belongs on the network platform;

[0079] The training module 502 is used to train the image classification model to be trained according to the main training task and the auxiliary training task to obtain a target image classification model.

[0080] Optionally, the auxiliary training task also includes at least one of the following: identifying the rotation angle of the image, perceiving the color information contained in the image, determining the aesthetic score of the image, and determining the probability of containing a sensitive human body in the image.

[0081] Optionally, the image classification model to be trained includes a shared layer and a network layer corresponding to each training task;

[0082] The training module 502 is specifically used for:

[0083] According to the first training sample corresponding to the main training task and the second training sample corresponding to the auxiliary training task, obtaining a first loss value of the to-be-trained image classification model corresponding to the main training task and a second loss value corresponding to the auxiliary training task;

[0084] Adjusting the model parameters in the shared layer together according to the first loss value and the second loss value;

[0085] According to the first loss value and the second loss value, the model parameters in the network layer corresponding to the main training task and the model parameters in the network layer corresponding to the auxiliary training task are adjusted respectively.

[0086] Optionally, the first training sample is obtained through the following modules:

[0087] The first acquisition module is used to acquire first sample images belonging to different sensitive image categories and label values ​​of the sensitive image categories to which they belong.

[0088] Optionally, the sensitive image category is used to characterize the type of image to which the sensitive image belongs; the sensitive image category is determined based on historical sensitive images under various categories contained in the network platform.

[0089] Optionally, the second training sample is obtained by the following module:

[0090] A second acquisition module is used to acquire second sample images belonging to different categories and the categories they belong to if the auxiliary training task is to identify the categories to which the images belong on the network platform;

[0091] A third acquisition module, configured to acquire third sample images belonging to different specified angles and the specified angles to which they belong if the auxiliary training task is to identify the rotation angle of the image;

[0092] A fourth acquisition module, configured to acquire a fourth sample image with color information and a corresponding grayscale image thereof if the auxiliary training task is to perceive the color information contained in the image;

[0093] a fifth acquisition module, configured to acquire a fifth sample image if the auxiliary training task is to determine the aesthetic score of the image, and to acquire the aesthetic score corresponding to the fifth sample image through a preset image aesthetic scoring model;

[0094] The sixth acquisition module is used to obtain a sixth sample image if the auxiliary training task is to determine the probability of containing a sensitive human body in the image, and to obtain a human classification label corresponding to the sixth sample image through a preset human recognition model; the human classification label represents the probability of containing a sensitive human body in the sixth sample image.

[0095] Optionally, the device 50 further includes:

[0096] A recognition module, used to recognize at least one auxiliary information of the image to be recognized through the target image classification model; the auxiliary information includes the category to which the image to be recognized belongs in the network platform;

[0097] A first determination module, configured to determine, based on the auxiliary information and the image to be identified, the probability that the image to be identified belongs to each sensitive image category through the target image classification model;

[0098] The second determination module is used to determine whether the image to be identified is a sensitive image according to the probability that the image to be identified belongs to each sensitive image category.

[0099] Optionally, the auxiliary information also includes at least one of the following: a rotation angle of the image to be identified, color information contained in the image to be identified, an aesthetic score of the image to be identified, and a probability that a sensitive human body is contained in the image to be identified.

[0100] In summary, the image classification model acquisition device provided by the embodiment of the present application can construct a main training task and at least one auxiliary training task for the image classification model to be trained; wherein the main training task is to identify the probability that the image belongs to each sensitive image category in combination with the auxiliary information obtained in the auxiliary training task; the auxiliary training task includes identifying the category to which the image belongs in the network platform, and training the image classification model to be trained according to the main training task and the auxiliary training task to obtain the target image classification model. In this way, the image classification model to be trained can learn how to identify the category to which the image belongs, and the ability to identify the probability of the image belonging to each sensitive image category in combination with the category to which the image belongs during the training process. In addition, it can be ensured that when the trained target image classification model is used for classification in the future, the target image classification model can take into account the presentation methods of sensitive images under different categories during the classification process, and classify in combination with the category to which the image belongs, thereby eliminating the influence of the presentation methods of sensitive images under different categories on the classification to a certain extent, reducing the probability of misjudgment and improving accuracy.

[0101] The present application also provides an electronic device, see Figure 6 , including: a processor 601, a memory 602, and a computer program 6021 stored in the memory and executable on the processor, wherein when the processor executes the program, the image classification model acquisition method of the aforementioned embodiment is implemented.

[0102] The present application also provides a readable storage medium. When the instructions in the storage medium are executed by a processor of an electronic device, the electronic device can execute the image classification model acquisition method of the aforementioned embodiment.

[0103] As for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0104] The algorithm and display provided herein are not inherently related to any particular computer, virtual system or other device. Various general purpose systems can also be used together with the teachings based on this. According to the above description, it is obvious that the structure required for constructing such systems. In addition, the application is not directed to any specific programming language either. It should be understood that various programming languages ​​can be utilized to realize the content of the application described herein, and the description of the specific language above is for the purpose of disclosing the best mode of implementation of the application.

[0105] In the description provided herein, a large number of specific details are described. However, it is understood that the embodiments of the present application can be practiced without these specific details. In some instances, well-known methods, structures and techniques are not shown in detail so as not to obscure the understanding of this description.

[0106] Similarly, it should be understood that in order to streamline the present application and help understand one or more of the various inventive aspects, in the above description of the exemplary embodiments of the present application, the various features of the present application are sometimes grouped together into a single embodiment, figure, or description thereof. However, the disclosed method should not be interpreted as reflecting the following intention: the claimed application requires more features than the features clearly stated in each claim. More specifically, as reflected in the claims below, the inventive aspects are less than all the features of the single embodiment disclosed above. Therefore, the claims following the specific embodiment are hereby expressly incorporated into the specific embodiment, wherein each claim itself serves as a separate embodiment of the present application.

[0107] Those skilled in the art will appreciate that the modules in the devices in the embodiments may be adaptively changed and arranged in one or more devices different from the embodiments. The modules or units or components in the embodiments may be combined into one module or unit or component, and further may be divided into a plurality of submodules or subunits or subcomponents. All features disclosed in this specification (including the accompanying claims, abstracts and drawings) and all processes or units of any method or device so disclosed may be combined in any combination, except that at least some of such features and / or processes or units are mutually exclusive. Unless otherwise expressly stated, each feature disclosed in this specification (including the accompanying claims, abstracts and drawings) may be replaced by an alternative feature that provides the same, equivalent or similar purpose.

[0108] The various component embodiments of the present application may be implemented in hardware, or in software modules running on one or more processors, or in a combination thereof. It should be understood by those skilled in the art that a microprocessor or a digital signal processor (DSP) may be used in practice to implement some or all of the functions of some or all of the components in the sorting device according to the present application. The present application may also be implemented as a device or apparatus program for executing part or all of the methods described herein. Such a program implementing the present application may be stored on a computer-readable medium, or may be in the form of one or more signals. Such a signal may be downloaded from an Internet website, or provided on a carrier signal, or provided in any other form.

[0109] It should be noted that the above embodiments illustrate the present application rather than limit the present application, and that those skilled in the art may design alternative embodiments without departing from the scope of the appended claims. In the claims, any reference symbol between brackets should not be constructed as a limitation to the claims. The word "comprising" does not exclude the presence of elements or steps not listed in the claims. The word "one" or "an" preceding an element does not exclude the presence of multiple such elements. The present application may be implemented by means of hardware including several different elements and by means of a suitably programmed computer. In a unit claim that lists several devices, several of these devices may be embodied by the same hardware item. The use of the words first, second, and third, etc. does not indicate any order. These words may be interpreted as names.

[0110] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0111] The above description is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present application should be included in the protection scope of the present application.

[0112] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

Claims

1. A method for obtaining an image classification model, It is characterized in that The method comprises: constructing a main training task and at least one auxiliary training task for a picture classification model to be trained; wherein the main training task is to identify the probability that a picture belongs to each sensitive picture category in combination with the auxiliary information obtained in the auxiliary training task; the auxiliary training task comprises identifying the category to which the picture belongs on the network platform; and according to the main training task and the auxiliary training task, training the picture classification model to be trained to obtain a target picture classification model; The image classification model to be trained includes a shared layer and a network layer corresponding to each training task; the training of the image classification model to be trained according to the main training task and the auxiliary training task includes: obtaining a first loss value of the image classification model to be trained corresponding to the main training task and a second loss value corresponding to the auxiliary training task according to a first training sample corresponding to the main training task and a second training sample corresponding to the auxiliary training task; jointly adjusting the model parameters in the shared layer according to the first loss value and the second loss value; and respectively adjusting the model parameters in the network layer corresponding to the main training task and the model parameters in the network layer corresponding to the auxiliary training task according to the first loss value and the second loss value.

2. The method according to claim 1, It is characterized in that The auxiliary training task also includes at least one of the following: identifying the rotation angle of the image, perceiving the color information contained in the image, determining the aesthetic score of the image, and determining the probability of containing a sensitive human body in the image.

3. The method according to claim 1, It is characterized in that The first training samples are obtained in the following manner: first sample images belonging to different sensitive image categories and label values ​​of the sensitive image categories to which they belong are obtained.

4. The method according to claim 1, 2 or 3, It is characterized in that The sensitive image category is used to characterize the type of image to which the sensitive image belongs; the sensitive image category is determined based on historical sensitive images under various categories contained in the network platform.

5. The method according to claim 2, It is characterized in that The second training sample is obtained in the following manner: if the auxiliary training task is to identify the category to which the image belongs on the network platform, then obtaining second sample images belonging to different categories and the categories to which they belong; If the auxiliary training task is the rotation angle of the identified image, then a third sample image belonging to different specified angles and the specified angles to which they belong are obtained; if the auxiliary training task is the color information contained in the perceived image, then a fourth sample image with color information and its corresponding grayscale image are obtained; if the auxiliary training task is the aesthetic score of the determined image, then a fifth sample image is obtained, and the aesthetic score corresponding to the fifth sample image is obtained through a preset image aesthetic scoring model; if the auxiliary training task is the probability of containing a sensitive human body in the determined image, then a sixth sample image is obtained, and the human body classification label corresponding to the sixth sample image is obtained through a preset human body recognition model; the human body classification label represents the probability of containing a sensitive human body in the sixth sample image.

6. The method according to claim 1, It is characterized in that The method also includes: identifying at least one auxiliary information of the image to be identified through the target image classification model; the auxiliary information includes the category to which the image to be identified belongs in the network platform; determining the probability that the image to be identified belongs to each sensitive image category through the target image classification model based on the auxiliary information and the image to be identified; and determining whether the image to be identified is a sensitive image based on the probability that the image to be identified belongs to each sensitive image category.

7. The method according to claim 6, It is characterized in that The auxiliary information also includes at least one of the following: a rotation angle of the image to be identified, color information contained in the image to be identified, an aesthetic score of the image to be identified, and a probability that the image to be identified contains a sensitive human body.

8. A device for acquiring a picture classification model, It is characterized in that The device includes: a construction module, which is used to construct a main training task and at least one auxiliary training task for the image classification model to be trained; wherein the main training task is to identify the probability that the image belongs to each sensitive image category in combination with the auxiliary information obtained in the auxiliary training task; the auxiliary training task includes identifying the category to which the image belongs in the network platform; a training module, which is used to train the image classification model to be trained according to the main training task and the auxiliary training task to obtain a target image classification model, wherein the image classification model to be trained includes a shared layer and a network layer corresponding to each training task; the training of the image classification model to be trained according to the main training task and the auxiliary training task includes: obtaining a first loss value of the image classification model to be trained corresponding to the main training task and a second loss value corresponding to the auxiliary training task according to a first training sample corresponding to the main training task and a second training sample corresponding to the auxiliary training task; jointly adjusting the model parameters in the shared layer according to the first loss value and the second loss value; and respectively adjusting the model parameters in the network layer corresponding to the main training task and the model parameters in the network layer corresponding to the auxiliary training task according to the first loss value and the second loss value.

9. An electronic device, It is characterized in that include: A processor, a memory, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the method for acquiring a picture classification model as described in one or more of claims 1-7 is implemented.

10. A readable storage medium, It is characterized in that When the instructions in the storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the image classification model acquisition method as described in one or more of method claims 1-7.

Citation Information

Patent Citations

  • Sensitive image recognition method and device

    CN107291737A

  • Picture processing method and device, mobile terminal and computer readable storage medium

    CN108805095A