Model training method, system, server and storage medium for identifying illegal content
By selecting and replacing sensitive and illegal content regions and training with enhanced models, the method addresses the challenge of inaccurate illegal content identification, achieving improved recognition accuracy.
Patent Information
- Application Number
- CN202111413529.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-25
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2041-11-25
AI Technical Summary
Existing image recognition technology is difficult to accurately distinguish between normal, sensitive and illegal content in the recognition of illegal content, especially due to the complex scene overlap and progressive relationship between these three categories, resulting in poor recognition results.
By obtaining sample image sets, filtering and replacing the content of sensitive and violation areas, generating corrected images, and using pre-trained models for training, enhance image feature association relationships, especially by adding global attention mechanism modules for training.
The generated violation content recognition model can more accurately identify violation content, enhance the correlation and distinctive features between image features, and improve the recognition accuracy.
Smart Images

Figure CN114241253B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of image recognition technology, and particularly to a model training method, system, server, and storage medium for identifying illegal content. Background Art
[0002] With the development of the Internet and 5G technology, a large number of original videos and live broadcasts are generated every day. Under regulatory requirements, no illegal content can appear in these video and live broadcast contents. Effectively identifying video illegal content and reducing manual intervention is a field worthy of exploration.
[0003] However, with the popularization of deep learning, existing image recognition all adopts a classification network based on a convolutional neural network; and in the process of training a convolutional neural network, in addition to the amount of training data and the classic network structure, effective training methods and adjustments to the network structure are also the key to improving the accuracy of image recognition; for the scenario of identifying illegal content, compared with other fields of image recognition, the recognition difficulty is higher; illegal content is usually divided into normal, sensitive, and illegal, and the scenarios of these three types of video content overlap with each other. Normal scenarios such as people, environments, and indoor scenes may all appear in sensitive and illegal scenarios, and sensitive and illegal are a classification situation with a progressive degree; based on this, conventional training methods and models often cannot achieve accurate results in identifying illegal content. Summary of the Invention
[0004] The purpose of the embodiments of the present application is to provide a model training method, system, server, and storage medium for identifying illegal content, so that the generated illegal content recognition model can accurately identify illegal content.
[0005] To solve the above technical problems, the embodiments of the present application provide a model training method for identifying illegal content, including: obtaining a sample image set, where the sample image set includes normal images, sensitive images, and illegal images; screening sensitive regions from the sensitive images, and replacing the region content in the sensitive regions according to a preset replacement rule to generate corrected normal images; screening illegal regions from the illegal images, and replacing the region content in the illegal regions according to the replacement rule to generate corrected sensitive images; inputting the sample image set, the corrected normal images, and the corrected sensitive images into a pre-trained model for training to generate an illegal content recognition model.
[0006] Embodiments of the present application also provide a system for identifying illegal content, including: a receiving module, the above-mentioned illegal content identification model, and a judgment module; wherein, the receiving module is configured to receive and decode the video content to be identified, and extract at least one frame of video image to be identified from the video content to be identified according to a preset extraction method; the illegal content identification model is configured to identify illegal content for each of the video images to be identified, and obtain the illegal results and illegal probabilities of each of the video images to be identified; the judgment module is configured to judge the identification result of the video content to be identified according to the illegal results and illegal probabilities of each of the video images to be identified.
[0007] Embodiments of the present application also provide a server, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the above-mentioned model training method for identifying illegal content.
[0008] Embodiments of the present application also provide a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the above-mentioned model training method for identifying illegal content is implemented.
[0009] In the embodiments of the present application, during the model training process for identifying illegal content, a sample image set is obtained, and the sample image set includes normal images, sensitive images, and illegal images; sensitive regions are screened out from the sensitive images, and the regional content in the sensitive regions is replaced according to a preset replacement rule to generate corrected normal images; illegal regions are screened out from the illegal images, and the regional content in the illegal regions is replaced according to the replacement rule to generate corrected sensitive images; the sample image set, the corrected normal images, and the corrected sensitive images are input into a pre-trained model for training to generate an illegal content identification model. This enables the pre-trained model of the present application to learn the progressive relationship between different categories of images from normal images, sensitive images, corrected normal images, illegal images, and corrected sensitive images, can enhance the correlation between image features, and can also better focus on the relevant features of illegal images and the differentiating features between illegal images and other images, so that the generated illegal content identification model of the present application can accurately identify illegal content. Description of the Drawings
[0010] One or more embodiments are illustrated by way of example with reference to the pictures in the corresponding drawings, and these exemplary illustrations do not limit the embodiments.
[0011] Figure 1 It is a flowchart of the model training method for identifying illegal content provided by the embodiments of the present application;
[0012] Figure 1a It is a schematic diagram of replacing the content of the illegal area of the illegal image provided by the embodiment of the present application;
[0013] Figure 1b It is a schematic structural diagram of the global attention mechanism module provided by the embodiment of the present application;
[0014] Figure 2 It is a flowchart of the model training method for illegal content recognition provided by the embodiment of the present application;
[0015] Figure 2a It is a schematic diagram of cropping normal images, sensitive images and illegal images provided by the embodiment of the present application;
[0016] Figure 3 It is a flowchart of the model training method for illegal content recognition provided by the embodiment of the present application;
[0017] Figure 4 It is a flowchart of the model training method for illegal content recognition provided by the embodiment of the present application;
[0018] Figure 5 It is a schematic structural diagram of the illegal content recognition system provided by the embodiment of the present application;
[0019] Figure 6 It is a schematic structural diagram of the server provided by the embodiment of the present application. Detailed implementation manners
[0020] To make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the following will elaborate on each embodiment of the present application in conjunction with the accompanying drawings. However, those of ordinary skill in the art can understand that in each embodiment of the present application, many technical details are proposed for the convenience of readers to better understand the present application. However, even without these technical details and various changes and modifications based on the following embodiments, the technical solutions required to be protected by the present application can still be implemented. The division of the following embodiments is for convenience of description and should not constitute any limitation to the specific implementation manner of the present application. The various embodiments can be combined and cross-referenced with each other on the premise of not conflicting.
[0021] The embodiment of the present application relates to a model training method for illegal content recognition, as Figure 1 shown, specifically including the following steps:
[0022] Step 101, obtain a sample image set, and the sample image set includes normal images, sensitive images and illegal images.
[0023] In an exemplary implementation, the sample images in the sample image set can be divided into three categories: normal images, sensitive images, and illegal images. Each sample image in the sample image set contains an image label, which is used to describe the image category of the sample image. The style of the image label can be (normal probability, sensitive probability, illegal probability). For example, (1, 0, 0) indicates that the sample image is a normal image, (0, 1, 0) indicates that the sample image is a sensitive image, and (0, 0, 1) indicates that the sample image is an illegal image; it can also be a text description of the relative image category, such as normal, sensitive, or illegal, etc.
[0024] In an exemplary implementation, taking pornographic and obscene content as the illegal content, the sensitive images in the sample image set can be sexy images, and the illegal images can be pornographic and obscene images.
[0025] Step 102: Screen out the illegal areas from the illegal images, and replace the area content in the illegal areas according to the replacement rule to generate corrected sensitive images.
[0026] In an exemplary implementation, as Figure 1a shown, to screen out the illegal areas from the illegal images, first, it is necessary to use a pre-set illegal prediction model to process the sensitive images to generate predicted illegal images. Then, use the Gradient-weighted Class Activation Mapping (abbreviated as Grad-CAM) method to process the predicted illegal images to generate the attention heat map of the predicted illegal images. Then, screen out the pixel points that meet the preset pixel conditions from the heat map according to the pixel values of each pixel point in the heat map. The pixel points that meet the pixel conditions form the first area, and the area in the illegal image corresponding to the first area is used as the illegal area of the illegal image.
[0027] In an exemplary implementation, after determining the illegal areas in the illegal images, it is necessary to replace the content in the illegal areas of the illegal images with the content in other areas of the illegal images to generate corrected sensitive images. The image category of the generated corrected sensitive images is sensitive images. However, since the corrected sensitive images are corrected from the illegal images, the image label of the corrected sensitive images can be (0, 0.95, 0.05) or corrected sensitive.
[0028] Step 103: Screen out the sensitive areas from the sensitive images, and replace the area content in the sensitive areas according to the preset replacement rule to generate corrected normal images.
[0029] In an exemplary implementation, the method of screening out sensitive regions from sensitive images is roughly the same as the method of screening out illegal regions from illegal images, and the methods of generating corrected normal images and corrected sensitive images are also roughly the same. The corrected normal images are corrected from the sensitive images. Therefore, the image label of the corrected normal images can be (0.95, 0.05, 0) or corrected normal.
[0030] Step 104: Input the sample image set, the corrected normal images, and the corrected sensitive images into a pre-trained model for training to generate an illegal content recognition model.
[0031] In an exemplary implementation, the pre-trained model is constructed based on the EfficientNet network architecture, and a global attention mechanism module (GC-Block module) is added at a specified position in the pre-trained model. The addition of the global attention mechanism module enables the pre-trained model to better focus on the correlation between image features during the training process. Among them, the pre-trained model used in this application can also be a recognition model that has been trained but has poor recognition effects on illegal content.
[0032] In an exemplary implementation, the traditional EfficientNet network architecture consists of a feature extraction layer, a classification layer, and a softmax activation layer. Taking the example that the feature extraction needs to go through 9 stages, the feature extraction process of the pre-trained model is shown in Table 1:
[0033] Table 1 Feature Extraction Process of the Pre-trained Model
[0034]
[0035] In an exemplary implementation, Conv refers to a convolutional block, and MBConv refers to an inverted residual module. A global attention mechanism module (GC-Block module) is added between the 3rd and 4th stages, the 4th and 5th stages, the 5th and 6th stages, the 6th and 7th stages, and the 7th and 8th stages of the feature extraction layer to enhance the correlation between image features. The structural schematic diagram of the GC-Block module is as Figure 1b shown.
[0036] In an exemplary implementation, after being processed by the feature extraction layer, the extracted features are input into the classification layer for classification processing. The classification layer consists of a global pooling layer and three fully connected layers, which are used to identify illegalities in the image based on the extracted features and obtain the illegal result and illegal probability of the image. The softmax activation layer is connected to the last fully connected layer of the classification layer and is used to normalize the illegal result and illegal probability output by the classification layer.
[0037] In an exemplary implementation, normal images, sensitive images, and illegal images in a sample image set, as well as corrected normal images and corrected sensitive images, are input into a pre-trained model. The pre-trained model processes each image, and can learn the discriminative features between normal images and sensitive images from the corrected normal images and their corresponding sensitive images. Similarly, it can also learn the discriminative features between illegal images and sensitive images from the corrected sensitive images and their corresponding illegal images.
[0038] In an exemplary implementation, to prevent the parameters of the pre-trained model from fluctuating violently due to the setting of the learning rate, a warmup method is adopted first. Let the model be trained in the first few epochs (such as 3 epochs), starting from a very low learning rate (such as 10e-6) and increasing linearly to the preset learning rate (such as 10e-2), and then the learning rate is decreased in a cosine curve manner for learning. At the same time, the optimizer uses the traditional Stochastic Gradient Descent combined with Momentum (parameter is 0.9) for training.
[0039] In the embodiments of the present application, during the training process of the illegal content recognition model, a sample image set is obtained. The sample image set includes normal images, sensitive images, and illegal images. The sensitive regions are screened out from the sensitive images, and the regional content in the sensitive regions is replaced according to the preset replacement rules to generate corrected normal images. The illegal regions are screened out from the illegal images, and the regional content in the illegal regions is replaced according to the replacement rules to generate corrected sensitive images. The sample image set, corrected normal images, and corrected sensitive images are input into the pre-trained model for training to generate an illegal content recognition model. This enables the pre-trained model of the present application to learn the progressive relationship between different categories of images from normal images, sensitive images, corrected normal images, illegal images, and corrected sensitive images, can enhance the correlation between image features, and can also better focus on the relevant features of illegal images and the discriminative features between illegal images and other images, so that the generated illegal content recognition model of the present application can accurately identify illegal content.
[0040] The embodiments of the present application relate to a method for training a model for identifying illegal content, as Figure 2 shown, specifically including the following steps:
[0041] Step 201, obtain a sample image set, where the sample image set includes normal images, sensitive images, and illegal images.
[0042] In an exemplary implementation, this step is substantially the same as step 101 in the embodiments of the present application, and will not be elaborated here one by one.
[0043] Step 202: Screen out the illegal areas from the illegal images, and replace the area content in the illegal areas according to the preset replacement rules to generate corrected sensitive images.
[0044] In an exemplary implementation, this step is substantially the same as step 102 in the embodiments of the present application, and will not be elaborated here one by one.
[0045] Step 203: Screen out the sensitive areas from the sensitive images, and replace the area content in the sensitive areas according to the replacement rules to generate corrected normal images.
[0046] In an exemplary implementation, this step is substantially the same as step 103 in the embodiments of the present application, and will not be elaborated here one by one.
[0047] Step 204: Perform scene recognition on the normal images, sensitive images, and illegal images. When the scenes of the normal images, sensitive images, or illegal images are the same, use the first cropping method to crop the normal images to obtain normal background images;
[0048] In an exemplary implementation, perform scene recognition on each sample image in the sample dataset in sequence, and monitor whether there are cases where the scenes of the normal images and sensitive images, normal images and illegal images, sensitive images and illegal images, and normal images, sensitive images, and illegal images are the same. When the scenes are the same, use the first cropping method to crop the normal images to obtain normal background images, and assign the normal image label (1, 0, 0) or the image label (normal background) to the normal images. Among them, as Figure 2a shown, the first cropping method refers to the normal cropping of images. The scene recognition method used in this application is some currently disclosed scene recognition methods, which will not be elaborated here.
[0049] Step 205: Use the second cropping method to crop the sensitive images or illegal images to obtain sensitive background images or illegal background images.
[0050] In an exemplary implementation, as Figure 2a shown, the second cropping method refers to a cropping method that only crops the background information around the image. Use the second cropping method to crop the sensitive background images or illegal background images from the sensitive images or illegal images, and assign the normal image label (0.95, 0.05, 0) or the image label (normal sensitive background) to the sensitive background images, and assign the normal image label (0.9, 0.05, 0.05) or the image label (normal illegal background) to the illegal background images; both the sensitive background images and illegal background images belong to the normal image category.
[0051] Step 206: Add normal image labels to the normal background images, sensitive background images, and illegal background images, and add them to the sample dataset.
[0052] In an exemplary implementation, the obtained normal background images, sensitive background images, and illegal background images are added to the sample dataset and input into the pre-trained model for training together.
[0053] Step 207: Input the sample image set, corrected normal images, and corrected sensitive images into the pre-trained model for training to generate an illegal content recognition model.
[0054] In an exemplary implementation, this step is substantially the same as step 104 in the embodiments of the present application, and will not be elaborated here one by one.
[0055] In the embodiments of the present application, on the basis of other embodiments, by learning the normal background images, sensitive background images, and illegal background images, the pre-trained model can learn the differences between background images and the features of recognition objects (such as sensitive parts) from such data.
[0056] The embodiments of the present application relate to a model training method for identifying illegal content, as Figure 3 shown, and specifically include the following steps:
[0057] Step 301: Obtain a sample image set, which includes normal images, sensitive images, and illegal images.
[0058] In an exemplary implementation, this step is substantially the same as step 101 in the embodiments of the present application, and will not be elaborated here one by one.
[0059] Step 302: Screen out the illegal regions from the illegal images, and replace the region content in the illegal regions according to the preset replacement rules to generate corrected sensitive images.
[0060] In an exemplary implementation, this step is substantially the same as step 102 in the embodiments of the present application, and will not be elaborated here one by one.
[0061] Step 303: Screen out the sensitive regions from the sensitive images, and replace the region content in the sensitive regions according to the replacement rules to generate corrected normal images.
[0062] In an exemplary implementation, this step is substantially the same as step 103 in the embodiments of the present application, and will not be elaborated here one by one.
[0063] Step 304: Split the sample dataset into multiple sample data subsets.
[0064] In an exemplary implementation, during the training process based on the sample dataset, the mini-batch training method can be adopted to split a large sample dataset into multiple sample data subsets.
[0065] Step 305: Input each sample data subset, the corrected normal images corresponding to the sensitive images in each sample data subset, and the corrected sensitive images corresponding to the illegal images in each sample data subset into each pre-trained model for training to generate each illegal content recognition sub-model.
[0066] In an exemplary implementation, this step is substantially the same as step 102 in the embodiment of the present application, and will not be elaborated here one by one.
[0067] In an exemplary implementation, for a sample data subset containing normal images, sensitive images, and illegal images, the number of illegal or sensitive images is relatively small compared to normal images, and the collection difficulty is relatively large; this leads to an imbalance among different types of images in a sample data subset. The illegal content recognition sub-model generated by training based on a sample data subset with an imbalance among different types of images has a relatively low accuracy in identifying illegal content. At this time, oversampling can be performed on the sensitive images and / or illegal images in a sample data subset, or on the basis of oversampling, a specified number of sensitive images and / or illegal images can be obtained from the sample data set and added to the sample data subset to increase the occurrence frequency of these samples with relatively less data in each iteration, and the sample data set after adding sensitive images and / or the illegal images is used to perform the next round of iterative training on the illegal content recognition sub-model, and so on, until the accuracy of the generated illegal content recognition sub-model in identifying illegal content meets the preset requirements.
[0068] Step 306: According to the preset fusion rule, fuse the model parameters of each illegal content recognition sub-model to generate an illegal content recognition model.
[0069] In an exemplary implementation, in step 305, each sample data subset, the corrected normal images corresponding to the sensitive images in each sample data subset, and the corrected sensitive images corresponding to the illegal images in each sample data subset are input into each pre-trained model for training. Since the network structures of each pre-trained model are the same but the training strategies are different, there are differences in the model parameters of each illegal content recognition sub-model obtained by training each pre-trained model. A fusion rule for model parameters can be preset to fuse the model parameters of each illegal content recognition sub-model to obtain the model parameters of a new illegal content recognition model, and the newly obtained illegal content recognition model is the finally generated one.
[0070] In the embodiment of the present application, on the basis of other embodiments, multiple illegal content recognition sub-models trained with the same network structure but different training strategies can be used, and the parameters of the multiple illegal content recognition sub-models can be fused to obtain the final model parameters, which not only does not increase the size of the model but also achieves the effect of fusing the illegal content recognition sub-models and improves the accuracy of the obtained illegal content recognition model.
[0071] An embodiment of the present application relates to a method for training a model for identifying illegal content, as Figure 4 shown, which specifically includes the following steps:
[0072] Step 401, obtain a sample image set, which includes normal images, sensitive images, and illegal images.
[0073] In an exemplary implementation, this step is substantially the same as the method of step 101 in the embodiment of the present application, and will not be elaborated here one by one.
[0074] Step 402, screen out the illegal regions from the illegal images, and replace the region content in the illegal regions according to the preset replacement rules to generate corrected sensitive images.
[0075] In an exemplary implementation, this step is substantially the same as the method of step 102 in the embodiment of the present application, and will not be elaborated here one by one.
[0076] Step 403, screen out the sensitive regions from the sensitive images, and replace the region content in the sensitive regions according to the replacement rules to generate corrected normal images.
[0077] In an exemplary implementation, this step is substantially the same as the method of step 103 in the embodiment of the present application, and will not be elaborated here one by one.
[0078] Step 404, input the sample image set, the corrected normal images, and the corrected sensitive images into a pre-trained model for training to generate an illegal content identification model.
[0079] In an exemplary implementation, this step is substantially the same as the method of step 104 in the embodiment of the present application, and will not be elaborated here one by one.
[0080] Step 405, perform model conversion on the illegal content identification model according to a preset model acceleration tool, and perform packaging processing on the converted illegal content identification model.
[0081] In the embodiment of the present application, a preset model acceleration tool (such as the TensorRT model) is used to perform model conversion on the generated illegal content identification model, and the converted illegal content identification model is encapsulated into an Application Program Interface (API) or a Software Development Kit (SDK).
[0082] In the embodiment of the present application, on the basis of other embodiments, model conversion can also be performed on the generated illegal content identification model to accelerate the inference speed and improve the practicability of the illegal content identification model.
[0083] The step division of the above various methods is only for clear description. When implemented, they can be combined into one step or some steps can be split into multiple steps. As long as the same logical relationship is included, it is within the protection scope of this patent. Making insignificant modifications to the algorithm or adding insignificant designs to the process, but without changing the core design of its algorithm and process, are all within the protection scope of this patent.
[0084] The embodiments of this application relate to a system for identifying illegal content. It is characterized in that the details of the system for identifying illegal content in this embodiment will be specifically described below. The following content is only implementation details provided for convenient understanding and is not necessary for implementing this example. Figure 5 It is a schematic diagram of the system for identifying illegal content in this embodiment, including: a receiving module 501, an illegal content recognition model 502, and a judgment module 503.
[0085] Among them, the receiving module 501 is used to receive and decode the video content to be recognized, and extract at least one frame of video image to be recognized from the video content to be recognized according to a preset extraction method.
[0086] The illegal content recognition model 502 is used to identify illegal content in each video image to be recognized, and obtain the illegal results and illegal probabilities of each video image to be recognized.
[0087] The judgment module 503 is used to judge the recognition result of the video content to be recognized according to the illegal results and illegal probabilities of each video image to be recognized.
[0088] In an exemplary implementation, the video content to be recognized received by the receiving model 501 is encapsulated. After receiving the video content to be recognized, it is necessary to perform a decompression operation on the video content to be recognized to obtain the video content to be recognized. Then, according to the preset extraction method, each frame of video image to be recognized is extracted from the video content to be recognized. For example, when the video content to be recognized is a live video, one frame can be extracted every 10s; when the video content to be recognized is an on-demand video, one frame can be extracted every 1s. Then, each frame of video image to be recognized extracted is input into the illegal content recognition model 502 for processing.
[0089] In an exemplary implementation, the illegal content recognition model 502 is encapsulated in the C++ SDK format through TensorRT. Its input is a video frame image array composed of each frame of video image to be recognized, and the output is the illegal result and illegal probability of each frame of video image to be recognized. Then, the illegal results and illegal probabilities of each frame of video image to be recognized are input into the judgment module 503 for processing.
[0090] In an exemplary implementation, the determination module 503 counts the proportion of illegal pictures in the entire video content to be recognized based on the illegal results and illegal probabilities of single-frame video images to be recognized, and determines the recognition result of the entire video content to be recognized according to a certain proportion; for each frame of video image to be recognized, those with an illegal probability above 0.85 can also be directly intercepted, and those below 0.85 are sent to humans for verification; for the sexy probability, it can be intercepted according to the specific requirements of the service. Additionally, when the probabilities of the three categories are all close, for example, 0.33 indicates that the model cannot make a good judgment, and such images are also sent to humans for verification.
[0091] It is not difficult to find that this embodiment is a system embodiment corresponding to the above method embodiment, and this embodiment can be implemented in cooperation with the above method embodiment. The relevant technical details and technical effects mentioned in the above embodiment are still valid in this embodiment. To avoid repetition, they are not elaborated here. Correspondingly, the relevant technical details mentioned in this embodiment can also be applied to the above embodiment.
[0092] An embodiment of the present application relates to a server, such as Figure 6 shown, including: at least one processor 601; and a memory 602 communicatively connected to the at least one processor 601; wherein, the memory 602 stores instructions executable by the at least one processor 601, and the instructions are executed by the at least one processor 601 to enable the at least one processor 601 to execute the model training method for identifying illegal content in the above embodiments.
[0093] Among them, the memory and the processor are connected by a bus. The bus can include any number of interconnected buses and bridges, and the bus connects various circuits of one or more processors and memories together. The bus can also connect various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art, so they will not be further described herein. The bus interface provides an interface between the bus and the transceiver. The transceiver can be one element or multiple elements, such as multiple receivers and transmitters, and provides a unit for communicating with various other devices on the transmission medium. The data processed by the processor is transmitted over the wireless medium through the antenna. Further, the antenna also receives data and transmits the data to the processor.
[0094] The processor is responsible for managing the bus and general processing, and can also provide various functions, including timing, peripheral interface, voltage regulation, power management, and other control functions. The memory can be used to store the data used by the processor when performing operations.
[0095] An embodiment of the present application relates to a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the method embodiments described above are implemented.
[0096] That is, those skilled in the art can understand that all or part of the steps in implementing the methods of the above embodiments can be completed by instructing relevant hardware through a program. The program is stored in a storage medium and includes several instructions to enable a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs.
[0097] Those of ordinary skill in the art can understand that the above embodiments are specific embodiments for implementing the present application, and in practical applications, various changes can be made in form and details without departing from the spirit and scope of the present application.
Claims
1. A method for training a model for identifying illegal content, characterized in that, The method includes: Obtaining a sample image set, where the sample image set includes normal images, sensitive images, and illegal images; Screening out illegal regions from the illegal images, and replacing the region content in the illegal regions according to a preset replacement rule to generate corrected sensitive images; Screening out sensitive regions from the sensitive images, and replacing the region content in the sensitive regions according to the replacement rule to generate corrected normal images; Inputting the sample image set, the corrected normal images, and the corrected sensitive images into a pre-trained model for training to generate an illegal content recognition model.
2. The method for training a model for identifying illegal content according to claim 1, wherein Before inputting the sample image set, the corrected normal images, and the corrected sensitive images into the pre-trained model for training, it further includes: Performing scene recognition on the normal images, the sensitive images, and the illegal images; When the scenes of the normal images, the sensitive images, or the illegal images are the same, using a first cropping method to crop the normal images to obtain normal background images; Using a second cropping method to crop the sensitive images or the illegal images to obtain sensitive background images or illegal background images; Adding normal image labels to the normal background images, the sensitive background images, and the illegal background images, and adding them to the sample image set.
3. The method for training a model for identifying illegal content according to claim 1, characterized in that, Inputting the sample image set, the corrected normal images, and the corrected sensitive images into the pre-trained model for training to generate an illegal content recognition model, including: Splitting the sample image set into multiple sample data subsets; Inputting each sample data subset, the corrected normal images corresponding to the sensitive images in each sample data subset, and the corrected sensitive images corresponding to the illegal images in each sample data subset into each pre-trained model for training to generate each illegal content recognition sub-model; Fusing the model parameters of each illegal content recognition sub-model according to a preset fusion rule to generate the illegal content recognition model.
4. The method for training a model for identifying illegal content according to claim 3, wherein, After inputting each sample data subset, the corrected normal images corresponding to the sensitive images in each sample data subset, and the corrected sensitive images corresponding to the illegal images in each sample data subset into each pre-trained sub-model of the pre-trained model for training to generate each illegal content recognition sub-model, it further includes: When the recognition effect of the illegal content recognition sub-model does not meet the preset requirements, obtaining a specified number of the sensitive images or the illegal images from the sample image set and adding them to the sample data subset, and using the sample data subset after adding the sensitive images or the illegal images to train the illegal content recognition sub-model.
5. The method for training a model for identifying illegal content according to claim 1, wherein Screening out illegal regions from the illegal images, including: Processing the illegal images using a preset illegal prediction model to generate predicted illegal images; Processing the predicted illegal images according to the gradient-based weighted activation mapping method to generate a heat map of the predicted illegal images; Taking the region corresponding to the first region composed of each pixel point in the illegal image that meets the preset pixel conditions in the heat map as the illegal region.
6. The method for training a model for identifying illegal content according to claim 1, wherein The pre-trained model is trained based on a learning strategy with learning rate warm-up and an optimizer using the Stochastic Gradient Descent with Momentum algorithm.
7. The method for training a model for identifying illegal content according to claim 1, wherein After inputting the sample image set, the corrected normal images, and the corrected sensitive images into the pre-trained model for training to generate a violation content recognition model, the method further includes: performing model conversion on the violation content recognition model according to a preset model acceleration tool, and performing encapsulation processing on the converted violation content recognition model.
8. An illegal content recognition system, characterized in that, The system includes: a receiving module, a violation content recognition model generated by using the violation content recognition model training method according to any one of claims 1-7, and a judging module; Wherein, the receiving module is configured to receive and decode the video content to be recognized, and extract at least one frame of video image to be recognized from the video content to be recognized according to a preset extraction method; The violation content recognition model is configured to perform violation content recognition on each of the video images to be recognized, and obtain the violation results and violation probabilities of each of the video images to be recognized; The judging module is configured to judge the recognition result of the video content to be recognized according to the violation results and violation probabilities of each of the video images to be recognized.
9. A server, characterized in that, Comprising: At least one processor; And, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the model training method for violation content recognition according to any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the model training method for violation content recognition according to any one of claims 1 to 7.
Citation Information
Patent Citations
Group behavior recognition model based on progressive relationship learning and training method thereof
CN110516599A
Method and system for detection of deformable structures in medical images
US20090010509A1