Multistage classification sensitive image recognition method, terminal and computer storage medium
Through the multi-level classification sensitive image recognition method, the combination of self-attention convolution neural network and global classification layer is used to solve the problem of insufficient classification in the prior art, and the detailed classification and recognition of sensitive images are achieved, with good scalability and adjustability.
Patent Information
- Application Number
- CN202410135767.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-30
- Publication Date
- 2025-08-01
AI Technical Summary
The existing sensitive image recognition methods are not classified in detail enough and cannot be flexibly adjusted according to the scene, resulting in false positives or missed inspections.
The multi-level classification sensitive image recognition method is adopted to extract image features through self-attention convolution neural network, and combine global classification layer and multi-level classification structure, and use parameter training of self-attention feature processing layer and multi-level classification layer, optimize the model using adversarial data amplification strategy, and merge classification results to improve recognition accuracy.
The detailed classification and recognition of sensitive images is realized, with good scalability and adjustability, and the types and ranges of sensitive images can be added or modified as needed, improving the accuracy and stability of recognition.
Smart Images

Figure CN120411580A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of neural networks, and particularly relates to a multi-level classification sensitive image recognition method, a terminal, and a computer storage medium. Background Art
[0002] The related technology of sensitive image recognition was born when the Internet was widely used. Researchers generally call such images NSFW images, which is the abbreviation of "not safe for work", and generally refers to images and videos that are not suitable for appearing in the workplace or public environment. Narrowly defined, sensitive images generally refer to "pornographic" images, and most open-source work mainly focuses on pornographic classification. With the development of the Internet and the complexity of Internet content, the defined scope of sensitive image recognition needs to be continuously expanded, and its classification will also include classifications such as violence and terrorism, politics, etc., which are also not suitable for work, advertising, and publicity, as well as some pictures used for illegal publicity on the Internet.
[0003] Deep learning technology (Deep Learning technology, or simply DL technology) has been used in the related field of sensitive image recognition for some time. In 2016, the earliest NSFW open-source model using DL technology was proposed in the prior art, and some papers also achieved good recognition rates on public datasets using related technologies. Most existing video-related websites have set up their own neural network recognition schemes. However, on the one hand, because these models are based on public datasets, many only perform recognition on the "pornographic" classification, and the scope covered by the datasets may not be applicable to the current complex content review requirements; at the same time, most current recognition systems are not detailed enough in classification and cannot be flexibly adjusted according to the scenario, resulting in false alarms or missed detections. Summary of the Invention
[0004] The purpose of the present invention is to provide a multi-level classification sensitive image recognition method, a terminal, and a computer storage medium, aiming to solve the problem that the classification of the existing sensitive image recognition method is not detailed enough and cannot be flexibly adjusted according to the scenario.
[0005] The technical solution adopted by the present invention to solve the technical problem is as follows:
[0006] The present invention provides a multi-level classification sensitive image recognition method, and the multi-level classification sensitive image recognition method includes:
[0007] Obtain a to-be-tested image, and extract the features of the to-be-tested image from the to-be-tested image;
[0008] Input the features of the image to be tested into the global classification layer to output the global classification result, and input the features of the image to be tested into the classification layers of each level of the multi-level classification structure respectively to output the classification results of each level of the multi-level classification;
[0009] Merge the classification results of each level of the multi-level classification to obtain the merged classification result, and fuse the merged classification result with the result of the global classification to form the probabilities of each category and output.
[0010] Further, extracting the features of the image to be tested from the image to be tested specifically includes:
[0011] Input the image to be tested into the self-attention convolutional neural network to output the preliminary features of the image to be tested;
[0012] Input the preliminary features of the image to be tested into the self-attention feature processing layer to output the features of the image to be tested.
[0013] Further, the parameters of the self-attention convolutional neural network, the attention feature processing layer, the classification layers of each level and the global classification layer are obtained through training, and the training includes:
[0014] Obtain the data sets of normal images and sensitive images, and label the data sets to form a validation set and a preliminary training set;
[0015] Perform data augmentation processing on the preliminary training set to form an enhanced training set;
[0016] Train the self-attention convolutional neural network, the attention feature processing layer, the classification layers of each level and the global classification layer according to the enhanced training set and the validation set.
[0017] Further, the data augmentation processing includes adversarial data augmentation processing, and the adversarial data augmentation processing specifically includes:
[0018] Select sensitive images;
[0019] Perform adversarial processing on the sensitive images according to the adversarial processing strategy of adversarial images;
[0020] Add the images after adversarial processing to the training set.
[0021] Further, the adversarial processing includes one or more of blurring, noise, texture, color inversion, occlusion and rotation deformation.
[0022] Further, the loss function of the training is:
[0023]
[0024] Among them, MSE is the mean square error loss function, CE is the cross entropy loss function, BCE is the binary cross entropy loss function, L is the total number of levels of the multi-level classification structure, λ is the weight value of the multi-level classification, i is the i-th level classification layer, P i is the predicted probability of the i-th level classification layer, P G is the predicted probability of the global classification layer, and label is the onthot vector.
[0025] Furthermore, the multi-level classification sensitive image recognition method further includes:
[0026] Calculate the F1 score of the classification result based on the probabilities of each category output multiple times;
[0027] The training parameters and training strategy of the training are improved according to the F1 score of the classification result.
[0028] Furthermore, the classification results of each level of the multi-level classification are merged to obtain a merged classification result, and the merged classification result is fused with the result of the global classification to form the probability of each category and output it. Specifically, the probability output of each category after fusion is calculated:
[0029] Output=(1-λ)Out G +λOut L
[0030] Where, Out G is the global classification layer probability output, Out L is the output form of the multi-class classification layer probability output, and λ is the weight value of the multi-class classification.
[0031] In addition, to achieve the above-mentioned purpose, the present invention also provides a terminal, which includes: a memory, a processor, and a multi-level classification sensitive image recognition program stored on the memory and run on the processor. When the multi-level classification sensitive image recognition program is executed by the processor, it controls the terminal to implement the steps of the multi-level classification sensitive image recognition method as described above.
[0032] In addition, to achieve the above-mentioned purpose, the present invention also provides a computer storage medium, which stores a multi-level classification sensitive image recognition program, and when the multi-level classification sensitive image recognition program is executed by the processor, it implements the steps of the multi-level classification sensitive image recognition method as described above.
[0033] The present invention adopts the above technical solution to achieve the following effects:
[0034] The present invention identifies sensitive images through the structure of a multi-level classification network. It can not only divide the categories of sensitive images into different levels according to the degree of detail, so as to define and classify sensitive images in detail, but also has good scalability and adjustability, and can add and modify the types and ranges of sensitive images according to needs. Brief Description of the Drawings
[0035] Figure 1 is a flowchart of the steps of the multi-level classification sensitive image recognition method in a preferred embodiment of the present invention;
[0036] Figure 2 is a schematic diagram of the model structure of the multi-level classification sensitive image recognition method in a preferred embodiment of the present invention;
[0037] Figure 3 is a diagram of the structural change of the self-attention convolutional neural network in a preferred embodiment of the present invention;
[0038] Figure 4 is a flowchart of the steps of training in a preferred embodiment of the present invention;
[0039] Figure 5 is a schematic diagram of the operating environment of a preferred embodiment of the terminal of the present invention. Detailed Embodiments
[0040] To make the objectives, technical solutions and advantages of the present invention clearer and more definite, the following further describes the present invention in detail with reference to the accompanying drawings and by way of examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0041] Embodiment 1
[0042] Please refer to Figure 1 and Figure 2 , Embodiment 1 of the present invention is a multi-level classification sensitive image recognition method, which includes the steps:
[0043] S1. Obtain a test image, and extract the features of the test image from the test image.
[0044] Specifically, the step S1 includes the steps:
[0045] S11. Input the test image into a neural network, and output the preliminary features of the test image.
[0046] In this embodiment, the neural network is specifically a self-attention convolutional neural network. In other alternative embodiments, the neural network can also be other neural network structures for feature extraction, such as a neural network based on the Transformer structure.
[0047] Specifically, the self-attention convolutional neural network is mainly composed of a convolutional layer and a self-attention layer. Please refer to Figure 3 As shown, the self-attention convolutional neural network uses the reparameterization technique. During training, the main structure behaves as a multi-branch model structure, and each convolutional layer contains multiple branch paths such as 3×3 convolution, 1×1 convolution, and shortcut. In the prediction mode after reparameterization, the main structure behaves as a single-path model structure. The overall network has no branch structure, only uses 3×3 convolution, and only uses ReLU as the activation function to achieve a very high computational density.
[0048] This is because the multi-branch structure has advantages during training, but has poor performance in actual inference. Therefore, in this embodiment, the multi-branch structure is used to train the convolutional layer during training to improve the training effect, and the single-path branch structure is used during use to ensure the performance during use.
[0049] The output of the self-attention convolutional neural network is a vector with a dimension of N×2560, where N is the batch size of the input image, that is, the number of samples, and 2560 is the dimension of the feature vector of a single image.
[0050] S12. Input the preliminary features of the to-be-tested image into the self-attention feature processing layer, and output the features of the to-be-tested image.
[0051] The input of the self-attention feature processing layer is a feature vector with a dimension of N×2560 calculated by the self-attention convolutional neural network, and the output is an embedding vector with a dimension of N×768. This module is used to adaptively distinguish the features required for different hierarchical classifications and encode them into corresponding embedding vectors. Embedding refers to the process of mapping high-dimensional data (such as text, images, audio) into a low-dimensional space. An embedding vector is an N-dimensional real-valued vector, usually a vector composed of real numbers, which represents the input data as a point in a continuous numerical space. The structure of the self-attention feature processing layer is a single-layer multi-head attention structure.
[0052] In this step, an Attention pooling is performed through the self-attention feature processing layer to adaptively distinguish the features required for different hierarchical classifications, so as to be able to judge the clothing, actions, and objects of the people in the image. At the same time, information such as the picture scene, the relationship between the background and the foreground, and the composition can also be considered, and the deep semantic information of the image can be recognized, improving the accuracy while meeting the rationality of classification in applications.
[0053] S2. Input the features of the image to be tested into the global classification layer to output the global classification result, and input the features of the image to be tested into the classification layers at all levels of the multi-level classification structure respectively to output the classification results at all levels of the multi-level classification.
[0054] In this embodiment, the sensitive image categories, sub-categories and classification criteria are defined, and according to relevant regulations and relevant academic research, the types and levels of sensitive images are clearly classified and defined.
[0055] The multi-level classification layer is constructed by using a multi-level classification structure of type HMCN-F, that is, a structure of a multi-level classification network, which is generally used for classification tasks with inclusion relationships and hierarchical structures. In the experiment, some classifications such as "sexy" and "pornographic" are not completely mutually exclusive in semantic features. Therefore, a classification task with a hierarchical structure needs to be designed, such as adding a parent classification of "porn-related" to improve the classification accuracy and semantic rationality.
[0056] The multi-level classification structure has the advantages of low computational cost and being suitable for extracting local information from the hierarchical structure. However, if only the multi-level classification structure is used, it is easy to cause overfitting because the classifier of the multi-level classification structure is more suitable for extracting local information from the regions of the hierarchical structure. Therefore, in this application, a global classification layer is designed so that the global classification layer and the multi-level classification layer can be combined later to integrate the advantages of global classification and multi-level classification.
[0057] Specifically, the multi-level classification specifically includes first-level classification, second-level classification and third-level classification. The first-level classification classifies sensitive images into multiple first-level categories, the second-level classification classifies sensitive images into multiple second-level categories, and the third-level classification classifies sensitive images into multiple third-level categories. Among them, the second-level categories are sub-categories of the first-level categories, and the third-level categories are sub-categories of the second-level categories. For example, the first-level categories can be divided into categories such as "porn-related", "violent" and "horrible", and the first-level category of "porn-related" includes one or more sub-classifications. For example, the first-level category of "porn-related" can include the sub-classification of "pornographic" as the second-level category, and the second-level category of "pornographic" can include, for example, "anime porn" as the third-level category.
[0058] Among them, the output of the first-level classification is a probability vector with the number of first-level categories, and its elements respectively correspond to the probabilities of each first-level category. Similarly, the output of the second-level classification is a probability vector with the number of second-level categories, and its elements respectively correspond to the probabilities of each second-level category. The output of the third-level classification is a probability vector with the number of third-level categories, and its elements respectively correspond to the probabilities of each third-level category.
[0059] The output of the global classification is a probability vector of the sum of the numbers of categories at all levels, and its elements respectively correspond to the probabilities of each category at each level of classification. For example, if there are 3 first-level categories in the first-level classification, 5 second-level categories in the second-level classification, and 10 third-level categories in the third-level classification, the global classification output is a probability vector of length 18.
[0060] Among them, the values of the elements of the probability vector range from 0 to 1.
[0061] S3. Combine the classification results of each level of the multi-level classification to obtain a combined classification result, and fuse the combined classification result with the result of the global classification to form the probabilities of each category and output them.
[0062] Specifically, in this embodiment, the fusion of the combined result and the result of the global classification to form the probabilities of each category for output is specifically to multiply the combined output of the first, second, and third-level classifications by a coefficient (usually set to 0.1), and add it to the global classification output to obtain the final classification result. The formula for weighted fusion is as follows:
[0063] Output=(1 - λ)Out G +λOut L ;
[0064] In the formula, Out G is the probability output of the global classification layer, and its output form is, for example: [0.5, 0.3, 0.3, 0.1, 0.0], Out L is the probability output of the multi-level classification layer, and its output form is, for example: [0.5, 0.3, 0.3, 0.1, 0.0], and λ is the artificially set weight parameter.
[0065] The self-attention convolutional neural network, attention feature processing layer, each level of classification layer, and global classification layer of this application are trained using the data set processed by data augmentation. Specifically, please refer to Figure 4 , and the training includes the steps:
[0066] A1. Obtain the data sets of normal images and sensitive images, and label the data sets to form a validation set and a preliminary training set.
[0067] Obtain normal image and sensitive image data through web crawlers and some public and private data sets. After labeling according to the data set definition, construct a sensitive image training set and a validation set. Among them, the data set definition is to define the sensitive image categories, sub-categories, and classification criteria. According to relevant regulations and relevant academic research, the classifications such as pornographic, violent and terrorist, politically sensitive, vulgar and disgusting, subculture, and Internet meme pictures are clearly defined, and the secondary classifications are clearly defined according to the definition.
[0068] A2. Perform data augmentation on the preliminary training set to form an enhanced training set.
[0069] Specifically, in this embodiment, general data augmentation and adversarial data augmentation are performed on the images of the training set. The general data augmentation includes rotation, random cropping, and mirror flipping. The adversarial data augmentation includes blurring, rotational deformation, color inversion, grayscale conversion, blocking key parts with color blocks, adding texture and noise, and splicing with normal images.
[0070] Among them, performing general data augmentation and adversarial data augmentation on the images of the training set specifically means randomly collecting some sensitive pictures and performing data augmentation on them.
[0071] It can be seen that in this embodiment, not only general data augmentation is performed on the images of the training set, but also an adversarial image data enhancement strategy is designed according to the characteristics of adversarial images on the network. As a result, the system can perform semantic recognition on images containing interference information, such as images with added blurring, noise, texture, color inversion, occlusion, and rotational deformation, and make accurate judgments, obtaining a multi-level classification sensitive image recognition algorithm with sufficiently accurate recognition, sufficiently stable system, sufficient defense against adversarial images, rapid response, and efficient calculation.
[0072] A3. Train the self-attention convolutional neural network, the attention feature processing layer, the classification layers at all levels, and the global classification layer according to the enhanced training set and the validation set.
[0073] Specifically, the loss function for training is:
[0074]
[0075] In the formula, MSE is the Mean Square Error loss function, CE is the CrossEntropy loss function, BCE is the Binary Cross Entropy loss function, L is the total number of levels of the multi-level classification structure, λ is the weight value, i is the i-th classification layer, P i is the prediction probability of the i-th classification layer, P G is the prediction probability of the global classification layer, Label is the onthot vector, and the onthot vector is a vector with only one element being 1 and the rest of the elements being 0. For example: [1, 0, 0, 0, 0].
[0076] In an alternative embodiment, the training is also evaluated based on the predicted F1 score to improve the corresponding training parameters and training strategies. The improvement of the corresponding training parameters includes adjusting the hyperparameters of the model, such as the learning rate. Improving the training strategy means selecting different training strategies, such as the cosine annealing decay strategy or the gradient descent strategy, etc.
[0077] Among them, the F1 score can be understood as the harmonic mean of precision and recall, taking into account both the accuracy of the prediction (precision, preferring to miss detections rather than make mistakes) and the proportion of correct predictions among the actual true ones (recall, trying to find every object that should be found, but there may be more false detections).
[0078] By using the F1 score as the evaluation metric for the trained model, a balance between precision and recall is achieved, so that the finally obtained model can strike a balance between false positives and false negatives.
[0079] Embodiment 2
[0080] Please refer to Figure 5 , based on the above method, the present invention also provides a terminal, which includes: a memory 10, a processor 20, and a multi-level classification sensitive image recognition program stored on the memory 10 and executable on the processor 20. When the multi-level classification sensitive image recognition program is executed by the processor 20, it controls the terminal to implement the steps of the multi-level classification sensitive image recognition method as described above.
[0081] In some embodiments, the memory 10 may be an internal storage unit of the terminal, such as the hard disk or memory of the terminal. In other embodiments, the memory 10 may also be an external storage device of the terminal, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the terminal. Further, the memory 10 may also include both the internal storage unit and the external storage device of the terminal. The memory 10 is used to store application software installed on the terminal and various types of data, such as the program code for installing the terminal. The memory 10 may also be used to temporarily store data that has been output or will be output. In one embodiment, a multi-level classification sensitive image recognition program is stored on the memory 10, and this multi-level classification sensitive image recognition program can be executed by the processor 20 to implement the multi-level classification sensitive image recognition method in this application.
[0082] In some embodiments, the processor 20 may be a central processing unit (CPU), a microprocessor, or other data processing chips, which are used to run the program code stored in the memory 10 or process data, such as executing the multi-level classification sensitive image recognition method, etc.
[0083] Embodiment III
[0084] This embodiment provides a storage medium. The computer storage medium stores a multi-level classification sensitive image recognition program. When the multi-level classification sensitive image recognition program is executed by a processor, it implements the steps of the multi-level classification sensitive image recognition method as described above.
[0085] In summary, the present invention identifies sensitive images through the structure of a multi-level classification network. It can not only divide the categories of sensitive images into different levels according to the degree of detail, so as to define and classify sensitive images in detail, but also has good scalability and adjustability, and can add and modify the types and ranges of sensitive images according to needs.
[0086] It should be noted that in this article, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or terminal including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or terminal. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or terminal including that element.
[0087] Of course, those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware (such as a processor, a controller, etc.) through a computer program. The program can be stored in a computer-readable storage medium, and when the program is executed, it can include the processes of the above method embodiments. The storage medium can be a memory, a magnetic disk, an optical disk, etc.
[0088] It should be understood that the application of the present invention is not limited to the above examples. For those of ordinary skill in the art, improvements or transformations can be made according to the above description, and all these improvements and transformations should fall within the protection scope of the appended claims of the present invention.
Claims
1. A multi-level classification sensitive image recognition method, characterized in that, The multi-level classification sensitive image recognition method includes: Obtain a test image, and extract the features of the test image from the test image; Input the features of the test image into the global classification layer to output the global classification result, and input the features of the test image into the classification layers of each level of the multi-level classification structure respectively to output the classification results of each level of the multi-level classification; Merge the classification results of each level of the multi-level classification to obtain a merged classification result, and fuse the merged classification result with the result of the global classification to form the probabilities of each category and output them.
2. The multi-level classification sensitive image recognition method according to claim 1, wherein, The step of extracting the features of the test image from the test image specifically includes: Input the test image into the self-attention convolutional neural network to output the preliminary features of the test image; Input the preliminary features of the test image into the self-attention feature processing layer to output the features of the test image.
3. The multi-level classification sensitive image recognition method according to claim 2, characterized in that The parameters of the self-attention convolutional neural network, the attention feature processing layer, the classification layers of each level, and the global classification layer are obtained through training. The training includes: Obtain a data set of normal images and sensitive images, and label the data set to form a validation set and a preliminary training set; Perform data augmentation processing on the preliminary training set to form an enhanced training set; Train the self-attention convolutional neural network, the attention feature processing layer, the classification layers of each level, and the global classification layer according to the enhanced training set and the validation set.
4. A multi-level classification sensitive image recognition method according to claim 3, characterized in that The data augmentation processing includes adversarial data augmentation processing. The adversarial data augmentation processing specifically includes: Select sensitive images; Perform adversarial processing on the sensitive images according to the adversarial processing strategy of adversarial images; Add the images after adversarial processing to the training set.
5. A multi-level classification sensitive image recognition method according to claim 4, characterized in that, The adversarial processing includes one or more of blurring, noise, texture, color inversion, occlusion, and rotational deformation.
6. A multi-level classification sensitive image recognition method according to claim 3, characterized in that, The loss function of the training is: Among them, MSE is the mean squared error loss function, CE is the cross-entropy loss function, BCE is the binary cross-entropy loss function, L is the total number of levels of the multi-level classification structure, λ is the weight value of multi-level classification, i is the i-th classification layer, and P i is the predicted probability of the i-th classification layer, and P G is the predicted probability of the global classification layer, and label is the onthot vector.
7. A multi-level classification sensitive image recognition method according to claim 3, characterized in that, The multi-level classification sensitive image recognition method further includes: Calculate the F1 score of the classification result according to the probabilities of each category output multiple times; Improve the training parameters and training strategy of the training according to the F1 score of the classification result.
8. A multi-level classification sensitive image recognition method according to claim 1, characterized in that The step of merging the classification results of each level of the multi-level classification to obtain a merged classification result, and fusing the merged classification result with the result of the global classification to form the probabilities of each category and output them is specifically to calculate the probability output Output of each category after fusion: Output=(1 - λ)Out G + λOut L Where, Out G is the probability output of the global classification layer, Out L is the probability output form of the multi-level classification layer, and λ is the weight value of the multi-level classification.
9. A terminal, characterized in that, The terminal includes: a memory, a processor, and a multi-level classification sensitive image recognition program stored on the memory and executable on the processor. When the multi-level classification sensitive image recognition program is executed by the processor, it controls the terminal to implement the steps of the multi-level classification sensitive image recognition method according to any one of claims 1-8.
10. A computer storage medium, characterized in that, The computer storage medium stores a multi-level classification sensitive image recognition program. When the multi-level classification sensitive image recognition program is executed by a processor, it implements the steps of the multi-level classification sensitive image recognition method according to any one of claims 1-8.
Citation Information
Patent Citations
Sensitive image authentication method and terminal system
CN109145979A
Multi-modal memetic graph emotion detection method
CN116563619A
Patent text multilevel classification method and device based on graph neural network
CN116932765A
Image recognition method and apparatus, image classification method and apparatus, electronic device, and storage medium
WO2020239015A1