Multi-level sensitive text classification method and system, terminal and storage medium
Through a multi-level sensitive text classification method, sensitive text data of historical Internet content is used for data cleaning and model training, and the coverage and accuracy problems of sensitive statement recognition in the existing technology are solved, achieving high-precision and fast sensitive text classification.
Patent Information
- Application Number
- CN202510141582.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-08
- Publication Date
- 2025-05-30
AI Technical Summary
In the recognition of sensitive statements, the classification coverage is not wide enough, and the text deep semantics and optical character recognition errors cannot be distinguished, resulting in false alarms or missed detection, and the response speed is slow.
A multi-level sensitive text classification method is adopted to clean and build a hierarchical label data set by obtaining sensitive text data of historical Internet content, model training is carried out to generate a sensitive text classification model, and the text data to be identified is input into the model to output classification results.
It has achieved the wide coverage of identification classification and high degree of subdivision, the ability to identify text meaning and reference, and consider text context information to improve classification accuracy, reduce resource requirements, and improve response speed.
Smart Images

Figure CN120067331A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of text classification, and particularly to a multi-level sensitive text classification method, system, terminal and computer-readable storage medium. Background Art
[0002] With the rapid growth of the amount of Internet content such as articles, the cost of manually reviewing Internet content security has risen sharply, and at the same time, it is difficult to meet the growing content demand in terms of response speed. Therefore, using artificial intelligence algorithms to identify sensitive sentences has become an important solution to ensure content security.
[0003] However, the method based on traditional syntactic analysis has limitations. For example, the coverage rate of the sensitive word library is insufficient, and the limitation of the modifier discrimination rule template makes it difficult to accurately identify sensitive sentences in a large number of scenarios. In addition, due to the diversity of the text changes of sensitive sentences, the method based on deep learning is also easily affected by sentence noise, resulting in a large deviation in the results. On the other hand, the existing sensitive text recognition methods often fail to fully consider the deep meaning of the text, and the classification system is not detailed enough, lacking the ability to flexibly adjust according to specific scenarios, which may lead to false alarms or missed detections. At the same time, with the increase in the application requirements of edge devices, it may be restricted by the memory of hardware devices during operation, resulting in problems such as slow response speed.
[0004] Therefore, the existing technology still needs to be improved and developed. Summary of the Invention
[0005] The main purpose of the present invention is to provide a multi-level sensitive text classification method, system, terminal and storage medium, aiming to solve the problems in the prior art that the classification coverage of sensitive sentences is not wide enough, the deep semantic meaning of the text and optical character recognition errors cannot be distinguished, resulting in false alarms or missed detections, low recognition accuracy, and slow response speed.
[0006] To achieve the above object, the present invention provides a multi-level sensitive text classification method, and the multi-level sensitive text classification method includes the following steps:
[0007] Obtain sensitive text data of historical Internet content, clean the sensitive text data to obtain target sensitive text data, and construct a hierarchical label data set according to the target sensitive text data;
[0008] Preprocess the hierarchical label data set to obtain a training data set, and perform model training according to the training data set to obtain a sensitive text classification model;
[0009] Obtain the text data to be recognized of the current Internet content, input the text data to be recognized into the sensitive text classification model, and output the sensitive text classification result.
[0010] Optionally, for the multi-level sensitive text classification method, wherein, obtaining the sensitive text data of historical Internet content, cleaning the sensitive text data to obtain target sensitive text data, and constructing a hierarchical label data set according to the target sensitive text data, specifically including;
[0011] Obtain the sensitive text data of historical Internet content, perform deduplication processing on the sensitive text data to obtain deduplicated sensitive text data, and perform annotation processing on the deduplicated sensitive text data to obtain target sensitive text data;
[0012] Set the sensitive label level, and perform grading processing on the target sensitive text data according to the sensitive label level to obtain a hierarchical label data set.
[0013] Optionally, for the multi-level sensitive text classification method, wherein, preprocessing the hierarchical label data set to obtain a training data set, specifically including:
[0014] Set a preset text length, compare the text lengths of all text data in the hierarchical label data set with the preset text length to obtain a comparison result;
[0015] If there is text data in the comparison result whose text length is less than the preset text length, obtain the corresponding text data, and perform character padding on the text data to obtain a target hierarchical label data set;
[0016] Perform data augmentation on the target hierarchical label data set to obtain text-augmented data, perform text conversion on the text-augmented data to obtain multiple text sequences, and combine all the text sequences to obtain a training data set.
[0017] Optionally, for the multi-level sensitive text classification method, wherein, performing data augmentation on the target hierarchical label data set to obtain text-augmented data, specifically including:
[0018] Perform homophone replacement on the target hierarchical label data set to obtain text replacement data, and perform random character deletion on the text replacement data to obtain target text replacement data;
[0019] Perform equivalent character replacement on the target text replacement data to obtain text-augmented data.
[0020] Optionally, in the multi-level sensitive text classification method, the model training based on the training data set to obtain a sensitive text classification model specifically includes:
[0021] Create a sensitive text classification training model and input a set of text data samples in the training data set into the sensitive text classification training model;
[0022] Extract information features from the text data samples to obtain label level information, and construct positive samples according to the label level information to obtain predicted sensitive text classification samples;
[0023] Calculate the loss value between the predicted sensitive text classification sample and the true sensitive text classification result corresponding to the text data sample to obtain a target loss value, and correct the parameters of the sensitive text classification training model according to the target loss value;
[0024] Input the next set of text data samples into the sensitive text classification training model until the training situation of the sensitive text classification training model meets the preset conditions to obtain a trained sensitive text classification model.
[0025] Optionally, in the multi-level sensitive text classification method, the calculation of the loss value between the predicted sensitive text classification sample and the true sensitive text classification result corresponding to the text data sample to obtain a target loss value, and the correction of the parameters of the sensitive text classification training model according to the target loss value specifically includes:
[0026] Determine a loss function according to the predicted sensitive text classification sample, and calculate the loss value between the predicted sensitive text classification sample and the true sensitive text classification result corresponding to the text data sample according to the loss function to obtain a target loss value;
[0027] Correct the parameters of the sensitive text classification training model according to the target loss value.
[0028] Optionally, in the multi-level sensitive text classification method, the expression of the loss function is:
[0029]
[0030] where L is the loss function, L C is the cross-entropy loss function, is the cross-loss function for generating positive samples, λ is the weight of the contrast learning loss, and L con is the contrast learning loss.
[0031] Optionally, in the multi - level sensitive text classification method, the multi - level sensitive text classification system includes:
[0032] A data processing module, configured to obtain sensitive text data of historical Internet content, perform data cleaning on the sensitive text data to obtain target sensitive text data, and construct a hierarchical label data set according to the target sensitive text data;
[0033] A model training module, configured to pre - process the hierarchical label data set to obtain a training data set, and perform model training according to the training data set to obtain a sensitive text classification model;
[0034] A sensitive classification module, configured to obtain text data to be recognized of current Internet content, input the text data to be recognized into the sensitive text classification model, and output a sensitive text classification result.
[0035] In addition, to achieve the above object, the present invention further provides a terminal, where the terminal includes: a memory, a processor, and a multi - level sensitive text classification program stored on the memory and executable on the processor. When the multi - level sensitive text classification program is executed by the processor, the steps of the multi - level sensitive text classification method as described above are implemented.
[0036] In addition, to achieve the above object, the present invention further provides a computer - readable storage medium, where the computer - readable storage medium stores a multi - level sensitive text classification program. When the multi - level sensitive text classification program is executed by a processor, the steps of the multi - level sensitive text classification method as described above are implemented.
[0037] In the present invention, sensitive text data of historical Internet content is obtained, the sensitive text data is subjected to data cleaning to obtain target sensitive text data, and a hierarchical label data set is constructed according to the target sensitive text data; the hierarchical label data set is pre - processed to obtain a training data set, and model training is performed according to the training data set to obtain a sensitive text classification model; text data to be recognized of current Internet content is obtained, the text data to be recognized is input into the sensitive text classification model, and a sensitive text classification result is output. The recognition and classification of the present invention have a wide coverage range and a high degree of subdivision, can also recognize the meaning and reference of the text, and can also consider the context information of the text, improve the classification accuracy, and require less resources and have a fast response speed during processing. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 is a flowchart of a preferred embodiment of the multi - level sensitive text classification method of the present invention;
[0039] Figure 2It is a schematic diagram of the model network architecture construction process of the multi-level sensitive text classification method of the present invention;
[0040] Figure 3 It is a schematic diagram of the model training process of the multi-level sensitive text classification method of the present invention;
[0041] Figure 4 It is a schematic diagram of the process of the model of the multi-level sensitive text classification method of the present invention for identifying and classifying sensitive texts;
[0042] Figure 5 It is the overall flowchart of the preferred embodiment of the multi-level sensitive text classification method of the present invention;
[0043] Figure 6 It is the structural diagram of the preferred embodiment of the multi-level sensitive text classification system of the present invention;
[0044] Figure 7 It is a schematic diagram of the operating environment of the preferred embodiment of the terminal of the present invention. Specific Embodiments
[0045] To make the objectives, technical solutions and advantages of the present invention clearer and more definite, the following further describes the present invention in detail with reference to the accompanying drawings and by way of examples. It should be understood that the specific examples described herein are only used to explain the present invention and are not used to limit the present invention.
[0046] It should be noted that if there are directional indications (such as up, down, left, right, front, back...) in the embodiments of the present invention, the directional indications are only used to explain the relative positional relationship and movement conditions between components in a specific posture (as shown in the drawings). If the specific posture changes, the directional indications will also change accordingly.
[0047] In addition, if there are descriptions such as "first" and "second" in the embodiments of the present invention, the descriptions of "first" and "second" are only for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In addition, the technical solutions between various embodiments can be combined with each other, but it must be based on the fact that those of ordinary skill in the art can implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection required by the present invention.
[0048] The multi-level sensitive text classification method described in the preferred embodiment of the present invention, as Figure 1 shown, the multi-level sensitive text classification method includes the following steps:
[0049] Step S10: Obtain the sensitive text data of historical Internet content, clean the sensitive text data to obtain the target sensitive text data, and construct a hierarchical label data set according to the target sensitive text data.
[0050] The step S10 includes:
[0051] Step S11: Obtain the sensitive text data of historical Internet content, perform duplicate removal processing on the sensitive text data to obtain deduplicated sensitive text data, and perform annotation processing on the deduplicated sensitive text data to obtain the target sensitive text data;
[0052] Step S12: Set the sensitive label level, and perform hierarchical processing on the target sensitive text data according to the sensitive label level to obtain a hierarchical label data set.
[0053] Specifically, in the embodiment of the present invention, in order to more accurately identify whether the text is sensitive, a neural network using a bidirectional attention mechanism is used to extract text features, a graph convolutional model is used to extract hierarchical label features, and contrast learning is used to integrate the hierarchical label features into the text information to train a neural network model with hierarchical label information; before model training, data for training needs to be obtained and cleaned. Specifically, the sensitive text data of historical Internet content is obtained, and the same text data in the sensitive text data needs to be removed, that is, duplicate removal processing is performed on the sensitive text data to obtain deduplicated sensitive text data, and then annotation processing is performed on the deduplicated sensitive text data to obtain the target sensitive text data; subsequently, hierarchical processing needs to be performed on the target sensitive text data. Specifically, the sensitive label level is set (for example, three levels of sensitive labels are set, where the first-level sensitive label is prohibited and non-prohibited), and the target sensitive text data is hierarchically processed according to the sensitive label level to obtain a hierarchical label data set.
[0054] Step S20: Preprocess the hierarchical label data set to obtain a training data set, and perform model training according to the training data set to obtain a sensitive text classification model.
[0055] The step S20 includes:
[0056] Step S21: Set a preset text length, compare the text lengths of all text data in the hierarchical label data set with the preset text length to obtain a comparison result;
[0057] Step S22: If there is text data in the comparison result whose text length is less than the preset text length, obtain the corresponding text data and perform character padding on the text data to obtain a target hierarchical label data set;
[0058] Step S23: Perform data augmentation on the target hierarchical label dataset to obtain text-augmented data, perform text transformation on the text-augmented data to obtain multiple text sequences, and combine all the text sequences to obtain a training dataset;
[0059] Step S24: Create a sensitive text classification training model and input a set of text data samples in the training dataset into the sensitive text classification training model;
[0060] Step S25: Extract information features from the text data samples to obtain label hierarchical information, and construct positive samples based on the label hierarchical information to obtain predicted sensitive text classification samples;
[0061] Step S26: Calculate the loss value between the predicted sensitive text classification samples and the true sensitive text classification results corresponding to the text data samples to obtain a target loss value, and correct the parameters of the sensitive text classification training model according to the target loss value;
[0062] Step S27: Input the next set of text data samples into the sensitive text classification training model until the training situation of the sensitive text classification training model meets the preset conditions to obtain a trained sensitive text classification model.
[0063] Specifically, after obtaining the hierarchical label dataset, it is necessary to obtain the corresponding training dataset. The specific processing process is as follows: First, corresponding data preprocessing is performed on the text data that needs to enter the subsequent Bert model (Bidirectional Encoder Representations from Transformers, a pre-trained language representation model based on context). Among them, the text data that needs to enter is the hierarchical label dataset used as the training set. The purpose of data preprocessing is to make the corresponding text data meet the conditions for entering the model, that is, a preset text length is set (for example, 128). The text lengths of all text data in the hierarchical label dataset are compared with the preset text length to obtain a comparison result. If there is text data in the comparison result whose text length is less than the preset text length, the corresponding text data is obtained, and character padding is performed on the text data to obtain the target hierarchical label dataset; at the same time, data augmentation needs to be performed on the target hierarchical label dataset, including random homophone replacement, random character deletion, and equivalent character replacement. The purpose is that the text recognized by the OCR (Optical Character Recognition) technology may have misspelled and missing characters, and this situation's impact on text classification accuracy can be alleviated through corresponding data augmentation methods. Specifically, homophone replacement is performed on the target hierarchical label dataset to obtain text replacement data, aiming to interfere with the semantics so that the model is more robust against such data; random character deletion is performed on the text replacement data to obtain the target text replacement data, aiming to enhance the model's anti-interference ability; and equivalent character replacement is performed on the target text replacement data to obtain text augmentation data, aiming to enable the model to learn different representations of numbers; finally, text conversion is performed on the text augmentation data to obtain multiple text sequences. In the present invention, the text augmentation data is converted into three text sequences: token_id, segment_id, and attention mask. Among them, token_id is an id sequence after the text is tokenized after passing through the tokenizer and mapped to the vocabulary. Each id value can query the corresponding phrase in the model vocabulary. After the mapping is completed, [CLS] and [SEP] identifiers are added at the start and end positions of the sequence, and then padding is performed at the end. While segment_id and attention mask are a set of opposite sequences. Among them, for the valid token part in token_id, segment_id is set to 0, and attention mask is set to 1. For the remaining padding characters, segment_id is set to 1, and attention mask is set to 0;Combine all the above text sequences to obtain a training dataset.
[0064] After obtaining the training dataset for model training, the training dataset can be used to train a model to obtain a sensitive text classification model. As Figure 2 shown, first, create a sensitive text classification training model, that is, a 4-layer tinybert model after distillation (tiny Bidirectional Encoder Representations from Transformers model). The purpose of using the 4-layer tinybert model after distillation is that the original bert model has a large number of parameters, high computing power requirements for inference after deployment, and slow model speed. Instead, the original large model is replaced with a tinybert model based on bert distillation to improve the model inference speed. Then, combine the tinybert model with the auxiliary training of hierarchical labels to build a complex network, that is, a sensitive text classification model. The process of training the sensitive text classification training model is as Figure 3 shown. Input a set of text data samples in the training dataset into the sensitive text classification training model. The sensitive text classification training model uses a graph convolutional neural network to extract information features from the text data samples to obtain label hierarchical information, and constructs positive samples using the adversarial attack idea under the guidance of the label hierarchical information to obtain predicted sensitive text classification samples. Among them, the use of the adversarial attack idea is to add FGM (Fast Gradient Method) adversarial training, so that the model has better robustness to texts with OCR recognition errors.
[0065] After that, by using the method of contrastive learning to minimize the loss function, it is necessary to calculate the loss value between the predicted sensitive text classification sample and the true sensitive text classification result corresponding to the text data sample. Specifically, determine the function according to the predicted sensitive text classification sample to obtain the loss function, and calculate the loss value between the predicted sensitive text classification sample and the true sensitive text classification result corresponding to the text data sample according to the loss function to obtain the target loss value. Adjust the regularization term in the loss function to constrain the output of the model and increase the probability that the confidence of the model output is greater than a fixed threshold. Among them, the expression of the loss function is: where L is the loss function, L C is the cross-entropy loss function, is the cross-loss function for generating positive samples, λ is the weight of the contrastive learning loss, L con is the contrastive learning loss, and L C 、 and L conThe definitions are as follows:
[0066]
[0067] Among them, y i,c is the true label of the c-th category in the i-th sample, is the predicted probability of the c-th category in the i-th sample, is the predicted probability of generating positive samples for the c-th category in the i-th sample, N is the size of the dataset, and C is the number of data included in the c-th category, is the contrastive learning loss of the m-th sample, and the nt-xent loss function (Normalized Temperature-scaled Cross Entropy Loss) is adopted.
[0068] And The expression of is: Among them, and x j are different augmented versions of the same input sample, and are both the similarities between different augmented versions of the same input sample, τ is a hyperparameter used to control the smoothness of the distribution, and j is any text data sample in the training dataset.
[0069] After that, the parameters of the sensitive text classification training model are corrected according to the calculated target loss value, so that the prediction error of the sensitive text classification training model gradually decreases during the training process; the next set of text data samples is input into the sensitive text classification training model, and the above process is repeated, which will not be elaborated here; until the training situation of the sensitive text classification training model meets the preset conditions, the preset conditions include that the target loss value meets the preset requirements or the number of training times reaches the preset number of times. The preset requirements can be determined according to the accuracy of the sensitive text classification model, which will not be elaborated here. The preset number of times can be the maximum number of training times of the sensitive text classification training model. For example, 2000 times, etc. Finally, a trained sensitive text classification model is obtained, and the sensitive text classification model can also be deployed to various platforms and devices to verify the performance.
[0070] Step S30, obtain the text data to be recognized of the current Internet content, input the text data to be recognized into the sensitive text classification model, and output the sensitive text classification result.
[0071] Specifically, in the embodiment of the present invention, after obtaining the trained sensitive text classification model, as Figure 4As shown, obtain the text data to be recognized of the current Internet content, input the text data to be recognized into the sensitive text classification model. The sensitive text classification model performs sensitive recognition and classification on the text data to be recognized through the multi-label classification layer, obtains the corresponding sensitive text classification result, and outputs the sensitive text classification result.
[0072] Furthermore, the overall flowchart of the preferred embodiment of the multi-level sensitive text classification method in the present invention is as Figure 5 shown, specifically:
[0073] Step S1: Obtain the sensitive text data of historical Internet content, perform data cleaning on the sensitive text data to obtain the target sensitive text data, and construct a hierarchical label data set according to the target sensitive text data;
[0074] Step S2: Preprocess the hierarchical label data set used as the training set, that is, perform character padding on the text data in the hierarchical label data set with a text length less than the preset text length to obtain the target hierarchical label data set, and perform data augmentation processing such as homophone random replacement, character random deletion, and equivalent character replacement on the target hierarchical label data set, and perform text conversion on the augmented data to obtain the training data set;
[0075] Step S3: Construct a sensitive text classification training model, that is, the distilled 4-layer tinybert model. Then, combine the tinybert model with the auxiliary training of hierarchical labels to build a complex network;
[0076] Step S4: Add FGM (Fast Gradient Method) adversarial training to the sensitive text classification training model, so that the sensitive text classification model has better robustness for texts with OCR recognition errors;
[0077] Step S5: Adjust the regularization term in the loss function to constrain the output of the sensitive text classification training model, so as to increase the probability that the output confidence of the sensitive text classification training model is greater than the fixed threshold;
[0078] Step S6: Train the sensitive text classification training model according to the training data set to obtain the sensitive text classification model;
[0079] Step S7: Deploy the trained sensitive text classification model to each platform and device to verify the performance, that is, obtain the text data to be recognized, input the text data to be recognized into the sensitive text classification model, and output the sensitive text classification result.
[0080] Furthermore, as Figure 6As shown in the figure, based on the above multi-level sensitive text classification method, the present invention also correspondingly provides a multi-level sensitive text classification system, wherein the multi-level sensitive text classification system includes:
[0081] A data processing module 51, configured to obtain sensitive text data of historical Internet content, clean the sensitive text data to obtain target sensitive text data, and construct a hierarchical label data set according to the target sensitive text data;
[0082] A model training module 52, configured to preprocess the hierarchical label data set to obtain a training data set, and perform model training according to the training data set to obtain a sensitive text classification model;
[0083] A sensitive classification module 53, configured to obtain text data to be recognized of current Internet content, input the text data to be recognized into the sensitive text classification model, and output a sensitive text classification result.
[0084] Further, as Figure 7 shown in the figure, based on the above multi-level sensitive text classification method, the present invention also correspondingly provides a terminal, and the terminal includes a processor 10, a memory 20, and a display 30. Figure 7 Only some components of the terminal are shown, but it should be understood that it is not required to implement all the shown components, and more or fewer components can be alternatively implemented.
[0085] In some embodiments, the memory 20 may be an internal storage unit of the terminal, such as a hard disk or memory of the terminal. In other embodiments, the memory 20 may also be an external storage device of the terminal, such as a plug-in hard disk equipped on the terminal, a SmartMedia Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Further, the memory 20 may also include both the internal storage unit and the external storage device of the terminal. The memory 20 is used to store application software installed on the terminal and various types of data, such as program codes installed on the terminal, etc. The memory 20 may also be used to temporarily store data that has been output or will be output. In one embodiment, a multi-level sensitive text classification program 40 is stored on the memory 20, and the multi-level sensitive text classification program 40 can be executed by the processor 10, so as to implement the multi-level sensitive text classification method in the present application.
[0086] In some embodiments, the processor 10 may be a central processing unit (CPU), a microprocessor, or other data processing chips, which are used to run the program code stored in the memory 20 or process data, such as executing the multi-level sensitive text classification method, etc.
[0087] In some embodiments, the display 30 may be an LED display, a liquid crystal display, a touch liquid crystal display, and an OLED (Organic Light-Emitting Diode) touch device, etc. The display 30 is used to display the information of the terminal and to display a visual user interface. The components of the terminal communicate with each other through a system bus.
[0088] In one embodiment, when the processor 10 executes the multi-level sensitive text classification program 40 in the memory 20, the following steps are implemented:
[0089] Obtain the sensitive text data of historical Internet content, clean the sensitive text data to obtain target sensitive text data, and construct a hierarchical label data set according to the target sensitive text data;
[0090] Preprocess the hierarchical label data set to obtain a training data set, and perform model training according to the training data set to obtain a sensitive text classification model;
[0091] Obtain the text data to be recognized of the current Internet content, input the text data to be recognized into the sensitive text classification model, and output the sensitive text classification result.
[0092] Among them, the obtaining the sensitive text data of historical Internet content, cleaning the sensitive text data to obtain target sensitive text data, and constructing a hierarchical label data set according to the target sensitive text data specifically includes;
[0093] Obtain the sensitive text data of historical Internet content, perform deduplication processing on the sensitive text data to obtain deduplicated sensitive text data, and perform annotation processing on the deduplicated sensitive text data to obtain target sensitive text data;
[0094] Set the sensitive label level, and perform grading processing on the target sensitive text data according to the sensitive label level to obtain a hierarchical label data set.
[0095] Among them, the preprocessing the hierarchical label data set to obtain a training data set specifically includes:
[0096] Set a preset text length, compare the text lengths of all text data in the hierarchical label dataset with the preset text length to obtain a comparison result;
[0097] If there is text data in the comparison result whose text length is less than the preset text length, obtain the corresponding text data and perform character padding on the text data to obtain a target hierarchical label dataset;
[0098] Perform data augmentation on the target hierarchical label dataset to obtain text-augmented data, perform text conversion on the text-augmented data to obtain multiple text sequences, and combine all the text sequences to obtain a training dataset.
[0099] Among them, the performing data augmentation on the target hierarchical label dataset to obtain text-augmented data specifically includes:
[0100] Perform homophone replacement on the target hierarchical label dataset to obtain text replacement data, and perform random character deletion on the text replacement data to obtain target text replacement data;
[0101] Perform equivalent character replacement on the target text replacement data to obtain text-augmented data.
[0102] Among them, the training a sensitive text classification model according to the training dataset specifically includes:
[0103] Create a sensitive text classification training model and input a set of text data samples in the training dataset into the sensitive text classification training model;
[0104] Extract information features from the text data samples to obtain label hierarchical information, and construct positive samples according to the label hierarchical information to obtain predicted sensitive text classification samples;
[0105] Calculate the loss value between the predicted sensitive text classification sample and the true sensitive text classification result corresponding to the text data sample to obtain a target loss value, and correct the parameters of the sensitive text classification training model according to the target loss value;
[0106] Input the next set of text data samples into the sensitive text classification training model until the training situation of the sensitive text classification training model meets the preset conditions to obtain a trained sensitive text classification model.
[0107] Among them, the calculating the loss value between the predicted sensitive text classification sample and the true sensitive text classification result corresponding to the text data sample to obtain a target loss value, and correcting the parameters of the sensitive text classification training model according to the target loss value specifically includes:
[0108] Determine a function based on the predicted sensitive text classification samples to obtain a loss function, and calculate the loss value between the predicted sensitive text classification samples and the true sensitive text classification results corresponding to the text data samples according to the loss function to obtain a target loss value;
[0109] Modify the parameters of the sensitive text classification training model according to the target loss value.
[0110] Among them, the expression of the loss function is:
[0111]
[0112] Among them, L is the loss function, L C is the cross-entropy loss function, is the cross-loss function for generating positive samples, λ is the weight of the contrast learning loss, L con is the contrast learning loss.
[0113] The present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a multi-level sensitive text classification program, and when the multi-level sensitive text classification program is executed by a processor, the steps of the multi-level sensitive text classification method described above are implemented.
[0114] In summary, the present invention provides a multi-level sensitive text classification method, system, terminal and storage medium. The method includes: obtaining sensitive text data of historical Internet content, cleaning the sensitive text data to obtain target sensitive text data, and constructing a hierarchical label data set according to the target sensitive text data; preprocessing the hierarchical label data set to obtain a training data set, and training a model according to the training data set to obtain a sensitive text classification model; obtaining text data to be recognized of current Internet content, inputting the text data to be recognized into the sensitive text classification model, and outputting a sensitive text classification result. The recognition and classification of the present invention have a wide coverage range and a high degree of subdivision, can also recognize the meaning and reference of the text, and can also consider the context information of the text, improve the classification accuracy, and require less resources and have a fast response speed during processing.
[0115] It should be noted that in this article, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or terminal including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or elements inherent to such a process, method, article or terminal. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or terminal including that element.
[0116] Of course, those of ordinary skill in the art can understand that all or part of the processes of implementing the above-described embodiment methods can be completed by instructing relevant hardware (such as a processor, a controller, etc.) through a computer program. The program can be stored in a computer-readable storage medium readable by a computer. When the program is executed, it can include the processes of the above-described method embodiments. The computer-readable storage medium can be a memory, a magnetic disk, an optical disk, etc.
[0117] It should be understood that the application of the present invention is not limited to the above examples. For those of ordinary skill in the art, improvements or transformations can be made according to the above description. All such improvements and transformations should fall within the protection scope of the appended claims of the present invention.
Claims
1. A multi-level sensitive text classification method, characterized in that: The multi-level sensitive text classification method includes: Acquire sensitive text data of historical Internet content, perform data cleaning on the sensitive text data to obtain target sensitive text data, and construct a hierarchical label data set based on the target sensitive text data; Preprocessing the hierarchical label data set to obtain a training data set, and performing model training based on the training data set to obtain a sensitive text classification model; The text data to be identified of the current Internet content is obtained, the text data to be identified is input into a sensitive text classification model, and a sensitive text classification result is output.
2. The multi-level sensitive text classification method according to claim 1 is characterized in that: The step of acquiring sensitive text data of historical Internet content, performing data cleaning on the sensitive text data to obtain target sensitive text data, and constructing a hierarchical label data set according to the target sensitive text data specifically includes: Acquire sensitive text data of historical Internet content, perform deduplication processing on the sensitive text data to obtain deduplication sensitive text data, and perform annotation processing on the deduplication sensitive text data to obtain target sensitive text data; A sensitivity label level is set, and the target sensitive text data is graded according to the sensitivity label level to obtain a hierarchical label data set.
3. The multi-level sensitive text classification method according to claim 1, characterized in that: The preprocessing of the hierarchical label data set to obtain a training data set specifically includes: Setting a preset text length, comparing the text lengths of all text data in the hierarchical label data set with the preset text length to obtain a comparison result; If the text length of the text data in the comparison result is less than the preset text length, the corresponding text data is obtained, and the text data is character-filled to obtain a target level label data set; The target level label data set is data enhanced to obtain text enhanced data, the text enhanced data is text converted to obtain multiple text sequences, and all the text sequences are combined to obtain a training data set.
4. The multi-level sensitive text classification method according to claim 3 is characterized in that: The data enhancement of the target level label data set to obtain text enhancement data specifically includes: Performing homophone replacement on the target level label data set to obtain text replacement data, and performing random character deletion on the text replacement data to obtain target text replacement data; The target text replacement data is replaced with equivalent words to obtain text enhancement data.
5. The multi-level sensitive text classification method according to claim 2, characterized in that: The performing model training according to the training data set to obtain a sensitive text classification model specifically includes: Creating a sensitive text classification training model, and inputting a set of text data samples in the training data set into the sensitive text classification training model; Extracting information features from the text data sample to obtain label level information, and constructing positive samples based on the label level information to obtain predicted sensitive text classification samples; Calculating the loss value of the predicted sensitive text classification sample and the real sensitive text classification result corresponding to the text data sample to obtain a target loss value, and modifying the parameters of the sensitive text classification training model according to the target loss value; The next set of text data samples is input into the sensitive text classification training model until the training status of the sensitive text classification training model meets the preset conditions, thereby obtaining a trained sensitive text classification model.
6. The multi-level sensitive text classification method according to claim 5, characterized in that: The calculating of the loss value of the predicted sensitive text classification sample and the real sensitive text classification result corresponding to the text data sample to obtain a target loss value, and correcting the parameters of the sensitive text classification training model according to the target loss value specifically includes: Performing function determination according to the predicted sensitive text classification sample to obtain a loss function, and calculating a loss value between the predicted sensitive text classification sample and a true sensitive text classification result corresponding to the text data sample according to the loss function to obtain a target loss value; The parameters of the sensitive text classification training model are modified according to the target loss value.
7. The multi-level sensitive text classification method according to claim 6, characterized in that: The expression of the loss function is: Among them, L is the loss function, L C is the cross entropy loss function, is the cross loss function for generating positive samples, λ is the weight of contrastive learning loss, L con is the contrastive learning loss.
8. A multi-level sensitive text classification system, characterized in that: The multi-level sensitive text classification system includes: A data processing module is used to obtain sensitive text data of historical Internet content, perform data cleaning on the sensitive text data to obtain target sensitive text data, and construct a hierarchical label data set based on the target sensitive text data; A model training module is used to preprocess the hierarchical label data set to obtain a training data set, and perform model training based on the training data set to obtain a sensitive text classification model; The sensitive classification module is used to obtain the text data to be identified in the current Internet content, input the text data to be identified into the sensitive text classification model, and output the sensitive text classification result.
9. A terminal, characterized in that: The terminal includes: a memory, a processor, and a multi-level sensitive text classification program stored in the memory and executable on the processor. When the multi-level sensitive text classification program is executed by the processor, the steps of the multi-level sensitive text classification method as described in any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a multi-level sensitive text classification program, and when the multi-level sensitive text classification program is executed by a processor, the steps of the multi-level sensitive text classification method as described in any one of claims 1-7 are implemented.