A method and system for detecting bad Chinese speech
By adopting multi-model consistency voting strategy, data enhancement, R-Drop regularization, dual-channel classification tasks and contrast learning module technical means in the detection of Chinese bad speech, the problem of keyword matching in the existing technology cannot handle context and implicit meaning, and a more efficient and accurate detection of Chinese bad speech is achieved.
Patent Information
- Application Number
- CN202411977154.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-20
- Estimated Expiration
- 2044-12-31
AI Technical Summary
The existing Chinese bad speech detection methods mainly rely on keyword matching, and cannot effectively deal with context and implicit meaning, which may cause misjudgment and misjudgment. Traditional models often lack understanding of the deep semantics of text, making it difficult to identify more subtle or variant forms of bad content, resulting in low detection efficiency.
A multi-model consistency voting strategy, data enhancement module, R-Drop regularization module, dual-channel classification task module and comparison learning module are used to build a Chinese bad speech detection model. Through these technical means, the accuracy and robustness of the model are improved, and the understanding and contrast learning ability of the deep semantics of the text are enhanced.
It improves the accuracy and reliability of Chinese bad speech detection, reduces the possible misjudgment caused by a single model, enhances the generalization ability of the model and the ability to identify subtle differences, and improves the sensitivity and accuracy of the overall detection.
Smart Images

Figure CN119377415B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of speech detection, and particularly to a method and system for detecting Chinese bad speech. Background Art
[0002] The method for detecting Chinese bad speech is mainly used to identify and filter Chinese texts containing inappropriate content, which combines keyword matching, machine learning classifiers, natural language processing technologies, and multi-modal information analysis, aiming to accurately identify bad speech by understanding the semantics and context of the text, as well as possible image and audio information. With the development of technology, these detection technologies are becoming more and more intelligent and automated, effectively helping to manage and filter bad information in the network environment.
[0003] However, the existing methods for detecting Chinese bad speech mainly rely on keyword matching, and cannot effectively handle context and implicit meanings, which may lead to misjudgment and missed judgment. Traditional models often lack the understanding of the deep semantics of the text and are difficult to identify more subtle or variant forms of bad content, resulting in low efficiency of detecting Chinese bad speech. Summary of the Invention
[0004] In order to solve the technical problems that the existing methods for detecting Chinese bad speech mainly rely on keyword matching, cannot effectively handle context and implicit meanings, may lead to misjudgment and missed judgment, and traditional models often lack the understanding of the deep semantics of the text and are difficult to identify more subtle or variant forms of bad content, resulting in low efficiency of detecting Chinese bad speech, the present invention provides a method and system for detecting Chinese bad speech.
[0005] The technical solutions provided by the embodiments of the present invention are as follows:
[0006] First aspect:
[0007] A method for detecting Chinese bad speech provided by an embodiment of the present invention includes:
[0008] S1: Obtain an initial tweet dataset containing bad speech;
[0009] S2: Preprocess the initial tweet dataset;
[0010] S3: Use a multi-model consistency voting strategy to classify and label the preprocessed initial tweet dataset to obtain a Chinese bad speech dataset;
[0011] S4: Construct a Chinese bad speech detection model, where the Chinese bad speech detection model includes a data augmentation module, an R-Drop regularization module, a dual-channel classification task module, and a contrast learning module connected in sequence;
[0012] S5: Input the Chinese offensive speech dataset into the Chinese offensive speech detection model for training;
[0013] S6: Obtain the real-time Chinese offensive speech dataset;
[0014] S7: Input the Chinese offensive speech dataset into the trained Chinese offensive speech detection model and output the Chinese offensive speech detection result.
[0015] Second aspect:
[0016] A Chinese offensive speech detection system provided by an embodiment of the present invention includes:
[0017] A processor;
[0018] A memory, on which computer-readable instructions are stored. When the computer-readable instructions are executed by the processor, the Chinese offensive speech detection method as in the first aspect is implemented.
[0019] Third aspect:
[0020] A computer-readable storage medium provided by an embodiment of the present invention, on which a computer program is stored. When the program is executed by the processor, the Chinese offensive speech detection method as in the first aspect is implemented.
[0021] The beneficial effects brought by the technical solutions provided by the embodiments of the present invention at least include:
[0022] In the present invention, by using the multi-model consistency voting strategy, the detection accuracy and reliability can be improved, which helps to balance the biases of different models and reduce the misjudgment that may be brought by a single model. Through the data augmentation module, the diversity of training samples can be expanded, and the R-Drop regularization module can achieve better generalization on the diverse data of the model, reducing overfitting. Through the dual-channel classification task module, the subtle differences in offensive speech can be effectively captured, and the differences can be deeply understood from different perspectives, enhancing the robustness of the model and improving the final classification effect, ensuring that the model can comprehensively capture the essential features of the input samples from multiple angles. Through the contrast learning module, the subtle differences between offensive speech and normal speech can be better learned and distinguished, improving the sensitivity and accuracy of the overall detection. Description of the Drawings
[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0024] Figure 1A flowchart of a Chinese bad speech detection method provided by an embodiment of the present invention;
[0025] Figure 2 A structural diagram of a Chinese bad speech detection model provided by an embodiment of the present invention;
[0026] Figure 3 A structural diagram of a Chinese bad speech detection system provided by an embodiment of the present invention. The technical solutions in the present invention will be described below with reference to the accompanying drawings.
[0027] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as an "example" in the present invention should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of the word "example" is intended to present concepts in a specific manner. In addition, in the embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one of the two.
[0028] In the embodiments of the present invention, "image" and "picture" can sometimes be used interchangeably. It should be noted that when the difference is not emphasized, their intended meanings are the same. "Of", "corresponding", and "corresponding" can sometimes be used interchangeably. It should be noted that when the difference is not emphasized, their intended meanings are the same.
[0029] In the embodiments of the present invention, sometimes a subscript such as W 1 may be written in a non-subscript form such as W1. When the difference is not emphasized, their intended meanings are the same.
[0030] To make the technical problems, technical solutions and advantages to be solved by the present invention clearer, the following will be described in detail with reference to the accompanying drawings and specific embodiments.
[0031] Refer to the attached Figure 1 illustrates a flowchart of a Chinese bad speech detection method provided by an embodiment of the present invention.
[0032] Refer to the attached Figure 2 illustrates a structural diagram of a Chinese bad speech detection model provided by an embodiment of the present invention.
[0033] As Figure 2, the Chinese bad speech detection model adopts a dual-channel BERT architecture. The two paths share parameters, extract features from the input text respectively, generate different representations, calculate the contrast loss through the contrast learning module to improve the model's recognition ability of the differences between features. At the same time, each of the two channels generates classification predictions to generate classification losses. Finally, by combining the contrast loss and the classification loss, the model is jointly optimized to more accurately classify and detect bad speech.
[0034] An embodiment of the present invention provides a Chinese bad speech detection method, and the method includes:
[0035] S1: Obtain an initial tweet dataset containing bad speech.
[0036] In a possible implementation manner, by obtaining an initial tweet dataset containing bad speech, the characteristics of bad speech in the actual social network can be better reflected.
[0037] In a possible implementation manner, S1 specifically includes:
[0038] S101: According to a preset list of bad speech sensitive words, search through the Twitter API, and filter out the tweets containing sensitive words.
[0039] Among them, the Twitter API is an application programming interface provided by the Twitter platform, which allows programs to interact with Twitter data, such as searching for tweets, obtaining user information, etc.
[0040] It should be noted that searching through the Twitter API according to the sensitive word list can quickly filter out the tweets containing sensitive content, quickly locate potential bad speech, save a large amount of time and human resources for subsequent data annotation and model training, and at the same time ensure that the preliminary data is targeted and reduce the noise of irrelevant data.
[0041] S102: Conduct preliminary manual annotation on the tweets to obtain annotated tweets.
[0042] Specifically, manual annotation can ensure that the meaning of sensitive words in a specific context is correctly understood, avoid misjudgment and missed judgment, thereby providing a high-quality data basis for subsequent model training and improving the reliability and accuracy of the model in dealing with complex semantics and implicit expressions.
[0043] S103: Record the usernames of the authors of the annotated tweets to form a list of high-risk users of bad speech.
[0044] S104: According to the list of high-risk users of bad speech, capture the historical tweets of high-risk users of bad speech through the Twitter API to form an initial tweet data.
[0045] S2: Preprocess the initial tweet dataset.
[0046] It should be noted that preprocessing the initial tweet dataset can clean up noisy data, improve data consistency and normalization, thereby enhancing the model's ability to extract valid information and reducing misjudgments caused by data quality differences.
[0047] In a possible implementation, S2 is specifically as follows:
[0048] Perform preprocessing on the initial tweet dataset including data cleaning and formatting.
[0049] Specifically, performing preprocessing including data cleaning and formatting specifically includes removing non-Chinese content in the tweets, including English characters, numbers, symbols, etc., and only retaining Chinese characters, cleaning up irrelevant information in the tweets, including deleting URL links, @ user information, and emojis, converting traditional Chinese in the tweets to simplified Chinese, unifying the language format, filtering short text content, and only retaining tweets with a length of 5 characters or more to form a structured high-quality tweet dataset.
[0050] S3: Use the multi-model consensus voting strategy to classify and label the preprocessed initial tweet dataset to obtain a Chinese bad speech dataset.
[0051] Among them, the multi-model consensus voting strategy refers to using multiple different models to predict the data at the same time, voting through the prediction results of each model, and when most models give the same classification result, using it as the final classification result to improve the accuracy and stability of the judgment.
[0052] It should be noted that using the multi-model consensus voting strategy to classify and label the preprocessed tweet dataset can effectively reduce the bias brought by a single model and improve the reliability of the classification result.
[0053] In a possible implementation, S3 specifically includes:
[0054] S301: Divide the preprocessed initial tweet dataset into a first dataset and a second dataset.
[0055] S302: Manually label the first dataset to form an initial labeled dataset.
[0056] S303: Select multiple Chinese pre-trained models, where the Chinese pre-trained models include the BERT-Base model, the RoBERTa-wwm-ext model, the MacBERT model, and the Chinese-ELECTRA model.
[0057] Specifically, the BERT-Base model is a pre-trained language model based on the bidirectional Transformer structure for capturing context information. The RoBERTa-wwm-ext model is an enhanced version of BERT that improves language understanding ability through the "whole word masking" strategy. The MacBERT model is an improved version of BERT that focuses on optimizing text matching and text classification tasks. The Chinese-ELECTRA model is a new type of pre-trained model that improves training efficiency by replacing the traditional generation task with a discriminative task.
[0058] S304: Based on the initial labeled dataset, fine-tune each Chinese pre-trained model for the binary classification task.
[0059] S305: Input the second dataset into each fine-tuned Chinese pre-trained model for automatic classification to generate the classification results and the consistency ratio of the second dataset.
[0060] S306: Determine whether the consistency ratio is greater than the preset ratio. If so, output the classification results as the automatic annotation results. Otherwise, perform manual annotation on the second dataset to form the final labeled dataset.
[0061] In actual operation, if 75% or more of the prediction results are consistent among the four models, directly adopt the consistent results as the annotation results of the tweet. For tweets with a consistency lower than 75%, submit them to the manual annotation team for manual annotation to ensure the accuracy of the annotation results.
[0062] S307: Add the final labeled dataset to the initial labeled dataset, and repeat steps S303 to S307 until the annotation of the high-quality tweet dataset is completed.
[0063] S308: Output the Chinese offensive speech dataset.
[0064] S4: Construct a Chinese offensive speech detection model, where the Chinese offensive speech detection model includes a data augmentation module, an R-Drop regularization module, a dual-channel classification task module, and a contrastive learning module connected in sequence.
[0065] It should be noted that by integrating the data augmentation module, the R-Drop regularization module, the dual-channel classification task module, and the contrastive learning module, the model is comprehensively improved in terms of sample richness, prediction stability, classification accuracy, and feature robustness, effectively enhancing the detection ability and adaptability to Chinese offensive speech.
[0066] S5: Input the Chinese offensive speech dataset into the Chinese offensive speech detection model for training.
[0067] It should be noted that by training the model, the model can fully learn the data features, thereby improving the accuracy and generalization ability of detection.
[0068] In one possible implementation, the data enhancement module includes a BERT unit, and the BERT unit includes a Dropout algorithm.
[0069] Specifically, the BERT unit is responsible for extracting contextual features from text to improve text understanding and processing capabilities. The Dropout algorithm is a regularization method that randomly discards some neurons during training to prevent overfitting and improve the generalization ability of the model. The Dropout algorithm is built into the Bert unit. By randomly "closing" some neurons, the model will not overly rely on certain specific neurons during training, thereby effectively reducing the risk of overfitting. By randomly "closing" neurons, different feature representations will be output for the same sample. The same sample is input into the BERT unit, and the output features are enhanced samples for each other.
[0070] S5 specifically includes:
[0071] S501: Based on the Chinese bad speech dataset, feature extraction is performed through the BERT unit to obtain data features.
[0072] It should be noted that by extracting features from the Chinese bad speech dataset through the BERT unit, we can deeply mine the contextual information and semantic features of the text, thereby obtaining high-quality data representation.
[0073] In a possible implementation manner, S501 specifically includes:
[0074] S5011: Load Chinese bad speech dataset:
[0075]
[0076] Where D represents the bad speech dataset, x i indicates the first in the bad speech dataset i input sentences, y i indicates x i The corresponding real label;
[0077] Specifically, y i The value of is a binary category, 0 represents "normal speech" and 1 represents "bad speech".
[0078] S5012: Perform text preprocessing on the Chinese bad speech dataset, including word segmentation, special marking, and conversion to the BERT format to obtain text data. Among them, the special marking includes adding a marker CLS at the beginning of the sentence in the Chinese bad speech dataset and adding a marker SEP at the end of the sentence in the Chinese bad speech dataset.
[0079] S5013: Divide the text data according to a preset size to generate batch data.
[0080] S5014: Input the batch data into the BERT unit, and extract the output vector containing CLS as the data feature:
[0081]
[0082] Among them, h CLS represents the data feature, f BERT represents the feature extraction function, input_ids represents the vocabulary index sequence of the input sentence, attention_mask represents the valid input marker, and token_type_ids represents the sentence segment distinction.
[0083] S502: Use the Dropout algorithm to perform data augmentation on the feature data to generate positive sample pairs.
[0084] In a possible implementation manner, S502 specifically includes:
[0085] S5021: Use the Dropout algorithm to perform double encoding on the data feature to generate the first feature representation and the second feature representation:
[0086]
[0087] Among them, represents the first feature representation, represents the second feature representation, x represents the input sentence, represents the first Dropout mask, represents the second Dropout mask.
[0088] S5022: Use the first feature representation and the second feature representation as different enhanced features of the same sentence in the Chinese bad speech dataset to form positive sample pairs for contrast learning. Among them, the first feature representation and the second feature representation are correlated.
[0089] S503: Input the positive sample pair into the classifier of the R-Drop regularization module, and output the predicted distribution of the positive sample pair:
[0090]
[0091] Among them, represents the first predicted distribution, represents the softmax function, represents the second predicted distribution, W represents the weight matrix of the classifier, b represents the bias term of the classifier, represents the first feature representation, represents the second feature representation.
[0092] S504: Calculate the KL divergence loss and the cross-entropy loss according to the predicted distribution.
[0093] Among them, the KL divergence is a metric for measuring the difference between two probability distributions, used to evaluate the matching degree between the model's predicted distribution and the target distribution. The cross-entropy loss is a loss function for measuring the difference between the predicted probability distribution and the true distribution (label), and is commonly used in classification tasks.
[0094] It should be noted that by combining the KL divergence loss and the cross-entropy loss, the accuracy and consistency of the predicted distribution can be optimized simultaneously, enhancing the classification stability and robustness of the model.
[0095] S505: Input the positive sample pair into the dual-channel classification task module to generate the final classification result.
[0096] In a possible implementation, S505 specifically includes:
[0097] S5051: Input the positive sample pair into the dual-channel classification task module to generate the first classification probability distribution and the second classification probability distribution.
[0098]
[0099] Among them, represents the first classification probability distribution, represents the second classification probability distribution, W 1 and W 2 respectively represent the weight matrices of the first-channel and second-channel classifiers, b 1 and b 2 respectively represent the bias terms of the first-channel and second-channel classifiers.
[0100] S5052: Combine the first classification probability distribution and the second classification probability distribution to generate the final classification result:
[0101]
[0102] where, represents the final classification result of the i -th category, represents the first classification probability distribution of the i -th category, represents the second classification probability distribution of the i -th category, represents the final classification result of the j -th category.
[0103] S506: Generate the contrastive learning loss through the ratio learning module:
[0104]
[0105] L cl represents the contrastive learning loss, sim represents the similarity between features, i = 1, 2, ···, n , n represents the number of input sentences participating in the loss calculation in the Chinese bad speech dataset τ represents the temperature hyperparameter, J= 1, 2, ···, N , N represents the total number of input sentences in the Chinese bad speech dataset, h j represents the negative sample.
[0106] S507: Combine the cross-entropy loss, KL divergence loss, and contrastive learning loss to construct the total loss function:
[0107]
[0108] where, L total represents the total loss function, L kl represents the KL divergence loss, L ce represents the cross-entropy loss represents the hyperparameter.
[0109] In a possible implementation, the KL divergence loss is specifically:
[0110]
[0111] where, Lkl represents the KL divergence loss, D KL represents the KL divergence, i = 1, 2, ···, n , n represents the number of input sentences in the Chinese bad speech dataset that participate in the loss calculation, P ( y ) represents the probability distribution P over the class y and the predicted probability value, Q ( y ) represents the probability distribution Q and the predicted probability value over the class y.
[0112] The cross-entropy loss is specifically:
[0113]
[0114] where, L ce represents the cross-entropy loss, m represents, x i represents the i th input sentence in the Chinese bad speech dataset, y i represents x i the corresponding true label, i = 1, 2, ···, n , n represents the number of input sentences in the Chinese bad speech dataset that participate in the loss calculation.
[0115] S508: Use the gradient descent optimization algorithm to adjust the parameters of the Chinese bad speech detection model until the total loss function value is less than the preset loss function value.
[0116] S6: Obtain the real-time Chinese bad speech dataset.
[0117] It should be noted that by obtaining the latest Chinese bad speech data in real time, it can ensure that the model can adapt to the dynamically changing language environment and improve the timeliness and coverage of detection.
[0118] S7: Input the Chinese bad speech dataset into the trained Chinese bad speech detection model and output the Chinese bad speech detection result.
[0119] The present invention performs BERT encoding twice on the same sentence, and each encoding uses Dropout randomness to generate different sentence embeddings, which are used as positive sample pairs for contrastive learning to achieve data-level enhancement. In the training process, the R-Drop strategy is adopted to effectively utilize the randomness of Dropout by minimizing the bidirectional KL divergence between the model output distributions after two independent Dropout operations, and significantly enhance the robustness and generalization ability of the model. By designing dual-channel classification task constraint feature learning, it is ensured that the model can learn consistent and effective classification information from different feature representations. The contrastive learning mechanism is introduced to bring positive sample pairs closer in the feature space and push away other negative samples, prompting the model to learn more essential and discriminative feature representations.
[0120] The beneficial effects of the technical solution provided by the embodiment of the present invention include at least:
[0121] In the present invention, the multi-model consistency voting strategy is used to improve the accuracy and reliability of detection, help balance the deviations of different models, and reduce the misjudgment that may be caused by a single model. The data enhancement module can expand the diversity of training samples. The R-Drop regularization module can achieve better generalization on the diverse data of the model and reduce overfitting. The dual-channel classification task module can effectively capture the subtle differences in bad speech, deeply understand the differences from different perspectives, enhance the robustness of the model, and improve the final classification effect. Ensure that the model can comprehensively capture the essential characteristics of the input samples from multiple angles. Through the comparative learning module, it can better learn and distinguish the subtle differences between bad speech and normal speech, and improve the sensitivity and accuracy of the overall detection.
[0122] Reference Manual Attached Figure 3 , shows a schematic diagram of the structure of a Chinese bad speech detection system provided by the present invention.
[0123] The present invention also provides a Chinese bad speech detection system 20, which is applied to the above-mentioned Chinese bad speech detection method, including:
[0124] Processor 201.
[0125] Memory 202, memory 202 stores computer-readable instructions, and when the computer-readable instructions are executed by processor 201, the Chinese bad speech detection method of the method embodiment is implemented.
[0126] The Chinese bad speech detection system 20 provided by the present invention can execute the above-mentioned Chinese bad speech detection method and achieve the same or similar technical effects. To avoid repetition, the present invention will not be described in detail.
[0127] The beneficial effects of the technical solution provided by the embodiment of the present invention include at least:
[0128] In the present invention, by using the multi-model consistency voting strategy, the accuracy and reliability of detection can be improved, which helps to balance the biases of different models and reduce the misjudgments that may be brought about by a single model. Through the data augmentation module, the diversity of training samples can be expanded. The R-Drop regularization module can achieve better generalization on the diverse data of the model and reduce overfitting. Through the dual-channel classification task module, the subtle differences in bad remarks can be effectively captured, the differences can be deeply understood from different perspectives, the robustness of the model can be enhanced, and the final classification effect can be improved, ensuring that the model can comprehensively capture the essential features of the input samples from multiple perspectives. Through the contrastive learning module, the subtle differences between bad remarks and normal remarks can be better learned and distinguished, and the sensitivity and accuracy of the overall detection can be improved.
[0129] It should be understood that the processor in the embodiments of the present invention may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0130] It should also be understood that the memory in the embodiments of the present invention may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0131] The above embodiments can be implemented in whole or in part by software, hardware (such as circuits), firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. A computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions according to the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wired (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that contains one or more collections of available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, or magnetic tape), an optical medium (such as a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.
[0132] It should be understood that the term "and / or" in this document is merely a description of the association relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B can be singular or plural. In addition, the character " / " in this document generally represents an "or" relationship between the associated objects before and after, but it may also represent an "and / or" relationship, which can be specifically understood by referring to the context.
[0133] In the present invention, "at least one" means one or more, and "a plurality" means two or more. "At least one of the following" or its similar expressions refer to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c can be single or multiple.
[0134] It should be understood that in various embodiments of the present invention, the magnitudes of the sequence numbers of the above processes do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0135] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in connection with the embodiments disclosed herein can be implemented in electronic hardware, or in a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Skilled artisans may use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0136] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the devices, apparatuses, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.
[0137] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.
[0138] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0139] In addition, the functional units in the various embodiments of the present invention can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.
[0140] If a function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0141] An embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the Chinese bad speech detection method as in the method embodiment.
[0142] The computer-readable storage medium provided by the present invention can implement the steps and effects of the Chinese bad speech detection method in the above method embodiment. To avoid repetition, the present invention will not elaborate further.
[0143] The beneficial effects brought by the technical solution provided by the embodiments of the present invention at least include:
[0144] In the present invention, by using the multi-model consistency voting strategy, the accuracy and reliability of detection can be improved, which helps to balance the biases of different models and reduce the misjudgment that may be caused by a single model. Through the data augmentation module, the diversity of training samples can be extended. The R-Drop regularization module can achieve better generalization on the diverse data of the model and reduce overfitting. Through the dual-channel classification task module, the subtle differences in bad speech can be effectively captured, and the differences can be deeply understood from different perspectives, enhancing the robustness of the model and improving the final classification effect, ensuring that the model can comprehensively capture the essential features of the input samples from multiple angles. Through the contrastive learning module, the subtle differences between bad speech and normal speech can be better learned and distinguished, improving the sensitivity and accuracy of the overall detection.
[0145] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
[0146] The following points need to be explained:
[0147] (1)The attached drawings of the embodiments of the present invention only relate to the structures involved in the embodiments of the present invention, and other structures can refer to the general design.
[0148] (2)For clarity, in the attached drawings used to describe the embodiments of the present invention, the thickness of layers or regions is enlarged or reduced, that is, these drawings are not drawn to actual scale. It can be understood that when an element such as a layer, film, region or substrate is referred to as being "on" or "under" another element, the element can be "directly" on or under the other element or there can be intermediate elements.
[0149] (3)Without conflict, the embodiments of the present invention and the features in the embodiments can be combined with each other to obtain new embodiments.
[0150] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. The protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. A method for detecting bad Chinese speech, characterized in that: include: S1: Get the initial tweet dataset containing bad words; S2: preprocessing the initial tweet dataset; S3: Use the multi-model consensus voting strategy to classify and annotate the preprocessed initial tweet dataset to obtain the Chinese bad speech dataset; S4: constructing a Chinese bad speech detection model, wherein the Chinese bad speech detection model includes a data enhancement module, an R-Drop regularization module, a dual-channel classification task module and a contrastive learning module connected in sequence; S5: inputting the Chinese bad speech data set into the Chinese bad speech detection model for training; S6: Obtain a real-time dataset of Chinese offensive speech; S7: inputting the Chinese bad speech data set into the trained Chinese bad speech detection model, and outputting the Chinese bad speech detection result; Wherein, the data enhancement module includes a BERT unit, and the BERT unit includes a Dropout algorithm; The S5 specifically includes: S501: Based on the Chinese bad speech dataset, feature extraction is performed by the BERT unit to obtain data features; S502: Using the Dropout algorithm, perform data enhancement on the data features to generate positive sample pairs; The S502 specifically includes: S5021: Using the Dropout algorithm, double-encode the data features to generate a first feature representation and a second feature representation, wherein: represents the first feature representation, represents the second feature representation; S5022: using the first feature representation and the second feature representation as different enhanced features of the same sentence in the Chinese bad speech dataset to form a positive sample pair for contrastive learning, wherein the first feature representation and the second feature representation are correlated; S503: Input the positive sample pair into the classifier of the R-Drop regularization module, and output the predicted distribution of the positive sample pair: ; in, represents the first predictive distribution, represents the normalized exponential function, represents the second predictive distribution, W represents the weight matrix of the classifier, b represents the bias term of the classifier, represents the first feature representation, represents the second feature representation; S504: Calculate KL divergence loss and cross entropy loss according to the predicted distribution; S505: Inputting the positive sample pair into the dual-channel classification task module to generate a final classification result; S506: Generate a contrastive learning loss through the contrastive learning module: ; Said L cl represents contrastive learning loss, sim represents the similarity between features, i =1,2,···, n , n Represents the number of input sentences involved in the loss calculation in the Chinese bad speech dataset τ represents the temperature hyperparameter, J= 1,2,···, N , N Represents the total number of input sentences in the Chinese bad speech dataset, h j represents negative samples; S507: Combining the cross entropy loss, the KL divergence loss and the contrastive learning loss, construct a total loss function: ; in, L total represents the total loss function, L kl represents the KL divergence loss, L ce represents the cross entropy loss, represents a hyperparameter; S508: Using a gradient descent optimization algorithm to adjust the parameters of the Chinese bad speech detection model until the total loss function value is less than a preset loss function value.
2. The method for detecting bad Chinese speech according to claim 1, characterized in that: The S1 specifically includes: S101: Based on the preset list of sensitive words for bad speech, search through the Twitter API to filter out tweets containing sensitive words; S102: Performing preliminary manual labeling on the tweets to obtain labeled tweets; S103: Record the user name of the author of the marked tweet to form a list of high-risk users of bad speech; S104: According to the list of high-risk users for bad speech, historical tweets of high-risk users for bad speech are captured through Twitter API to form initial tweet data.
3. The method for detecting bad Chinese speech according to claim 1, characterized in that: The S2 is specifically: The initial tweet dataset is preprocessed including data cleaning and formatting.
4. The method for detecting bad Chinese speech according to claim 1, characterized in that: The S3 specifically includes: S301: Divide the preprocessed initial tweet dataset into a first dataset and a second dataset; S302: manually annotating the first data set to form an initial annotated data set; S303: Select multiple Chinese pre-training models, wherein the Chinese pre-training models include a BERT-Base model, a RoBERTa-wwm-ext model, a MacBERT model, and a Chinese-ELECTRA model; S304: fine-tuning each Chinese pre-training model for a binary classification task based on the initial annotated data set; S305: Inputting the second data set into each fine-tuned Chinese pre-training model for automatic classification, and generating a classification result and consistency ratio of the second data set; S306: Determine whether the consistency ratio is greater than a preset ratio, if so, output the classification result as an automatic annotation result, otherwise, manually annotate the second data set to form a final annotated data set; S307: adding the final annotated data set to the initial annotated data set, and repeating steps S303 to S307 until the preprocessed initial tweet data set is annotated; S308: Outputting the Chinese offensive speech dataset.
5. The method for detecting bad Chinese speech according to claim 1, characterized in that: The S501 specifically includes: S5011: Load the Chinese bad speech dataset: ; Among them, D represents the Chinese bad speech dataset, x i Indicates the first i Input sentences, y i express x i The corresponding true label; S5012: Performing text preprocessing including word segmentation, special marking and conversion to BERT format on the Chinese bad speech dataset to obtain text data, wherein the special marking includes adding a mark [ CLS ] and adding tags to the end of sentences in the Chinese bad speech dataset [ SEP ]; S5013: Divide the text data into pieces according to a preset size to generate batch data; S5014: Input the batch data into the BERT unit, extract the data including [ CLS ] The output vector of the mark is used as the data feature: ; in, h CLS Represents data characteristics, f BERT Represents the feature extraction function, input_ids represents the vocabulary index sequence of the input sentence, attention_mask represents the valid input token, and token_type_ids represents the distinction between sentence segments.
6. The method for detecting bad Chinese speech according to claim 1, characterized in that: The S502 specifically includes: S5021: Using the Dropout algorithm, double-encode the data features to generate a first feature representation and a second feature representation: ; in, represents the first feature representation, represents the second feature representation, x represents the input sentence, represents the first Dropout mask, Represents the second Dropout mask.
7. The method for detecting bad Chinese speech according to claim 1, characterized in that: The S505 specifically includes: S5051: Input the positive sample pair into the dual-channel classification task module to generate a first classification probability distribution and a second classification probability distribution: ; in, represents the first classification probability distribution, represents the second classification probability distribution, W 1 and W 2 represents the weight matrix of the first channel and the second channel classifier respectively, b 1 and b 2 represents the bias terms of the first channel and the second channel classifier respectively; S5052: Fusing the first classification probability distribution and the second classification probability distribution to generate a final classification result; ; in, Indicates i The final classification result of categories is Indicates i The first classification probability distribution of categories, Indicates i The second classification probability distribution of categories, Indicates j The final classification results of the categories.
8. The method for detecting bad Chinese speech according to claim 1, characterized in that: The KL divergence loss is specifically: ; in, L kl represents the KL divergence loss, D KL represents the KL divergence, i =1,2,···, n , n Represents the number of input sentences involved in the loss calculation in the Chinese bad speech dataset, P ( y ) represents the probability distribution P In category y The predicted probability value on Q ( y ) represents the probability distribution Q The predicted probability value on category y; The cross entropy loss is specifically: ; in, L ce represents the cross entropy loss, m express, x i Indicates the first i Input sentences, y i express x i The corresponding true label, i =1,2,···, n , n Represents the number of input sentences involved in the loss calculation in the Chinese bad speech dataset.
9. A Chinese bad speech detection system, characterized in that: include: processor; A memory having computer-readable instructions stored thereon, wherein when the computer-readable instructions are executed by the processor, the method for detecting bad Chinese speech as claimed in any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Specific clothing image identification and detection method based on machine learning
CN107437099A
Social network aggressive speech detection method based on multi-task learning
CN116244441A