Hatred speech detection method and device based on comparative learning and computer equipment
By constructing a hate speech detection model based on contrast learning, and using preset sensitive keywords and dialogue generation models to construct the data set, the problem of low detection accuracy of hate speech in the existing technology is solved, and higher detection accuracy and data set diversity are achieved.
Patent Information
- Application Number
- CN202411927008.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-25
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2044-12-25
AI Technical Summary
The prior art cannot effectively detect hate speech, especially in short texts, semantic information cannot be fully mined, and the difference between positive and negative samples of hate speech is small, resulting in a low detection accuracy rate.
Using a contrast learning-based method, a hate speech detection model is constructed, including word embedding layer, semantic extraction layer, contrast mapping layer, contrast loss layer and output layer, and a hatred dialogue speech data set is constructed using preset sensitive keywords and dialogue generation models to train the model to improve detection accuracy.
By introducing contrast learning technology, the semantic differences between positive and negative samples of hate speech are expanded, the accuracy of detecting hate speech is improved, and the diversity and sample quality of the data set are enhanced.
Smart Images

Figure CN120030158A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data classification, in particular to the technical field of network speech governance, and in particular to a hate speech detection method, device and computer equipment based on contrastive learning. Background Art
[0002] As the number of users on social platforms continues to grow, hate speech has not only increased in number, but also spread more quickly, posing a huge challenge to social stability and personal mental health. Hate speech refers to words related to insults, hatred, prejudice, discrimination, crimes, sensitive topics, and physical harm, usually targeting a specific race, religion, gender, ethnicity, sexual orientation, or other group. Hate speech can be expressed in a variety of ways, from obvious insults and threats to implicit offensive speech, but an overly strict hate speech review mechanism will greatly reduce the response speed of the big model and the scope of interaction between users and the big model.
[0003] In order to prevent the negative impact of hate speech on individuals and society, the detection of hate speech has become a research focus. There are many methods for hate speech detection, including rule-based sensitive word filtering methods, machine learning-based methods (such as support vector machines (SVMs), decision trees, random forests, etc.), and deep learning-based methods (such as hate speech detection based on pre-trained models). However, these methods cannot fully mine the semantic information in short texts, and the difference between positive and negative samples of hate speech is very small, resulting in low accuracy of hate speech detection. Summary of the invention
[0004] In order to solve the above problems, the embodiments of the present invention provide a method, apparatus and computer device for detecting hate speech based on contrastive learning, which improve the accuracy of detecting hate speech.
[0005] In a first aspect, an embodiment of the present invention provides a method for detecting hate speech based on contrastive learning, comprising:
[0006] Construct a hate speech dataset based on preset sensitive keywords and dialogue generation model;
[0007] Constructing a hate speech detection model, and using the hate speech data set to train the hate speech detection model; wherein the hate speech detection model sequentially comprises a word embedding layer, a semantic extraction layer, a contrast mapping layer, a contrast loss layer, and an output layer;
[0008] The speech to be detected is input into the trained hate speech detection model, and the detection result of the speech to be detected is output.
[0009] Optionally, constructing a hate speech data set based on preset sensitive keywords and a dialogue generation model includes:
[0010] Get different types of hate speech as positive samples;
[0011] Input the preset sensitive keywords into the dialogue generation model to generate negative samples;
[0012] Based on the positive samples and the negative samples, the hate speech dataset is constructed.
[0013] Optionally, constructing a hate speech detection model includes:
[0014] The Roberta-wwm model is used as the word embedding layer; the MLP layer is used as the contrast mapping layer; and the Bi-LSTM or Bi-GRU is used as the semantic extraction layer.
[0015] Optionally, constructing a hate speech detection model includes:
[0016] splicing label information of each category before the text to be input to obtain the input text; wherein the categories include abnormal and normal;
[0017] Input the input text into the word embedding layer, and output a context vector, a label vector sequence of each category, and a word embedding vector matrix;
[0018] Inputting the label vector sequence of each category and the word embedding vector matrix into the semantic extraction layer, and outputting a plurality of semantic feature vectors; the semantic feature vectors include label vectors output by the label vectors of each category;
[0019] Input the context vector, the label vector of the abnormal category, and the label vector of the normal category into the contrast mapping layer, and output a first feature vector, a second feature vector, and a third feature vector respectively;
[0020] Calculate the similarity between the first feature vector and the second feature vector or the third feature vector of each category, and output the category corresponding to the highest similarity value by the output layer to obtain a prediction vector;
[0021] The prediction vector, the first feature vector and the second feature vector are input into the contrast loss layer for contrast learning, and a loss value is output.
[0022] Optionally, using the hate speech dataset to train the hate speech detection model includes:
[0023] The hate speech data set is processed in batches to obtain a plurality of sample sets including a plurality of positive samples and negative samples; wherein the category of the positive samples is abnormal; and the category of the negative samples is normal;
[0024] The sample set is input into the hate speech detection model, and in the contrast loss layer, the predicted vector and the true label vector of the category are cross-entropy calculated to obtain a first loss value; the first feature vector of each positive sample output by the contrast mapping layer is similar to the second feature vector of the other positive samples to obtain a second loss value; the second feature vector of each positive sample output by the contrast mapping layer is similar to the first feature vector of the other positive samples to obtain a third loss value;
[0025] Summing the first loss value, the second loss value and the third loss value to obtain the loss value;
[0026] When the loss value meets a preset training termination condition, the training is completed to obtain the trained hate speech detection model.
[0027] Optionally, the second eigenvector and the third eigenvector are obtained by the following method:
[0028] When the label name of the abnormal category is extreme and the label name of the normal category is normal, the word embedding layer outputs two label vector sequences; wherein the label vector sequences are both composed of a first vector and a second vector;
[0029] Inputting the first vector and the second vector of the label vector with the label name of "extreme" into the semantic extraction layer and performing average pooling to obtain the label vector with the label name of "extreme"; and inputting the label vector into the contrast mapping layer to obtain the second feature vector;
[0030] The first vector and the second vector of the label vector with a normal label name are input into the semantic extraction layer and average pooled to obtain the label vector with a normal label name; and the label vector is input into the contrast mapping layer to obtain the third feature vector.
[0031] In a second aspect, an embodiment of the present invention further provides a hate speech detection device based on contrastive learning, comprising:
[0032] A data construction module is used to construct a hate speech dataset based on preset sensitive keywords and a dialogue generation model;
[0033] A model building module, used to build a hate speech detection model, and train the hate speech detection model using the hate speech dialogue data set; wherein the hate speech detection model includes a word embedding layer, a semantic extraction layer, a contrast mapping layer, a contrast loss layer, and an output layer in sequence;
[0034] The detection module is used to input the speech to be detected into the trained hate speech detection model and output the detection result of the speech to be detected.
[0035] In a third aspect, an embodiment of the present invention further provides a computing device, including a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, it implements any of the above-mentioned methods for detecting hate speech based on contrastive learning.
[0036] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed in a computer, the computer is caused to execute any of the above-mentioned methods for detecting hate speech based on contrastive learning.
[0037] In a fifth aspect, an embodiment of the present invention further provides a computer program product, comprising computer instructions, which, when executed by a processor, implement the steps of the method described in any first aspect of this specification.
[0038] The embodiment of the present invention provides a method, apparatus and computer device for detecting hate speech based on contrastive learning. The method first generates a hate speech data set using preset sensitive keywords and a dialogue generation model, then constructs a hate speech detection model based on a word embedding layer, a semantic extraction layer, a contrast mapping layer, a contrast loss layer and an output layer, and trains the model using the hate speech data set to obtain a trained hate speech detection model. By inputting the speech to be detected into the trained hate speech detection model, the detection result of the speech to be detected can be obtained. In this way, the negative samples of the hate speech data set constructed by the present invention also contain preset sensitive keywords, so that the positive and negative samples are more balanced, and the diversity and sample quality of the data set are enhanced; at the same time, a contrast loss layer based on contrastive learning is introduced into the hate speech detection model, which further expands the semantic difference between the positive and negative samples of hate speech and improves the accuracy of detecting hate speech. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0040] Figure 1 is a flowchart of a method for detecting hate speech based on contrastive learning provided by an embodiment of the present invention;
[0041] Figure 2 is a schematic diagram of the architecture of a hate speech detection model provided by an embodiment of the present invention;
[0042] Figure 3 is a schematic diagram of the architecture of another hate speech detection model provided by an embodiment of the present invention;
[0043] Figure 4 is a hardware architecture diagram of a computing device provided by an embodiment of the present invention;
[0044] Figure 5 It is a structural diagram of a hate speech detection device based on contrastive learning provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0045] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0046] Existing methods for detecting hate speech cannot fully exploit the semantic information in short texts, and the differences between positive and negative samples of hate speech are very small. The specific manifestations and reasons are as follows:
[0047] First, the current deep learning model cannot effectively distinguish between positive and negative samples. The reason is that the vector sequence extracted by the pre-trained model often reduces the semantic feature distance between positive and negative samples after the semantic features are extracted by GRU or LSTM. For example, the two sentences "I don't like violent incidents" and "I like violent incidents" are respectively input into the recurrent neural network (RNN), long short-term memory network (LSTM) and gated recurrent unit (GRU) after the word embedding matrix is extracted by the pre-trained model to obtain the final text semantic features. However, the semantic features of these two sentences are often very close in Euclidean distance, making the current hate speech model unable to effectively distinguish this type of text;
[0048] Second, traditional deep learning models usually only regard labels as independent target variables and directly map them to categories, lacking a deep understanding of the complex relationship between labels and input text.
[0049] Therefore, based on the above reasons, the present invention provides a hate speech detection method based on contrastive learning.
[0050] The following is the concept of the present invention: Figure 1 As shown, an embodiment of the present invention provides a method for detecting hate speech based on contrastive learning, the method comprising:
[0051] Step 100, constructing a hate speech dataset based on preset sensitive keywords and a dialogue generation model;
[0052] Step 102, constructing a hate speech detection model, and using the hate speech dataset to train the hate speech detection model; wherein the hate speech detection model includes a word embedding layer, a semantic extraction layer, a contrast mapping layer, a contrast loss layer, and an output layer in sequence;
[0053] Step 104: input the speech to be detected into the trained hate speech detection model, and output the detection result of the speech to be detected.
[0054] In an embodiment of the present invention, a hate speech data set is first generated using preset sensitive keywords and a dialogue generation model, and then a hate speech detection model is constructed based on a word embedding layer, a semantic extraction layer, a contrast mapping layer, a contrast loss layer, and an output layer, and the model is trained using the hate speech data set to obtain a trained hate speech detection model, and the detection result of the speech to be detected can be obtained by inputting the speech to be detected into the trained hate speech detection model. In this way, the negative samples of the hate speech data set constructed by the present invention also contain preset sensitive keywords, so that the positive and negative samples are more balanced, and the diversity and sample quality of the data set are enhanced; at the same time, a contrast loss layer based on contrast learning is introduced into the hate speech detection model, which further expands the semantic difference between the positive and negative samples of hate speech and improves the accuracy of detecting hate speech.
[0055] It should be noted that, in accordance with the requirements of the Interim Measures for the Administration of Generative Artificial Intelligence Services jointly promulgated by the Cyberspace Administration of China and relevant departments, the definition of hate speech in this invention includes but is not limited to the following three points: (1) Strong emotional opposition: hate speech often carries strong offensive, insulting or derogatory emotions; (2) Semantic complexity: hate speech includes direct insults, as well as expressing an unreasonable point of view through a preset malicious stance; (3) Context-dependent: the same sentence may have different emotional colors in different contexts. For example, "These people really deserve to die" may be an expression of anger in a certain situation, but it may also be a clear hate speech against other groups in certain circumstances.
[0056] Described below Figure 1 How the various steps are performed.
[0057] In step 100, a hate speech dataset is constructed based on preset sensitive keywords and a dialogue generation model, including:
[0058] Get different types of hate speech as positive samples;
[0059] Input the preset sensitive keywords into the dialogue generation model to generate negative samples;
[0060] Based on positive and negative samples, a hate speech dataset is constructed.
[0061] It should be noted that the positive samples are collected from existing datasets widely used in large-scale language model evaluation, which include three types of hate speech, including swear words, insults, prejudice, discrimination, and crimes. In this way, the positive samples cover different types of malicious speech, and fully consider the diversity and complexity of hate speech in the sample annotation process, providing more comprehensive and representative positive samples for subsequent training. The Hate Conversation Speech Dataset is also called CHCD (Chinese Hostile Conversation Dataset), or CHCD dataset for short.
[0062] In the existing hate speech data set, the content of negative samples (i.e., normal speech) is often too simple, and the topics involved are relatively single, lacking in-depth discussion of sensitive issues, which leads to the problem of imbalance between positive and negative samples in the existing data set, thereby affecting the training results of the hate speech detection model. In an embodiment of the present invention, by improving the construction method of the hate speech data set, taking the attitude expression of preset sensitive keywords as the core, the dialogue generation model is used to ensure that the text maintains the naturalness and consistency of the dialogue in semantics, while avoiding the infiltration of malicious speech. In this way, these generated negative samples of non-hate speech containing preset sensitive keywords not only cover a variety of common normal social scenes, but also can effectively reflect the general values of the general public, thereby enhancing the diversity and authenticity of the data set, while also reducing the quality control problems caused by crawling social platform data to obtain negative samples.
[0063] In a preferred embodiment, the dialogue generation model adopts the ChatGLM4-9B model.
[0064] In an embodiment of the present invention, the ChatGLM4-9B model has stronger reasoning performance, longer context processing capability, multi-language capability, multi-modal capability and efficient pre-training technology.
[0065] In step 102, a hate speech detection model is constructed, including:
[0066] The Roberta-wwm model is used as the word embedding layer; the MLP layer is used as the contrast mapping layer; and the Bi-LSTM or Bi-GRU is used as the semantic extraction layer.
[0067] In the embodiment of the present invention, the model uses the Roberta-wwn model, the Bi-LSTM model and contrastive learning to gradually realize the automatic detection and classification of hate speech, and obtains an efficient and accurate hate speech detection model.
[0068] In an embodiment of the present invention, the Roberta-wwm model achieves more efficient semantic modeling in the Chinese context through the full-word masking technology, effectively avoiding the inaccuracy of Chinese word segmentation. Since the Roberta-wwm model can better capture the contextual information of words, especially for processing polysemous words, synonyms and complex sentence structures, it can learn a large amount of contextual information of Chinese text through pre-training in the hate speech detection task, so that the representation of each word not only contains the grammatical information of the word itself, but also incorporates its semantics in a specific context. By inputting text, the Roberta-wwm model can generate an embedding vector for each word. These embedding vectors form a highly dispersed representation in the semantic space, which can reflect the multi-dimensional semantic characteristics of words.
[0069] In the embodiment of the present invention, since the bidirectional long short-term memory network Bi-LSTM and the bidirectional gated recurrent unit Bi-GRU both adopt a bidirectional structure, they can process the input sequence from front to back or from back to front, respectively, thereby capturing the context information of the previous text and the context information of the subsequent text in the input sequence at the same time. Therefore, applying Bi-LSTM or Bi-GRU in the hate speech detection task can effectively capture these semantic nuances through its bidirectional context modeling ability, and improve the ability to understand complex texts, thereby improving the detection accuracy of the hate speech detection model.
[0070] In an embodiment of the present invention, contrastive learning can construct positive and negative sample pairs, allowing the model to learn how to shorten the distance between similar samples in the feature space and push the distance between dissimilar samples away, thereby effectively distinguishing positive and negative samples, thereby effectively improving the discrimination and generalization capabilities of the hate speech detection model.
[0071] In a preferred embodiment, construct Figure 2 Hate speech detection models shown, including:
[0072] The label information of each category is concatenated before the text to be input to obtain the input text; wherein the categories include abnormal and normal;
[0073] Input the input text into the word embedding layer, and output the context vector, the label vector sequence of each category, and the word embedding vector matrix;
[0074] Input the label vector sequence and word embedding vector matrix of each category into the semantic extraction layer, and output a number of semantic feature vectors; the semantic feature vector includes the label vector output by the label vector of each category;
[0075] The context vector, the label vector of the non-normal category, and the label vector of the normal category are input into the contrast mapping layer, and the first feature vector, the second feature vector, and the third feature vector are output respectively;
[0076] The first eigenvector is respectively calculated with the second eigenvector or the third eigenvector of each category, and the category corresponding to the highest similarity value is outputted by the output layer to obtain a prediction vector;
[0077] The prediction vector, the first eigenvector and the second eigenvector are input into the contrast loss layer for contrast learning, and the loss value is output.
[0078] It should be noted that in the embodiment of the present invention, the label name of the abnormal category is extreme, that is, the label information is extreme; the label name of the normal category is normal, that is, the label information is normal. Specifically, the context vector integrates the context information and the label information; the label vector is obtained by converting the label information into a vector form, with a total of two label vectors; the word embedding layer of the input text to be input is converted into a vector form of the feature space to obtain a word embedding vector matrix, and each vector of the matrix corresponds to each word of the text to be input, that is, the number of vectors in the matrix is the same as the number of words in the text to be input.
[0079] For the label vector, for example, if the label information is "hate", which consists of a single word, the word embedding layer will output a vector, which is a label vector sequence of the abnormal category. For example, if the label information is "biased", which consists of two words, the word embedding layer will output two vectors. At this time, the two vectors need to be fused (i.e., average pooling) to obtain a one-dimensional vector, which is a label vector sequence of the abnormal category.
[0080] Specifically, in the contrast mapping layer, the context vector is input into the contrast mapping layer, and the first feature vector is output; the label vector of the abnormal category is input into the contrast mapping layer, and the second feature vector is output; the label vector of the normal category is input into the contrast mapping layer, and the third feature vector is output.
[0081] In the embodiment of the present invention, by splicing the label information to the beginning of the text to be input, and then inputting the text into the word embedding layer, the word embedding form of the text is obtained, and the context vector, the label vector sequence of each category included in the hate speech data set, and the word embedding vector matrix are output, thereby using the features of the label information and the CLS features of the word embedding layer to expand the distance between positive and negative samples. At the same time, in the semantic extraction layer, the complex relationship between the label information and the text to be input is deeply understood, and the understanding ability is improved. Therefore, the present invention enhances the accuracy of hate speech text classification by making full use of the label information.
[0082] In the embodiment of the present invention, multiple experiments have proved that the contrast mapping layer improves the learning effect of the hate speech detection model. Since the context vector and the semantic feature vector are respectively input into the contrast mapping layer MLP layer, the feature vector is mapped to an optimized feature space, thereby promoting the objective function of contrast learning to be learned and optimized more effectively.
[0083] In step 102, a hate speech detection model is trained using the hate speech dataset, including:
[0084] The hate speech dataset is processed in batches to obtain several sample sets containing positive samples and negative samples. Among them, the category of positive samples is abnormal; the category of negative samples is normal.
[0085] The sample set is input into the hate speech detection model. In the contrast loss layer, the cross entropy of the predicted vector and the true label vector of the category is calculated to obtain a first loss value. The similarity between the first feature vector of each positive sample output by the contrast mapping layer and the second feature vector of other positive samples is calculated to obtain a second loss value. The similarity between the second feature vector of each positive sample output by the contrast mapping layer and the first feature vector of other positive samples is calculated to obtain a third loss value.
[0086] The first loss value, the second loss value, and the third loss value are summed to obtain a loss value;
[0087] When the loss value meets the preset training termination condition, the training is completed and a trained hate speech detection model is obtained.
[0088] In an embodiment of the present invention, in order to further expand the spatial distance between hate speech (i.e., positive samples) and normal speech (i.e., negative samples) in the feature space, the hate speech dialogue dataset is processed in batches to construct a sample set containing positive samples and negative samples.
[0089] Specifically, the loss function contained in the contrast loss layer consists of three parts. The first loss value represents the cross entropy loss Loss between the prediction vector predicted by the model and the true label vector of the category (i.e., one-hot form).CE , obtained by the following formula:
[0090]
[0091] like Figure 3 As shown, the second loss value represents the first feature vector e of each positive sample CLS The label semantic vector (i.e., the second eigenvector) of the same sample POS Loss CLS , obtained by the following formula:
[0092]
[0093] In formula (2), the first eigenvector e of each positive sample is CLS The second eigenvector e of other similar samples POS Perform dot product, and then divide by the temperature parameter to scale the dot product value to calculate the similarity score, and then get the similarity matrix logits 1 Among them, the similarity matrix logits 1 The value of each row in represents the similarity score between the sample and other similar samples, and then the similarity matrix logits 1 Normalize and sum each row to get the sum of similarities of each sample to other similar samples.
[0094] like Figure 3 As shown, the third loss value represents the second feature vector e of each positive sample POS The first eigenvector e of the same sample CLS Loss TAG , obtained by the following formula:
[0095]
[0096] The calculation process of formula (3) is the same as that of formula (2), wherein, in the above formulas (1) to (3), P i is the set of positive samples in the i-th sample set; label p P i A positive sample in context i is the first eigenvector e of the target sample CLS ; A i is the set of positive and negative samples in the i-th sample set; I is the set of all samples in the hate speech dataset, N is the number of all samples; K is the class label set, k is one of the class labels; τ is the temperature parameter;
[0097] Then the loss value Loss = Loss CE+Loss CLS +Loss TAG .
[0098] It should be noted that Figure 3 The output layer is not shown, and the semantic extraction layer only shows a semantic feature vector; and when the input is a negative sample, the output of the mapping layer is compared NEG Represents the third eigenvector; when the input is a positive sample, the output of the mapping layer is compared POS represents the second eigenvector.
[0099] In the embodiment of the present invention, the loss function of the hate speech detection model shortens the distance between positive samples through the above three parts, and effectively expands the distance between positive samples and negative samples, so that the hate speech detection model can learn more detailed features, thereby improving the classification ability of unknown texts.
[0100] In a preferred embodiment, the second eigenvector and the third eigenvector are obtained by the following method:
[0101] When the label name of the abnormal category is extreme and the label name of the normal category is normal, the word embedding layer outputs two label vector sequences; wherein the label vector sequences are composed of the first vector and the second vector;
[0102] The first vector and the second vector of the label vector with the label name of "biased" are input into the semantic extraction layer and average pooled to obtain the label vector with the label name of "biased"; and the label vector is input into the contrast mapping layer to obtain the second feature vector;
[0103] The first vector and the second vector of the normal label vector are input into the semantic extraction layer and average pooled to obtain the normal label vector; and the label vector is input into the contrast mapping layer to obtain the third feature vector.
[0104] Since the label name is in Chinese, the label may be composed of multiple characters, so the word embedding sequence of each label needs to be averaged by pooling. In the data set of the present invention, there are two categories of labels, one label name is "extreme" and the other label name is "normal". Taking "extreme" as an example, "extreme" is input into the word embedding layer to output a label vector sequence including the first vector V1 and the second vector V2. After passing through the semantic extraction layer and average pooling, the label vector V is obtained. p , the label vector V p After the contrast mapping layer, the second feature vector e is output POS The third eigenvector is obtained in the same way.
[0105] In an embodiment of the present invention, conversation speech is taken as a starting point, preset sensitive keywords are introduced, and a hate speech conversation data set of a user's conversation with a large model is constructed; at the same time, a hate speech detection model is built based on Bi-LSTM or Bi-GRU, which effectively mines the deep semantic information in the text, and uses the pre-trained model Roberta-wwm and contrastive learning technology to expand the distance between positive and negative samples, effectively extracting the semantic information between hate speech and non-hate speech, thereby effectively detecting hate speech in user conversations and improving the accuracy of hate speech detection.
[0106] After actual testing, the mainstream model and the hate speech detection model of the present invention were used to detect hate speech, and the test results of detection accuracy (%) were obtained as shown in Table 1 and Table 2. Table 1 uses the existing commonly used hate speech data set for training, and Table 2 uses the hate speech dialogue data set constructed by the above step 100 for training; the two data sets have the same topics and the same ratio of positive and negative samples, but the negative samples in the data set used in Table 1 are too simple and the sensitive topics involved are also relatively single.
[0107] Table 1
[0108]
[0109] Table 2
[0110]
[0111] It should be noted that in Table 1, in the first row, when the existing model includes a word embedding layer, FNN, and cross entropy is used in the loss function, the accuracy of the validation data set is 92.02%, and the accuracy of the test data set is 82.4%; when the existing model includes a word embedding layer, Bi-GRU, and cross entropy is used in the loss function, the accuracy of the validation data set is 91.88%, and the accuracy of the test data set is 82.25%; when the existing model includes a word embedding layer, Bi-LSTM, and cross entropy is used in the loss function, the accuracy of the validation data set is 92.07%, and the accuracy of the test data set is 82.27%. In the second row, the architecture of the existing model includes a word embedding layer, FNN, and when cross entropy + basic contrastive learning is used in the loss function, the accuracy of the validation data set is 92.4%, and the accuracy of the test data set is 82.74%; the architecture of the existing model includes a word embedding layer, Bi-GRU, and when cross entropy + basic contrastive learning is used in the loss function, the accuracy of the validation data set is 92.07%, and the accuracy of the test data set is 82.4%; the architecture of the existing model includes a word embedding layer, Bi-LSTM, and when cross entropy + basic contrastive learning is used in the loss function, the accuracy of the validation data set is 92.32%, and the accuracy of the test data set is 82.83%. It should be noted that in the existing models in the first and second rows, the inputs are all texts to be input.
[0112] However, the detection accuracy of the above existing models for hate speech is still low. Therefore, in order to improve the existing model, in the third row, an experimental model A including a word embedding layer, a contrast mapping layer (MLP layer), an FNN, and a cross entropy + contrast learning framework of the present invention in the loss function is designed, and the accuracy of the verification data set is 92.49%, and the accuracy of the test data set is 82.83%; and an experimental model B including a word embedding layer, a semantic extraction layer (Bi-GRU), a contrast mapping layer (MLP layer), and a cross entropy + contrast learning framework of the present invention in the loss function is designed, and the accuracy of the verification data set is 92.41%, and the accuracy of the test data set is 82.96%; and an experimental model C including a word embedding layer, a semantic extraction layer (Bi-LSTM), a contrast mapping layer (MLP layer), and a cross entropy + contrast learning framework of the present invention in the loss function is designed, and the accuracy of the verification data set is 92.29%, and the accuracy of the test data set is 83.13%. Obviously, the experimental models B and C provided in the embodiments of the present invention have higher detection accuracy for hate speech. Among them, in the model of the third row, its input is the text to be input with labels of various categories. It can be seen from Table 2 that after training with the hate speech data set constructed by the present invention, the detection accuracy of the experimental model B and the experimental model C provided by the embodiment of the present invention for hate speech is further improved, wherein the accuracy of the verification data set is as high as 95.05%, and the accuracy of the test data set is 94.98%.
[0113] It can be seen from Table 1 and Table 2 that the detection accuracy of the hate speech detection model of the present invention is higher than that of the mainstream model. The hate speech detection method based on contrastive learning proposed in the present invention has the following advantages:
[0114] (1) By using the contrastive learning method, the model can learn more detailed features by comparing the similarities between positive and negative samples, thereby improving the classification ability of unseen samples. In the hate speech detection task, although the surface features of hate speech and normal speech may be similar, through contrastive learning, the model can better distinguish subtle semantic differences, thereby improving the detection accuracy, which is about 2% higher than the mainstream model;
[0115] (2) The hate speech dataset contains a variety of positive examples of hate speech. This paper effectively expands the types and forms of hate speech by collecting positive examples of three types of hate speech, including insults, prejudice, discrimination, and crimes. These positive examples are not only highly offensive and inflammatory, but also fully reflect the complexity of hate speech, ensuring that the training data is richer and more diverse.
[0116] (3) Generate negative samples based on preset sensitive keywords and a dialogue generation model, effectively avoiding quality control problems in the crawled data and ensuring the high quality and consistency of the negative sample content.
[0117] As Figure 4 , Figure 5 shown, an embodiment of the present invention provides a hate speech detection device based on contrastive learning. The device embodiment can be implemented by software, or by hardware or a combination of software and hardware. At the hardware level, as Figure 4 shown, it is a hardware architecture diagram of a computing device where a hate speech detection device based on contrastive learning provided by an embodiment of the present invention is located. In addition to Figure 4 the shown processor, memory, network interface, and non-volatile memory, the computing device where the device is located in the embodiment usually may also include other hardware, such as a forwarding chip responsible for processing packets, and so on. Taking software implementation as an example, as Figure 5 shown, as a logically meaningful device, it is formed by the CPU of its computing device reading the corresponding computer program in the non-volatile memory into the memory and running. A hate speech detection device based on contrastive learning provided in this embodiment includes:
[0118] A data construction module 500, configured to construct a hate dialogue speech dataset based on preset sensitive keywords and a dialogue generation model;
[0119] A model construction module 502, configured to construct a hate speech detection model and train the hate speech detection model using the hate dialogue speech dataset; wherein, the hate speech detection model sequentially includes a word embedding layer, a semantic extraction layer, a contrast mapping layer, a contrast loss layer, and an output layer;
[0120] A detection module 504, configured to input the speech to be detected into the trained hate speech detection model and output the detection result of the speech to be detected.
[0121] In some specific implementation manners, the data construction module 500 may be used to execute the above step 100, the model construction module 502 may be used to execute the above step 102, and the detection module 504 may be used to execute the above step 104.
[0122] In some specific implementation manners, the data construction module 500 is further used to perform the following operations:
[0123] Obtain different types of hate speech as positive samples;
[0124] Input the preset sensitive keywords into the dialogue generation model to generate negative samples;
[0125] Based on the positive samples and negative samples, construct a hate dialogue speech dataset.
[0126] In some specific implementations, the model building module 502 is further configured to perform the following operations:
[0127] The Roberta-wwm model is used as the word embedding layer; the MLP layer is used as the contrast mapping layer; and the Bi-LSTM or Bi-GRU is used as the semantic extraction layer.
[0128] In some specific implementations, the model building module 502 is further configured to perform the following operations:
[0129] The label information of each category is concatenated before the text to be input to obtain the input text; wherein the categories include abnormal and normal;
[0130] Input the input text into the word embedding layer, and output the context vector, the label vector sequence of each category, and the word embedding vector matrix;
[0131] Input the label vector sequence and word embedding vector matrix of each category into the semantic extraction layer, and output a number of semantic feature vectors; the semantic feature vector includes the label vector output by the label vector of each category;
[0132] The context vector, the label vector of the non-normal category, and the label vector of the normal category are input into the contrast mapping layer, and the first feature vector, the second feature vector, and the third feature vector are output respectively;
[0133] The first eigenvector is respectively calculated for similarity with the second eigenvector or the third eigenvector of each category, and the category corresponding to the highest similarity value is outputted by the output layer to obtain a prediction vector;
[0134] The prediction vector, the first eigenvector and the second eigenvector are input into the contrast loss layer for contrast learning, and the loss value is output.
[0135] In some specific implementations, the model building module 502 is further configured to perform the following operations:
[0136] The hate speech dataset is processed in batches to obtain several sample sets containing positive samples and negative samples. The positive samples are classified as abnormal and the negative samples are classified as normal.
[0137] The sample set is input into the hate speech detection model. In the contrast loss layer, the cross entropy of the predicted vector and the true label vector of the category is calculated to obtain a first loss value. The similarity between the first feature vector of each positive sample output by the contrast mapping layer and the second feature vector of other positive samples is calculated to obtain a second loss value. The similarity between the second feature vector of each positive sample output by the contrast mapping layer and the first feature vector of other positive samples is calculated to obtain a third loss value.
[0138] The first loss value, the second loss value, and the third loss value are summed to obtain a loss value;
[0139] When the loss value meets the preset training termination condition, the training is completed and a trained hate speech detection model is obtained.
[0140] In some specific implementations, the model building module 502 is further configured to perform the following operations:
[0141] When the label name of the abnormal category is extreme and the label name of the normal category is normal, the word embedding layer outputs two label vector sequences; wherein the label vector sequences are composed of the first vector and the second vector;
[0142] The first vector and the second vector of the label vector with the label name of "biased" are input into the semantic extraction layer and average pooled to obtain the label vector with the label name of "biased"; and the label vector is input into the contrast mapping layer to obtain the second feature vector;
[0143] The first vector and the second vector of the normal label vector are input into the semantic extraction layer and average pooled to obtain the normal label vector; and the label vector is input into the contrast mapping layer to obtain the third feature vector.
[0144] It is to be understood that the structure illustrated in the embodiment of the present invention does not constitute a specific limitation on a hate speech detection device based on contrastive learning. In other embodiments of the present invention, a hate speech detection device based on contrastive learning may include more or fewer components than shown in the figure, or combine some components, or split some components, or arrange the components differently. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.
[0145] The information interaction, execution process and other contents between the modules in the above-mentioned device are based on the same concept as the embodiment of the method of the present invention. For the specific contents, please refer to the description in the embodiment of the method of the present invention, and no further description is given here.
[0146] An embodiment of the present invention further provides a computing device, including a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, a method for detecting hate speech based on contrastive learning in any embodiment of the present invention is implemented.
[0147] An embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the processor executes a method for detecting hate speech based on contrastive learning in any embodiment of the present invention.
[0148] An embodiment of the present application also provides a computer program product, which includes a computer program. A processor of a computer device reads the computer program from a computer-readable storage medium, and the processor executes the computer program, so that the computer device executes a hate speech detection method based on contrastive learning as described in any of the above embodiments.
[0149] Specifically, a system or device equipped with a storage medium can be provided, on which software program code that implements the functions of any of the above-mentioned embodiments is stored, and a computer (or CPU or MPU) of the system or device can read and execute the program code stored in the storage medium.
[0150] In this case, the program code itself read from the storage medium can realize the function of any one of the above-mentioned embodiments, and thus the program code and the storage medium storing the program code constitute a part of the present invention.
[0151] The storage medium embodiments for providing the program code include a floppy disk, a hard disk, a magneto-optical disk, an optical disk (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RAM, DVD-RW, DVD+RW), a magnetic tape, a non-volatile memory card, and a ROM. Alternatively, the program code can be downloaded from a server computer by a communication network.
[0152] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, which carry computer-readable program code. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Computer-readable signal media may also be any computer-readable medium other than a computer-readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, system, or device.
[0153] The program code embodied on the computer readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0154] Computer program code for performing the operations of the present invention may be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0155] In addition, it should be clear that the functions of any of the above embodiments can be implemented not only by executing the program code read by the computer, but also by enabling an operating system operating on the computer to complete part or all of the actual operations based on instructions from the program code.
[0156] In addition, it can be understood that the program code read from the storage medium is written to a memory provided in an expansion board inserted into the computer or to a memory provided in an expansion module connected to the computer, and then based on the instructions of the program code, a CPU installed on the expansion board or expansion module is enabled to perform part or all of the actual operations, thereby realizing the functions of any of the above-mentioned embodiments.
[0157] It should be noted that, in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the statement "comprise a ..." do not exclude the presence of other identical factors in the process, method, article or device including the elements.
[0158] A person of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps of the above method embodiments; and the aforementioned storage medium includes: ROM, RAM, magnetic disk or optical disk, etc., various media that can store program codes.
[0159] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for detecting hate speech based on contrastive learning, characterized in that: include: Construct a hate speech dataset based on preset sensitive keywords and dialogue generation model; Constructing a hate speech detection model, and using the hate speech data set to train the hate speech detection model; wherein the hate speech detection model sequentially comprises a word embedding layer, a semantic extraction layer, a contrast mapping layer, a contrast loss layer, and an output layer; The speech to be detected is input into the trained hate speech detection model, and the detection result of the speech to be detected is output.
2. The method according to claim 1, characterized in that The method of constructing a hate speech data set based on preset sensitive keywords and a dialogue generation model includes: Get different types of hate speech as positive samples; Input the preset sensitive keywords into the dialogue generation model to generate negative samples; Based on the positive samples and the negative samples, the hate speech dataset is constructed.
3. The method according to claim 1, characterized in that The constructing of a hate speech detection model includes: The Roberta-wwm model is used as the word embedding layer; the MLP layer is used as the contrast mapping layer; and the Bi-LSTM or Bi-GRU is used as the semantic extraction layer.
4. The method according to any one of claims 1 to 3, characterized in that: The constructing of a hate speech detection model includes: splicing label information of each category before the text to be input to obtain the input text; wherein the categories include abnormal and normal; Input the input text into the word embedding layer, and output a context vector, a label vector sequence of each category, and a word embedding vector matrix; Input the label vector sequence of each category and the word embedding vector matrix into the semantic extraction layer, and output a plurality of semantic feature vectors; the semantic feature vectors include label vectors output by the label vectors of each category; Input the context vector, the label vector of the abnormal category, and the label vector of the normal category into the contrast mapping layer, and output a first feature vector, a second feature vector, and a third feature vector respectively; Calculate the similarity between the first feature vector and the second feature vector or the third feature vector of each category, and output the category corresponding to the highest similarity value by the output layer to obtain a prediction vector; The prediction vector, the first feature vector and the second feature vector are input into the contrast loss layer for contrast learning, and a loss value is output.
5. The method according to claim 4, characterized in that Using the hate speech dataset to train the hate speech detection model includes: The hate speech data set is processed in batches to obtain a plurality of sample sets including positive samples and negative samples; wherein the category of the positive samples is abnormal; and the category of the negative samples is normal; The sample set is input into the hate speech detection model, and in the contrast loss layer, the predicted vector and the true label vector of the category are cross-entropy calculated to obtain a first loss value; the first feature vector of each positive sample output by the contrast mapping layer is similar to the second feature vector of the other positive samples to obtain a second loss value; the second feature vector of each positive sample output by the contrast mapping layer is similar to the first feature vector of the other positive samples to obtain a third loss value; Summing the first loss value, the second loss value and the third loss value to obtain the loss value; When the loss value meets a preset training termination condition, the training is completed to obtain the trained hate speech detection model.
6. The method according to claim 4, characterized in that The second eigenvector and the third eigenvector are obtained by the following method: When the label name of the abnormal category is extreme and the label name of the normal category is normal, the word embedding layer outputs two label vector sequences; wherein the label vector sequences are both composed of a first vector and a second vector; Inputting the first vector and the second vector of the label vector with the label name of "extreme" into the semantic extraction layer and performing average pooling to obtain the label vector with the label name of "extreme"; and inputting the label vector into the contrast mapping layer to obtain the second feature vector; The first vector and the second vector of the label vector with a normal label name are input into the semantic extraction layer and average pooled to obtain the label vector with a normal label name; and the label vector is input into the contrast mapping layer to obtain the third feature vector.
7. A hate speech detection device based on contrastive learning, characterized in that: include: A data construction module is used to construct a hate speech dataset based on preset sensitive keywords and a dialogue generation model; A model building module, used to build a hate speech detection model, and train the hate speech detection model using the hate speech dialogue data set; wherein the hate speech detection model includes a word embedding layer, a semantic extraction layer, a contrast mapping layer, a contrast loss layer, and an output layer in sequence; The detection module is used to input the speech to be detected into the trained hate speech detection model and output the detection result of the speech to be detected.
8. A computing device, comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the method according to any one of claims 1 to 6 is implemented.
9. A computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, causes the computer to execute the method according to any one of claims 1 to 6.
10. A computer program product, characterized in that The method comprises computer instructions, which, when executed by a processor, implement the steps of the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Social network sensitive gensis detection method based on supervised learning
CN116257698A
Hidden hatred speech detection and fine classification method and device based on combined comparative learning
CN116303995A
Hatred speech detection method and system based on semantic enhancement and storage medium
CN118520070A
Contrastive learning method based on implications for detecting implicit hate expression, apparatus and computer program for performing the method
US20240135243A1