Implicit hate speech detection and classification method and device based on joint contrastive learning

By applying a joint contrast learning method in social media, combining pre-trained language models and fully connected networks, using the contrast learning framework and common sense knowledge base ATOMIC for knowledge enhancement, the problem of difficult detection and classification of implicit hate speech is solved, and higher detection accuracy and classification accuracy are achieved, providing effective decision support for speech governance in cyberspace.

CN116303995BActive Publication Date: 2025-05-23NAT UNIV OF DEFENSE TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310349777.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-04
Publication Date
2025-05-23
Estimated Expiration
2043-04-04

AI Technical Summary

Technical Problem

It is difficult for the prior art to effectively detect and subcategorize implicit hate speech, especially in social media. Implicit hate speech is difficult to accurately identify and classify by existing text classification models due to its obscureness and diversity.

Method used

Using a joint contrast learning method, by constructing an implicit hate speech detection model and classification model, combining pre-trained language model and fully connected network, a comparison learning framework and common sense knowledge base ATOMIC are used to enhance knowledge, and the intentions of the text publisher and the psychological reactions of readers are reasoned, so as to conduct more accurate implicit hate speech detection and subclassification.

Benefits of technology

It significantly improves the detection accuracy, recall rate and F1 value of implicit hate speech, which can more effectively identify and classify implicit hate speech, and provides better decision-making support for speech governance in cyberspace.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116303995B_ABST
    Figure CN116303995B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and device for detecting and sub-classifying implicit hate speech based on joint contrastive learning. The method and device can detect and sub-classify hate speech in social media, provide decision support for speech governance in cyberspace, collect text data in online social media, build and train an implicit hate speech detection model, obtain implicit hate speech, perform knowledge enhancement based on a common sense knowledge base ATOMIC, infer the intention of the text publisher, the subject reaction and the object reaction of the reader; build and optimize an implicit hate speech classification model, respectively use implicit hate speech, intention, subject reaction and object reaction of the same category as positive samples in contrastive learning to calculate contrastive learning loss, perform weighted summation to obtain joint contrastive learning loss, sum it with cross entropy loss to obtain a total loss function, optimize the total loss function, obtain the classification result of implicit hate speech through the trained implicit hate speech classification model, and use it for speech governance in social media.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of online public opinion analysis, data mining, and social media speech governance, and is a method and device for detecting and subdividing implicit hate speech based on joint contrastive learning. Background Art

[0002] The widespread use of social media and its anonymity have led to the widespread spread of hate speech, which has caused actual harm to people. Currently, many studies focus on hate speech governance, but most of them focus on coarse-grained detection and discovery of explicit hate speech, while ignoring fine-grained classification of implicit hate speech. Implicit hate speech mainly refers to "the use of coded or indirect language, such as sarcasm, metaphors, and beating around the bush, to devalue certain groups or individuals, or to convey prejudice and harmful views against them." Implicit here mainly refers to hate speech that is not directly and explicitly expressed.

[0003] The development of social media technology has greatly facilitated the dissemination of information, but it has also provided an opportunity for the spread of harmful speech. Compared with explicit hate speech, implicit hate speech is more obscure and difficult to detect. If we fail to fully understand the harmfulness of relevant implicit hate speech and take effective measures to manage it, it will pose a great hidden danger in the long run. The detection and sub-classification of implicit hate speech is of great significance for accurately countering hate speech and locating and combating hate groups in real life. However, there are still few researchers who have conducted research on implicit hate speech. Summary of the invention

[0004] In response to the above problems, the present invention discloses a method and device for detecting and subdividing implicit hate speech based on joint contrastive learning, which can detect hate speech in social media and subdivide implicit hate speech, thereby providing decision support for the governance of speech in cyberspace.

[0005] The technical solution is as follows: a method for detecting and subdividing implicit hate speech based on joint contrastive learning, characterized in that it includes the following steps:

[0006] Step 1: Collect text data from online social media, and perform data cleaning and data preprocessing on the collected text data;

[0007] Step 2: constructing an implicit hate speech detection model, wherein the implicit hate speech detection model includes a pre-trained language model module and a fully connected network module, wherein the pre-trained language model module is used to extract text features, and the fully connected network module is used to classify the extracted text features, and the implicit hate speech detection model is trained using a labeled data set, wherein the trained implicit hate speech detection model can distinguish non-hate speech, explicit hate speech, and implicit hate speech in text data, and the collected text data is input into the trained implicit hate speech detection model to obtain implicit hate speech;

[0008] Step 3: For the detected implicit hate speech, knowledge enhancement is performed on the implicit hate text based on the common sense knowledge base ATOMIC to infer the intention of the text author, the subjective reaction of the author, and the objective reaction of the reader;

[0009] Step 4: constructing an implicit hate speech classification model, wherein the implicit hate speech classification model includes a pre-trained language model module and a fully connected network module, wherein the pre-trained language model module is used to extract features, and the fully connected network module is used to classify the extracted features;

[0010] Using the contrastive learning framework, implicit hate speech, intention, subject reaction, and object reaction of the same category are used as positive samples in contrastive learning to calculate the contrastive learning loss, and the four contrastive learning losses are jointly weighted and summed to obtain the joint contrastive learning loss;

[0011] The total loss function of the implicit hate speech classification model is obtained by summing the joint contrastive learning loss and the cross entropy loss, and the total loss function of the implicit hate speech classification model is optimized to obtain a trained implicit hate speech classification model;

[0012] Step 5: Input the detected implicit hate speech into the trained implicit hate speech classification model, and output the classification results of the implicit hate speech. The classification results of the implicit hate speech are used for social media speech governance.

[0013] Furthermore, in step 1, the data preprocessing includes removing emoticons, removing Chinese and English characters, removing null values ​​and words without actual meaning, and lowercasing characters for English.

[0014] Furthermore, in step 2, the pre-trained language model module of the implicit hate speech detection model is constructed based on the pre-trained language model PLMs, and the pre-trained language model PLMs is used to obtain the given text x i The characteristic representation is expressed as:

[0015] h(x i )=PLMs(x i )

[0016] Among them, h(x i ) is the text feature obtained by pre-training language model PLMs;

[0017] The fully connected network module of the implicit hate speech detection model uses a non-linear activation function to classify text into non-hate speech, explicit hate speech, and implicit hate speech, expressed as:

[0018]

[0019] Among them, sigmoid is a nonlinear activation function. is the detection result of the implicit hate speech detection model, none represents non-hate speech, explicit represents explicit hate speech, and implicit represents implicit hate speech.

[0020] Furthermore, in step 3, the following steps are specifically included:

[0021] Implicit hate speech text identified by the implicit hate speech detection model Based on the common sense knowledge base ATOMIC, the intention and reaction of the person who sent the implicit hate speech text, as well as the reaction of the person who read this implicit hate speech text,

[0022] The standard triples of the common sense knowledge base ATOMIC are<event,relation,event> ,The nodes in the common sense knowledge base ATOMIC are events, the edges between nodes are relationship types, and ,relation includes subject intention xIntent, subject reaction xReact and object reaction oReact;

[0023] Use the SentenceBERT model to encode implicit hate speech and all event nodes in the ATOMIC knowledge graph to obtain their text feature representation:

[0024] h SBERT (x i )=SentenceBERT(x i )

[0025]

[0026] Among them, h SBERT (x i ) indicates implicit hate speech, Represents the feature representation of event nodes in the ATOMIC knowledge graph;

[0027] Use the cosine similarity function to calculate the similarity between implicit hate text and event nodes in ATOMIC, and select the top K events in similarity as the event source for implicit hate speech reasoning;

[0028] According to implicit hate speech And the corresponding events with the top K similarity ranking in ATOMIC, infer the corresponding K possible subject intentions xintent, subject reactions xreact, and object reactions oreact, and connect their respective K texts, expressed as:

[0029]

[0030]

[0031]

[0032] Among them, xIntent(x i )、xReact(x i )、oReact(x i ) are implicit hate speech texts (x i ) is a collection of the subject intention xintent, subject reaction xreact, and object reaction oreact. They are implicit hate speech texts (x i )’s kth subject intention xintent, subject reaction xreact, and object reaction oreact.

[0033] Furthermore, in step 4, the following steps are specifically included:

[0034] Set the subject intent xIntent(x i ) as a positive sample for contrastive learning, and calculate the intention contrast loss of implicit hate speech, expressed as:

[0035]

[0036] 1 of them [k≠i] represents the indicator function, the function value is 1 if and only if k≠i, τ is a hyperparameter, f(u,v)=sim(u,v)=u T v / ‖u‖‖v‖ is the similarity calculation function, h(xIntent(x i )) represents a positive sample, Indicates the batch size set during training;

[0037] The main reaction xReact(x i ) is used as a positive sample for contrastive learning, and the subject response contrast loss is calculated, which is expressed as:

[0038]

[0039] 1 of them [k≠i] represents the indicator function, the function value is 1 if and only if k≠i, τ is a hyperparameter, f(u,v)=sim(u,v)=u T v / ‖u‖‖v‖ is the similarity calculation function, h(xReact(x i )) represents a positive sample, Indicates the batch size set during training;

[0040] The object reaction oReact(x i ) is used as a positive sample for contrastive learning and the object response contrast loss is calculated, which is expressed as:

[0041]

[0042] 1 of them [k≠i] represents the indicator function, the function value is 1 if and only if k≠i, τ is a hyperparameter, f(u,v)=sim(u,v)=u T v / ‖u‖‖v‖ is the similarity calculation function, h(oReact(x i )) represents a positive sample, Indicates the batch size set during training;

[0043] The same category of implicit hate speech (x j ) as a positive sample for contrastive learning and calculate the supervised contrast loss:

[0044]

[0045] 1 of them [k≠i] , represents the indicator function, the function value is 1 if and only if k≠i, and if and only if y i =y j When the function value is 1, τ is a hyperparameter, f(u,v)=sim(u,v)=u T v / ‖u‖‖v‖ is the similarity calculation function, h((x j )) represents a positive sample, Indicates the batch size set during training;

[0046] The weighted sum of the contrast losses is used as the joint contrast loss, expressed as:

[0047]

[0048] Among them, γ xI , γ xR , γ oR, γ implicit are weights respectively;

[0049] When training the implicit hate speech classification model, the implicit hate speech classification model is also optimized through a dataset with implicit hate speech classification labels. The cross entropy loss is calculated according to the labels predicted by the implicit hate speech classification model. The cross entropy loss function is expressed as follows:

[0050]

[0051] in Denotes the predicted label, y i represents the true label, N b Indicates the size of the data batch, represents the cross entropy loss function;

[0052] The cross entropy loss is added to the joint contrastive learning loss as the total loss function for the implicit hate speech classification model:

[0053]

[0054] in, Total loss function for the implicit hate speech classification model.

[0055] A computer device, characterized in that it comprises: a processor, a memory and a program;

[0056] The program is stored in the memory, and the processor calls the program stored in the memory to execute the above-mentioned implicit hate speech detection and sub-classification method based on joint contrastive learning.

[0057] A computer-readable storage medium, characterized in that: the computer-readable storage medium is used to store a program, and the program is used to execute the above-mentioned implicit hate speech detection and sub-classification method based on joint contrastive learning.

[0058] The present invention has the following beneficial effects:

[0059] In view of the fact that implicit hate speech may have great differences in language expression and is more obscure and difficult to detect, the existing text classification pre-trained language model is not ideal for the detection and fine-grained classification of implicit hate speech. In the present invention, the psychological state involved in implicit hate speech includes the speaker's intention, and the reader's psychological reaction is usually the same. Therefore, by using the knowledge enhancement method, based on the inference knowledge base ATOMIC, the subject intention and subject psychological reaction of the poster involved in implicit hate speech, as well as the objective psychological reaction of the reader are inferred. The method of the present invention can express the implicit meaning of implicit hate speech to a certain extent, which is conducive to assisting the pre-trained language model to understand its semantics and perform more accurate fine-grained classification.

[0060] The present invention designs a new joint contrastive learning framework. Based on the classic contrastive learning loss function, the subject intention, subject psychological reaction, object psychological reaction and samples of the same category of the inferred implicit hate speech are respectively used as positive samples in the contrastive learning loss, which can make the representations of samples with similar intentions and psychological reactions more similar in the representation space, and make the different ones more distinguishable. The present invention takes the weighted sum of the above losses and adds them to the cross-entropy loss as the optimization target of the implicit hate speech classification model. Compared with the existing text classification pre-trained language models such as BERT, RoBERTa, HateBERT and BERTweet that only use the cross-entropy loss function, the classification accuracy, recall rate and F1 value are significantly improved.

[0061] The method of the present invention can be used in social media speech governance and data mining, and can be especially used for the discovery and precise attack of harmful speech in cyberspace, and has broad application prospects. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] Figure 1 Schematic diagram of the steps of the implicit hate speech detection and sub-classification method based on joint contrastive learning in the embodiment;

[0063] Figure 2 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0064] The main categories of implicit hate speech and their corresponding definitions are as follows:

[0065] Grievance: This type of speech involves frustration with minority privilege and casting majority groups as the true victims of racism. This type of language is associated with extremist behavior and support for violence.

[0066] Incitement to Violence: This category includes statements that flaunt the unity and power of a group and further the ideology of a known hate group, primarily extremist statements.

[0067] Inferiority: Inferiority speech usually implies that a group or individual is inferior to another group, and mainly includes dehumanization (denying a person's humanity) and contempt (comparing the subject to diseases, insects, and animals). Inferiority speech is also related to certain language that violates human dignity, dominates, and promotes the superiority of the group. For example, "It is no coincidence that the best places to live are mostly inhabited by a certain race of people."

[0068] Irony is the use of sarcasm, humor, and ridicule to attack or demean a protected class or individual. Irony is often used by modern online hate groups to cover up their aggressive rhetoric and extremism.

[0069] Stereotypes: Statements that associate a protected class with negative attributes, such as criminality or terrorism. For example: All people of a certain group are violent.

[0070] Threatening: This type of speech conveys the speaker's claim of suffering, harm, loss, or violation of rights to the target. Example: All non-white immigration should be ended.

[0071] As described in the background technology and the main categories and corresponding definitions of implicit hate speech, it can be understood that implicit hate speech may have great differences in language expression and is more obscure and difficult to detect, but the psychological state involved includes the speaker's intention and the reader's psychological reaction is usually the same. However, current researchers have conducted little research on the intention and subject and object reactions involved in implicit hate speech. Therefore, in order to promote research in this field, the present invention provides a method for detecting and subdividing implicit hate speech based on joint contrastive learning.

[0072] See Figure 1 The present invention provides a method for detecting and classifying implicit hate speech based on joint contrastive learning, comprising the following steps:

[0073] Step 1: Collect text data from online social media, and perform data cleaning and data preprocessing on the collected text data;

[0074] Step 2: Construct an implicit hate speech detection model. The implicit hate speech detection model includes a pre-trained language model module and a fully connected network module. The pre-trained language model module is used to extract text features, and the fully connected network module is used to classify the extracted text features. The implicit hate speech detection model is trained using a labeled data set. The trained implicit hate speech detection model can distinguish non-hate speech, explicit hate speech, and implicit hate speech in text data. The collected text data is input into the trained implicit hate speech detection model to obtain implicit hate speech.

[0075] Step 3: For the detected implicit hate speech, knowledge enhancement is performed on the implicit hate text based on the common sense knowledge base ATOMIC to infer the intention of the text author, the subjective reaction of the author, and the objective reaction of the reader;

[0076] Step 4: Construct an implicit hate speech classification model. The implicit hate speech classification model includes a pre-trained language model module and a fully connected network module. The pre-trained language model module is used to extract features, and the fully connected network module is used to classify the extracted features.

[0077] Using the contrastive learning framework, implicit hate speech, intention, subject reaction, and object reaction of the same category are used as positive samples in contrastive learning to calculate the contrastive learning loss, and the four contrastive learning losses are jointly weighted and summed to obtain the joint contrastive learning loss;

[0078] The total loss function of the implicit hate speech classification model is obtained by summing the joint contrastive learning loss and the cross entropy loss, and the total loss function of the implicit hate speech classification model is optimized to obtain a trained implicit hate speech classification model;

[0079] Step 5: Input the detected implicit hate speech into the trained implicit hate speech classification model and output the classification results of implicit hate speech. The classification results of implicit hate speech are used for social media speech governance, including targeted processing by online social media companies and refined management by relevant management departments to avoid social conflicts caused by various types of implicit hate speech.

[0080] Specifically, in one embodiment of the present invention, in step 1, news text data related to social events is obtained from a news database, text data from online social media such as Sina Weibo and Twitter is collected, and data cleaning and necessary data preprocessing are performed, mainly including removing emoticons, removing Chinese and English characters, removing null values ​​and words without actual meaning, and lowercasing English characters.

[0081] Specifically, in step 2, an implicit hate speech detection model is constructed. The implicit hate speech detection model includes a pre-trained language model module and a fully connected network module. The pre-trained language model module of the implicit hate speech detection model is constructed based on the pre-trained language model PLMs. The pre-trained language model PLMs is used to obtain a given text x i The characteristic representation is expressed as:

[0082] h(x i )=PLMs(x i )

[0083] Among them, h(x i ) is the text feature obtained by pre-training language model PLMs;

[0084] The fully connected network module of the implicit hate speech detection model uses a non-linear activation function to perform supervised parameter fine-tuning training on the pre-trained language model PLMs and the non-linear activation function according to the labeled data set, so that it has the ability to detect non-hate, explicit hate, and implicit hate speech, which is expressed as:

[0085]

[0086] Among them, sigmoid is a nonlinear activation function. is the detection result of the implicit hate speech detection model, none represents non-hate speech, explicit represents explicit hate speech, and implicit represents implicit hate speech.

[0087] Finally, the collected text data is input into the trained implicit hate speech detection model to obtain implicit hate speech.

[0088] In step 3 of one embodiment of the present invention, the implicit hate speech text determined by the implicit hate speech detection model in step 2 is Based on the common sense knowledge base ATOMIC, the intentions and reactions of the person who issued the implicit hate speech text, as well as the reactions of the person who read this implicit hate speech text, are inferred. ATOMIC is an if-then structured knowledge graph in which the nodes are events and the edges between the nodes are relationship types. The relationship type mainly indicates the direction of reasoning, that is, the content of the reasoning.

[0089] The standard triples of the common sense knowledge base ATOMIC are<event,relation,event> ,The nodes in the common sense knowledge base ATOMIC are events, and the edges between nodes are relationship types, among which there are nine types of reasoning relationship types to choose from. This method mainly uses the subject intention xIntent, the subject reaction xReact and the object reaction oReact.

[0090] For example, when the event "X abuses Y" occurs, it can be inferred that the subject intention is xIntent: X wants to verbally attack someone, the subject reaction is xReact: X feels angry, and the object reaction is oReact: Y feels dissatisfied. The above can be formalized into three different triples, including: <X abuses Y, xIntent, X wants to verbally attack someone>, <X abuses Y, xReact: X feels angry>, <X abuses Y, oReact, Y feels dissatisfied>.

[0091] To perform corresponding reasoning for implicit hate speech, in the embodiment, the SentenceBERT model is used to encode the implicit hate speech and all event nodes in the ATOMIC knowledge graph to obtain their text feature representations:

[0092] h SBERT (x i ) = SentenceBERT(x i )

[0093]

[0094] Among them, h SBERT (x i ) represents that of the implicit hate speech, represents the feature representation of the event nodes in the ATOMIC knowledge graph;

[0095] After that, the cosine similarity function is used to calculate the similarity between the implicit hate text and the event nodes in ATOMIC, and the top K events with the highest similarity are selected as the event sources for implicit hate speech reasoning;

[0096] Then, based on the implicit hate speech and its corresponding top K events with the highest similarity in ATOMIC, the corresponding K possible subject intentions xintent, subject reactions xreact, and object reactions oreact are respectively inferred, and the K texts of each are concatenated and expressed as:

[0097]

[0098]

[0099]

[0100] Among them, xIntent(x i )、xReact(x i )、oReact(x i ) are respectively the implicit hate speech text (x i) is a collection of the subject intention xintent, subject reaction xreact, and object reaction oreact. They are implicit hate speech texts (x i )’s kth subject intention xintent, subject reaction xreact, and object reaction oreact.

[0101] In step 4 of one embodiment of the present invention, a new joint contrastive learning framework is designed. Contrastive learning aims to bring the representations of similar samples closer in the representation space and move the representations of different samples farther apart in the representation space. Based on the classic contrastive learning loss function, the present invention uses the subject intention, subject psychological reaction, object psychological reaction and samples of the same category obtained by reasoning of implicit hate speech as positive samples in the contrastive learning loss, which can make the representations of samples with similar intentions and psychological reactions more similar in the representation space and make the different ones more distinguishable. Then, the weighted sum of the above losses is taken and added to the cross entropy loss as the optimization target of the model.

[0102] The joint contrastive learning framework in the embodiment includes:

[0103] 401: Connect the K subject intents xIntent(x i ) as a positive sample for contrastive learning, and calculate the intention contrast loss of implicit hate speech, expressed as:

[0104]

[0105] 1 of them [k≠i] represents the indicator function, the function value is 1 if and only if k≠i, τ is a hyperparameter, f(u,v)=sim(u,v)=u T v / ‖u‖‖v‖ is the similarity calculation function, h(xIntent(x i )) represents a positive sample, Indicates the batch size set during training;

[0106] This step is mainly based on the results of previous studies, that is, implicit hate speech of the same category usually has similar intentions. For example, implicit hate speech of the "threat" category usually has the intention of the poster to intimidate and attempt to harm others.

[0107] 402: Connect the K entities to react xReact(x i ) is used as a positive sample for contrastive learning, and the subject response contrast loss is calculated, which is expressed as:

[0108]

[0109] 1 of them [k≠i]represents the indicator function, the function value is 1 if and only if k≠i, τ is a hyperparameter, f(u,v)=sim(u,v)=u T v / ‖u‖‖v‖ is the similarity calculation function, h(xReact(x i )) represents a positive sample, Indicates the batch size set during training;

[0110] The purpose of this step is similar to that of step 401, that is, implicit hate speech of the same category usually has similar subject reactions. For example, for implicit hate speech of the "threat" category, the subject reaction of the poster is usually self-satisfied, excited, etc.

[0111] 403: Connect K objects to react oReact(x i ) is used as a positive sample for contrastive learning and the object response contrast loss is calculated, which is expressed as:

[0112]

[0113] 1 of them [k≠i] represents the indicator function, the function value is 1 if and only if k≠i, τ is a hyperparameter, f(u,v)=sim(u,v)=u T v / ‖u‖‖v‖ is the similarity calculation function, h(oReact(x i )) represents a positive sample, Indicates the batch size set during training;

[0114] The purpose of this step is similar to that of step 401, that is, implicit hate speech of the same category usually has similar object reactions, for example, for implicit hate speech of the "threat" category, the reader's object reaction is usually fear, etc.

[0115] 404: The same category of implicit hate speech (x j ) as a positive sample for contrastive learning and calculate the supervised contrast loss:

[0116]

[0117] 1 of them [k≠i] , represents the indicator function, the function value is 1 if and only if k≠i, and if and only if y i =y j When the function value is 1, τ is a hyperparameter, f(u,v)=sim(u,v)=u T v / ‖u‖‖v‖ is the similarity calculation function, h((x j )) represents a positive sample, Indicates the batch size set during training;

[0118] The purpose of this step is to further bring implicit hate speech of the same category closer together in the representation space, and push samples of different categories apart in the representation space. Implicit hate speech of the same category mainly refers to samples that are both sarcastic or threatening, that is, samples of the same category.

[0119] 405: The weighted sum of the contrast loss is taken as the joint contrast loss, expressed as:

[0120]

[0121] Among them, γ xI , γ xR , γ oR , γ implicit They represent the weights of the corresponding losses respectively.

[0122] 406: When training the implicit hate speech classification model, the implicit hate speech classification model is also optimized through a data set with implicit hate speech classification labels. The implicit hate speech classification labels include but are not limited to grievance, incitement to violence, inferiority, irony, stereotype, and threat. The cross entropy loss is calculated based on the labels predicted by the implicit hate speech classification model. The cross entropy loss function is expressed as follows:

[0123]

[0124] in Denotes the predicted label, y i represents the true label, N b Indicates the size of the data batch, represents the cross entropy loss function;

[0125] 407: Add the cross entropy loss to the joint contrastive learning loss as the total loss function for the implicit hate speech classification model:

[0126]

[0127] in, Total loss function for the implicit hate speech classification model.

[0128] The implicit hate speech classification model is trained iteratively using the data from the training set until the model converges and has good discrimination.

[0129] In step 5, the detected implicit hate speech is input into the trained implicit hate speech classification model, and the classification results of the implicit hate speech are output, and fine-grained labels are given, mainly including categories such as grievance, incitement to violence, inferiority, sarcasm, stereotype, threat, etc. The classification results of the implicit hate speech obtained are used for targeted processing by online social media companies and / or refined management by management departments to avoid social conflicts caused by various types of implicit hate speech.

[0130] The implicit hate speech detection and sub-classification method based on joint contrastive learning in the above embodiment is experimentally demonstrated:

[0131] The classification performance of the implicit hate speech detection and sub-classification method BERTweet (joint-cl) based on joint contrastive learning in the present invention and the pre-trained language models including BERT, RoBERTa, HateBERT, and BERTweet using only cross entropy loss (CE) were compared in two implicit hate speech benchmark data IHC and DYNAHATE. The indicators used included precision, recall, F1 value, and accuracy.

[0132] The experimental data is shown in Table 1 below. In the test results, the implicit hate speech detection and sub-classification method BERTweet (joint-cl) based on joint contrastive learning in the present invention has better test precision, recall, F1 value and accuracy. The test results show that the method of the present invention has obvious superiority.

[0133]

[0134] Table 1

[0135] In view of the fact that implicit hate speech may have great differences in language expression and is more obscure and difficult to detect, the existing text classification pre-trained language model is not ideal for the detection and fine-grained classification of implicit hate speech. In the present invention, the psychological state involved in implicit hate speech includes the speaker's intention, and the reader's psychological reaction is usually the same. Therefore, by using the knowledge enhancement method, based on the inference knowledge base ATOMIC, the subject intention and subject psychological reaction of the poster involved in implicit hate speech, as well as the objective psychological reaction of the reader are inferred. The method of the present invention can express the implicit meaning of implicit hate speech to a certain extent, which is conducive to assisting the implicit hate speech classification model to understand its semantics and perform more accurate fine-grained classification.

[0136] The present invention designs a new joint contrastive learning framework when optimizing the implicit hate speech classification model. Based on the classic contrastive learning loss function, the subject intention, subject psychological reaction, object psychological reaction and samples of the same category of the inferred implicit hate speech are respectively used as positive samples in the contrastive learning loss, which can make the representations of samples with similar intentions and psychological reactions more similar in the representation space, and make the different ones more distinguishable. The present invention takes the weighted sum of the above losses and adds them to the cross-entropy loss as the optimization target of the implicit hate speech classification model. Compared with the existing text classification pre-trained language models that only use the cross-entropy loss function, such as BERT, RoBERTa, HateBERT and BERTweet, the classification accuracy, recall rate and F1 value are significantly improved.

[0137] Internet content security supervision is a huge project that requires various means to fight extremists. The method of the present invention can be used in social media speech governance and data mining, and can especially be used for the discovery and precise attack of harmful speech in cyberspace, and has broad application prospects.

[0138] In an embodiment of the present invention, a computer device is also provided, comprising: a processor, a memory and a program;

[0139] The program is stored in the memory, and the processor calls the program stored in the memory to execute the above-mentioned implicit hate speech detection and sub-classification method based on joint contrastive learning.

[0140] The computer device may be a terminal, and its internal structure diagram may be as follows: Figure 2 As shown. The computer device includes a processor, a memory, a network interface, a display screen and an input device connected by a bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, an implicit hate speech detection and sub-classification method based on joint contrast learning is implemented. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a key, trackball or touchpad provided on the housing of the computer device, or an external keyboard, touchpad or mouse, etc.

[0141] The memory may be, but is not limited to, a random access memory (RAM), a read only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable read-only memory (EEPROM), etc. The memory is used to store programs, and the processor executes the programs after receiving the execution instruction.

[0142] The processor can be an integrated circuit chip with signal processing capabilities. The above-mentioned processor can be a general-purpose processor, including a central processing unit (Central Processing Unit, referred to as: CPU), a network processor (Network Processor, referred to as: NP), etc. The processor can also be other general-purpose processors, digital signal processors (Digital Signal Processor, DSP), application-specific integrated circuits (Application Specific Integrated Circuit, ASIC), field programmable gate arrays (Field-Programmable Gate Array, FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The various methods, steps and logic block diagrams disclosed in the embodiments of the present application can be implemented or executed. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc.

[0143] Those skilled in the art will understand that Figure 2 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0144] In an embodiment of the present invention, a computer-readable storage medium is further provided, and the computer-readable storage medium is used to store a program, and the program is used to execute the above-mentioned implicit hate speech detection and sub-classification method based on joint contrastive learning.

[0145] Those skilled in the art will appreciate that the embodiments of the embodiments of the present invention may be provided as methods, computer devices, or computer program products. Therefore, the embodiments of the present invention may take the form of complete hardware embodiments, complete software embodiments, or embodiments combining software and hardware. Moreover, the embodiments of the present invention may take the form of computer program products implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.

[0146] The embodiments of the present invention are described with reference to flowcharts of methods, computer devices, or computer program products according to the embodiments of the present invention. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device generate a device for implementing the functions specified in the flowchart.

[0147] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing terminal device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured product including an instruction device that implements the functions specified in the flowchart.

[0148] The above is a detailed introduction to the application of the implicit hate speech detection and subclassification method based on joint contrastive learning, computer device, and computer-readable storage medium provided by the present invention. Specific examples are used in this article to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea; at the same time, for those skilled in the art, according to the idea of ​​the present invention, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be understood as limiting the present invention.

Claims

1. A joint contrastive learning-based implicit hate speech detection and classification method. It is characterized in that The following steps are involved: Step 1: Collect text data from online social media, and perform data cleaning and data preprocessing on the collected text data; Step 2: constructing an implicit hate speech detection model, wherein the implicit hate speech detection model includes a pre-trained language model module and a fully connected network module, wherein the pre-trained language model module is used to extract text features, and the fully connected network module is used to classify the extracted text features, and the implicit hate speech detection model is trained using a labeled data set, wherein the trained implicit hate speech detection model can distinguish non-hate speech, explicit hate speech, and implicit hate speech in text data, and the collected text data is input into the trained implicit hate speech detection model to obtain implicit hate speech; Step 3: For the detected implicit hate speech, knowledge enhancement is performed on the implicit hate text based on the common sense knowledge base ATOMIC to infer the intention of the text author, the subjective reaction of the author, and the objective reaction of the reader; Step 4: constructing an implicit hate speech classification model, wherein the implicit hate speech classification model includes a pre-trained language model module and a fully connected network module, wherein the pre-trained language model module is used to extract features, and the fully connected network module is used to classify the extracted features; Using the contrastive learning framework, implicit hate speech, intention, subject reaction, and object reaction of the same category are used as positive samples in contrastive learning to calculate the contrastive learning loss, and the four contrastive learning losses are jointly weighted and summed to obtain the joint contrastive learning loss; The total loss function of the implicit hate speech classification model is obtained by summing the joint contrastive learning loss and the cross entropy loss, and the total loss function of the implicit hate speech classification model is optimized to obtain a trained implicit hate speech classification model; Step 5: Input the detected implicit hate speech into the trained implicit hate speech classification model, and output the classification results of the implicit hate speech. The classification results of the implicit hate speech are used for social media speech governance.

2. According to claim 1, a method for detecting and classifying implicit hate speech based on joint contrastive learning, Features: In step 1, data cleaning and data preprocessing include removing emoticons, Chinese and English characters, removing null values ​​and words without actual meaning, and lowercasing English characters.

3. According to the method for detecting and classifying implicit hate speech based on joint contrastive learning according to claim 1, Features: In step 2, the pre-trained language model module of the implicit hate speech detection model is built based on the pre-trained language model PLMs, and the pre-trained language model PLMs is used to obtain the given text x i The characteristic representation is expressed as: h(x i )=PLMs(x i ) Among them, h(x i ) is the text feature obtained by pre-training language model PLMs; The fully connected network module of the implicit hate speech detection model uses a non-linear activation function to classify text into non-hate speech, explicit hate speech, and implicit hate speech, expressed as: Among them, sigmoid is a nonlinear activation function. is the detection result of the implicit hate speech detection model, none represents non-hate speech, explicit represents explicit hate speech, and implicit represents implicit hate speech.

4. According to claim 1, a method for detecting and classifying implicit hate speech based on joint contrastive learning, Features: In step 3, the following steps are specifically included: Implicit hate speech text identified by the implicit hate speech detection model Based on the common sense knowledge base ATOMIC, the intention and reaction of the person who sent the implicit hate speech text, as well as the reaction of the person who read this implicit hate speech text, The standard triples of the common sense knowledge base ATOMIC are<event,relation,event> ,The nodes in the common sense knowledge base ATOMIC are events, the edges between nodes are relationship types, and ,relation includes subject intention xIntent, subject reaction xReact and object reaction oReact; Use the SentenceBERT model to encode implicit hate speech and all event nodes in the ATOMIC knowledge graph to obtain their text feature representation: h SBERT (x i )=SentenceBERT(x i ) Among them, h SBERT (x i ) indicates implicit hate speech, Represents the feature representation of event nodes in the ATOMIC knowledge graph; Use the cosine similarity function to calculate the similarity between implicit hate text and event nodes in ATOMIC, and select the top K events in similarity as the event source for implicit hate speech reasoning; According to implicit hate speech And the corresponding events with the top K similarity ranking in ATOMIC, infer the corresponding K possible subject intentions xintent, subject reactions xreact, and object reactions oreact, and connect their respective K texts, expressed as: Among them, xIntent(x i )、xReact(x i )、oReact(x i ) are implicit hate speech texts (x i ) is a collection of the subject intention xintent, subject reaction xreact, and object reaction oreact. They are implicit hate speech texts (x i )’s kth subject intention xintent, subject reaction xreact, and object reaction oreact.

5. According to claim 1, a method for detecting and classifying implicit hate speech based on joint contrastive learning, Features: In step 4, the following steps are specifically included: Set the subject intent xIntent(x i ) as a positive sample for contrastive learning, and calculate the intention contrast loss of implicit hate speech, expressed as: 1 of them [k≠i] represents the indicator function, the function value is 1 if and only if k≠i, τ is a hyperparameter, is the similarity calculation function, h(xIntent(x i )) represents a positive sample, Indicates the batch size set during training; Take the main reaction xReact(x i ) as a positive sample for contrastive learning, and calculate the main reaction contrastive loss, expressed as: where 1 [k≠i] represents an indicator function that has a value of 1 if and only if k ≠ i, and τ is a hyperparameter, is a similarity calculation function, h(xReact(x i )) represents a positive sample, represents the batch size set during the training process; The object reaction oReact(x i ) is used as a positive sample for contrastive learning and the object response contrast loss is calculated, which is expressed as: 1 of them [k≠i] represents the indicator function, the function value is 1 if and only if k≠i, τ is a hyperparameter, is the similarity calculation function, h(oReact(x i )) represents a positive sample, Indicates the batch size set during training; The same category of implicit hate speech (x j ) as a positive sample for contrastive learning and calculate the supervised contrast loss: 1 of them [k≠i] , represents the indicator function, the function value is 1 if and only if k≠i, and if and only if y i =y j When the function value is 1, τ is a hyperparameter, is the similarity calculation function, h((x j )) represents a positive sample, Indicates the batch size set during training; The weighted sum of the contrast losses is used as the joint contrast loss, expressed as: Among them, γ xI , γ xR , γ oR , γ implicit are weights respectively.

6. According to claim 5, a method for detecting and classifying implicit hate speech based on joint contrastive learning, Features: In step 4, when training the implicit hate speech classification model, the implicit hate speech classification model is also optimized through a data set with implicit hate speech classification labels, and the cross entropy loss is calculated according to the labels predicted by the implicit hate speech classification model. The cross entropy loss function is expressed as follows: in Denotes the predicted label, y i represents the true label, N b Indicates the size of the data batch, represents the cross entropy loss function; The cross entropy loss is added to the joint contrastive learning loss as the total loss function for the implicit hate speech classification model: in, Total loss function for the implicit hate speech classification model.

7. The method for detecting and classifying implicit hate speech based on joint contrastive learning according to claim 6, Features: Implicit hate speech classification labels include, but are not limited to, grievance, incitement to violence, inferiority, sarcasm, stereotype, and threat.

8. A computer device, It is characterized in that It includes: a processor, a memory and a program; The program is stored in the memory, and the processor calls the program stored in the memory to execute the implicit hate speech detection and sub-classification method based on joint contrastive learning according to claim 1.

9. A computer-readable storage medium, Features: The computer-readable storage medium is used to store a program, and the program is used to execute the implicit hate speech detection and sub-classification method based on joint contrastive learning according to claim 1.

Citation Information

Patent Citations

  • Self-supervised public opinion comment viewpoint object classification method based on comparative learning

    CN114548321A

  • BERT model depolarization method and system based on comparative learning

    CN115860083A