Invisible hatred language detection method based on interpretation generation and multi-agent voting

By generating explanations and combining multi-agent voting strategies, the accuracy and reliability of implicit hatred language detection are solved, and the effective recognition and generalization ability of implicit hatred language is improved.

CN120256632APending Publication Date: 2025-07-04YANGZHOU UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510311911.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The prior art has limitations in detecting implicit hatred language, which is difficult to accurately identify and generate effective explanations, and large language models are prone to hallucinations.

Method used

By generating explanations and combining multi-angle prompt information to build large language model input, design internal and external double-layer multi-agent voting strategies, use interpretation generation and multi-agent voting methods to perform implicit hatred language detection, reduce hallucination phenomena and improve accuracy.

Benefits of technology

It significantly improves the accuracy and reliability of implicit hatred language detection, enhances the ability to generalize different languages ​​and cultural backgrounds, and reduces the randomness and outliers of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120256632A_ABST
    Figure CN120256632A_ABST
Patent Text Reader

Abstract

The invention discloses a hidden hatred language detection method based on interpretation generation and multi-agent voting, which comprises the following steps of: 1) constructing the input of a large language model LLMs in combination with prompt information of multi-angle design and text content of network comments; 2) guiding the large language model LLMs through prompt to generate explanatory output and category labels; 3) executing a single agent for multiple times, and performing internal voting based on multiple results of the agent to determine a final decision of the agent; and 4) summarizing the internal voting results of the plurality of agents in the step 3) through external voting, comprehensively making a decision, and finally determining a category label of a given network comment text to finish hidden hatred language detection. By guiding LLMs generation and explanation, the context and implicit information behind network comments are revealed, so that the hidden hatred language and the dominant hatred language are distinguished. In addition, a double-layer intelligent agent voting strategy is designed, so that the hallucination phenomenon of LLMs is reduced, and the accuracy and reliability of hidden hatred language detection are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of research on Chinese implicit hate speech and short text classification, and particularly to a method for detecting implicit hate speech based on explanation generation and multi-agent voting. Background Art

[0002] In recent years, the method of combining deep learning with dictionaries has been widely used in the field of hate speech detection. Such methods perform excellently in identifying explicit hate speech (usually containing obvious insulting or hate words), but there are still limitations in detecting implicit hate speech. The research on detecting implicit hate speech has experienced a process from early methods relying on feature engineering to the application of neural network models, and has further developed into methods based on pre-trained language models (PLMs) in recent years, such as BERT and HateBERT. These PLM-driven technologies have shown excellent performance improvement in the hate speech detection task.

[0003] Due to the particularity of implicit hate speech, recent research has begun to explore providing underlying explanations to assist detection. Some research efforts are dedicated to generating keyword-based explanations for detection, but such methods often fail to capture the implicit hate content that is not explicitly expressed in the text. Other research uses trained generative models to detect through free-text reasoning written manually. Although these methods have certain potential, they are limited by the logical inconsistencies in annotation reasoning, resulting in less than ideal detection and explanation effects. There is also research that utilizes external knowledge sources, task decomposition, and knowledge injection, treating hate speech detection as a few-shot learning task to improve generalization ability and detection performance. However, the explanations generated by these methods are not directly used for detection, and at the same time, the introduced large language models are prone to hallucination phenomena. Summary of the Invention

[0004] The purpose of the present invention is to overcome the defects of the prior art and provide a method for detecting implicit hate speech based on explanation generation and multi-agent voting. By generating explanations to reveal the implicit context behind the speech, it can effectively distinguish implicit hate speech, explicit hate speech, and non-hate speech; use these generated explanations to guide the LLMs for detection, and design an internal and external double-layer multi-agent voting strategy to reduce the hallucination phenomena that may occur in the LLMs, significantly improving the accuracy and reliability of classification.

[0005] The purpose of the present invention is achieved as follows: A method for detecting implicit hate speech based on explanation generation and multi-agent voting, comprising the following steps:

[0006] 1) Combine the prompt information designed from multiple perspectives with the text content of online reviews to construct the input of the large language model LLMs;

[0007] 2) Guide large language models (LLMs) to generate explanatory outputs and category labels through prompts;

[0008] 3) A single agent executes multiple times and conducts internal voting based on its own multiple results to determine the final decision of the agent;

[0009] 4) Aggregate the internal voting results of multiple agents in step 3) through external voting, make comprehensive decisions, and finally determine the category label of the given network comment text to complete the detection of implicit hate language.

[0010] As a further limitation of the present invention, the step 1) specifically comprises:

[0011] Step 1.1) Design prompts to guide large language models (LLMs). Consider multiple perspectives, including sentiment, themes, context, user background, historical background, and the definition of implicit hate language, and integrate these perspectives into prompts. The specific perspectives of online comment analysis include:

[0012] Sentiment analysis: Analyze whether the post content contains hateful sentiment by evaluating the post’s subject, content, and emotional tone;

[0013] Theme and Topic Analysis: Analyze the intent and nature of posts by evaluating their subject, content, and relevant context;

[0014] Contextual and situational analysis: Identifying seemingly neutral speech that may contain hidden hate by analyzing the social and cultural context in which the speech is expressed;

[0015] User background and historical background analysis: Identify the hateful tendencies hidden in the user's speech patterns by analyzing the user's historical behavior and past speech;

[0016] Definition of implicit hate language: By defining the characteristics of implicit hate language, LLMs are guided to identify words or texts that appear neutral or normal on the surface, but actually contain negative emotions, prejudice or discrimination against specific groups, individuals or identities, and express hatred or hostility in a subtle and obscure way through suggestion, irony, exaggeration and metaphor;

[0017] Step 1.2) The prompt word set P constructed from the above perspectives = {p1, p2, ..., p out}, p i is the i-th prompt word prompt in the prompt word set, and the number of prompt words out is set to an odd number.

[0018] As a further limitation of the present invention, the step 2) specifically comprises:

[0019] Step 2.1) Each prompt word p i Guide the LLMs to interpret the online comments and determine whether the comment is malicious based on the interpretations obtained from the LLMs; given the prompt word p i The j-th detection result is determined jointly by the generated interpretation and the label as follows:

[0020]

[0021] Step 2.2) Add JSON format constraints to the output, specifying that "interpretation" and "label" are complete, i.e., add a constraint statement at the end of each prompt word prompt, and the "interpretation" and "label" in the answer are enclosed in square brackets "[]", and neither the interpretation nor the label can be empty.

[0022] As a further limitation of the present invention, step 3) specifically includes:

[0023] Step 3.1) Internal voting is carried out within a single agent; internal voting can smooth randomness and reduce errors in a single output. For each prompt word p i , generate in detection results R i :

[0024]

[0025] The number of internal voting rounds in is set to an odd number;

[0026] Step 3.2) The final result V of the internal voting i is obtained by calculating the most frequently occurring category in R i as follows:

[0027]

[0028] where the Mode function is used to determine the value that appears most frequently in the set R i , filtering out the influence of random errors and outliers; r represents an element in the set R i , and Frequency(r) represents the number of times r appears in R i .

[0029] As a further limitation of the present invention, step 4) specifically includes:

[0030] Step 4) External voting is carried out among different agents; based on the internal voting results, external voting further aggregates the detection results of different agents;

[0031] The internal voting results V of all agentsi After the out-round external voting, the calculation method is as follows:

[0032] V final = Mode({V1, V2, …, V out [[ID=8}]}) (4)

[0033] The final detection result V final is the category with the highest frequency among the internal voting results V i of all agents, so as to balance with the voting results of other agents when there are defects in the design of a certain agent or prompt word; the number of external voting rounds out is set to an odd number.

[0034] Adopting the above technical solutions, compared with the prior art, the beneficial effects of the present invention are as follows: (1) By generating and filling in the humanistic history, cultural stories and their true intentions behind the remarks, the present invention provides richer context information, enabling LLMs to more accurately detect implicit hate speech; (2) The present invention designs a double-layer voting mechanism combining internal voting of agents and cross-agent voting, effectively reducing the hallucination phenomenon that may occur in the detection task of LLMs, and is more robust than a single detection model; (3) The present invention adopts the zero-shot method, enabling the model to perform implicit hate speech detection without additional labeled data, enhancing the generalization ability for different languages, cultural backgrounds and new types of implicit hate speech. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 The overall framework diagram of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0036] As Figure 1 shown in the implicit hate speech detection method based on explanation generation and multi-agent voting, includes the following steps:

[0037] 1) Combine the prompt information designed from multiple angles with the text content of the online comment to construct the input of the large language model LLMs;

[0038] Step 1.1) Design the prompts for guiding the large language model LLMs, by comprehensively considering from multiple angles, including aspects such as emotion, theme and topic, context and situation, user background and historical background, and the definition of implicit hate speech, and integrating these angles into the prompts; the specific angles of online comment analysis involved in the present invention include:

[0039] Sentiment analysis: By evaluating the theme, content and emotional color of the post, analyze whether the post content contains hate emotions; even if it is seemingly neutral, sentiment analysis can identify hidden biases or hostilities, helping to judge whether the speech has negative emotions;

[0040] Subject and Topic Analysis: Analyze the intention and nature of a post by evaluating its subject, content, and relevant background;

[0041] Context and Situation Analysis: Identify seemingly neutral but potentially implicitly hateful remarks by analyzing the social and cultural environment in which the speech is made;

[0042] User Background and Historical Background Analysis: Identify implicit hate tendencies in a user's speech pattern by analyzing the user's historical behavior and past remarks; Historical data can reveal the user's cultural and social stance and help judge potential biases in their speech;

[0043] Definition of Implicit Hate Language: By defining the characteristics of implicit hate language, guide LLMs to identify speech or text that seemingly appears neutral or normal on the surface but actually harbors negative emotions, biases, or discrimination against specific groups, individuals, or identities, and expresses hatred or hostility in a subtle and implicit way through means such as implication, irony, exaggeration, and metaphor;

[0044] Step 1.2) The set of prompts P = {p1, p2, …, p out} constructed from the above perspectives, where p i is the i-th prompt prompt in the set of prompts, and the number out of prompts is set to be odd;

[0045] In this embodiment, out is set to 5. The specific content of the set of prompts is shown in the following table:

[0046]

[0047]

[0048]

[0049] 2) Guide the large language model LLMs to generate explanatory outputs and class labels through prompts;

[0050] Step 2.1) Each prompt p i guides LLMs to interpret the online comment and determine whether the comment is malicious based on the interpretation obtained from LLMs; Specifically, under the given prompt p i , the j-th detection result is jointly determined by the generated interpretation and the label , and the calculation method is as follows:

[0051]

[0052] Step 2.2) Add JSON format constraint output, specifying that "explanation" and "label" are complete, i.e., add constraint statement output at the end of each prompt. In the answer, "explanation" and "label" are enclosed in square brackets "[]", and neither the explanation nor the label can be empty.

[0053] The detailed format example is as follows:

[0054]

[0055] 3) A single agent executes multiple times and conducts an internal vote based on its own multiple results to determine the final decision of the agent;

[0056] Step 3.1) The internal vote is conducted within a single agent; the internal vote can smooth randomness and reduce errors in a single output. For each prompt p i , generate in detection results R i :

[0057]

[0058] The number of internal voting rounds in is set to an odd number;

[0059] Step 3.2) The final result V of the internal vote i is obtained by calculating the most frequently occurring category in R i as follows:

[0060]

[0061] Among them, the Mode function is used to determine the value that appears most frequently in the set R i , filtering out the influence of random errors and outliers; r represents an element in the set R i , and Frequency(r) represents the number of times r appears in R i .

[0062] 4) Aggregate the internal vote results of multiple agents in step 3) through external voting, make a comprehensive decision, and finally determine the category label of the given network review text to complete the detection of implicit hate speech;

[0063] The external vote is conducted among different agents; based on the internal vote results, the external vote further aggregates the detection results of different agents;

[0064] The internal vote results V of all agents i go through out rounds of external voting, and the calculation method is as follows:

[0065] V final = Mode({V1, V2,..., V out) (4)

[0066] The final detection result V final is the internal voting result V of all agents i which is the category with the highest frequency of occurrence in V. Thus, when there are defects in the design of a certain agent or prompt word, the voting results of other agents are used for balance; the number of external voting rounds out is set to an odd number.

[0067] To verify the performance of the present invention in implicit hate speech detection, experiments were conducted on four datasets; these datasets were all divided into two categories: "friendly" and "malicious";

[0068] The datasets include two well-known English datasets SBIC and Latent Hatred (LHd), and two Chinese datasets ToxiCN and ProsCons. Data preprocessing and statistical analysis were performed on the above four datasets, and the results are summarized in Table 1.

[0069] Table 1 Dataset Statistics

[0070]

[0071] To ensure that the detection effect can be truly measured, two representative evaluation metrics, Accuracy and F1, were selected, and their specific definitions are as follows:

[0072]

[0073] Here,

[0074]

[0075] where tp is the number of malicious comments correctly predicted by the algorithm, fp is the number of comments predicted as malicious but actually friendly, fn is the number of comments predicted as friendly but actually malicious, tn is the number of friendly comments correctly predicted, N is the total number of network comments predicted, and the F1 score is the harmonic mean of precision and recall.

[0076] To demonstrate the performance of the test results, other traditional baseline methods are adopted on four datasets for comparison with the proposed implicit hate speech detection method based on the explanation generation-based two-layer multi-agent voting mechanism of the present invention. These baseline methods include 1) deep neural network-based methods: hate speech detection method based on sentiment features (SKS); 2) pre-trained language model-based methods: HateBERT, contrastive learning method for generating machine statements (ConPrompt); 3) prompt learning-based methods: prompt learning (PL), soft template prompt fine-tuning (Soft), knowledge-based prompt learning from external knowledge (KPT), upgraded knowledge-based prompt learning from external knowledge (KPT++); 4) large language model-based methods: GPT-3.5, interpretive hate speech detection method (Fr-HARE).

[0077] In the experiment, the zero-shot method is adopted for the method of the present invention and the large model interpretive method Fr-HARE. For the few-shot scenarios of the model prompt learning method and the large language model method GPT-3.5, 5 positive and negative samples are randomly selected from the training set. For the positive and negative sample training numbers of the deep neural network and the pre-trained language model, they are 400, 400, 400, and 200 on the datasets SBIC, LHd, ToxiCN, and ProsCons respectively. The test results of the datasets are shown in Table 2. It can be seen from Table 2 that the method of the present invention outperforms other methods in all four metrics on the four datasets.

[0078] Table 2 Experimental Results

[0079]

[0080] The present invention proposes an implicit hate speech detection method based on explanation generation and multi-agent voting. Explanations are generated through LLMs to identify the human history, cultural stories, and true intentions behind the posts. By designing a two-layer multi-agent voting strategy, internal voting and external voting are carried out within a single agent and between different agents respectively, thus effectively suppressing the hallucination phenomenon that may occur in large models. A large number of experiments have verified the excellent performance and remarkable effectiveness of this method on four datasets.

[0081] The present invention is not limited to the above embodiments. Based on the disclosed technical solutions of the present invention, those skilled in the art can make some substitutions and deformations to some technical features without creative labor according to the disclosed technical content, and these substitutions and deformations are all within the protection scope of the present invention.

Claims

1. A method for detecting implicit hate speech based on explanation generation and multi-agent voting, characterized in that Including the following steps: 1) Combine the prompt information designed from multiple perspectives with the text content of online reviews to construct the input for the large language model LLMs; 2) Guide the large language model LLMs to generate explanatory outputs and category labels through prompts; 3) Execute a single agent multiple times and conduct internal voting based on its own multiple results to determine the final decision of the agent; 4) Aggregate the internal voting results of multiple agents in step 3) through external voting, make a comprehensive decision, and finally determine the category label of the given online review text to complete the detection of implicit hate speech.

2. The implicit hate speech detection method based on explanation generation and multi-agent voting according to claim 1, characterized in that The specific content of step 1) includes: Step 1.1) Design prompts for guiding the large language model LLMs. Through comprehensive consideration from multiple perspectives, including sentiment, theme and topic, context and situation, user background and historical background, and the definition of implicit hate speech, integrate these perspectives into the prompts; The specific analysis perspectives of online reviews involved include: Sentiment analysis: Analyze whether the post content contains hate sentiment by evaluating the theme, content, and emotional color of the post; Theme and topic analysis: Analyze its intention and nature by evaluating the theme, content, and relevant background of the post; Context and situation analysis: Identify words that seem neutral on the surface but may contain implicit hate by analyzing the social and cultural environment in which the speech is made; User background and historical background analysis: Identify the implicit hate tendency in the user's speech pattern by analyzing the user's historical behavior and past remarks; Definition of implicit hate speech: By defining the characteristics of implicit hate speech, guide the LLMs to identify words or texts that seemingly appear neutral or normal on the surface but actually hide negative emotions, prejudices, or discrimination against specific groups, individuals, or identities, and express hatred or hostility in a subtle and implicit way through means such as implication, irony, exaggeration, and metaphor; Step 1.2) In the set of prompts P = {p1, p2, …, p out} constructed through the above angles, p i is the i-th prompt in the set of prompts, where the number out of the prompts is set to be odd.

3. The implicit hate speech detection method based on explanation generation and multi-agent voting according to claim 1, characterized in that The specific content of step 2) includes: Step 2.1) Each prompt p i guides the LLMs to interpret the online comments and determines whether the comment is malicious based on the interpretations obtained from the LLMs; given the prompt p i the j-th detection result is determined by the generated interpretation and the label jointly, and the calculation method is as follows: Step 2.2) Add JSON format constraints to the output, specifying that "explanation" and "label" are complete, that is, add a constraint statement output at the end of each prompt. The "explanation" and "label" in the answer are enclosed in square brackets "[]", and neither the explanation nor the label can be empty.

4. The implicit hate speech detection method based on explanation generation and multi-agent voting according to claim 1, characterized in that The specific content of step 3) includes: Step 3.1) The internal voting is carried out within a single agent; the internal voting can smooth randomness and reduce errors in a single output. For each prompt word p i , in detection results R are generated i : Set the number of internal voting rounds in to an odd number; Step 3.2) The final result V of the internal voting i Obtained by calculating the most frequently occurring category in R i as follows: Among them, the Mode function is used to determine the value that appears most frequently in the set R i and filter out the influence of random errors and outliers; r represents an element in the set R i and Frequency(r) represents the number of times r appears in R i among them.

5. The implicit hate speech detection method based on interpretation generation and multi-agent voting according to claim 1, characterized in that The specific content of step 4) includes: Step 4) External voting is carried out among different agents; Based on the internal voting results, external voting further aggregates the detection results of different agents; The internal voting results V of all agents i After out rounds of external voting, the calculation method is as follows: V final = Mode({V1, V2, …, V out}) (4) The final detection result V final is the internal voting result V of all agents i with the highest frequency of occurrence in the category, so as to balance with the voting results of other agents when there are defects in the design of a certain agent or prompt word; the number of external voting rounds out is set to an odd number.

Citation Information

Cited By

  • Commercial customer service system based on large model emotion recognition labeling and correction

    CN120973950A

  • Knowledge distillation-based online game dialogue malicious language interpretable detection method

    CN121524350A