Harmful knowledge graph construction and harmful information identification method based on large language model
By building a harmful knowledge graph and combining the semantic understanding ability of large language models, the missed and mis-checked problems of harmful information recognition in the existing technology are solved, and the recognition effect and generalization are improved.
Patent Information
- Application Number
- CN202510310645.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-06-20
AI Technical Summary
The prior art has missed checks, high false check rates, and lack of in-depth semantic understanding and background knowledge in the identification of harmful information, resulting in unsatisfactory recognition results.
The harmful knowledge graph construction method based on the large language model is adopted, and harmful text data is collected and preprocessed, and the large language model is used for in-text learning and triple extraction, and the external knowledge graph is used for identification.
It significantly reduces the false positive rate of harmful information identification, improves the recognition performance of implicit harmful information, and improves the generalization and adaptability of the model.
Smart Images

Figure CN120179831A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of natural language processing and knowledge graph mining, and in particular to a method for constructing a harmful knowledge graph and identifying harmful information based on a large language model. Background Art
[0002] Social media has become the main source of information for people all over the world. At the same time, social media also brings the risk of spreading harmful information. To address this risk, traditional methods are to find harmful information from a vast amount of text data by establishing filtering rules and improving the sensitive word list, so as to purify the network environment. This rule-based method is simple, fast, and can quickly update the sensitive word list to adapt to different public opinion environments.
[0003] With the development of technology, the technology of using machine learning to identify harmful information has gradually emerged. Representative methods include using word vectorization methods such as word2vec to obtain text representations, and then training classifiers for classification and other methods. The machine learning method considers more context semantic information than the rule method, and has a certain degree of generalization and universality.
[0004] With the development of natural language processing technology, the classification method based on the pre-trained language model BERT has gradually become the mainstream. BERT can capture the deep semantic information and complex context relationships in the text through bidirectional pre-training on a vast amount of text data. At the same time, because BERT has a relatively small number of parameters, it can be easily fine-tuned to adapt to the identification of harmful texts in different fields.
[0005] Deficiencies of the prior art: I) The rule-based method has a high rate of missed checks and false checks, a high cost of maintaining the sensitive word library, and lacks context semantic information Traditional rule-based methods often predefine a sensitive word library and then use methods such as regular expressions for matching. This method first requires manual maintenance of a sensitive word library, which has a relatively high maintenance cost and needs to be updated as the social media language is updated. At the same time, because the sensitive word library cannot cover all sensitive words, this method has a high missed check rate; and because rule-based matching ignores context semantic information, it is easy to cause text that is not originally harmful information to be mischecked just because two words happen to be connected together and match a certain sensitive word. Generally speaking, the recognition effect of this method is relatively poor.
[0006] II) The machine learning-based method lacks the ability of in-depth semantic understanding and is difficult to identify implicit harmful texts Machine learning-based methods rely on text vectorization methods such as word2vec and TF-IDF. However, these methods are all developed based on the bag-of-words model. Although they can understand the meaning of individual words, they cannot capture the connections between contexts. This also leads to the fact that although these methods can capture more semantic information, they still cannot handle implicit harmful information such as metaphors and irony.
[0007] III) The method based on the BERT model lacks background knowledge related to harmful information and has insufficient semantic understanding and expression ability Although the method based on the BERT model has undergone bidirectional pre-training, enabling the model to learn the associations between contexts, since harmful information often involves knowledge in many aspects of the real world such as history, geography, society, and humanities, the BERT model lacks the understanding and learning of this type of knowledge, resulting in unsatisfactory results in text judgments involving relevant backgrounds. Summary of the Invention
[0008] The purpose of the present invention is to provide a method for constructing a harmful knowledge graph and identifying harmful information based on a large language model. The harmful knowledge graph mines background knowledge related to harmful information, filling the gap in the lack of this aspect of knowledge in the large language model; harmful information identification utilizes the semantic understanding ability of the large language model and the external knowledge of the knowledge graph to obtain better generalization ability and adaptability, and can improve the effect of identifying harmful information on social media.
[0009] The specific technical solution for achieving the purpose of the present invention is as follows: A method for constructing a harmful knowledge graph and identifying harmful information based on a large language model, the method comprising the following steps: Step 1: Collect harmful text information data All kinds of remarks involving gender discrimination, racial discrimination, regional discrimination, and personal attacks on network platforms belong to harmful text information data; a large number of existing studies have collected network text data and carried out manual annotation according to whether the text is harmful and what kind of harm it belongs to. Collect multiple such data sets as initial data; then perform data preprocessing, screen out harmful content from the data sets, and unify all labels marked as harmful, offensive, and insulting in different data sets as harmful. Step 2: Input the initial data obtained in Step 1 into the large language model. By constructing multi-example prompt words for the harmful text explanation task, show the large language model how to explain the toxicity source of harmful information, so as to utilize the in-text learning ability of the large language model to enable the large language model to learn this ability and output the toxicity source of harmful information. Step 3: Integrate the harmful text processed in Step 1 and the explanations output by the large language model in Step 2. By constructing multi-instance prompt words for the triple (i.e., the text structure of subject, predicate, and object) extraction task, show the large language model how to perform triple extraction. Also, utilize the in-text learning ability of the large language model to enable it to understand this task and master this ability, so that the large language model outputs the core triples of the toxicity source of harmful information; Step 4: Take the triples output in Step 3 as input. First, use regular expressions to check their formats, delete the outputs that do not meet the format requirements, and then construct multi-instance prompt words for the triple checking task. Let the large language model perform self-checking to determine whether the content of the output triples meets the requirements of being toxic and harmful; Step 5: Take the head and tail entity nodes of the triples after checking as input, use the BERT pre-trained model to obtain their word representations, and aggregate different names of the same entity together through a clustering algorithm as the nodes of the knowledge graph. Take the predicates connecting the two nodes in the triples as the edges in the knowledge graph; Step 6: Use the large language model and design multi-instance prompt words for the entity extraction task to enable the large language model to understand the entity extraction task, and then extract the core entities from the text to be judged; Step 7: Use the extracted core entities to perform retrieval on the knowledge graph to find the corresponding entities on the knowledge graph; Step 8: Arrange the retrieved entities in pairs, search for the shortest path on the knowledge graph with them as the starting point and the ending point, and take all the triples on the searched shortest path as candidate triples; finally, organize the candidate triples into text form and return; Step 9: Input the candidate triple text returned in Step 8 into the SentenceBERT pre-trained model to obtain the representations, filter and sort the similarity between its representations and the representations of the text to be judged. In the order of similarity from high to low, provide the triples to the large language model as supplementary knowledge, and make a final judgment by extracting the prediction probability of the large language model output being toxic or non-toxic.
[0010] Furthermore, the data preprocessing described in Step 1 includes: 1.1: Filter the texts labeled as non-toxic, and only retain the harmful information texts for subsequent knowledge graph generation; 1.2: Align the text labels of different data sources with different granularities, and unify the fine-grained classifications of hatred against race and hatred against region into harmful; 1.3: Align the text labels of different data sources, and unify all the labels with different expressions: offensive, hateful, and harmful into harmful.
[0011] Further, the construction of multi-example prompt words for the harmful text interpretation task described in step 2 specifically includes: 2.1: Task definition, requiring the large language model to output the reasons why this text is judged to be harmful; 2.2: Output format, requiring the large language model to output the reasons in text format without outputting other irrelevant content; 2.3: Multi-example prompt, leveraging the in-context learning ability of the large language model, and improving the performance of the large language model in the task through multiple artificially constructed example prompts.
[0012] Further, the construction of multi-example prompt words for the triple extraction task described in step 3 specifically includes: 3.1: Task definition, requiring the large language model to extract the core triples from the text and the explanation, and this triple contains the key information leading to the harmfulness of the text; and defines the format of the triple as the subject, predicate, and object format; 3.2: Output format, requiring the large language model to output the triples in JSON format without outputting other irrelevant content; 3.3: Multi-example prompt, leveraging the in-context learning ability of the large language model, and improving the performance of the large language model in the task through multiple artificially constructed example prompts.
[0013] Further, the construction of multi-example prompt words for the triple checking task described in step 4 specifically includes: 4.1: Task definition, requiring the large language model to screen out the triples that summarize the harmful information from the given text and triples, and not output irrelevant and harmless triples; 4.2: Output format, requiring the large language model to output the triples in the same JSON format as in step 3 without outputting other irrelevant content; 4.3: Multi-example prompt, leveraging the in-context learning ability of the large language model, and improving the performance of the large language model in the task through multiple artificially constructed example prompts.
[0014] Further, step 5 specifically includes: 5.1: Extract the head and tail entity nodes of the triple, and obtain their corresponding word vector representations through the BERT pre-trained model; 5.2: Cluster the word vector representations, and cluster the entity names with high similarity into the same cluster; 5.3: There are multiple names pointing to the same entity in the same cluster, and use the most frequently occurring entity name as the name of the entity pointed to in the cluster.
[0015] Further, the design of multi-example prompt words for the entity extraction task described in step 6 includes: 6.1: Task definition, requiring the large language model to extract the entities involved from the given text; 6.2: Output format, requiring the large language model to output the entities in a list format without any additional content; 6.3: Multi-example prompt, leveraging the in-context learning ability of the large language model, and enhancing the performance of the large language model in the task through multiple artificially constructed example prompts; Furthermore, the retrieval on the knowledge graph described in step 7 specifically includes: 7.1: Obtain the semantic representations of the nodes on the knowledge graph and the entities extracted in step 6 through a pre-trained BERT model; 7.2: Store the node representations on the knowledge graph into the Faiss vector database to accelerate querying; 7.3: Use the semantic representations of the entities extracted in step 6 to perform similarity retrieval in the Faiss vector database to find the most similar entities; Furthermore, step 9 specifically includes: 9.1: Use a pre-trained SentenceBERT model to obtain the sentence representations of the text to be judged, and at the same time obtain the sentence representations of the candidate triple texts obtained in step 8; 9.2: Calculate the similarity between the sentence representation of the text to be judged and the sentence representations of the candidate triples; 9.3: Sort the candidate triples in descending order of similarity to the sentence representation of the text to be judged, and discard the candidate triples with similarity lower than the threshold; 9.4: Embed the finally obtained triples as supplementary knowledge into the prompt for the final judgment. The final prompt includes: 9.4.1: Task definition, requiring the model to judge whether the given text is toxic with reference to the provided triple supplementary knowledge; 9.4.2: Return format, asking the model to choose between toxic and non-toxic, and the model only needs to return one option; 9.4.3: Multi-example prompt words, by providing multiple artificially constructed examples, enabling the model to learn how to utilize the provided triple supplementary knowledge to improve the model's performance in this task; 9.5: Extract the probabilities of the model choosing two different options, and select the option with the higher probability as the model's judgment result; at the same time, the model generates the basis and reasons for the judgment; make the final judgment by extracting the predicted probabilities of the large language model outputting toxic or non-toxic.
[0016] Advantages compared to existing methods: 1) Significantly reduced the false positive rate of harmful information identification: By injecting external knowledge into the large language model and guiding its thinking direction, the large language model is prevented from being overly sensitive to some entities, thus avoiding the false positive problem and protecting freedom of speech on social media.
[0017] 2) Improved the identification performance of implicit harmful information: By leveraging the semantic understanding ability of the large language model and the supplement of external knowledge, the large language model can understand and identify the content expressed by relatively implicit harmful information.
[0018] 3) Enhanced the generalization of the model: By continuously and expandably mining the harmful knowledge graph, the present invention can adapt to the update and iteration of social media language, and achieve good generalization and universality. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 is a flowchart for constructing the harmful knowledge graph of the present invention; Figure 2 is a flowchart for identifying harmful information using a large language model and a harmful knowledge graph. DETAILED DESCRIPTION OF THE INVENTION
[0020] The present invention will be described in detail below with reference to the accompanying drawings and embodiments.
[0021] Refer to Figure 1 and Figure 2 The present invention specifically includes: 1) Data construction Data definition: Harmful information refers to various remarks on the Internet that involve gender discrimination, racial discrimination, regional discrimination, and personal attacks.
[0022] Data collection: Since a large number of existing studies have recruited data annotators to annotate harmful information, three already annotated datasets were directly collected. They have different levels and types of annotation information.
[0023] Data preprocessing: In order to unify the different annotation granularities and types of different datasets, the annotation granularities of all datasets were aligned to the relatively coarse-grained annotations of "harmful" and "non-toxic and harmless". At the same time, all annotations of "toxic", "hateful", and "offensive" were unified into the "harmful" category for subsequent processing. For verification, the dataset was divided in a ratio of 9:1, and one-tenth of the data was reserved for testing.
[0024] 2) Knowledge graph construction Reason explanation: Since harmful information often may be conveyed through relatively implicit semantic information, directly extracting triples may lead to poor results. Therefore, a step for reason explanation is designed. By informing the model that the given text is harmful and designing a multi-instance prompt template, the large language model can infer from the given text the reasons why the text is labeled as hateful, thus assisting the subsequent triple extraction step.
[0025] Triple extraction: This step mainly relies on the original text and supplements it with the reasons generated in the previous step, and uses the large language model to extract triples expressing hatred from it. Since the large language model may have hallucination problems, a self-check mechanism is designed. First, use regular expressions to check the output format of the large language model and delete those that do not meet the format requirements; then input the remaining triples into the large language model again and require it to screen out the triples expressing hatred information. Through this mechanism, the hallucination of the model is suppressed, and the noise that may be introduced in the construction of the knowledge graph is reduced.
[0026] Entity resolution: In the knowledge graph constructed in the previous step, some entity nodes actually point to the same entity, but different names are used to express it. This step can reduce the noise in the knowledge graph caused by inconsistent entity names, and at the same time can reduce the space occupancy of the knowledge graph and improve the retrieval efficiency on the knowledge graph.
[0027] 3) Knowledge graph retrieval enhances harmful information recognition In order to enable the large language model to obtain sufficient and accurate supplementary knowledge when identifying harmful information, the following five steps are designed: Entity extraction: For the new text to be identified, first, the large language model needs to be used to extract potentially relevant entities from the text. Through the design of multi-instance prompts, the large language model can accurately extract relevant entities from it.
[0028] Node query: Since the entity names extracted by the large language model may not be the standard names in the knowledge graph, directly using the names for matching will lead to a decline in performance. The nodes in the knowledge graph are pre-generated with node representations through a pre-trained BERT model and stored in a vector database to support subsequent efficient queries. Then, when identifying harmful information, only need to use the pre-trained BERT model to obtain the entity representations extracted in the previous step and then retrieve them in the vector database. To avoid retrieving irrelevant entities, a threshold is set for control.
[0029] Path Retrieval: Through the node query in the previous step, multiple relevant nodes in the knowledge graph can be obtained. Since the text is related to these nodes simultaneously, a shortest path retrieval method is designed to obtain the triples related to the text. The specific method is as follows: First, combine the retrieved entity nodes in pairs, and then for each pair, find the shortest path between these two nodes on the knowledge graph.
[0030] Path Textualization: The paths retrieved in the previous step are returned in the form of triples, but large language models are difficult to directly understand triples and need to be converted into text form to assist understanding. Since the triples are already in the form of subject-predicate-object, the corresponding text form can be obtained by simply linking them in order.
[0031] Similarity Ranking and Filtering: Since there may be triples on the path that are not relevant to the original text, the path texts obtained in the previous step and the query text are encoded through a pre-trained language model to obtain their corresponding representations. Then, the similarity between the representation of the query text and the representation of the path text is calculated, and the path texts are sorted in descending order of similarity to the query text representation. The path texts below the set threshold are filtered out to obtain the final knowledge text for enhancing the large prediction model.
[0032] 4) Model Input and Prediction The knowledge content supplemented to the large language model is obtained in the previous step, which needs to be incorporated into the multi-example prompt template and input to the large language model. To avoid errors in the large language model's instruction compliance and reduce resource consumption during the large language model's generation, the task is designed as a multiple-choice question in the prompt template, using "a" to represent harmful information and "b" to represent non-toxic and harmless information. At the same time, there is no need for the model to output text. Only the probabilities of the model's last word being predicted as "a" and "b" are taken, and the size relationship between the two is compared, and the relatively larger one is taken as the final answer.
[0033] 5) Method Evaluation Two methods of directly using the large language model for judgment (naive LLM) and supplementing knowledge to the large language model through direct retrieval-enhanced generation (RAG) are selected as baselines. To evaluate the effectiveness and generalization of this method, two types of experiments are designed: To verify the effectiveness of the method, it is tested on the dataset for constructing the knowledge graph, so that the knowledge graph and the test data come from the same source; To verify the generalization of the method, a knowledge graph is constructed on one dataset and then tested on the test set of another dataset, so that the knowledge graph and the test data come from different scenarios.
[0034] The evaluation metrics include accuracy (Acc.), F1-score (F1), area under the precision-recall curve (AUC), and false positive rate (FPR). And Qwen2.5-14B-Instruct and Llama3.1-8B-Instruct are selected as the base models. Finally, the experimental results obtained by testing on the same-source dataset are shown in Table 1, and the test results on different datasets are shown in Table 2. The optimal results are shown in bold. The method of the present invention is superior to the baseline method in various scenarios, significantly improving the generalization ability and significantly reducing the false positive rate, and avoiding the harm of harmful information identification to freedom of speech.
[0035] Table 1 Performance comparison of the present invention with other large language model-based methods under the same-source dataset
[0036] Table 2 Performance comparison of the present invention with other large language model-based methods under different-source datasets
Claims
1. A method for constructing harmful knowledge graph and identifying harmful information based on a large language model, characterized in that: The method comprises the following steps: Step 1: Collect harmful text message data All kinds of speech involving gender discrimination, racial discrimination, regional discrimination and personal attacks on online platforms are harmful text information data; a large number of studies have collected online text data and manually annotated the text according to whether it is harmful and what kind of harm it is, and collected multiple such data sets as initial data; Then, data preprocessing was performed to filter out harmful content from the dataset, and all labels marked as harmful, offensive, and insulting in different datasets were unified as harmful; Step 2: Input the initial data obtained in step 1 into the large language model, and show the large language model how to explain the toxicity source of the harmful information by constructing multiple example prompt words for the harmful text interpretation task, so as to use the in-text learning ability of the large language model to enable the large language model to learn this ability and output the toxicity source of the harmful information; Step 3: Integrate the harmful text processed in step 1 with the explanation output by the large language model in step 2. By constructing multiple example prompt words for the triple extraction task (i.e., text with subject, predicate, and object structure), the large language model is shown how to extract triples. The large language model is also enabled to understand the task and master the ability to learn in text by using its in-text learning capability, so that the large language model can output the core triples of the source of the toxicity of the harmful information. Step 4: Take the triples outputted in step 3 as input, first use regular expressions to check their format, delete the outputs that do not meet the format requirements, then construct multiple example prompt words for the triple checking task, and use the large language model to perform self-checking to determine whether the content of the output triples meets the requirements of being toxic or harmful; Step 5: Take the head and tail entity nodes of the checked triples as input, use the BERT pre-trained model to get their word representations, use the clustering algorithm to aggregate different names of the same entity together as nodes of the knowledge graph, and use the predicates that link the two nodes in the triples as edges in the knowledge graph; Step 6: Use a large language model and design multiple example prompt words for the entity extraction task, so that the large language model can understand the entity extraction task and then extract the core entities from the text to be judged; Step 7: Use the extracted core entities to search on the knowledge graph and find the corresponding entities on the knowledge graph; Step 8: Arrange the retrieved entities in pairs and use them as the starting point and end point to search for the shortest path on the knowledge graph. All triples on the shortest path searched are used as candidate triples. Finally, the candidate triples are organized into text form and returned. Step 9: Input the candidate triple text returned in step 8 into the SentenceBERT pre-trained model to obtain a representation, filter and sort its representation and the representation similarity of the text to be judged, and provide the triples as supplementary knowledge to the large language model in descending order of similarity, and make the final judgment by extracting the prediction probability of toxicity or non-toxicity output by the large language model; where: The data preprocessing described in step 1 includes: 1.1: Filter the texts marked as non-toxic and keep only the harmful information texts for subsequent knowledge graph generation; 1.2: Align text labels of different granularities from different data sources, and unify the fine-grained classification of racial hatred and regional hatred as harmful; 1.3: Align text labels from different data sources and unify labels of different expressions: offensive, hateful, and harmful as harmful; Step 2 describes the construction of multiple example prompt words for the harmful text interpretation task, specifically including: 2.1: Task definition, requiring the large language model to output the reason why this text is judged to be harmful; 2.2: Output format: The large language model is required to output the reasons in text format without outputting other irrelevant content. 2.3: Multiple example prompts, using the in-text learning ability of the large language model, through multiple manually constructed example prompts, to improve the performance of the large language model in the task; Step 3 describes constructing multiple example prompt words for the triple extraction task, specifically including: 3.1: Task definition: The large language model is required to extract the core triples from the text and explanation, and this triple contains the key information that causes the text to be harmful; and the format of the triple is defined as subject, predicate and object format; 3.2: Output format: The large language model is required to output triples in JSON format and not output other irrelevant content; 3.3: Multiple example prompts, using the in-text learning ability of the large language model, through multiple manually constructed example prompts, to improve the performance of the large language model in the task; Step 4 constructs multiple example prompt words for the triple inspection task, specifically including: 4.1: Task definition: The large language model is required to filter out the triples that summarize harmful information from the given text and triples, and not output irrelevant and harmless triples; 4.2: Output format: The large language model is required to output triples in the same JSON format as step 3, and no other irrelevant content is output; 4.3: Multiple example prompts, using the in-text learning ability of the large language model, through multiple manually constructed example prompts, to improve the performance of the large language model in the task; The step 5 specifically includes: 5.1: Extract the head and tail entity nodes of the triple and obtain their corresponding word vector representations through the BERT pre-trained model; 5.2: Cluster the word vector representations and cluster entity names with high similarity into the same cluster; 5.3: If there are multiple names pointing to the same entity in the same cluster, use the most frequently appearing entity name as the name of the entity pointed to in the cluster; Step 6 describes designing multiple example prompt words for the entity extraction task, including: 6.1: Task definition, requiring the large language model to extract the entities involved from the given text; 6.2: Output format, requiring the large language model to output entities in list format, without outputting other additional content; 6.3: Multiple example prompts, using the in-text learning ability of the large language model, through multiple manually constructed example prompts, to improve the performance of the large language model in the task; Step 7 describes searching on the knowledge graph, specifically including: 7.1: Obtain the semantic representation of the nodes on the knowledge graph and the entities extracted in step 6 through the pre-trained BERT model; 7.2: Store node representations on the knowledge graph into the Faiss vector database to speed up queries; 7.3: Use the semantic representation of the entity extracted in step 6 to perform similarity search in the Faiss vector database to find the most similar entity; The step 9 specifically includes: 9.1: Use the pre-trained SentenceBERT model to obtain the sentence representation of the text to be judged, and at the same time obtain the sentence representation of the candidate triple text obtained in step 8; 9.2: Calculate the similarity between the sentence representation of the text to be judged and the candidate triple sentence representation; 9.3: Sort the candidate triples from high to low according to their similarity with the sentence representation of the text to be judged, and discard the candidate triples whose similarity is lower than the threshold; 9.4: The final triples are used as supplementary knowledge and embedded into the final judgment prompt words. The final prompt words include: 9.4.1: Task definition, requiring the model to refer to the given triples of supplementary knowledge to determine whether the given text is toxic; 9.4.2: Return format, let the model choose between toxic and non-toxic, only need the model to return one option; 9.4.3: Multiple example prompts: By giving multiple manually constructed examples, the model can learn how to use the given triples to supplement knowledge and improve the performance of the model on this task. 9.5: Extract the probability of the model selecting two different options, and take the option with a greater probability as the model's judgment result; at the same time, the model generates and outputs the basis and reasons for the judgment; the final judgment is made by extracting the large language model to output the predicted probability of toxicity or non-toxicity.
Citation Information
Cited By
Harmful cue word visual analysis system based on risk perception
CN120995040A
Toxic text collection method and system based on retrieval enhancement generation
CN121071167A
A Large Model-Based Construction Method for Online Harmful Information Sample Libraries
NL4000863A