Method and system for generating entity recognition data set based on large model

By exporting and sampling entities from the knowledge graph database and using big models to generate and verify text, the problem of unrealistic and reliable entity words and noise in the data set generation in the prior art are solved, and high-quality data set construction is achieved.

CN120218213APending Publication Date: 2025-06-27ANHUI UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510365902.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The prior art cannot ensure that the entity words marked in the text are authentic and reliable when generating entity recognition data sets, and there is a problem of noise in the built data set.

Method used

By exporting entities from the knowledge graph database in the vertical field, sampling entities, and using the big model to generate text containing these entities, matching entities in the annotation text to obtain labels, generating data sets, and verifying the data sets through the big model to filter out irregular data.

Benefits of technology

Ensure that the entities in the data set truly exist in the real vertical field scenarios, ensure that the entity words marked in the text are authentic and reliable, avoid noise in the data set, and improve the overall quality of the data set.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120218213A_ABST
    Figure CN120218213A_ABST
Patent Text Reader

Abstract

The invention discloses a method and system for generating an entity recognition data set based on a large model, and the method comprises the steps: exporting all entities from a knowledge graph database in a vertical field, and generating an entity list; sampling a plurality of entities in the entity list; generating a text containing the sampled entities by using the large model; matching entities in the tagged text to obtain a tag, and generating a data set by using the text and the tag; verifying the data set by using a large model, and filtering out non-standard data in the data set; the method has the advantages that the entity words marked in the text are real and reliable, and no noise exists in the constructed data set.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of entity recognition, and particularly to a method and system for generating an entity recognition data set based on a large model. Background Art

[0002] Named Entity Recognition (NER) aims to identify the required entities and their types from unstructured text, and is a basic text task. Entities in text are important key information and are crucial for applications such as machine translation and knowledge graph question answering, so it has great research significance and application value. A high-performance entity recognition model is indispensable for high-quality data sets. However, in vertical fields, the cost of manually annotating data sets is high and the data quality is difficult to guarantee. Therefore, how to automatically generate high-quality data sets has also become the focus of current research.

[0003] Traditional methods for generating entity recognition data sets use a pre-trained masked model to annotate existing text to obtain entity labels. For example, a named entity recognition method based on a pre-trained language model disclosed in Chinese Patent Publication No. CN113806494A. The essence of this method is knowledge transfer, but the meanings of entity words in vertical fields are special, resulting in inaccurate annotation results. Traditional methods for automatically constructing entity recognition data sets can be divided into two categories: distant supervision and weakly supervised learning. Distant supervision is a method that uses existing knowledge bases (such as Freebase, Wikipedia, etc.) to align with unannotated text to achieve automatic entity annotation. The distant supervision method is easily affected by noisy data and annotation inconsistencies because entity mentions in text do not always exactly correspond to entities in the knowledge base. Weakly supervised learning methods use a small amount of annotated data and a large amount of unannotated data to train an initial masked model. Weakly supervised learning methods have achieved remarkable results in reducing the annotation cost, but still face challenges such as unstable data quality and insufficient model generalization ability.

[0004] We are currently in the era of large language models. Research shows that large language models have demonstrated excellent capabilities in various downstream natural language processing tasks, including question answering, mathematical reasoning, and code generation. In terms of generating data sets for tasks such as text classification and question answering inference, compared with traditional methods, large language models can generate semantically coherent and diverse samples by virtue of their extensive knowledge and context learning ability, thus meeting the needs of small-supervised model training. However, for information extraction tasks, due to the limitations of large language models in dealing with relationships between entities, generating entity recognition data sets remains a huge challenge. Existing technologies focus on directly annotating existing text with large language models to output the entities contained in the text. This method also requires original corpus. In addition, large language models have hallucinated text, resulting in noisy output results.

[0005] In summary, when these methods are applied in vertical fields, they cannot guarantee the authenticity and reliability of the entity words marked in the text, and there is noise in the constructed dataset. Therefore, manual inspection is generally required during the actual application process, although this also brings inevitable annotation costs. Summary of the Invention

[0006] The technical problem to be solved by the present invention is that the existing methods for generating entity recognition datasets cannot guarantee the authenticity and reliability of the entity words marked in the text, and there is noise in the constructed dataset.

[0007] The present invention solves the above technical problems through the following technical means: a method for generating an entity recognition dataset based on a large model, including:

[0008] S1. Export all entities from the knowledge graph database in the vertical field to generate an entity list;

[0009] S2. Sample a number of entities from the entity list;

[0010] S3. Use the large model to generate text containing the sampled entities;

[0011] S4. Match the entities in the annotated text to obtain labels, and generate a dataset using the text and labels;

[0012] S5. Use the large model to verify the dataset and filter out the non-standard data in the dataset.

[0013] The present invention exports entities and samples entities, and uses the large model to generate text containing these entity words, so that the position (label) of the entity words in the text can be determined through string search, and a dataset is generated, thereby ensuring that the entities in the dataset actually exist in the real vertical field scenario, guaranteeing the authenticity and reliability of the entity words marked in the text, and also using the large model to verify the dataset and filter out the non-standard data in the dataset, so that there is no noise in the constructed dataset.

[0014] Further, S1 includes:

[0015] Use the neo4j driver package to read all entity names and relationship names in the neo4j graph database, and export them in the form of a list and store them in a local file.

[0016] Further, S2 includes:

[0017] Define how many entities are to be extracted for each entity type. For each entity type, continuously extract entities from the entity list of that type until the number of extracted entities reaches the sampling setting requirement, and fill the extracted entities into the user prompt words of the large model.

[0018] Furthermore, S3 includes:

[0019] Select the ChatGLM-4 Chinese-English bilingual large language model to output a text containing sampled entities line by line according to the instruction requirements, and use line breaks to split out these texts.

[0020] Furthermore, S4 includes:

[0021] S41. Define the label prefix of the first character of the entity as B, the label prefix of the last character of the entity as E, and the label prefix of the middle characters of the entity as I; the label of other characters that are not entities is O;

[0022] S42. Initialize the labels of each character in the text to O;

[0023] S43. Search for the positions where the entity appears in the text, and write the labels of the characters at these positions as B-entity type;

[0024] S44. Based on the number of characters of the entity, deduce the position of the last character of the entity, and write the label of this character as E-entity type, and write the labels of the middle characters of the entity as I-entity type.

[0025] Furthermore, S43 also includes: Search for the positions where the entity appears in the text. According to the requirements, each entity appears only once in the text. If the entity does not appear or appears multiple times, it means that the large model has hallucinations and the generated text does not meet the specifications, so it is discarded.

[0026] Furthermore, S5 includes:

[0027] S51. Compile the instructions for the large model to let the large model identify the entities of the preset type in the text of the data set;

[0028] S52. If the identified entity appears an entity other than the pre-sampled entity, then discard the generated text.

[0029] The present invention also provides a system for generating an entity recognition data set based on a large model, including:

[0030] An entity list construction module, which is used to export all entities from the knowledge graph database in the vertical domain and generate an entity list;

[0031] An entity sampling module, which is used to sample a number of entities in the entity list;

[0032] A text generation module, which is used to generate a text containing the sampled entities by using the large model;

[0033] A data set generation module, which is used to match the entities in the labeled text to obtain labels, and generate a data set by using the text and the labels;

[0034] A dataset verification module, which is used to verify the dataset using a large model and filter out the non-standard data in the dataset.

[0035] Furthermore, the entity list construction module is also used for:

[0036] Using the neo4j driver package, read all entity names and relationship names in the neo4j graph database, export them in the form of a list, and store them in a local file.

[0037] Furthermore, the entity sampling module is also used for:

[0038] Define how many entities are to be extracted for each entity type. For each entity type, continuously extract entities from the entity list of that type until the number of extracted entities reaches the sampling requirement, and fill the extracted entities into the user prompt words of the large model.

[0039] Furthermore, the text generation module is also used for:

[0040] Select the ChatGLM-4 Chinese-English bilingual large language model to output a text containing the sampled entities line by line according to the instruction requirements, and split these texts using line breaks.

[0041] Furthermore, the dataset generation module is also used for:

[0042] S41. Define the label prefix of the first character of the entity as B, the label prefix of the last character of the entity as E, and the label prefix of the middle characters of the entity as I; the label of other characters that are not entities is O;

[0043] S42. Initialize the labels of each character in the text to O;

[0044] S43. Search for the positions where the entity appears in the text, and write the labels of the characters at these positions as B - entity type;

[0045] S44. Based on the number of characters of the entity, infer the position of the last character of the entity, and write the label of this character as E - entity type, and write the labels of the middle characters of the entity as I - entity type.

[0046] Furthermore, S43 also includes: Search for the positions where the entity appears in the text. According to the requirements, each entity should appear only once in the text. If the entity does not appear or appears multiple times, it means that the large model has hallucinations and the generated text does not meet the specifications, so it is discarded.

[0047] Furthermore, the dataset verification module is also used for:

[0048] S51. Write instructions for the large model to identify entities of a preset type in the text of the data set;

[0049] S52. If the identified entity is an entity other than the pre-sampled entity, then discard the generated text.

[0050] The advantages of the present invention are as follows:

[0051] (1) The present invention exports entities and samples entities, and allows the large model to generate text containing these entity words. This can determine the position (label) of the entity words in the text through string search, generate a data set, thereby ensuring that the entities in the data set actually exist in the real vertical domain scenario, ensuring the authenticity and reliability of the entity words marked in the text, and also using the large model to verify the data set and filter out the non-standard data in the data set, so that the constructed data set has no noise.

[0052] (2) Problems such as entity label misplacement or incorrect entity type annotation may occur during the construction of the data set. The present invention uses the large language model to output entities of predefined types in the text, and judges whether the sample is incorrect by comparing with the actual label, thereby avoiding incorrect generated data sets and improving the overall quality of the data set. Description of the Drawings

[0053] Figure 1 It is a flowchart of the method for generating an entity recognition data set based on a large model disclosed in an embodiment of the present invention;

[0054] Figure 2 It is a flowchart of matching annotations in the method for generating an entity recognition data set based on a large model disclosed in an embodiment of the present invention;

[0055] Figure 3 It is a flowchart of filtering out non-standard text in the method for generating an entity recognition data set based on a large model disclosed in an embodiment of the present invention. Detailed Embodiments

[0056] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments of the present invention fall within the protection scope of the present invention.

[0057] Embodiment 1

[0058] As Figure 1 shown, Embodiment 1 of the present invention provides a method for generating an entity recognition data set based on a large model, including:

[0059] S1. Export all entities from the knowledge graph database in the vertical domain to generate an entity list. The specific process is as follows:

[0060] The task objective of generating the entity recognition dataset is to construct a generator. Input the entity word t i as the entity type, and t i as the entity word of type t. With the help of a large model, generate text containing these entity words, and finally label the actual positions where these entity words E appear in the text according to the BIO format.

[0061] With the help of the neo4j driver package, read all entity names and relationship names in the neo4j graph database, and export them in the form of a list and store them in a local file. neo4j is a tool package specifically for reading and writing neo4j databases. It can connect to the database and perform operations such as adding, deleting, modifying, and querying data in the database. The export task is actually to let the neo4j driver execute a query statement to query all entity names and relationship names.

[0062] The results returned by the above query statement are lists, and write them into a text file. For the entity list, it also needs to be aggregated according to the entity type before writing the file. Each entity type has an entity list, and write them into files respectively. The file name is the entity type.

[0063] The entity types and relationship types in the database may not be self-explanatory, which is not conducive to the large model's understanding. Therefore, a mapping needs to be defined to indicate the meanings of entity types and relationship types.

[0064] For example:

[0065] # Entity type: People.name -> person, Course.code -> course code

[0066] # Relationship type: belong -> belong to, superior_department -> superior unit

[0067] S2. Sample several entities from the entity list. The specific process is as follows:

[0068] Randomly select entities from the entity list. Ultimately, it is to fill the entities into the large model prompt template to form a complete prompt instruction, so that the large model generates text containing these entity words. The specific method of entity sampling and forming the large model prompt is as follows:

[0069] Define sampling settings: how many entities to extract for each entity type. Because large models may cause hallucinations and not output strictly according to the prompt, the number of entities is controlled between 1 and 3 to reduce the difficulty of the task.

[0070] For each entity type, entities are continuously extracted from the entity list of that type until the number of extracted entities reaches the set requirements for sampling.

[0071] The extracted entities fill the large model user_prompt. The entity types are indicated by self-explanatory words. The same type of entity words are separated by ',', and different types of entity words are separated by ','.

[0072] Here is an example:

[0073] #Entity sampling results: {'People.name':['张磊'],'Major.name':['大数据']}

[0074] #user_prompt: Input: Person: Zhang Lei, Subject code: 0854, 0934, Person: Zhao Shu, Major: Big Data. \nOutput:\n

[0075] S3, using the large model to generate text containing the sampled entities;

[0076] Nowadays, large language models have received widespread attention as a breakthrough technology in the AI ​​community, and have demonstrated excellent performance in processing large-scale text information, context understanding, and natural language generation. Using the generalization generation ability and extensive parameter knowledge of large language models to generate data sets is one of the current popular applications of large language models. Existing technologies based on large models are limited to generating data sets for tasks such as text classification and instruction tuning, and it is difficult to generate high-quality entity recognition data sets. The gap can be attributed to the shortcomings of large models in processing the relationship between entities. The data set generated by the present invention is Chinese corpus, so the ChatGLM-4 Chinese-English bilingual large language model is selected to improve the quality of generating Chinese data sets. This step inputs the prompt into the large model, allowing it to output content, and finally filters out other characters such as line breaks, serial numbers, etc., and parses the output content to obtain text. The specific approach is as follows:

[0077] (1) The large model outputs the following content according to the prompt command: one text per line containing entity words. First, these texts are separated according to the line break character to obtain Text_List = [text1, text2, text3, ...], text i is the i-th text.

[0078] (2) Since the text output by the large model will have labels, it is still necessary to filter out the preceding labels for each text.

[0079] S4. Match the entities in the labeled text to obtain labels, and generate a dataset using the text and labels;

[0080] Entity annotation adopts the BIO (Beginning, Inside, Outside) format, which can accurately represent the entity types and boundaries in the sequence text. For example, "B-PER" represents the start of a person name entity, "I-PER" represents the internal part of a person name entity, and "O" represents the non-entity part. The large model generates text containing sampled entity words, and the positions where the entity words appear still need to be determined through string matching. Match the entity words one by one and label the text according to the BIO format: the label prefix of the first character of the entity word is B, the label prefix of the last character is E, and the label prefixes of the middle characters are I; the labels of other characters are O. As Figure 2 shown, the specific method is as follows:

[0081] The original text is a character sequence, that is, text = [c1, c2, c3...], c i is the i-th character.

[0082] S41. Initialize the label of each character in the original text to O, Tag = [t1, t2, t3,...], t i = O.

[0083] S42. Search for the positions where the entity words appear in the original text. According to the requirements, each entity word appears only once in the text. If the entity word does not appear or appears multiple times, it means that the large model has hallucinations and the generated text does not meet the specifications, and it needs to be discarded.

[0084] S43. Write the labels of the characters at these positions as B-entity type.

[0085] S44. Based on the number of characters of the entity word, deduce the position of the last character of the entity word, and write the label of this character as E-entity type, and write the labels of the middle characters of the entity word as I-entity type.

[0086] S5. Due to its own hallucination problem, the large language model may not generate text that only contains the given entity words as required, resulting in the actual generated text lacking entity words or containing extra entity words. The data that does not meet the specifications has noise, which reduces the overall quality of the dataset. This step aims to verify whether there is noise in the dataset through the large model, so as to filter out the non-standard data, mainly by using the large model to verify the dataset and filter out the non-standard data in the dataset. As Figure 3 shown, the specific method is as follows:

[0087] S51. Write a large model prompt to enable the large model to identify entities of a preset type in the text of the data set;

[0088] S52. If entities other than the pre-sampled entities appear in the identified entities, then discard the generated text to further improve the overall quality of the data set.

[0089] Through the above technical solutions, the present invention ensures the generation of a high-quality data set and the real existence of entities therein in the real scenario by generating text containing entity words in the knowledge graph and verifying the standardized data with the help of a large model. There may be problems such as entity label misplacement or incorrect entity type annotation during the construction of the data set. The present invention uses a large language model to output entities of predefined types in the text and determines whether the sample is incorrect by comparing with the actual label.

[0090] Embodiment 2

[0091] Based on Embodiment 1, Embodiment 2 of the present invention further provides a system for generating an entity recognition data set based on a large model, including:

[0092] An entity list construction module for exporting all entities from the knowledge graph database in the vertical domain to generate an entity list;

[0093] An entity sampling module for sampling a number of entities in the entity list;

[0094] A text generation module for using the large model to generate text containing the sampled entities;

[0095] A data set generation module for matching the entities in the labeled text to obtain labels and generating a data set using the text and labels;

[0096] A data set verification module for using the large model to verify the data set and filtering out non-standard data in the data set.

[0097] Specifically, the entity list construction module is further used for:

[0098] Using the neo4j driver package, reading all entity names and relationship names in the neo4j graph database, and exporting them in the form of a list and storing them in a local file.

[0099] Specifically, the entity sampling module is further used for:

[0100] Defining how many entities are to be extracted for each entity type. For each entity type, continuously extract entities from the entity list of that type until the number of extracted entities reaches the sampling setting requirement, and fill the extracted entities into the user prompt of the large model.

[0101] Specifically, the text generation module is further configured to:

[0102] Select the ChatGLM-4 Chinese-English bilingual large language model to output a text containing sampled entities line by line according to the instruction requirements, and split these texts using line breaks.

[0103] Specifically, the dataset generation module is further configured to:

[0104] S41. Define the label prefix of the first character of the entity as B, the label prefix of the last character of the entity as E, and the label prefix of the middle characters of the entity as I; the label of other characters that are not entities is O;

[0105] S42. Initialize the labels of each character in the text to O;

[0106] S43. Search for the positions where the entity appears in the text, and write the labels of the characters at these positions as B-entity type;

[0107] S44. Based on the number of characters of the entity, deduce the position of the last character of the entity, and write the label of this character as E-entity type, and write the labels of the middle characters of the entity as I-entity type.

[0108] More specifically, S43 further includes: Search for the positions where the entity appears in the text. According to the requirements, each entity should appear only once in the text. If the entity does not appear or appears multiple times, it means that the large model has hallucinations and the generated text does not meet the specifications, so it should be discarded.

[0109] Specifically, the dataset verification module is further configured to:

[0110] S51. Compile the instructions for the large model to let the large model identify the entities of the preset type in the text of the dataset;

[0111] S52. If the identified entity appears an entity other than the pre-sampled entity, then discard the generated text.

[0112] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for generating an entity recognition dataset based on a large model, characterized in that: include: S1. Export all entities from the knowledge graph database in the vertical field and generate an entity list; S2, sampling several entities in the entity list; S3, using the large model to generate text containing the sampled entities; S4, matching entities in the annotated text to obtain labels, and using the text and labels to generate a dataset; S5. Use the large model to verify the data set and filter out irregular data in the data set.

2. The method for generating an entity recognition dataset based on a large model according to claim 1, characterized in that: S1 includes: Use the neo4j driver package to read all entity names and relationship names in the neo4j graph database, export them into a list, and store them in a local file.

3. The method for generating an entity recognition dataset based on a large model according to claim 1, characterized in that S2 include: Define how many entities to extract for each entity type. For each entity type, continuously extract entities from the entity list of that type until the number of extracted entities reaches the sampling setting requirements, and fill the extracted entities into the user prompt words of the large model.

4. The method for generating an entity recognition dataset based on a large model according to claim 1, characterized in that S3 include: Select the ChatGLM-4 Chinese-English bilingual language model to output a text containing the sampled entity per line according to the instructions, and use line breaks to separate the text.

5. The method for generating an entity recognition dataset based on a large model according to claim 1, characterized in that: S4 includes: S41, define the label prefix of the first word of the entity as B, the label prefix of the last word as E, the label prefix of the words in between as I; the labels of other words that are not entities are O; S42, initialize the label of each word of the text to 0; S43, searching for the positions where the entity appears in the text, and writing the labels of the words at these positions as B-entity types; S44. According to the number of characters in the entity, the position of the last character of the entity is deduced, and the label of the character is written as E-entity type, and the labels of the middle characters of the entity are all written as I-entity type.

6. The method for generating an entity recognition dataset based on a large model according to claim 5, characterized in that: S43 also includes: searching for the position where the entity appears in the text. According to the requirements, each entity appears only once in the text. If the entity does not appear or appears multiple times, it means that the large model has hallucinations and the generated text does not meet the specifications and is discarded.

7. The method for generating an entity recognition dataset based on a large model according to claim 1, characterized in that S5 include: S51. Write instructions for the big model to enable the big model to identify entities of a preset type in the text of the data set; S52. If the recognized entities include entities other than the pre-sampled entities, the generated text is discarded.

8. A system for generating entity recognition datasets based on a large model, characterized in that: include: The entity list construction module is used to export all entities from the knowledge graph database in the vertical field and generate an entity list; The entity sampling module is used to sample several entities in the entity list; A text generation module, used to generate text containing the sampled entities using the large model; The dataset generation module is used to match entities in the annotated text to obtain labels and generate datasets using text and labels; The dataset verification module is used to verify the dataset using the large model and filter out irregular data in the dataset.

9. The system for generating entity recognition dataset based on a large model according to claim 8, characterized in that: The Entity List building block is also used to: Use the neo4j driver package to read all entity names and relationship names in the neo4j graph database, export them into a list, and store them in a local file.

10. The system for generating entity recognition dataset based on a large model according to claim 8, characterized in that: The entity sampling module is also used to: Define how many entities to extract for each entity type. For each entity type, continuously extract entities from the entity list of that type until the number of extracted entities reaches the sampling setting requirements, and fill the extracted entities into the user prompt words of the large model.

Citation Information

Patent Citations

  • Named entity recognition method based on pre-training language model

    CN113806494A