Entity labeling method and device based on large language model

Through the automation method based on the large language model, efficient entity recognition and labeling of text data is achieved, and the problems of cumbersome and low efficiency in the prior art are solved, and processing efficiency and accuracy are improved.

CN120257988APending Publication Date: 2025-07-04BEIJING GONGJUN TECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510196393.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The current entity recognition and labeling process mainly relies on manual operations, resulting in cumbersome and inefficient.

Method used

An automated method based on a large language model is adopted to process entity information by obtaining text data, preprocessing, segmenting, and inputting preset language models, and determining the labeling method according to the target field to achieve automatic entity annotation.

Benefits of technology

It reduces manual intervention, improves processing efficiency, and can quickly identify and label entity information in large amounts of text data, solving the problems of cumbersome and inefficient manual operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120257988A_ABST
    Figure CN120257988A_ABST
Patent Text Reader

Abstract

The invention discloses an entity labeling method and device based on a large language model, and relates to the technical field of entity labeling. The method comprises the following steps: acquiring to-be-processed text data; processing the to-be-processed text data to obtain text information; segmenting the text information according to a preset rule to obtain multiple pieces of sub-text information; obtaining target sub-text information from the multiple pieces of sub-text information, and inputting the target sub-text information into a preset language model for processing to obtain entity information; obtaining a target field corresponding to the target sub-text information, and determining a first labeling mode according to the target field; and marking entity information in the target sub-text information based on the first marking mode to complete entity marking of the to-be-processed text data. By implementing the technical scheme provided by the invention, the problems of tedious manual operation and low efficiency in the current entity identification and labeling process are effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of entity annotation, and particularly to an entity annotation method and device based on a large language model. Background Art

[0002] In the vast field of natural language processing (NLP), entity recognition and entity annotation form the cornerstone of information extraction and understanding. These two tasks play an irreplaceable role in mining key information from text data.

[0003] Entity recognition, as one of the core links in NLP, its core function is to accurately extract key information units such as person names, place names, and organization names from text. These information units are not only the core components of text content but also the basis for understanding and analyzing text. Through entity recognition, we can transform originally unstructured text data into structured data models, such as knowledge graphs or databases, thus greatly facilitating subsequent information retrieval, analysis, and application. On the basis of entity recognition, entity annotation further marks the recognized entities and assigns specific labels to them, such as "person name", "place name", etc. This process not only enhances the readability of text data but also enables unstructured text data to be transformed into structured information, providing great convenience for subsequent analysis and processing. However, the current process of entity recognition and annotation still mainly relies on manual operations, that is, usually including manually selecting entity content from long texts and then having proofreaders review to ensure the accuracy of the selected content. But the entire screening process is cumbersome, resulting in the problem of low efficiency of manual screening.

[0004] Therefore, there is an urgent need for an entity annotation method and device based on a large language model that can solve the above technical problems. Summary of the Invention

[0005] This application provides an entity annotation method and device based on a large language model, which effectively solves the problems of cumbersome manual operations and low efficiency in the current entity recognition and annotation process.

[0006] First aspect, this application provides an entity annotation method based on a large language model. The method includes: obtaining text data to be processed; processing the text data to be processed to obtain text information; segmenting the text information according to a preset rule to obtain multiple sub-text information; obtaining target sub-text information from the multiple sub-text information, inputting the target sub-text information into a preset language model for processing to obtain entity information; obtaining the target domain corresponding to the target sub-text information, and determining a first annotation method according to the target domain; annotating the entity information in the target sub-text information based on the first annotation method to complete the entity annotation of the text data to be processed.

[0007] By adopting the above technical solution, the text data to be processed is obtained through an automated means, and the text data to be processed is preliminarily processed to obtain text information, thereby reducing the need for manual intervention. Then, the text information is segmented according to a preset rule, and the segmentation helps to reduce the workload of a single processing unit to improve the processing efficiency; the target sub-text information is screened out from the multiple sub-text information, avoiding the cumbersome manual selection; the target sub-text information is input into a preset language model for processing, and the entity information can be automatically identified and extracted; according to the target domain corresponding to the target sub-text information, the first annotation method is determined, and the entity information is annotated based on the first annotation method, and the annotation task can be automatically completed without manual intervention, effectively solving the problems of cumbersome manual operation and low efficiency in the current entity recognition and annotation process.

[0008] Optionally, before inputting the target sub-text information into a preset language model for processing to obtain entity information, it is necessary to construct a preset language model, which specifically includes: obtaining a sample data set; using an initial language model to train the sample data set until the loss value between the output result and the actual result meets the convergence condition, ending the training, and taking the initial language model at the end of the training as the preset language model, where each sample data set in the sample data set includes a marked preset entity, and the preset entity includes a person name, a place name, and an organization name.

[0009] By adopting the above technical solution, the sample data set is the basis for training the model. By continuously providing the sample data set to the model and allowing the initial language model to learn how to extract useful information from these data, when the loss value meets the convergence condition, the training process ends, and the initial language model at the end of the training is determined as the preset language model. Once the preset language model is trained, it can quickly process a large amount of text data and automatically identify and annotate the preset entities therein.

[0010] Optionally, after obtaining the target sub-text information from multiple sub-text information, the method further includes: extracting multiple entities to be monitored from the target sub-text information; determining whether all the multiple entities to be monitored exist in the entity database; when all the multiple entities to be monitored exist in the entity database, confirming to input the target sub-text information into a preset language model for processing.

[0011] By adopting the above technical solution, before inputting the target sub-text information into the preset language model, the entities to be monitored therein are first extracted, and it is checked whether these entities already exist in the entity database. If all the entities to be monitored exist in the database, then the target sub-text information can be directly input into the preset language model for processing. Pre-screening the entities to be monitored can ensure that only the target sub-text information containing known entities will be sent into the preset language model for processing.

[0012] Optionally, after determining whether all the multiple entities to be monitored exist in the entity database, the method further includes: if there is an entity to be monitored among the multiple entities to be monitored that does not exist in the entity database, outputting the target entity to be monitored, where the target entity to be monitored is the entity to be monitored that does not exist in the entity database; inputting the target entity to be monitored into a preset entity model for processing to obtain a target value; determining whether the target value is greater than or equal to a preset value; when the target value is greater than or equal to the preset value, confirming to store the target entity to be monitored in the sample data set.

[0013] By adopting the above technical solution, when it is found that there is an entity among the entities to be monitored that does not exist in the entity database, the entity information that does not exist in the entity database is extracted as the target entity to be monitored. Inputting the target entity to be monitored into the preset entity model for processing can obtain a target value, and the target value reflects the recognition degree or confidence of the entity model for the target entity to be monitored. Adding the target entity to be monitored that meets the conditions to the sample data set can continuously enrich and optimize the data set.

[0014] Optionally, after determining whether the target value is greater than or equal to the preset value, the method further includes: when the target value is less than the preset value, confirming to classify the target entity to be monitored into the data set to be reviewed; generating a prompt message from the data set to be reviewed and sending the prompt message to the reviewer.

[0015] By adopting the above technical solution, when the target value obtained after processing the target entity to be monitored by the preset entity model is lower than the preset value, it means that there may be uncertainty in the entity information. Classifying the target entity to be monitored into the data set to be reviewed and having it manually reviewed by the reviewer can further improve the accuracy of the data and ensure the reliability of subsequent processing or analysis results.

[0016] Optionally, generate a prompt message for the dataset to be audited and send the prompt message to the auditors, specifically including: obtaining the audit quantity corresponding to the dataset to be audited; determining whether the audit quantity is less than or equal to a preset quantity; when the audit quantity is less than or equal to the preset quantity, confirming that the number of auditors corresponding to the dataset to be audited is the first quantity, and sending the prompt message to the auditors of the first quantity; when the audit quantity is greater than the preset quantity, confirming that the number of auditors corresponding to the dataset to be audited is the second quantity, and sending the prompt message to the auditors of the second quantity.

[0017] By adopting the above technical solution, the number of auditors to be audited is flexibly adjusted according to the audit quantity of the dataset to be audited. When the audit quantity is small, the number of auditors to be audited can be reduced to avoid waste of human resources; when the audit quantity is large, the number of auditors to be audited is increased to ensure that the audit task can be completed in time. This dynamic adjustment helps to optimize the allocation of human resources and improve the overall work efficiency.

[0018] Optionally, generate a prompt message for the dataset to be audited and send the prompt message to the auditors, specifically including: obtaining the audit quantity corresponding to the dataset to be audited; determining whether the audit quantity is less than or equal to a preset quantity; when the audit quantity is less than or equal to the preset quantity, confirming that the number of auditors corresponding to the dataset to be audited is the first quantity, and sending the prompt message to the auditors of the first quantity; when the audit quantity is greater than the preset quantity, confirming that the number of auditors corresponding to the dataset to be audited is the second quantity, and sending the prompt message to the auditors of the second quantity.

[0019] By adopting the above technical solution, the requirement information corresponding to the text data to be processed can be obtained, the goals and key points of the annotation can be clarified, and according to the second annotation method determined by the requirement information, the second annotation method is selected and the entity information in the target sub-text information is accurately annotated, which can ensure the accuracy of the annotation result.

[0020] In the second aspect of the present application, an entity annotation device based on a large language model is provided. The device includes an acquisition unit, a processing unit, and an annotation unit; the acquisition unit acquires the text data to be processed; the processing unit processes the text data to be processed to obtain text information; the text information is segmented according to a preset rule to obtain a plurality of sub-text information; the target sub-text information is obtained from the plurality of sub-text information, and the target sub-text information is input into a preset language model for processing to obtain entity information; the target field corresponding to the target sub-text information is obtained, and the first annotation method is determined according to the target field; the annotation unit annotates the entity information in the target sub-text information based on the first annotation method to complete the entity annotation of the text data to be processed.

[0021] Optionally, the obtaining unit is used to obtain a sample data set; the processing unit is used to train the sample data set by using an initial language model until the loss value between the output result and the actual result meets the convergence condition, end the training, and use the initial language model at the end of the training as the preset language model, where each sample data set in the sample data set includes a pre-labeled preset entity, and the preset entity includes a person name, a place name, and an organization name.

[0022] Optionally, the obtaining unit is used to extract multiple entities to be monitored from the target sub-text information; the processing unit is used to determine whether all the multiple entities to be monitored exist in the entity database; when all the multiple entities to be monitored exist in the entity database, confirm to input the target sub-text information into the preset language model for processing.

[0023] Optionally, the processing unit is used to, if there is an entity to be monitored among the multiple entities to be monitored that does not exist in the entity database, output the target entity to be monitored, where the target entity to be monitored is the entity to be monitored that does not exist in the entity database; input the target entity to be monitored into the preset entity model for processing to obtain a target value; determine whether the target value is greater than or equal to a preset value; when the target value is greater than or equal to the preset value, confirm to store the target entity to be monitored in the sample data set.

[0024] Optionally, the processing unit is used to, when the target value is less than the preset value, confirm to classify the target entity to be monitored into the data set to be reviewed; generate a prompt message from the data set to be reviewed and send the prompt message to the reviewer.

[0025] Optionally, the obtaining unit is used to obtain the review quantity corresponding to the data set to be reviewed; the processing unit is used to determine whether the review quantity is less than or equal to a preset quantity; when the review quantity is less than or equal to the preset quantity, confirm that the number of reviewers corresponding to the data set to be reviewed is the first quantity, and send a prompt message to the reviewers of the first quantity; when the review quantity is greater than the preset quantity, confirm that the number of reviewers corresponding to the data set to be reviewed is the second quantity, and send a prompt message to the reviewers of the second quantity.

[0026] Optionally, the obtaining unit is used to obtain the requirement information corresponding to the text data to be processed; the processing unit is used to determine a second annotation method according to the requirement information; the annotation unit is used to annotate the entity information in the target sub-text information based on the second annotation method.

[0027] In a third aspect of the present application, an electronic device is provided. The electronic device includes a processor, a memory, a user interface, and a network interface. The memory is used to store instructions. The user interface and the network interface are used to communicate with other devices. The processor is used to execute the instructions stored in the memory, so that the electronic device executes the method of any one of the above in the present application.

[0028] In a fourth aspect of the present application, a computer-readable storage medium is provided. The computer-readable storage medium stores instructions, and when the instructions are executed, the method of any one of the above in the present application is executed.

[0029] In summary, one or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages: 1. By means of automation, the text data to be processed is obtained, and the text data to be processed is preliminarily processed to obtain text information, thus reducing the need for manual intervention. Then, the text information is segmented by preset rules. The segmentation helps to reduce the workload of a single processing unit to improve the processing efficiency; the target sub-text information is screened out from multiple sub-text information, avoiding the cumbersome manual selection; the target sub-text information is input into a preset language model for processing, and entity information can be automatically identified and extracted; according to the target field corresponding to the target sub-text information, the first annotation method is determined, and the entity information is annotated based on the first annotation method, which can automatically complete the annotation task without manual intervention, effectively solving the problems of cumbersome manual operation and low efficiency in the current entity recognition and annotation process.

[0030] 2. The sample data set is the basis for training the model. By continuously providing the model with the sample data set and allowing the initial language model to learn how to extract useful information from these data, when the loss value meets the convergence condition, the training process ends. The initial language model at the end of the training is determined as the preset language model. Once the preset language model is trained, it can quickly process a large amount of text data and automatically identify and annotate the preset entities therein. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 is a schematic flowchart of a method for entity annotation based on a large language model provided by an embodiment of the present application; Figure 2 is a schematic structural diagram of an apparatus for entity annotation based on a large language model provided by an embodiment of the present application; Figure 3 is a schematic structural diagram of an electronic device disclosed by an embodiment of the present application.

[0032] Explanation of reference numerals: 201, acquisition unit; 202, processing unit; 203, labeling unit; 300, electronic device; 301, processor; 302, memory; 303, user interface; 304, network interface; 305, communication bus. DETAILED DESCRIPTION

[0033] In order to enable technicians in this field to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the drawings in the embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of this application, not all of the embodiments.

[0034] In the description of the embodiments of the present application, words such as "for example" or "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design described as "for example" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "for example" or "for example" is intended to present related concepts in a specific way.

[0035] In the description of the embodiments of the present application, the meaning of the term "multiple" refers to two or more. For example, multiple systems refer to two or more systems, and multiple screen terminals refer to two or more screen terminals. In addition, the terms "first" and "second" are used for descriptive purposes only and cannot be understood as indicating or implying relative importance or implicitly indicating the indicated technical features. Thus, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. The terms "include", "comprise", "have" and their variations all mean "including but not limited to", unless otherwise specifically emphasized.

[0036] In the vast field of natural language processing (NLP), entity recognition and entity annotation constitute the cornerstones of information extraction and understanding. These two tasks play an irreplaceable role in mining key information from text data.

[0037] Entity recognition, as one of the core aspects of NLP, its core function lies in accurately extracting key information units such as person names, place names, and organization names from text. These information units are not only the core components of the text content but also the basis for understanding and analyzing the text. Through entity recognition, we can transform the originally unstructured text data into a structured data model, such as a knowledge graph or a database, thus greatly facilitating subsequent information retrieval, analysis, and application. On the basis of entity recognition, entity annotation further marks the recognized entities and assigns specific labels to them, such as "person name", "place name", etc. This process not only enhances the readability of the text data but also enables the transformation of unstructured text data into structured information, providing great convenience for subsequent analysis and processing. However, the current process of entity recognition and annotation still mainly relies on manual operations, that is, usually including manually selecting entity content from long texts and then having proofreaders conduct a second check to ensure the accuracy of the selected content. But the entire screening process is cumbersome, resulting in the problem of low efficiency of manual screening.

[0038] Therefore, how to solve the problems of cumbersome manual operations and low efficiency in the current entity recognition and annotation process. An entity annotation method based on a large language model provided by an embodiment of this application is applied to a server. The server of this application can be a platform that provides entity annotation services. Figure 1 It is a schematic flowchart of an entity annotation method based on a large language model provided by an embodiment of this application. Refer to Figure 1 and this method includes the following steps S101 - step S106.

[0039] S101: Obtain the text data to be processed.

[0040] In the above S101, the server screens the data from multiple data sources and then obtains the text data to be processed. The data sources include texts, database records, web page contents, and user inputs, etc. According to the data sources, corresponding tools or programming languages (such as the pandas library in Python, SQL query statements, web crawlers, etc.) are used to scrape or read the data. The obtained data is integrated to obtain the text data to be processed.

[0041] S102: Process the text data to be processed to obtain text information.

[0042] In the above S102, after obtaining the text data to be processed, the text data to be processed is processed, and the processing includes cleaning and removing stop words. Cleaning data refers to removing irrelevant characters (such as HTML tags, special symbols, extra spaces, etc.) in the text data to be processed, and unifying the text format (such as case conversion, encoding conversion). Removing stop words is to remove words that appear frequently in the text but are meaningless to the analysis (such as "的", "是", "in", "the", etc.). If the text data to be processed is in English, the stems can be extracted to restore the words to their basic form, such as restoring "running" to "run", which helps to reduce vocabulary diversity. The processed information is then integrated to obtain text information. At this time, the text information refers to the text information that has been cleaned and stop words have been removed.

[0043] S103: Segment the text information according to a preset rule to obtain a plurality of sub-text information.

[0044] In the above S103, after obtaining the text information, it is necessary to set the segmentation rules according to the characteristics and needs of the text, such as segmentation by sentence, paragraph, specific mark (such as line break, specific keyword). The text information is divided into multiple sub-text information according to the segmentation rules. After segmentation, it is necessary to ensure that each sub-text information after segmentation is semantically complete and can express an independent meaning. The segmentation rules should be consistent to avoid different segmentation standards in the same text information. In addition to segmenting the text information according to the above-mentioned set rules, it can also be based on statistical learning methods or segmentation methods in specific fields. The statistical learning-based methods include hidden Markov models (HMM), conditional random fields (CRF) and maximum entropy models. In specific fields, such as biomedicine, finance, etc., named entity recognition (NER) can identify and segment entities with specific meanings, such as names of people, places, names of organizations, names of diseases, names of drugs, etc. It can also be based on syntactic analysis to identify and segment sentence components in the text, such as subjects, predicates, objects, etc. through syntactic analysis trees. The specific method for segmenting the text information can be selected based on the actual application scenario, and no further limitation is given here. For example, an original text information is segmented to obtain multiple sub-text information.

[0045] S104: Obtain target sub-text information from the plurality of sub-text information, input the target sub-text information into a preset language model for processing, and obtain entity information.

[0046] In the above S104, after obtaining multiple sub-text information, one sub-text information can be obtained from the multiple sub-text information, and this sub-text information represents the target sub-text information. Before inputting the target sub-text information into a preset language model for processing to obtain entity information, it is necessary to construct the preset language model, which specifically includes: obtaining a sample data set; using an initial language model to train the sample data set until the loss value between the output result and the actual result meets the convergence condition, ending the training, and taking the initial language model at the end of the training as the preset language model. Among them, each sample data set in the sample data set includes labeled preset entities, and the preset entities include personal names, place names, and organization names. Specifically, sample data can be collected first. When collecting sample data, it is necessary to ensure that the collected sample data all contains text data with labeled preset entities (personal names, place names, organization names). Text data can be collected from various sources (such as web pages, news articles, social media, public databases, etc.). Ensure the diversity and representativeness of the data to cover texts in different fields and styles. Manually annotate the collected text data to mark the personal names, place names, and organization names in it. Professional annotation tools or platforms (such as doccano, Prodigy, etc.) can also be used to improve the annotation efficiency and accuracy. Clean the annotated data to remove irrelevant characters, duplicate content, etc. Convert the data into a format suitable for model training, such as word segmentation, removing stop words, etc. Divide the processed data into a training set, a validation set, and a test set according to a certain proportion. The training set is used to train the model, the validation set is used to adjust the model parameters and select the best model, and the test set is used to evaluate the performance of the model. Then summarize the training set, the validation set, and the test set to obtain the sample data set. According to the task requirements and resource conditions, select a suitable initial language model. A pre-trained model based on the Transformer architecture (such as BERT, GPT series, ERNIE, etc.) can be selected, and these models perform well in natural language processing tasks. Set the parameters of the model, such as the number of layers, the number of hidden units, the learning rate, etc. According to the hardware resources and time cost, select a suitable batch size and the number of training epochs. Select a suitable loss function to measure the difference between the model output result and the actual result. For the named entity recognition task, common loss functions include cross-entropy loss, etc. Input the training set into the model and adjust the model parameters through the backpropagation algorithm to minimize the loss function. During the training process, use the validation set to monitor the performance of the model to avoid overfitting. Use the test set to evaluate the performance of the model and calculate metrics such as accuracy, recall rate, F1 score, etc. According to the evaluation results, adjust the model parameters, training strategy, or select other models for training. During the training process, monitor the change trend of the loss value. When the loss value fluctuates within a certain range and no longer decreases significantly, it is considered that the model has converged. According to the actual situation, specific convergence conditions can be set (such as the loss value is less than a certain threshold, the change in the loss value in consecutive multiple epochs is less than a certain range, etc.).When the model meets the convergence condition, save the model parameters at the end of training. Train the model in the above manner to obtain a preset language model. At this time, the preset language model has been trained with a large amount of text data and can recognize and understand the semantic information in the text. Then, input the target sub-text into the language model, and use the built-in entity recognition function of the model (such as named entity recognition NER) to extract entity information in the text, such as person names, place names, organization names, etc.

[0047] In addition, in order to continuously update the sample data set in the preset language model so that the preset language model can recognize more entity information, before inputting the target sub-text information into the preset language model for processing, it is necessary to judge the target sub-text information to determine whether there is known entity information in the target sub-text information. If there is, it is confirmed that the target sub-text information is input into the preset language model for recognition operations, which specifically includes: extracting multiple entities to be monitored from the target sub-text information; judging whether all the multiple entities to be monitored exist in the entity database; when all the multiple entities to be monitored exist in the entity database, it is confirmed that the target sub-text information is input into the preset language model for processing. Specifically, the target sub-text information is loaded into the memory as a string. The string is cleaned to remove irrelevant characters, such as HTML tags, comments, CSS styles, etc., and the valid text content is retained. The text is tokenized, splitting the text into tokens (such as words, phrases, or characters). Since sub-text information refers to splitting the text information into paragraphs or sentences to obtain multiple sub-text information, the target sub-text information represents a paragraph or a complete sentence. Then, the paragraph or sentence is tokenized, and the sentence is split into multiple phrases according to the tokenization rules, and one phrase is an entity to be monitored. After extracting multiple entities to be monitored from the target sub-text information, the multiple entities to be monitored are compared with the entity database in turn. The entity database is a database that stores known entity information, and these entity information can be names of people, places, organizations, etc. The entity database can be built based on a relational database (such as MySQL, PostgreSQL, etc.) or a NoSQL database (such as MongoDB, etc.). The entity information in the database is obtained through manual input or other means. For each entity to be monitored, check whether there is a matching entity in the entity database. The matching process can be implemented based on methods such as string matching, fuzzy matching, or semantic matching. If the entity to be monitored exactly matches or meets a certain matching threshold with an entity in the database, it is considered that the entity exists. If all the entities to be monitored are found to have matching items in the entity database, the judgment result is "all exist". At this time, it has been confirmed that there is entity information in the target sub-text information, so the target sub-text information needs to be input into the preset language model for processing to obtain entity information. According to the above method of processing the target sub-text information, the other sub-text information in the multiple sub-text information is processed in turn, and then the sub-text information that meets the requirements is input into the preset language model for processing.

[0048] Further, if there is monitored entity information among multiple pieces of monitored entity information that is not in the entity database, output the target monitored entity information, where the target monitored entity information is the monitored entity information that does not exist in the entity database; input the target monitored entity information into a preset entity model for processing to obtain a target value; determine whether the target value is greater than or equal to a preset value; when the target value is greater than or equal to the preset value, confirm to store the target monitored entity information in the sample dataset. Specifically, in the previous steps, multiple pieces of monitored entity information have been matched with the entity database. At this time, it is necessary to traverse the list of monitored entity information to check whether each entity information has a matching item in the entity database. If it is found that a certain piece of monitored entity information does not have a matching item in the entity database, it is marked as the target monitored entity information. The marked target monitored entity information can be output in the form of text, list, or other forms suitable for subsequent processing. According to the task requirements, select a pre-trained preset entity model, which can be constructed based on machine learning or deep learning methods and is used to evaluate the probability that the monitored entity information belongs to entity information. Before inputting the target monitored entity information into the preset entity model, some preprocessing operations may be required, such as text cleaning, word segmentation, vectorization, etc. Input the preprocessed target monitored entity information into the preset entity model. The model processes the input data and outputs a target value. This value can be in the form of probability or score, etc., and is used to represent the correlation degree between the target monitored entity information and the entity information rules. Based on the correlation degree of the entity information rules, set a preset value, which is used to compare with the target value to determine whether the target monitored entity information meets the entity information standard. Compare the target value with the preset value. If the target value is greater than or equal to the preset value, it is considered that the target monitored entity information meets the entity information standard. When the target value is greater than or equal to the preset value, add the target monitored entity information to the sample dataset. At this time, the target monitored entity information can be understood as newly emerging entity information, entity information that has not appeared before. The updated sample dataset can be used for subsequent training, testing, or evaluation tasks. Adding the target monitored entity information to the sample dataset facilitates subsequent training of the preset language model using the sample dataset, helps improve the training effect and performance of the preset language model, and provides high-quality sample data for subsequent tasks. For example, if the target monitored entity information is A, input A into the preset entity model for processing to obtain a target value, the target value is 0.8, and the preset value is 0.7. At this time, the target value is greater than the preset value, and it is confirmed that the monitored entity information A is included in the sample dataset.

[0049] Furthermore, when the target value is less than the preset value, confirm that the target entity information to be monitored is summarized into the dataset to be reviewed; generate a prompt message based on the dataset to be reviewed and send the prompt message to the reviewer. Specifically, in the previous steps, the target value of the target entity information to be monitored has been calculated and compared with the preset value. If the target value is less than the preset value, it means that the target entity information to be monitored may not meet the entity information standard and further review is required. When the target value is less than the preset value, the target entity information to be monitored is separated from the current processing flow and summarized into a dedicated dataset to be reviewed. The dataset to be reviewed is a temporary storage area for storing entity information that needs further review or processing. During the process of summarizing into the dataset to be reviewed, some additional marking or management operations can be performed on the target entity information to be monitored so that the subsequent reviewers can more easily identify and process this information. For example, a unique identifier can be assigned to each entity information to be reviewed, some descriptive information can be added, or it can be classified into specific categories. Generate a corresponding prompt message according to the entity information in the dataset to be reviewed. The prompt message should contain sufficient information so that the reviewer can understand the basic situation of the entity information to be reviewed and the key points to be reviewed. The format of the prompt message can be in the form of text, table, chart, etc., and is selected according to actual needs. The reviewer refers to the person responsible for analyzing the entity information in the text. Send the generated prompt message to the reviewer through an appropriate channel. For example, if the target entity information to be monitored is b, b is input into the preset entity model for processing to obtain the target value. If the target value is 0.5 and the preset value is set to 0.7, at this time the target value is less than the preset value, confirm that the target entity information b is summarized into the dataset to be reviewed for subsequent sending to the reviewer for review.

[0050] At this time, prompt information is generated for the dataset to be reviewed and sent to the reviewers, specifically including: obtaining the review quantity corresponding to the dataset to be reviewed; determining whether the review quantity is less than or equal to a preset quantity; when the review quantity is less than or equal to the preset quantity, confirming that the number of reviewers corresponding to the dataset to be reviewed is the first quantity, and sending prompt information to the reviewers of the first quantity; when the review quantity is greater than the preset quantity, confirming that the number of reviewers corresponding to the dataset to be reviewed is the second quantity, and sending prompt information to the reviewers of the second quantity. Specifically, it is necessary to count the quantity of entity information to be reviewed in the dataset to be reviewed, that is, the review quantity. This can be achieved by traversing the dataset to be reviewed and counting the entity information therein. During the process of calculating the review quantity, data verification may be required to ensure the accuracy of the calculation. For example, it can be checked whether there are duplicate or invalid entity information in the dataset to be reviewed and excluded from the count. A preset quantity is set according to task requirements, the processing capacity of a single reviewer, or other relevant factors. The preset quantity is used to determine whether the size of the dataset to be reviewed exceeds the processing capacity of a single reviewer. The calculated review quantity is compared with the preset quantity. If the review quantity is less than or equal to the preset quantity, it is considered that the size of the dataset to be reviewed is within the processing capacity of a single reviewer; if the review quantity is greater than the preset quantity, it is considered that the size of the dataset to be reviewed exceeds the processing capacity of a single reviewer. When the review quantity is less than or equal to the preset quantity, confirm that the number of reviewers corresponding to the dataset to be reviewed is the first quantity. The first quantity can be set according to the actual situation. For example, it can be 1, specifically depending on the complexity and urgency of the review task. When the review quantity is greater than the preset quantity, confirm that the number of reviewers corresponding to the dataset to be reviewed is the second quantity. The second quantity is usually greater than the first quantity to ensure that there are enough reviewers to process the larger dataset to be reviewed. According to the determined number of reviewers, the corresponding reviewers are selected from the reviewer list. The selection process can be based on factors such as the professional skills, experience, and workload of the reviewers. Send prompt information to the selected reviewers. The prompt information should include key contents such as the basic information of the dataset to be reviewed, review requirements, and deadline. The channel for sending prompt information can be email, text message, instant messaging tool, etc., specifically depending on the preferences and actual situation of the reviewers. During the review process, the review progress and results can be tracked. If the reviewers review the entity information to be monitored in the dataset to be reviewed and store the entity information to be monitored that is confirmed as entity information in the sample dataset, and delete the entity information to be monitored that is confirmed not to be entity information.

[0051] S105: Obtain the target domain corresponding to the target subtext information and determine the first annotation method according to the target domain.

[0052] In the above S105, based on the content of the target sub-text, use a domain classification algorithm or manual judgment to determine its domain (such as medical, legal, financial, etc.). According to the domain characteristics and annotation requirements, select or design an appropriate annotation method. Different domains may have different requirements for entity annotation. For example, the medical domain may require detailed annotation of disease types, symptoms, etc., and different disease types can be annotated with different colors.

[0053] S106: Annotate the entity information in the target sub-text information based on the first annotation method to complete the entity annotation of the text data to be processed.

[0054] In the above S106, according to the selected first annotation method, annotate the entity information in the target sub-text. This can be simple tag addition (such as person name - PER, place name - LOC), or more complex structured information annotation (such as disease name - type - symptom). After annotation, in order to facilitate quick positioning of the annotated content, the same disease can be annotated with the same color for subsequent quick query of information related to the disease. Automatically verify the annotation results to ensure the accuracy and consistency of the annotation. Save the annotated text data in the required format, such as XML, JSON files with annotation information, or directly insert annotation markers in the original text.

[0055] In addition to annotating the entity information in the target sub-text information through the first annotation method, other annotation methods can also be used to annotate the entity information, specifically including: obtaining the requirement information corresponding to the text data to be processed; determining the second annotation method according to the requirement information; and annotating the entity information in the target sub-text information based on the second annotation method. Specifically, conduct a requirement analysis on the collected text data to clarify the goals, scope, and requirements of the annotation. This may include determining the types of entities to be annotated (such as person names, place names, organization names, etc.), the accuracy and format requirements of the annotation, etc. Organize the analyzed requirement information to form a clear and definite requirement specification document. According to the requirement information, select a suitable annotation method as the second annotation method. The annotation method may include semi-automatic annotation (such as using annotation tools for assistance) or automatic annotation (such as prediction based on machine learning models). Configure the selected annotation method to meet specific requirements. This includes setting the parameters of the annotation tool, training the machine learning model, etc. Develop detailed annotation specifications to ensure the consistency and accuracy of the annotation. Extract the target sub-text information to be annotated from the text data to be processed. Identify the entity information to be annotated in the target sub-text information. According to the second annotation method and the annotation specifications, annotate the identified entity information. The annotation process may involve steps such as inserting specific marker symbols in the text, highlighting in color, creating annotation labels, etc. The effective annotation of the entity information in the text data is achieved. This helps to improve the processing efficiency and accuracy of the text data and provides strong support for subsequent analysis and mining.

[0056] Through the above method, automatically obtain the text data to be processed from various data sources, then process and segment the text data to obtain multiple sub-text information, and then sequentially input each sub-text information into a preset language model for processing to obtain entity information, and then annotate the entity information in each sub-text information based on the selected annotation method, thereby realizing the entity annotation of the text data. When providing high-quality sample data for the training of the preset language model, the entity information to be monitored can be compared with the entity database to screen the entity information to be monitored, and the newly emerging entity information is screened and stored in the sample data set so that the preset language model can be retrained with the updated sample data set, improving the training effect and performance of the preset language model.

[0057] The embodiment of the present application also provides an entity annotation device based on a large language model. Figure 2 is a schematic structural diagram of an entity annotation device based on a large language model provided by the embodiment of the present application. Refer to Figure 2 , the device includes an acquisition unit 201, a processing unit 202, and an annotation unit 203.

[0058] The acquisition unit 201 acquires the text data to be processed.

[0059] The processing unit 202 processes the text data to be processed to obtain text information; divides the text information according to a preset rule to obtain a plurality of sub-text information; obtains target sub-text information from the plurality of sub-text information, inputs the target sub-text information into a preset language model for processing to obtain entity information; obtains the target field corresponding to the target sub-text information, and determines the first annotation method according to the target field.

[0060] The annotation unit 203 annotates the entity information in the target sub-text information based on the first annotation method to complete the entity annotation of the text data to be processed.

[0061] In a possible implementation manner, the acquisition unit 201 is used to acquire a sample data set; the processing unit 202 is used to train the sample data set by using an initial language model until the loss value between the output result and the actual result satisfies the convergence condition, ends the training, and uses the initial language model at the end of the training as the preset language model, where each sample data set in the sample data set includes a preset entity with a label, and the preset entity includes a person name, a place name, and an organization name.

[0062] In a possible implementation manner, the acquisition unit 201 is used to extract a plurality of entities to be monitored from the target sub-text information; the processing unit 202 is used to determine whether all the plurality of entities to be monitored exist in the entity database; when all the plurality of entities to be monitored exist in the entity database, it is confirmed that the target sub-text information is input into the preset language model for processing.

[0063] In a possible implementation manner, the processing unit 202 is used to, if there is an entity to be monitored among the plurality of entities to be monitored that does not exist in the entity database, output the target entity to be monitored, where the target entity to be monitored is the entity to be monitored that does not exist in the entity database; input the target entity to be monitored into a preset entity model for processing to obtain a target value; determine whether the target value is greater than or equal to a preset value; when the target value is greater than or equal to the preset value, it is confirmed that the target entity to be monitored is stored in the sample data set.

[0064] In a possible implementation manner, the processing unit 202 is used to, when the target value is less than the preset value, confirm that the target entity to be monitored is summarized into the data set to be audited; generate a prompt message from the data set to be audited, and send the prompt message to the auditor.

[0065] In a possible implementation, the obtaining unit 201 is configured to obtain the review quantity corresponding to the dataset to be reviewed; the processing unit 202 is configured to determine whether the review quantity is less than or equal to a preset quantity; when the review quantity is less than or equal to the preset quantity, confirm that the number of reviewers corresponding to the dataset to be reviewed is a first quantity, and send a prompt message to the reviewers of the first quantity; when the review quantity is greater than the preset quantity, confirm that the number of reviewers corresponding to the dataset to be reviewed is a second quantity, and send a prompt message to the reviewers of the second quantity.

[0066] In a possible implementation, the obtaining unit 201 is configured to obtain the requirement information corresponding to the text data to be processed; the processing unit is configured to determine a second annotation method according to the requirement information; the annotation unit 202 is configured to annotate the entity information in the target sub-text information based on the second annotation method.

[0067] It should be noted that when the device provided in the above embodiment implements its functions, only the above division of each functional module is used for illustration. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the device and method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be seen in the method embodiment, which will not be repeated here.

[0068] This application also discloses an electronic device. Refer to Figure 3 , Figure 3 This is a schematic structural diagram of an electronic device provided by an embodiment of the present application. The electronic device 300 may include: at least one processor 301, at least one network interface 304, a user interface 303, a memory 302, and at least one communication bus 305.

[0069] Among them, the communication bus 305 is used to realize the connection and communication between these components.

[0070] Among them, the user interface 303 may include a display screen (Display) and a camera (Camera). Optionally, the user interface 303 may further include a standard wired interface and a wireless interface.

[0071] Among them, the network interface 304 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface).

[0072] Among them, the processor 301 may include one or more processing cores. The processor 301 uses various interfaces and circuits to connect various parts within the entire server. By running or executing instructions, programs, code sets, or instruction sets stored in the memory 302, and by calling the data stored in the memory 302, it performs various functions of the server and processes data. Optionally, the processor 301 may be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). The processor 301 may integrate a combination of one or several of a central processing unit (CPU), a graphics processing unit (GPU), and a modem, etc. Among them, the CPU mainly processes the operating system, user interface, and application requests, etc.; the GPU is responsible for rendering and drawing the content to be displayed on the display screen; the modem is used to process wireless communications. It can be understood that the above-mentioned modem may not be integrated into the processor 301 and may be implemented separately by a single chip.

[0073] Among them, the memory 302 may include random access memory (RAM) and may also include read-only memory. Optionally, the memory 302 includes a non-transitory computer-readable storage medium. The memory 302 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 302 may include a program storage area and a data storage area. Among them, the program storage area can store instructions for implementing the operating system, instructions for at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the above-mentioned various method embodiments, etc.; the data storage area can store the data involved in the above-mentioned various method embodiments. Optionally, the memory 302 may also be at least one storage device located far from the aforementioned processor 301.

[0074] As Figure 3 shown, the memory 302, as a computer storage medium, may include an operating system, a network communication module, a user interface module, and an application program based on entity annotation of the large language model.

[0075] In Figure 3In the electronic device 300 shown, the user interface 303 is mainly used to provide an interface for the user to input and obtain the data input by the user; and the processor 301 can be used to call the application program stored in the memory 302 and annotated with entities based on the large language model. When executed by one or more processors, the electronic device is enabled to execute one or more of the methods as described in the foregoing embodiments.

[0076] It should be noted that, for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0077] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0078] In the several embodiments provided by this application, it should be understood that the disclosed device can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some service interfaces. The indirect couplings or communication connections of the devices or units can be in electrical or other forms.

[0079] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0080] In addition, in each embodiment of this application, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0081] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The aforementioned memory includes various media that can store program codes, such as USB flash drives, mobile hard disks, magnetic disks, or optical discs.

[0082] The above are only exemplary embodiments of the present disclosure, and the scope of the present disclosure cannot be limited thereby. That is, any equivalent changes and modifications made in accordance with the teachings of the present disclosure still fall within the scope covered by the present disclosure. After considering the specification and the disclosure of the practical truth, those skilled in the art will easily think of other implementation schemes of the present disclosure. This application aims to cover any variations, uses, or adaptive changes of the present disclosure, and these variations, uses, or adaptive changes follow the general principles of the present disclosure and include common general knowledge or conventional technical means in the technical field not recorded in the present disclosure.

Claims

1. An entity annotation method based on a large language model, characterized in that, The method includes: Obtain the text data to be processed; Process the text data to be processed to obtain text information; Segment the text information according to a preset rule to obtain multiple sub-text information; Obtain the target sub-text information from the multiple sub-text information, and input the target sub-text information into a preset language model for processing to obtain entity information; Obtain the target field corresponding to the target sub-text information, and determine the first annotation method according to the target field; Annotate the entity information in the target sub-text information based on the first annotation method to complete the entity annotation of the text data to be processed.

2. The method according to claim 1, characterized in that Before inputting the target sub-text information into a preset language model for processing to obtain entity information, it is necessary to construct the preset language model, which specifically includes: Obtain a sample data set; Use the initial language model to train the sample data set until the loss value between the output result and the actual result meets the convergence condition, end the training, and use the initial language model at the end of the training as the preset language model, where each sample data set in the sample data set includes a preset entity with a label, and the preset entity includes a person name, a place name, and an organization name.

3. The method according to claim 1, wherein After obtaining the target sub-text information from the multiple sub-text information, the method further includes: Extract multiple entities to be monitored from the target sub-text information; Determine whether all the multiple entities to be monitored exist in the entity database; When all the multiple entities to be monitored exist in the entity database, confirm to input the target sub-text information into the preset language model for processing.

4. The method according to claim 3, wherein After determining whether all the multiple entities to be monitored exist in the entity database, the method further includes: If there is an entity to be monitored that does not exist in the entity database among the multiple entities to be monitored, output the target entity to be monitored, where the target entity to be monitored is the entity to be monitored that does not exist in the entity database; Input the target entity to be monitored into a preset entity model for processing to obtain a target value; Determine whether the target value is greater than or equal to a preset value; When the target value is greater than or equal to the preset value, confirm to store the target entity to be monitored in the sample data set.

5. The method according to claim 4, wherein After determining whether the target value is greater than or equal to a preset value, the method further includes: When the target value is less than the preset value, confirm to classify the target entity to be monitored into the data set to be reviewed; Generate a prompt message from the data set to be reviewed and send the prompt message to the reviewer.

6. The method according to claim 5, characterized in that, The generating a prompt message from the data set to be reviewed and sending the prompt message to the reviewer specifically includes: Obtain the review quantity corresponding to the data set to be reviewed; Determine whether the review quantity is less than or equal to a preset quantity; When the number of reviews is less than or equal to the preset number, confirm that the number of reviewers corresponding to the dataset to be reviewed is the first number, and send the prompt message to the reviewers of the first number; When the number of reviews is greater than the preset number, confirm that the number of reviewers corresponding to the dataset to be reviewed is the second number, and send the prompt message to the reviewers of the second number.

7. The method according to claim 1, characterized in that Before determining the first annotation method according to the target field in the target field corresponding to the obtained target sub-text information, the method further includes: Obtain the requirement information corresponding to the text data to be processed; Determine the second annotation method according to the requirement information; Annotate the entity information in the target sub-text information based on the second annotation method.

8. An entity annotation device based on a large language model, characterized in that, The device includes an acquisition unit (201), a processing unit (202), and an annotation unit (203); The acquisition unit (201) acquires the text data to be processed; The processing unit (202) processes the text data to be processed to obtain text information; divides the text information according to a preset rule to obtain a plurality of sub-text information; obtains target sub-text information from the plurality of sub-text information, inputs the target sub-text information into a preset language model for processing to obtain entity information; obtains the target field corresponding to the target sub-text information, and determines the first annotation method according to the target field; The annotation unit (203) annotates the entity information in the target sub-text information based on the first annotation method to complete the entity annotation of the text data to be processed.

9. An electronic device, characterized in that, It includes a processor (301), a memory (302), a user interface (303), and a network interface (304). The memory (302) is used to store instructions. The user interface (303) and the network interface (304) are used to communicate with other devices. The processor (301) is used to execute the instructions stored in the memory (302) so that the electronic device (300) executes the method according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores instructions, and when the instructions are executed, the method according to any one of claims 1-7 is executed.

Citation Information

Cited By

  • Information labeling method and device, electronic equipment and computer readable storage medium

    CN121117956A