Entity standardization method and model based on large language model retrieval enhancement
By converting named entity standardization tasks into question-and-answer tasks, using large language model retrieval enhancement methods, combining BERT and TF-IDF similarity screening candidate entities, the problems of large resource consumption and lack of context information in the prior art are solved, and efficient entity standardization is achieved.
Patent Information
- Application Number
- CN202510721003.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-08-12
AI Technical Summary
In the prior art, the rules-based entity standardization technology consumes a lot of resources and poor generalization capabilities, while machine learning-based methods lack context information and poor model interpretability.
By constructing a prompt template, the named entity standardized task is converted into a question-and-answer task, the large language model retrieval enhancement method is used, candidate entities are filtered in combination with BERT representation vector and TF-IDF similarity, and a query prompt template is constructed, and the generator is input to generate the answer.
Effectively simulate the interaction between entity mentions and standard entities, context information and candidate entities, improving the accuracy of entity standardization and generalization capabilities of the model.
Smart Images

Figure CN120471054A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of large language model retrieval, and in particular to an entity standardization method and model based on large language model retrieval enhancement. Background Art
[0002] Large language models have demonstrated powerful contextual understanding capabilities in various natural language processing tasks in general domains.
[0003] Among the related technologies, rule-based entity standardization technology: This type of technology uses pre-defined rules to match entities in the text. The rule construction consumes resources and has poor generalization ability. Machine learning-based entity standardization technology: This type of method uses supervised learning models such as support vector machines and random forests for entity standardization. It does not use contextual information and the model has poor interpretability. Summary of the Invention
[0004] The purpose of the present invention is to solve the above problems and provide an entity standardization method based on large language model retrieval enhancement. By constructing a prompt template, the named entity standardization task is converted into a question-answering task, where the prompt template is a query about entity mentions, including entity mentions, context and a list of candidate entities.
[0005] In order to solve the above problems, the present invention provides the following technical solutions: On the one hand, a method for entity standardization based on large language model retrieval enhancement includes Compare entity mentions with the standard entity dictionary to obtain multiple candidate entities and form a candidate entity set; Get the context of an entity mention; A query prompt template is constructed based on entity mentions, context, and candidate entity sets, and is input into a generator to generate an answer, where the answer option corresponds to a standard entity in the candidate entity set.
[0006] In a related embodiment, comparing the entity mention with a standard entity dictionary to obtain a plurality of candidate entities includes: Calculate the BERT representation vector and TF-IDF score of entity mentions and standard entities in the standard entity dictionary respectively; Calculate the BERT similarity and TF-IDF similarity between entity mentions and standard entities; According to the similarity between entity mentions and standard entities, the k standard entities with the highest similarity are selected as candidate entities.
[0007] In a related embodiment, the BERT similarity and TF-IDF similarity calculation process between entity mentions and standard entities is as follows: Among them, the entity mentioned and standard entities BERT representation vector and TF-IDF score , BERT similarity between entity mentions and standard entities and TF-IDF similarity , f represents the cosine similarity function, R is a mathematical symbol representing the real number field, and the meaning in the formula is that the calculated similarity is a real number.
[0008] In a related embodiment, a query prompt template is constructed based on entity mentions, context, and candidate entity sets, including context, questions for selecting standard entities for entity mentions, candidate entities as options in multiple candidate entity sets, and seeking answers, and the candidate entities include their entity description information.
[0009] In a related embodiment, an answer is generated in an input generator, wherein the answer option corresponds to a standard entity in a candidate entity set and includes: Build a key-value database. Given an entity mention that needs to be standardized, calculate its BERT representation vector combined with the distance metric, and query the database for the entity mention with the closest cosine distance. training samples; After obtaining multiple neighbor training samples, a prompt text with an answer is constructed according to the query prompt template. The prompt texts of all samples and the prompt text corresponding to the given entity mention are concatenated and input into the generator together to obtain the answer screening result.
[0010] In a related embodiment, building a key-value database includes: The i-th entity mention in the training set is ,calculate BERT representation vector ; Calculated by the retriever k candidate entities , and 、 、 Context The answer options corresponding to the label of the sample (the true standard entity) Forming a quad ; Will BERT representation vector As a key, Quadruple as value, forming key-value pair ; After constructing key-value pairs for each sample in the training set according to the above process, these key-value pairs are stored in a list as a training sample database.
[0011] In the second aspect, a retrieval enhancement generation model based on a large language model is characterized in that it includes a retriever and a generator; the retriever is based on an entity standardization method for retrieval enhancement based on a large language model, and the generator is used to implement an entity standardization method for retrieval enhancement based on a large language model.
[0012] In a related embodiment, the loss function of the retriever is a contrastive learning loss function, and the generator is trained using maximum likelihood estimation, with the goal of maximizing the conditional probability of the target sequence.
[0013] In a third aspect, an electronic device includes a memory and a processor, wherein the memory stores computer-readable instructions, and when the computer-readable instructions are executed by the processor, the processor performs an entity standardization method based on large language model retrieval enhancement.
[0014] In a fourth aspect, a computer-readable storage medium stores a computer program, which, when executed on a computer, enables the computer to execute an entity standardization method based on large language model retrieval enhancement.
[0015] Compared with the prior art, the present invention has the following beneficial effects: This paper transforms the named entity standardization task into a question-answering task by constructing a prompt template. The prompt template is a query about entity mentions, including the entity mention, context, and a list of candidate entities. This method effectively simulates the interactions between entity mentions and standard entities, entity mentions and contextual information, and candidate entities. In addition, to ensure consistency between entity mentions and candidate entity contextual information, entity description information of the candidate entities is added to the prompt template. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the specific embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. Those skilled in the art can also derive other drawings based on these drawings without inventive work, among which: Figure 1 The framework of this embodiment Figure 1 .
[0017] Figure 2 The framework of this embodiment Figure 2 .
[0018] Figure 3 This is an example diagram of the retriever.
[0019] Figure 4 This is a schematic diagram of the prompt template.
[0020] Figure 5 This is an example diagram of the generator.
[0021] Figure 6 Filter example graphs for K-nearest neighbor training samples.
[0022] Figure 7 This is a flow chart of this embodiment. DETAILED DESCRIPTION
[0023] In order to make the purpose, technical solutions and advantages of the present invention clearer, the following Figures 1 to 7 The present invention is further described in detail. The described embodiments should not be regarded as limiting the present invention. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0024] like Figure 1 and 7 As shown, large language models demonstrate powerful contextual understanding capabilities in a variety of natural language processing (NLP) tasks across general domains. To leverage this contextual learning capability, this embodiment proposes an entity normalization model based on large language model retrieval enhancement. This model transforms the named entity normalization task into a question-answering task by constructing a prompt template. The prompt template is a query about an entity mention, including the entity mention, context, and a list of candidate entities.
[0025] This method can effectively simulate the interaction between entity mentions and standard entities, entity mentions and contextual information, and candidate entities. In addition, to ensure the consistency of entity mentions and candidate entity context information, the entity description information of the candidate entity is added to the prompt template.
[0026] like Figure 1 As shown, for example: the entity mention input is mild mania, the entity mention is compared with the standard dictionary to obtain multiple candidate entities, such as mood disorder, mania, depression, bipolar disorder, etc., and then the candidate entities are calculated.
[0027] An inquiry prompt template is constructed based on entity mentions, context, and candidate entity sets, including context, questions for selecting a standard entity for entity mentions, candidate entities as options in multiple candidate entity sets, and seeking answers. The candidate entities include their entity description information. In this embodiment, for example, the context includes evidence that bipolar disorder is linked to chromosome 18 and has a parent-of-origin effect. A standard entity is selected from the candidate list to represent the entity "bipolar disorder" mentioned in the context. Candidate list: A. Bipolar disorder: A major mood disorder characterized by severe mood swings and a tendency to remit and relapse; B. Depression: usually refers to a moderately severe depressive state, which is different from the severe depression in neurotic and psychotic disorders; C....; Answer: A.
[0028] Entity standardization method based on large language model retrieval enhancement Retrieval-enhanced generation methods enhance the capabilities of generative models by retrieving external knowledge sources. Their basic framework consists of a retriever and a generator. The retriever retrieves relevant information from an external knowledge base, while the generator generates text based on the retrieval results.
[0029] like Figure 2 As shown, in the entity standardization method based on large language model retrieval enhancement proposed in this embodiment, the standard entity dictionary corresponds to an external knowledge source, and a retriever filters out candidate entities from the standard entity dictionary. For example, the entity mention input is mild mania, and the entity mention is compared with the vocabulary to obtain multiple candidate entities, such as mood disorders, mania, depression, bipolar disorder, etc. The screening basis is the distance between the representation of the entity mention and the representation of the standard entity in the vocabulary in the feature space. After the retriever generates the candidate entity, it constructs the entity mention, context and candidate entity list together into a query. The specific form of the query is a prompt template. The generator takes a prompt text of a query as input and outputs the predicted answer.
[0030] For example, in context: Bipolar disorder is characterized by episodes of mania or hypomania alternating with depression; A standard entity is selected from the candidate list to represent the entity "hypomania" mentioned in the context; Candidate entity list: A. Mood disorders: Psychological disorders characterized by mood disturbances; B. Mania: A state of excitement with excessive activity, sometimes accompanied by psychotic symptoms; C. Bipolar disorder: ...; D. Depression: ...; E. Anxiety disorders: ...; Answer: Finally, the generator outputs the answer A.
[0031] like Figure 3As shown in the figure, the retriever includes a BERT-based semantic encoder and a TF-IDF-based shape encoder to represent entity mentions and standard entities in the vocabulary, such as mild mania. The retriever selects candidate entities corresponding to entity mentions based on the cosine distance between the representation vectors, such as mood disorders, mania, depression, and bipolar disorder. The model constructs a prompt template based on the entity mention, context, and corresponding candidate entities, and inputs the prompt text into the generator, which ultimately outputs an answer option corresponding to a standard entity.
[0032] Retriever The retriever functions similarly to the candidate generation stage, consisting of a BERT-based encoder and a TF-IDF model to calculate entity mentions. and standard entities BERT representation vector and TF-IDF score BERT is an encoder that calculates the representation vector of a word, while TF-IDF is a statistical algorithm used to calculate the frequency of a word (or phrase) in the entire corpus. The essential difference between the two is that BERT reflects the semantic information of a word, while TF-IDF scores reflect the appearance information of a word. Combining the two to calculate the similarity between words can consider the similarity between entity mentions and standard entities from two perspectives; thus, the BERT similarity between entity mentions and standard entities can be calculated. and TF-IDF similarity , the calculation process is as follows: in, Represents the cosine similarity function, R is a mathematical symbol representing the real number domain, and its meaning in the formula is that the calculated similarity is a real number.
[0033] The retriever selects entities based on the similarity between entity mentions and standard entities. The standard entities with the highest similarity are selected as candidate entities.
[0034] Prompt construction By entity mention , the context of entity mentions and retriever generated candidate entities and its corresponding entity description , construct a prompt as the input of the generator.
[0035] like Figure 4As shown, the prompt template contains instructions, contextual information, and a list of candidate entities. To ensure consistency between entity mentions and candidate entities, each candidate in the candidate entity list is accompanied by its corresponding entity description, which serves as the candidate entity's contextual information. By constructing prompts, this method transforms the named entity standardization problem into a multi-question question-answering task: selecting the correct answer from a set of answer options—candidate entities—and predicting the true standard entity.
[0036] Generator like Figure 5 As shown, the generator includes a large language model, such as a T5 model, which takes the constructed prompt text as input, answers the questions in the prompt text, and outputs the expected answer, that is, the predicted standard entity.
[0037] In this embodiment, each entity mention, contextual information, and the corresponding answer option symbol for each candidate entity are input into the generator as a single prompt template. This ensures interaction between entity mentions and candidate entities, entity mentions and contextual information, and candidate entities and candidate entities. Furthermore, the use of option symbols, such as "A," in the prompt templates allows the generator to generate single characters rather than specific entities, thus preventing invalid output.
[0038] Furthermore, research has shown that the performance of current language models in question-answering tasks is somewhat dependent on the order of answer options. To mitigate the impact of answer option order on model performance, this embodiment employs a data augmentation strategy: when training the generator, an additional prompt example is added to each training example, in which the order of the options is randomly swapped, while the rest of the text remains unchanged. This strategy allows the generator to better understand the relationship between the answer option symbols and the answer (the candidate entity itself), rather than simply memorizing the answer's position.
[0039] K-nearest neighbor training sample screening To improve the generalization ability of the model, this example introduces a K-nearest neighbor training sample screening module. This module uses the KNN algorithm to find training samples similar to the current query from the entire training set as examples for generator inference. To run the KNN algorithm, a key-value pair database based on the entire training set must be constructed. The construction process is as follows: ·Remember the first Entities mentioned as ,calculate BERT representation vector .
[0040] Calculated by the retriever of candidate entities , and 、 、 Context The answer options corresponding to the label of the sample (the true standard entity) Forming a quad .
[0041] ·Will BERT representation vector As a key, Quadruple as value, forming key-value pair .
[0042] After constructing key-value pairs for each sample in the training set according to the above process, these key-value pairs are stored in a list as the training sample database.
[0043] After constructing the above key-value pair database, given an entity mention that needs to be standardized , this module will calculate its BERT representation vector , and then based on the KNN algorithm, query the The closest cosine distance training samples , Top-N will correspond to the first N samples sorted by cosine distance from small to large, representing the first N samples with the highest similarity to the given entity mention. The process is as follows Figure 6 As shown. Since entity mentions are added when constructing the database The corresponding key-value pairs need to be The corresponding samples from Remove from .
[0044] get After generating K-nearest neighbor training samples, this module constructs prompt text with answers based on the prompt template. It then concatenates the prompt text from all samples, along with the prompt text corresponding to the given entity mention, and finally inputs them into the generator. Figure 2 shows the results of a K-nearest neighbor training sample screening. In this process, the prompt text generated based on the K-nearest neighbor training samples provides reference clues for the generator to answer questions, thereby enhancing the model's generalization ability.
[0045] Model training For the retriever, this embodiment adopts contrastive learning method for training. The goal of contrastive learning is to make entity mentions Similar standard entities The distance in the feature space is close, while making it similar to the standard entity The contrastive learning loss function is as follows: in, is the similarity function (cosine similarity), is the temperature parameter, Is an entity mention The negative sample set.
[0046] For the generator, this embodiment uses maximum likelihood estimation for training, with the goal of maximizing the conditional probability of the target sequence. The loss function is as follows: in, is the target sequence length, is the input sequence, is the first tokens, Before the target sequence tokens; P represents the model when the input sequence is x and the output is t-1 tokens before Under the condition of , the output of token t is The conditional probability of .
[0047] Experimental results and analysis Dataset and Experimental Settings Datasets: This paper evaluates the entity normalization method based on a large language model proposed in this example using three public datasets: NCBI, BC5CDR, and COMETA. The statistics for each dataset are shown in Table 1.
[0048] NCBI is a disease corpus containing 792 medical abstract texts. The entity mentions in each document were manually annotated and mapped to the CUI in the Medical Subject Thesaurus (MeSH) or the Human Genetics Thesaurus (OMIM).
[0049] BC5CDR is an evaluation dataset for the BioCreative V CDR-disease relation extraction task. It consists of two sub-datasets: BC5CDR-Disease and BC5CDR-Chemical. BC5CDR-Disease includes only samples where entity mentions are disease names. These entity mentions have been manually annotated and mapped to the MEDIC medical dictionary. BC5CDR-Chemical includes only samples where entity mentions are chemical substances. These entity mentions have been manually annotated and mapped to the chemical dictionary of the Comparative Toxicogenomics Database.
[0050] COMETA is a corpus of 20,000 biomedical entity mentions from the Reddit social platform, with expert annotations mapping the entity mentions to the systematic medical nomenclature - clinical terminology (SNOMED CT).
[0051] Table 1: Statistics of the data set used in the experiment of this example Dataset Number of entities in the training set Number of entities in the validation set Number of test set entities NCBI 5134 787 960 BC5CDR 9285 9515 9654 COMETA 13489 2176 4350 Experimental setup: This example experiment uses the T5-base model as the base model of the generator. The number of training rounds is 10 and the batch size is 8. In contrastive learning, the temperature coefficient Set to 0.01. The model training uses the Adam optimizer, and the learning rate of BERT related parameters is set to , the learning rates of other parameters are set to The number of candidate entities generated by the retriever 5, the number of similar samples generated by the K nearest neighbor sample screening module is 3.
[0052] Comparative experimental results and analysis The LLMRAG-BioEL model proposed in this example was evaluated on three public datasets: NCBI, BC5CDR, and COMETA. The benchmark models compared with the model in this example were divided into three categories: 1. Discriminative methods: These methods use encoders to retrieve relevant entities, including the BioSYN model proposed by Sung et al., the ResCNN model proposed by Lai et al., the Prompt-BioEL model proposed by Xu et al., and the IA-BioSYN model proposed by Peng et al. 2. Generative methods: These methods directly generate standard entities, including the BioBART model and GenBioEL model proposed by Yuan et al., and the kNN-BioEL model and BioELQA model proposed by Lin et al. 3. Large model-based methods: This type of method constructs prompts and uses large language models to generate standard entities. The large models involved include GPT-3.5, LLaMA-2 and Claude.
[0053] Table 2: Comparative experimental results (Acc@1) Model NCBI BC5CDR COMETA Average performance BioSYN-SapBERT 91.1 - 71.3 74.8 ResCNN 92.4 - 80.1 82.3 Prompt-BioEL 92.6 93.7 83.7 90.7 IA-BioSYN-SapBERT 93.1 93.7 84.6 91.0 BioBART 89.3 93.0 81.8 89.5 GenBioEL 91.1 93.3 81.4 89.7 kNN-BioEL 92.8 93.7 <![CDATA[ 85.7 ]]> <![CDATA[ 91.3 ]]> BioELQA <![CDATA[ 92.9 ]]> <![CDATA[ 93.8 ]]> 84.9 91.2 GPT-3.5 52.2 54.9 43.5 51.4 LLaMA-2-13B 59.2 66.5 40.7 58.5 Claude-2 70.2 78.0 53.3 70.3 LLMRAG-BioEL 92.6 93.8 85.8 91.4 The comparative experimental results are shown in Table 2, where the optimal model performance is bolded and the suboptimal model performance is underlined. The experimental results show that the LLMRAG-BioEL model proposed in this example performs worse than the discriminant IA-BioSYN model and the two generative methods on the NCBI dataset, but outperforms the other models on the BC5CDR and COMETA datasets. The model proposed in this example achieves the best average performance.
[0054] First, the LLMRAG-BioEL model significantly outperforms discriminative methods, demonstrating the effectiveness of the retrieval-augmented generation approach itself. Furthermore, most discriminative methods fail to utilize the contextual information of entity mentions, considering only the interaction between entity mentions and candidate entities. This also demonstrates the effectiveness of contextual information in entity standardization tasks.
[0055] Secondly, the performance of the LLMRAG-BioEL model is very similar to that of the compared generative methods. The main difference between the model of this embodiment and the compared models is still related to contextual information.
[0056] Finally, the LLMRAG-BioEL model far outperforms methods based on large language models. This is because large language models lack sufficient expertise in specific domains and still have limitations in specific natural language processing tasks. This embodiment model utilizes a retrieval-augmented generation framework to supplement the large model with exogenous knowledge. Specifically, the retriever filters candidate entities from the vocabulary, thereby improving the large model's prediction accuracy.
[0057] In summary, the Acc@1 of the LLMRAG-BioEL model on the NCBI, BC5CDR and COMETA datasets reached 92.6%, 93.8% and 85.8% respectively, and the average Acc@1 on the three datasets reached 91.4%, which is higher than other benchmark models.
[0058] Ablation experiment results and analysis This example uses ablation experiments to verify the effectiveness of each module in the LLMRAG-BioEL model proposed in this example. The ablation experiment group tests include: w / o context: The model does not use contextual information, including entity mention context and entity description information of candidate entities; w / o entity description: The model does not use the entity description information of the candidate entity; w / o random options: No data augmentation is performed to randomly shuffle the order of answer options; w / o TF-IDF retriever: The retriever only uses BERT to encode the entity representation; w / o KNN Sample filter: No KNN training sample filtering is performed, and the prompt text input to the generator contains only a single query; with random instances: Each query uses a random training sample as an example; w / o option character: Output the entity directly instead of the answer option symbol. Table 3: Ablation experiment results (Acc@1) Model NCBI BC5CDR COMETA LLMRAG-BioEL 92.6 93.8 85.8 w / o context 92.1 93.6 85.0 w / o entity description 92.6 93.7 85.6 w / o random options 92.4 93.6 84.5 w / o TF-IDF retriever 91.8 93.5 84.6 w / o KNN sample filter 90.1 92.8 82.7 with random examples 90.3 92.3 83.2 w / o option character 70.1 89.2 79.3 Table 3 lists the evaluation results of various ablation groups for LLMRAG-BioEL. Across all ablation groups, model performance degraded most significantly when generating meaningful entity names directly without using option characters as output (without option characters), with an average accuracy drop of 11.2% across all three datasets. In this ablation group, the generative model was prone to generating invalid text, leading to performance degradation. Without the K-nearest neighbor (KNN) sample filter and with random examples, the average model performance dropped by 2.2% and 2.1%, respectively. This indicates that the similar examples generated by the K-nearest neighbor (KNN) sample filter provide effective predictive cues for the generator. Without context, the average model performance dropped by 0.5%, demonstrating that contextual information is still effective when performing entity normalization in large language models. Experimental results demonstrate that each module of the LLMRAG-BioEL model has a beneficial effect on the entity normalization task within the retrieval-augmented generation framework.
[0059] Summary of this embodiment This embodiment introduces a large language model-based entity standardization method. It employs a large-model-based retrieval-enhanced generation framework to transform the entity standardization task into a question-answering task. It incorporates contextual information from entity mentions into prompt text and introduces a K-nearest neighbor training sample screening module to filter samples similar to the entity mentions in the query across the entire training set as clues for the large-model to generate standard entities. This paper validates the proposed LLMRAG-BioEL model on three public datasets in the field of entity standardization. The experimental results demonstrate that the model proposed in this embodiment effectively utilizes contextual information. With the assistance of the K-nearest neighbor training sample screening module, the model can obtain effective predictive prompt information when generating standard entities, thereby improving the model's performance on entity standardization tasks.
[0060] An electronic device comprises a memory, a processor, and a computer program stored in the memory and executable on the processor; when the processor executes the computer program, the processor implements the steps of an entity standardization method based on large language model retrieval enhancement.
[0061] The electronic device may be a desktop computer, a laptop, a PDA, a cloud server, or other electronic device. The electronic device may include, but is not limited to, a processor and a memory. Those skilled in the art will appreciate that the figures are merely examples of electronic devices and do not limit the scope of the electronic device. The electronic device may include more, fewer, or different components than shown.
[0062] The processor can be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.
[0063] Memory can be an internal storage unit of an electronic device, such as its hard drive or memory. It can also be an external storage device, such as a plug-in hard drive, SmartMedia Card (SMC), Secure Digital (SD) card, or Flash Card. Memory can also include both internal storage units and external storage devices. Memory is used to store computer programs and other programs and data required by the electronic device.
[0064] In the several embodiments provided in this application, it should be understood that the disclosed methods and models can also be implemented in other ways. The model embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of the methods, models, and computer program products according to multiple embodiments of the present invention. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or part of a code, which contains one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions marked in the boxes can also occur in an order different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, as well as the combination of boxes in the block diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified functions or actions, or can be implemented using a combination of dedicated hardware and computer instructions.
[0065] In addition, the functional modules in the various embodiments of the present invention may be integrated together to form an independent part, or each module may exist independently, or two or more modules may be integrated to form an independent part.
[0066] If the functions are implemented as software modules and sold or used as standalone products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product, stored in a storage medium, includes instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods described in various embodiments of the present invention. It should be noted that, in this document, relational terms such as first and second, etc., are used solely to distinguish one entity or operation from another, and do not necessarily require or imply any actual relationship or order between these entities or operations. Furthermore, the terms "comprise," "include," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or device comprising a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. Without further limitation, the phrase "comprising a..." does not preclude the presence of additional identical elements in the process, method, article, or device comprising the elements.
[0067] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Various modifications and variations are readily apparent to those skilled in the art. Any modifications, equivalent substitutions, improvements, and the like made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention. It should be noted that similar reference numerals and letters denote similar items in the following figures. Therefore, once an item is defined in one figure, it need not be further defined or explained in subsequent figures.
[0068] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.
Claims
1. A method for entity standardization based on large language model retrieval enhancement, characterized in that: Compare entity mentions with the standard entity dictionary to obtain multiple candidate entities and form a candidate entity set; Get the context of an entity mention; A query prompt template is constructed based on entity mentions, context, and candidate entity sets, and is input into a generator to generate an answer, where the answer option corresponds to a standard entity in the candidate entity set.
2. The entity standardization method based on large language model retrieval enhancement according to claim 1 is characterized in that: Comparing entity mentions with standard entity dictionaries yields several candidate entities, including: Calculate the BERT representation vector and TF-IDF score of entity mentions and standard entities in the standard entity dictionary respectively; Calculate the BERT similarity and TF-IDF similarity between entity mentions and standard entities; According to the similarity between entity mentions and standard entities, the k standard entities with the highest similarity are selected as candidate entities.
3. The entity standardization method based on large language model retrieval enhancement according to claim 2 is characterized in that: The BERT similarity and TF-IDF similarity calculation process between entity mentions and standard entities is as follows: Among them, the entity mentioned and standard entities BERT representation vector and TF-IDF score , BERT similarity between entity mentions and standard entities and TF-IDF similarity , f represents the cosine similarity function, R is a mathematical symbol representing the real number field, and the meaning in the formula is that the calculated similarity is a real number.
4. The entity standardization method based on large language model retrieval enhancement according to claim 1 is characterized in that: An inquiry prompt template is constructed based on entity mentions, context, and candidate entity sets, including context, questions for selecting standard entities for entity mentions, candidate entities as options in multiple candidate entity sets, and seeking answers. The candidate entities contain their entity description information.
5. The entity standardization method based on large language model retrieval enhancement according to claim 1 is characterized in that: The query prompt template is input into the generator to generate an answer, and the answer option corresponds to a standard entity in the candidate entity set, including: Build a key-value database. Given an entity mention that needs to be standardized, calculate its BERT representation vector and combine it with the distance metric to query the entity mention from the database with the closest cosine distance. training samples; After obtaining multiple neighbor training samples, a prompt text with an answer is constructed according to the query prompt template. The prompt texts of all samples and the prompt text corresponding to the given entity mention are concatenated and input into the generator together to obtain the answer screening result.
6. The entity standardization method based on large language model retrieval enhancement according to claim 5 is characterized in that: Building a key-value database includes: The i-th entity mention in the training set is ,calculate BERT representation vector ; Calculated by the retriever k candidate entities , and 、 、 Context The answer options corresponding to the label of the sample Forming a quadruple ; Will BERT representation vector As a key, Quadruple as value, forming key-value pair ; After constructing key-value pairs for each sample in the training set according to the above process, these key-value pairs are stored in a list as a training sample database.
7. A retrieval enhancement generation model based on a large language model, characterized by: It includes a retriever and a generator; the retriever is used to implement the entity standardization method based on large language model retrieval enhancement as described in any one of claims 1-6, and the generator is used to implement the entity standardization method based on large language model retrieval enhancement as described in any one of claims 1-6.
8. The retrieval enhancement generation model based on a large language model according to claim 7 is characterized in that: The loss function of the retriever is the contrastive learning loss function, and the generator is trained using maximum likelihood estimation, with the goal of maximizing the conditional probability of the target sequence.
Citation Information
Cited By
Depression analysis scoring method based on knowledge distillation and retrieval enhancement generation
CN120656733A
Entity identification method and device for nutritional and healthy food field
CN120706427A
Nutritionally healthy food domain entity recognition method and apparatus
CN120706427B