Entity linking method and system based on large language model

By converting entity linking tasks into question-and-answer questions for single-choice questions, and using multiple rounds of dialogue technology of large language models, the problem of ignoring entity associations within documents in the existing technology is solved, and better document-level entity linking performance and generalization are achieved.

CN120181077AActive Publication Date: 2025-06-20INST OF COMPUTING TECH CHINESE ACAD OF SCI

Patent Information

Application Number
CN202510639279.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-19
Publication Date
2025-06-20
Estimated Expiration
2045-05-19

AI Technical Summary

Technical Problem

When handling document-level entity linking tasks, existing entity linking technologies ignore potential semantic associations between individual entity mentions within the document, resulting in poor performance.

Method used

Using multiple rounds of dialogue technology based on large language model, the entity linking task is converted into a single-choice question question question. By enhancing the context, generating candidate entity lists and description information, building reference diagrams, and using a random walk algorithm with restart, the correlation between entity referents is calculated, and the linking task of all entity referents in the document is processed in the same dialogue context of the large model.

Benefits of technology

Effectively capture potential connections between entity reference items, improves document-level entity link task performance, reduces maintenance costs, and has good generalization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120181077A_ABST
    Figure CN120181077A_ABST
Patent Text Reader

Abstract

The invention discloses an entity linking method and system based on a large language model, and the method comprises the steps: carrying out the enhancement of the context of a given entity reference item in an entity document through a large language model, generating a candidate entity list for the entity reference item, and generating description information for each candidate entity in the candidate entity list; constructing a question and answer task of a single choice question for each entity reference item; constructing a reference graph by utilizing the entity reference items and the candidate entity list to obtain association degrees among the entity reference items; all the entity reference items are sorted, question and answer pairs of the preset number of entity reference items with the highest association degree of the current entity reference items are selected as dialogue contexts according to the sorting result, single choice questions are sequentially input into a large language model for answering, and the entity linking result of each entity reference item is obtained. According to the method, the potential relation between the entity reference items can be effectively captured, and the method is good in performance on the entity link task of the document level.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an entity linking method and system, and particularly to an entity linking method and system based on a large language model. Background Art

[0002] Entity linking refers to correctly linking the given entity mentions in a document to the entities in a knowledge base, which is a basic task in natural language processing and plays a key role in fields such as information extraction, knowledge graph construction, and question system construction.

[0003] Existing entity linking techniques are divided into two categories. The first category is rule-based methods, which rely on manual rules, dictionary matching, or probability models to match entity mentions, performing well in specific domains but having high maintenance costs and difficulty in handling complex semantics. The second category is deep learning-based methods, which automatically learn the representations of entities and contexts through neural networks to achieve end-to-end entity linking, capable of automatically learning representations and handling complex semantics, but requiring a large amount of domain-specific labeled data and having poor generalization ability.

[0004] With the rapid development of large language models (LLMs), relying on their powerful language understanding ability and zero-shot learning ability, the entity linking task is gradually shifting towards a new paradigm centered around large language models. This paradigm demonstrates good generalization performance in various application scenarios and has received extensive attention. However, current mainstream methods often independently process each entity mention in a document, ignoring the potential semantic associations between entity mentions within the document, resulting in poor performance in document-level entity linking tasks. To address this issue, some studies have attempted to model entity linking as a comprehensive question-answering task, integrating all entity mentions and their candidate entity information in the same document to achieve unified disambiguation. However, such methods still have a significant gap in actual effects and have not shown sufficient performance advantages. Therefore, how to make full use of large language models for document-level joint disambiguation remains a key challenge in current entity linking research. Summary of the Invention

[0005] To solve the above problems, this paper proposes an entity linking method and system based on multi-turn dialogue technology of large language models.

[0006] In a first aspect, an embodiment of the present application provides an entity linking method based on a large language model, the method comprising:

[0007] Enhancing the context of the given entity mentions in the entity document using a large language model, generating a candidate entity list for the entity mentions, and generating description information for each candidate entity in the candidate entity list;

[0008] Using the enhanced context, the candidate entity list, and the description information of the candidate entities, construct a multiple-choice question and answer task for each entity mention;

[0009] Using the entity mentions in the document and the candidate entity list generated for the entity mentions, construct a reference graph, and obtain the correlation degree between entity mentions through the random walk algorithm with restart;

[0010] Sort all entity mentions. According to the sorting result, sequentially input the multiple-choice questions corresponding to the entity mentions into the large language model for answering. When inputting, select the question and answer pairs of entity mentions whose correlation degree with the current entity mention exceeds the threshold as the dialogue context, and obtain the entity linking results of each entity mention.

[0011] In the embodiment of the present invention, in the above step of enhancing the context of the entity mentions given in the document by using the large language model, it includes:

[0012] Use the entity mention and the preset number of characters before and after the entity mention as the context, and complete the context enhancement through the form of question and answer of the large language model according to the question template. Among them, the question template includes: the text content and the enhanced context of the entity mention output by the large language model, and the text content includes the context of the entity mention.

[0013] In the embodiment of the present invention, in the above step of generating a list of candidate entities for the entity mentions, it includes:

[0014] According to the enhanced context, use the entity linking model to filter the entities in the knowledge base, generate a list of candidate entities including the most likely entities to be linked as the candidate entity list, obtain the candidate entity list of the entity mentions and the matching scores of each candidate entity, and select a preset number of candidate entities as the candidate entity list.

[0015] In the embodiment of the present invention, in the above step of generating description information for each candidate entity in the list of candidate entities, it includes:

[0016] The description information includes: the first part of the content is one of the content of the corresponding entity description information in the knowledge base, and one of the content is the abstract description of the current entity; and the second part of the content is the content with the strongest relevance to the context recalled from the corresponding entity description information in the knowledge base, and the recall uses the vector recall method;

[0017] After splitting the entity document into multiple blocks according to a fixed character length, use the text embedding model to generate embeddings for each block, calculate the cosine similarity between the embedding of the enhanced context and the embeddings of all blocks, and use the block with the highest score as the second part of the description information of the candidate entity; splice the first part and the second part of the description as the description information of the candidate entity.

[0018] In the embodiments of the present invention, in the step of constructing a multiple-choice question and answer task for each entity reference item, it includes:

[0019] For the entity reference item, use the enhanced context obtained as the context of the multiple-choice question, use the obtained list of candidate entities as options, and use the obtained candidate entity description information as the description information of each option;

[0020] The template of the multiple-choice question includes: context, question, options, and the last option is none of the above.

[0021] In the embodiments of the present invention, in the step of constructing a reference graph and obtaining the correlation degree between entity reference items, it includes:

[0022] Construct a weighted graph including two types of nodes: entity reference items and candidate entities, where the weight between the entity reference item and the candidate entity is the matching score of the obtained candidate entity and the entity reference item. The weight between candidate entities is the semantic correlation degree between candidate entities. The semantic correlation degree is obtained by calculating the number of in-link entities shared by two entities in the knowledge base.

[0023] Calculate the correlation degree between entity reference items through a random walk algorithm with restart. Starting from each entity reference item to be disambiguated, explore its propagation path in the reference graph, and then measure its correlation with other entity reference items.

[0024] In the embodiments of the present invention, in the step of sorting all entity reference items, it includes:

[0025] Use a heuristic method to estimate the link difficulty of entity reference items and sort them. Calculate the highest matching score among the candidate entities of the entity reference item as the matching difficulty score of the entity reference item, and sort according to the score.

[0026] In the embodiments of the present invention, in the step of sequentially inputting the multiple-choice questions corresponding to the entity reference items into a large language model for answering according to the sorting result, and selecting the question and answer pairs of entity reference items with a correlation degree exceeding the threshold with the current entity reference item as the dialogue context when inputting, and obtaining the entity linking result of each entity reference item, it includes:

[0027] When asking a multiple-choice question about one entity reference item, select at most N link information of entity reference items with the highest correlation degree with the current entity reference item from the previously processed entity reference items, including the multiple-choice questions and the answers made by the large model as the dialogue context; where N is the number of entity reference items used as the dialogue context;

[0028] The question templates of the large language model include: task definition, instruction description, and examples. Among them, the instruction description includes: context, question, options, and answers.

[0029] In a second aspect, an embodiment of the present application provides an entity linking system based on a large language model, which adopts the entity linking method based on the large language model as described above. The system includes:

[0030] Context enhancement module: used to enhance the context of the entity referent using the large language model;

[0031] Candidate entity generation module: used to generate a list of candidate entities for the entity referent;

[0032] Description information recall module: used to generate description information for each candidate entity in the list of candidate entities;

[0033] Single-choice question construction module: used to construct a single-choice question in a fixed format for each entity referent using the enhanced context, candidate entity list, and description information of the candidate entity;

[0034] Entity referent relevance calculation module: constructs a reference graph using the entity referents in the document and the list of candidate entities generated for the entity referents, and obtains the relevance between entity referents through the random walk algorithm with restart;

[0035] Sorting module: used to sort all entity referents;

[0036] Entity linking module: According to the sorting result, sequentially input the single-choice questions corresponding to the entity referents into the large language model for answering. When inputting, select the Q&A pairs of entity referents with a relevance exceeding the threshold to the current entity referent as the dialogue context, and obtain the entity linking result of each entity referent.

[0037] In a third aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the steps of the entity linking method based on the large language model.

[0038] In a fourth aspect, an embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps of the entity linking method as described above.

[0039] Compared with the related prior art, it has the following outstanding beneficial effects:

[0040] The method of the present invention converts the entity linking task into a multiple-choice question and answer task, makes full use of the context within the document and the rich world knowledge inherent in the large language model for reasoning, without the need to pre-define rules, a large amount of labeled data, and train or fine-tune the model, with low maintenance costs and good generalization;

[0041] The method of the present invention processes the linking tasks of all entity mentions within the document in the same dialogue context of the large model, can effectively capture the potential connections between entity mentions, and performs well in document-level entity linking tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:

[0043] Figure 1 is a schematic diagram of the entity linking method based on the large language model of the present invention;

[0044] Figure 2 is a schematic diagram of the entity linking system based on the large language model of the present invention;

[0045] Figure 3 is a schematic diagram of the computer hardware of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0046] In the present invention, "at least one" means one or more, and "a plurality" means two or more. "At least one (item)" or its similar expression refers to any combination of these items, including any combination of single (item) or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, c can be single or multiple.

[0047] It should also be understood that the term "and / or" in this article is only a description of the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Among them, A and B can be singular or plural. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after, but may also represent an "and / or" relationship, which can be specifically understood with reference to the context before and after.

[0048] It should also be understood that in various embodiments of the present invention, the magnitude of the serial numbers of the above processes does not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.

[0049] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling, direct coupling, or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the devices or units can be in electrical, mechanical, or other forms.

[0050] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0051] In addition, in each embodiment of the present invention, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0052] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention. The foregoing storage medium includes: USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks, and other media that can store program codes.

[0053] To make the above features and effects of the present invention more clearly understandable, the following specific examples are given and detailed descriptions are made in conjunction with the accompanying drawings of the specification as follows. This specification discloses one or more embodiments incorporating the features of the present invention. The disclosed embodiments are only for illustrative purposes. The protection scope of the present invention is not limited to the disclosed embodiments, and the present invention is defined by the appended claims.

[0054] The following is a system embodiment corresponding to the method embodiment above. This embodiment can be implemented in cooperation with the above embodiment. The relevant technical details mentioned in the above embodiment are still valid in this embodiment. To avoid repetition, they will not be elaborated here. Correspondingly, the relevant technical details mentioned in this embodiment can also be applied to the above embodiment.

[0055] The method of the present invention aims to propose an entity linking method and system based on the multi-turn dialogue technology of large language models. The method of the present invention converts the entity linking task into a multiple-choice question and answer task, makes full use of the context within the document and the rich world knowledge inherent in the large language model for reasoning, without the need to pre-define rules, a large amount of labeled data, and train or fine-tune the model, with low maintenance costs and good generalization; the method of the present invention processes the linking tasks of all entity mentions within the document in the same dialogue context of the large model, can effectively capture the potential connections between entity mentions, and performs well in document-level entity linking tasks.

[0056] The following will be a detailed description of the method of the embodiment of the present application in combination with specific embodiments:

[0057] Embodiment 1

[0058] As Figure 1 shown, the embodiment of the present application proposes an entity linking method based on a large language model, and the method includes:

[0059] Step S1: Use the large language model to enhance the context of the given entity mention in the entity document; in the specific embodiment of the present invention, step S1 can summarize and generalize the description of the entity mention in the original document, reduce the context length and increase the information density; in addition, the world knowledge of the large language model itself can be used to supplement descriptive information for the entity mention.

[0060] Step S2: Generate a candidate entity list for the entity mention; in the specific embodiment of the present invention, the number of entities to be matched in step S2 is in the millions, and the length of the dialogue context of the large language model is limited. Step S2 can effectively filter out most irrelevant entities and only retain a small number of entities that may be linked as candidate entities;

[0061] Step S3: Generate descriptive information for each candidate entity in the candidate entity list; in the specific embodiment of the present invention, in step S3, since the names of the candidate entities are usually very similar, the descriptive information can help the large language model effectively distinguish these entities

[0062] Step S4: Using the enhanced context, the candidate entity list, and the description information of the candidate entities, construct a multiple-choice question and answer task for each entity reference; in a specific embodiment of the present invention, the entity linking task in step S4 is a typical task with complex, detailed, and strict requirements. Although large language models are good at understanding and generating natural language, they may have difficulties in accurately following very specific and complex task instructions. By converting the entity linking task into a multiple-choice question and answer task, it is more in line with the characteristics of text generation by large language models and can effectively improve performance; in addition, by restricting the large language model to only select from the given options, hallucinations can be avoided;

[0063] Step S5: Using the entity reference and the candidate entity list to construct a reference graph to obtain the correlation degree between entity references; in an embodiment of the present invention, the random walk algorithm with restart is used for calculation;

[0064] Step S6: Sort all entity references. The sorting in step S5 refers to estimating the linking difficulty of entity references through a heuristic method. The entity references that are more easily linked are processed earlier in step S7;

[0065] Step S7: According to the sorting result, sequentially input the multiple-choice questions corresponding to the entity references into the large language model for answering. When inputting, select the question and answer pairs of entity references whose correlation degree with the current entity reference exceeds the threshold as the dialogue context to obtain the entity linking result of each entity reference;

[0066] In a specific embodiment of the present invention, when processing the linking of subsequent entity references, according to the correlation degree information obtained in step S5, from the entity references processed previously, select at most N entity reference linking information with the highest correlation degree with the current entity reference, including multiple-choice questions and the answers made by the large model as the dialogue context, thereby indirectly considering the connection between them; due to the sorting in step S5, the error accumulation caused by incorrect previous links is avoided, and because the entities that are more difficult to match are processed later, more context information can be utilized and the accuracy is higher. By restricting the number and correlation degree of the dialogue context, the overhead of the dialogue with the LLM is effectively reduced.

[0067] In an embodiment of the present invention, in the above step S1 of enhancing the context of the given entity references in the document using a large language model, it includes:

[0068] Taking the entity reference and the preset number of characters before and after the entity reference as the context, and completing the context enhancement through the form of asking questions and answering according to the question template by the large language model. The question template includes: the text content and the enhanced context of the entity reference output by the large language model, where the text content includes the context of the entity reference.

[0069] More specifically, in the embodiments of the present invention, a large language model is used to enhance the context of entity referring terms. The context of an entity referring term is the text content near the entity referring term in the document. Specifically, in this embodiment, the entity referring term and 150 characters before and after the entity referring term are used as the context. The context enhancement is completed in the form of asking questions to the large language model, and the large language model here includes but is not limited to Wenxin Yiyan, Tongyi Qianwen, DeepSeek, etc. The specific question template is as follows: Consider the following text content:

[0070] Text content: {Context of the entity referring term}

[0071] Please provide more descriptive information about {entity referring term} based on the above given text. {Entity referring term} must be included in your answer. Answer as follows:

[0072] The result output by the large language model is the entity referring term

[0073] The enhanced context .

[0074] In the embodiments of the present invention, in step S2 of generating a list of candidate entities for the entity referring term, it includes:

[0075] According to the enhanced context, use an entity linking model to filter the entities in the knowledge base, generate a list of multiple entities that are most likely to be linked as the candidate entity list, obtain the candidate entity list of the entity referring term and the matching score of each candidate entity, and take a preset number of candidate entities as the candidate entity list.

[0076] More specifically, in the embodiments of the present invention, a list of candidate entities is generated for the entity referring term. Since the number of entities to be matched for the entity referring term is often in the millions, and the dialogue context length of the large language model is limited, it is necessary to first filter the entities in the knowledge base to generate the most likely to be linked number of entities as the candidate entity list . In this embodiment, the traditional entity linking model BLINK is used for filtering. Specifically, for the entity referring term , the enhanced context obtained in step S1 is input to the BLINK model to obtain the candidate entity list of the entity referring term and the matching score of each candidate entity, and the first k candidate entities are taken as the candidate entity list . In this embodiment, k = 10.

[0077] ​In the embodiments of the present invention, in step S3 of generating description information for each candidate entity in the candidate entity list, it includes:

[0078] The description information in step S3 includes two parts. The first part is the first sentence of the corresponding entity description information in the knowledge base, and this sentence is usually the abstract description of the current entity; the second part is the sentence with the strongest context relevance recalled from the corresponding entity description information in the knowledge base by using the vector recall method.

[0079] The description information includes: the first part of the content is one of the content of the corresponding entity description information in the knowledge base, and one of the content is the abstract description of the current entity; and the second part of the content is the content with the strongest context relevance recalled from the corresponding entity description information in the knowledge base, wherein the recall adopts the vector recall method.

[0080] After splitting the entity document into multiple blocks according to a fixed character length, use the text embedding model to generate embeddings for each block, calculate the cosine similarity between the embedding of the enhanced context and the embeddings of all blocks, and the block with the highest score is used as the second part of the candidate entity description information; concatenate the first part and the second part of the description as the description information of the candidate entity.

[0081] More specifically, in the embodiments of the present invention, description information is generated for each candidate entity in the candidate entity list. Since the names of candidate entities are usually very similar, such as apple, Apple Inc., apple tree, etc., adding description information to the candidate entities can help the large language model to distinguish. The description information includes two parts. Assume that the description document about the entity is , then the first part of the description information is the first sentence of this description document, and this sentence is usually the document abstract, which is the most comprehensive and concise introduction to this entity, that is . The second part is the sentence recalled by the vector recall method that is most similar to the enhanced context obtained in step S1.

[0082] Specifically, split the description document into multiple blocks according to a fixed character length, . Then use the text embedding model to generate embeddings for each block . In order to obtain the description information that best matches the context, calculate the cosine similarity between the embedding of the enhanced context obtained in step S1 and the embeddings of all blocks of this entity document, and the block with the highest score is used as the second part of the candidate entity description information, that is:

[0083]

[0084] Finally, in the embodiments of the present invention, the descriptions of the two parts are spliced as the candidate entity description information . In this embodiment, the character length is selected as 150, the overlapping character length between blocks is 20, and the text embedding model selects Qwen2.5-32B-Instruct.

[0085] In the embodiments of the present invention, in step S4 of constructing a multiple-choice question and answer task for each entity reference item, it includes:

[0086] For the entity reference item, the enhanced context obtained is used as the context of the multiple-choice question, the obtained candidate entity list is used as the options, and the obtained candidate entity description information is used as the description information of each option;

[0087] The template of the multiple-choice question includes: context, question, options, and the last option is none of the above.

[0088] More specifically, in the embodiments of the present invention, a multiple-choice question is constructed for each entity reference item. Since the entity linking task is a typical task with complex, detailed, and strict requirements, and large language models have difficulties in accurately following very specific and complex task instructions and perform poorly. Therefore, converting the entity linking task into a multiple-choice question and answer task is more in line with the characteristics of large language model text generation. Specifically, for the entity reference item , the enhanced context obtained in step S1 is used as the context of the question, the candidate entity list obtained in step S2 is used as the options, and the obtained in step S3 is used as the description information of each option. Particularly, the last option is "none of the above". The specific multiple-choice question template is as follows:

[0089] Context: {Enhanced context of the entity reference item}

[0090] Question: Which of the following entities is the {entity reference item} described above?

[0091] Options:

[0092] Candidate entity 1: Description information of candidate entity 1

[0093] Candidate entity 2: Description information of candidate entity 2

[0094] …

[0095] (k + 1) None of the above.

[0096] In the embodiments of the present invention, a reference graph including entity mentions and candidate entities is constructed, and the association degree between entity mentions is calculated by a random walk algorithm with restart. In step S5, it includes:

[0097] Construct a weighted graph including two types of nodes: entity mentions and candidate entities, where the weight between an entity mention and a candidate entity is the matching score between the candidate entity and the entity mention obtained in step 2. The weight between candidate entities is the semantic association degree between candidate entities. The semantic association degree is obtained by calculating the number of in-link entities shared by two entities in the knowledge base.

[0098] Calculate the association degree between entity mentions by a random walk algorithm with restart. Starting from each entity mention to be disambiguated, explore its propagation path in the reference, and then measure its correlation with other entity mentions.

[0099] Construct a reference graph using all entity mentions and all candidate entities recalled for the entity mentions in step S2. This graph is a weighted graph, where the weight between an entity mention and a candidate entity is the matching score obtained in step S2, and the weight between candidate entities is the semantic association degree between two entities.

[0100] The semantic association degree is obtained by calculating the number of in-link entities shared by two entities in the knowledge base. Its calculation method is:

[0101]

[0102] Among them, a and b represent two candidate entities, A and B are respectively the sets of entities linked to a and b in the knowledge base, |A ∩ B| represents the number of entities linked to both, and |W| is the total number of entities in the knowledge base. If two entities share more in-link entities (i.e., co-occur frequently in the context), their semantic relevance SR(a, b) is higher.

[0103] Calculating the association degree between entity mentions by a random walk algorithm with restart means starting from the entity mention m i and exploring its propagation path in the semantic graph, and then measuring its correlation with other entity mentions. Specifically, first define a normalized transition matrix A according to the graph structure, and the initial distribution p (0) is set to take the value of 1 on the node m i and 0 for the rest. Each step of the random walk is iteratively updated as follows:

[0104]

[0105] Iterate until convergence to obtain the steady-state distribution p * where each dimension p *(v) represents the access probability p of node v relative to the starting point m i of (v). The access probabilities p of all other entity referents m * (v). Take the access probability p of all other entity referents m j as a measure of the semantic relatedness between m * (m j ) and m i and m j . Finally, during the collective entity linking process, the disambiguation decisions within the entire document can be jointly optimized based on the association information between these entity referents.

[0106] In the embodiment of the present invention, in the above-mentioned step S6 of sorting all entity referents, it includes: sorting all entity referents means using a heuristic method to evaluate the linking difficulty of entity referents. The lower the evaluation difficulty of an entity referent, the higher its score, and it will be processed earlier in step S7.

[0107] Use a heuristic method to estimate the linking difficulty of entity referents and sort them. Calculate the matching score of one of the candidate entities of the entity referent as the matching difficulty score of the entity referent, and sort according to the score.

[0108] More specifically, in the embodiment of the present invention, all entity referents are sorted. Since the linking difficulties of different entity referents are different, and even in the same document, the information density and quality of their contexts are also different, a heuristic method is first used to estimate the linking difficulty of entity referents and sort them. In this embodiment, the matching score of the first candidate entity of the entity referent obtained in step S2 is used as the matching difficulty score of the entity referent. The higher the score, the easier it is to process, and they are sorted from high to low according to the score.

[0109] In the embodiment of the present invention, in the above-mentioned step S7 of inputting the multiple-choice questions corresponding to the entity referents into the large language model for answering in sequence according to the sorting result, and selecting the Q&A pairs of entity referents with an association degree exceeding the threshold with the current entity referent as the dialogue context when inputting to obtain the entity linking result of each entity referent, it includes:

[0110] When asking a multiple-choice question about one entity referent, select at most N link information of entity referents with the highest association degree with the current entity referent from the previously processed entity referents, including the multiple-choice questions and the answers made by the large model as the dialogue context; where N is the number of referents used as the dialogue context;

[0111] In a specific embodiment of the present invention, select at most N entity referents from the entity referents that meet two conditions, and use their Q&A pairs as the dialogue context. The two conditions include:

[0112] 1. The entity reference item has completed the linking;

[0113] 2. The relevance of the entity reference item to the currently processed entity reference item exceeds the threshold.

[0114] More specifically, in the embodiments of the present invention, according to the sorting result, the multiple-choice questions corresponding to the entity reference items are sequentially input into the large language model for answering. When inputting, the question-and-answer pairs of the up to N entity reference items with the highest relevance to the current entity reference item are used as the dialogue context to obtain the entity linking result of each entity reference item. Specifically, when asking the multiple-choice question Q i of the i-th entity reference item m i , the question-and-answer pairs of the associated entity reference items are passed to the large language model as the dialogue context; the associated entity reference items refer to the up to N entity reference items with the highest relevance to the current entity reference item, that is, the entity reference items obtained in step S5 with the highest relevance to the current entity reference item and exceeding the threshold δ, and the entity linking has been completed before processing the current entity reference item. The up to N entity reference items are selected in descending order of relevance and denoted as R i , and the question-and-answer pairs of the associated entity reference items are the multiple-choice questions and the answers made by the large model when processing these reference items , when processing the current entity reference item, the input and answer of the large language model are:

[0115]

[0116] Specifically, the so-called according to the sorting result here means that the later the ranking in step S7, that is, the more difficult it is estimated to be processed, the later it is processed in this step, so that the richer the dialogue context information it can utilize and the higher the accuracy can be improved.

[0117] For example: Suppose the entity reference items sorted in ascending order of matching difficulty in the document are m1, m2, m3, m4, m5, and each entity reference item is processed in turn.

[0118] 1. Process m1, input the multiple-choice question for processing m1 into the LLM to get an answer.

[0119] 2. Process m2, judge whether the relevance between m1 and m2 exceeds the threshold. If it exceeds the threshold, when inputting the multiple-choice question for processing m2 into the LLM, input the question-and-answer pair of m1 as the dialogue context together; ...

[0121] 5. Processor m5 selects, from the already processed m1, m2, m3, and m4, at most N entity mentions with the highest correlation with m5 and exceeding the threshold as associated entity mentions, and inputs the Q&A pairs of these associated mentions, together with the multiple-choice questions for processing m5, into the LLM;

[0122] Specifically, the Prompt here refers to the question template, which consists of a task definition, an instruction description, and examples, as follows:

[0123] Task definition:

[0124] You are an assistant highly proficient in natural language processing. You will complete entity linking through multiple-choice. Your task is to select the most matching entity from the given list of candidate entities after the given context. If none match, select "None of the above".

[0125] Instruction description:

[0126] Context: Please carefully analyze the given text.

[0127] Question: Determine the meaning of the entity mention in the sentence based on the context.

[0128] Options: Select the entity that most conforms to the meaning of the entity mention. If none, select "None of the above".

[0129] Answer: First give the option number and the selected entity name, such as (1) Apple. Then give your explanation of why you chose this option rather than the others,

[0130] Example:

[0131] Context: Apple is a company that produces mobile phones, tablets, and computers.

[0132] Question: Which of the following entities is the Apple described in the above text?

[0133] Options:

[0134] (1) Apple: A round, edible fruit.

[0135] (2) Apple (tech company): Apple Inc. is an American multinational technology company.

[0136] (3) None of the above.

[0137] Answer: (2) Apple (tech company). Because the Apple mentioned in the text is a company that can produce high-tech products, not a fruit.

[0138] In the embodiments of the present invention, by having the large model make selections from given options instead of generating text in an open-ended manner, the hallucination problem of the large language model can be effectively avoided. Performing entity linking for the entire document within the same dialogue context can make full use of the association relationships between entity referring terms in the document, improving the accuracy of document-level entity linking.

[0139] As described above, the method of the present invention can be preferably implemented.

[0140] Embodiment Two

[0141] As Figure 2 shown, an entity linking system 200 based on a large language model is provided in an embodiment of the present application. Using the entity linking method based on the large language model as described above, the system 200 includes:

[0142] A context enhancement module 210: for enhancing the context of entity referring terms using the large language model;

[0143] A candidate entity generation module 220: for generating a list of candidate entities for entity referring terms;

[0144] A description information recall module 230: for generating description information for each candidate entity in the list of candidate entities;

[0145] A single-choice question construction module 240: for constructing a single-choice question in a fixed format for each entity referring term using the enhanced context, the candidate entity list, and the description information of the candidate entities;

[0146] An entity referring term correlation degree calculation module 250: constructing a reference graph using the entity referring terms in the document and the list of candidate entities generated for the entity referring terms, and obtaining the correlation degree between entity referring terms through a random walk algorithm with restart;

[0147] A sorting module 260: for sorting all entity referring terms,

[0148] An entity linking module 270: according to the sorting result, sequentially inputting the single-choice questions corresponding to the entity referring terms into the large language model for answering. When inputting, selecting the Q&A pairs of entity referring terms with a correlation degree exceeding the threshold with the current entity referring term as the dialogue context, and obtaining the entity linking result of each entity referring term.

[0149] Embodiment Three

[0150] An embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the steps of the entity linking method based on the large language model are implemented.

[0151] Embodiment Four

[0152] An embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the entity linking method based on a large language model as described are implemented.

[0153] In addition, Figure 1 The entity linking method based on a large language model described in the embodiments of the present application can be implemented by an electronic device, such as a computer device. Figure 3 FIG. is a schematic diagram of the hardware structure of a computer device according to an embodiment of the present application.

[0154] In some of these embodiments, the computer device may further include a communication interface 83 and a bus 80. Among them, as Figure 3 shown, the processor 81, the memory 82, and the communication interface 83 are connected through the bus 80 to complete mutual communication.

[0155] Specifically, the above-mentioned processor 81 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.

[0156] The memory 82 can be used to store or cache various data files required for processing and / or communication, as well as possible computer program instructions executed by the processor 81.

[0157] The processor 81 reads and executes the computer program instructions stored in the memory 82 to implement any one of the entity linking methods based on a large language model in the above embodiments.

[0158] The technical features of the above-mentioned embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0159] The above-mentioned embodiments only represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.

Claims

1. An entity linking method based on a large language model, characterized in that: The method comprises: Using a large language model to enhance the context of an entity reference term given in an entity document, generating a candidate entity list for the entity reference term, and generating description information for each candidate entity in the candidate entity list; Using the enhanced context, the candidate entity list and the description information of the candidate entity, construct a multiple-choice question answering task for each entity reference item; Using the entity referents and the candidate entity list to construct a reference graph, and obtain the association between the entity referents; All entity reference items are sorted, and according to the sorting results, the multiple-choice questions corresponding to the entity reference items are input into the large language model for answering in turn. When inputting, the question-answer pairs of entity reference items whose correlation with the current entity reference item exceeds a threshold are selected as the dialogue context to obtain the entity linking results of each entity reference item.

2. The entity linking method based on a large language model according to claim 1, characterized in that: The step of enhancing the context of the entity reference term given in the entity document by using the large language model includes: The entity reference term and the characters of a preset length before and after the entity reference term are taken as context, and the context enhancement is completed by asking and answering the large language model according to a question template, wherein the question template includes: text content and the enhanced context of the entity reference term output by the large language model, wherein the text content includes the context of the entity reference term.

3. The entity linking method based on a large language model according to claim 1, characterized in that: The step of generating a candidate entity list for the entity referent includes: According to the enhanced context, the entity linking model is used to filter the entities in the knowledge base, and a candidate entity list containing multiple entities that are most likely to be linked is generated to obtain a candidate entity list of the entity referent and a matching score for each candidate entity, and a preset number of the candidate entities are taken as the candidate entity list.

4. The entity linking method based on a large language model according to claim 1, characterized in that: The step of generating description information for each candidate entity in the candidate entity list includes: The description information includes: a first part of the content is one of the contents of the corresponding entity description information in the knowledge base, one of the contents is a summary description of the current entity; and a second part of the content is the content most closely associated with the context recalled from the corresponding entity description information in the knowledge base, wherein the recall adopts a vector recall method; After dividing the entity document into multiple blocks according to a fixed character length, an embedding is generated for each block using a text embedding model, and the cosine similarity between the enhanced embedding of the context and the embeddings of all the blocks is calculated. The block with the highest score is used as the second part of the candidate entity description information; the first part and the second part descriptions are concatenated as the description information of the candidate entity.

5. The entity linking method based on a large language model according to claim 1, characterized in that: The question-answering task step of constructing a single-choice question for each entity reference term includes: For the entity referent, the obtained enhanced context is used as the context of the single-choice question, the obtained candidate entity list is used as options, and the obtained candidate entity description information is used as the description information of each option; The template of the multiple-choice question includes: context, question, options, and the last option is none of the above.

6. The entity linking method based on a large language model according to claim 1, characterized in that: The step of sorting all entity reference items includes: A heuristic method is used to estimate the link difficulty of entity reference terms and to sort them, a matching score of one of the candidate entities of the entity reference term is calculated as the matching difficulty score of the entity reference term, and the entities are sorted according to the scores.

7. The entity linking method based on a large language model according to claim 1, characterized in that: According to the sorting result, the single-choice questions corresponding to the entity referents are input into the large language model in turn for answering, and when inputting, the question-answer pairs of the entity referents whose relevance to the current entity referent exceeds a threshold are selected as the dialogue context, and the step of obtaining the entity linking result of each entity referent includes: When asking a multiple-choice question about one of the entity referents, select at most N entity referent link information with the highest correlation with the current entity referent from previously processed entity referents, including the multiple-choice question and the answer made by the large model as the dialogue context; wherein N is the number of entity referents used as the dialogue context; The question template of the large language model includes: task definition, instruction description and examples, wherein the instruction description includes: context, question, options and answers.

8. An entity linking system based on a large language model, using the entity linking method based on a large language model as claimed in any one of claims 1 to 7, characterized in that: The system comprises: Context enhancement module: used to enhance the context of entity referents using a large language model; Candidate entity generation module: used to generate a candidate entity list for entity referents; Description information recall module: used to generate description information for each candidate entity in the candidate entity list; A single-choice construction module: used to construct a single-choice question of a fixed format for each entity reference item by using the enhanced context, the candidate entity list and the description information of the candidate entity; Entity referent relevance calculation module: constructs a reference graph using the entity referent and the candidate entity list to obtain the relevance between entity referents; Sorting module: used to sort all entity reference items; Entity linking module: According to the sorting results, the multiple-choice questions corresponding to the entity reference items are input into the large language model for answering in turn. When inputting, the question-answer pairs of entity reference items whose correlation with the current entity reference item exceeds the threshold are selected as the dialogue context to obtain the entity linking results of each entity reference item.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the entity linking method based on a large language model described in any one of claims 1 to 7 are implemented.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps of the entity linking method based on a large language model are implemented as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Integrated entity linking method and system based on deep learning

    CN111062214A

  • Entity linking method for Chinese knowledge graph question-answering system

    CN111563149A

  • Large language model enhanced question and answer generation method

    CN118013051A

  • Intelligent question and answer method, system and device, computer equipment and readable storage medium

    CN118377881A

  • Intelligent agent method for complex question and answer reasoning of figure knowledge graph based on large model

    CN119513330A

Cited By

  • Information source retrieval method and system including alias identification and homonymous entity disambiguation, and storage medium

    CN120849457A

  • Information source retrieval methods, systems, and storage media including alias identification and same-name entity disambiguation

    CN120849457B

  • Knowledge graph exercise automatic mounting method and system for intelligent education

    CN121937262A

  • A smart education-oriented knowledge graph exercise automatic mounting method and system

    CN121937262B