Entity disambiguation and forgetting method and system based on large language model
By combining large language models and contrast learning, the problem of feature engineering time-consuming and complex semantic capture in the physical disambiguation of classical machine learning models is solved, and more efficient and accurate physical disambiguation and forgetting capabilities are achieved, which is suitable for a variety of application scenarios.
Patent Information
- Application Number
- CN202411932523.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-26
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2044-12-26
AI Technical Summary
Existing entity disambiguation methods based on classic machine learning models have problems such as feature engineering time-consuming and labor-intensive, difficulty capturing complex semantics and contextual information, and poor performance in the face of complex ambiguity.
The entity disambiguation and forgetting method based on large language models is adopted, and the LLaMA3 model is combined with comparison learning, and the ability to bring different context representations of the same entity closer and different entity representations farther away, enhance the disambiguation ability, and realize the forgetting of specific entity information by dynamically learning new entity characteristics and adapting to new contexts.
It significantly improves the accuracy of entity disambiguation and the robustness of the model, can more effectively distinguish different entities, and realizes effective forgetting of specific entity information. It is suitable for applications such as search engines, knowledge graph construction, question-and-answer systems and virtual assistants.
Smart Images

Figure CN120011534A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of entity disambiguation in natural language processing, and in particular to an entity disambiguation and forgetting method and system based on a large language model. Background Art
[0002] With the explosive growth of Internet information, text data contains a large number of ambiguous entities. Accurate entity disambiguation is crucial for applications such as information retrieval, question-answering systems, and text summarization. Modern applications, such as search engines, knowledge graph construction, question-answering systems, and virtual assistants, all rely on accurate identification and disambiguation of entities to provide users with accurate information. Entity disambiguation has become an urgent need for models to accurately understand and remember entities. However, the disambiguation task is challenging, especially when faced with homonymous entities (such as company names and place names) and polysemous words (the same word has different meanings). In addition, during the entity disambiguation process, the model may process entities containing personal privacy information (such as names, geographic locations, etc.). According to data privacy regulations, users have the right to request the deletion of their personal data. The entity disambiguation and forgetting method based on the large language model provides an efficient solution for identifying and distinguishing entities with multiple meanings in a large amount of text and forgetting specific entity information. It is particularly important in search engines, knowledge graph construction, question-answering systems, and virtual assistants to provide users with the information they need, ensure its accuracy and completeness, and comply with data privacy regulations. It has a very broad application scenario.
[0003] In the past, most of the entity disambiguation methods based on machine learning used machine learning models such as SVM and decision trees to encode entities and their contexts as features and classify them. The pain points of this type of method are: 1. Classic machine learning models usually require a lot of manual feature engineering to extract effective features for training. This is not only time-consuming and labor-intensive, but the design quality of the features directly affects the performance of the model. Manual features are difficult to capture complex semantics and contextual information, which limits the expressive power of the model. 2. Compared with modern large language models, classic machine learning models have obvious deficiencies in understanding and utilizing contextual information, and it is difficult to capture long-distance dependencies and complex contextual relationships. 3. For those entities with high ambiguity, classic models often find it difficult to effectively distinguish their specific meanings. This is because classic models rely on predefined features and lack an understanding of the deep semantics of entities, resulting in poor performance when faced with complex ambiguities. Summary of the invention
[0004] In order to solve the above problems, the purpose of the present invention is to provide an entity disambiguation and forgetting technology based on a large language model, with the LLaMA3 model as the basic architecture, which fundamentally solves the many defects exposed by the previous classic machine learning models (SVM, decision tree), and the present invention combines contrastive learning, the model automatically learns how to bring different contextual representations of the same entity closer and push the representations of different entities farther away, and the disambiguation ability is significantly enhanced. In addition, contrastive learning enhances the model's distinction between these entities by constructing positive and negative sample pairs, can dynamically learn new entity features, adapt to new contexts, and achieve effective forgetting of specific entity information.
[0005] In order to achieve the above technical objectives, the present application provides an entity disambiguation and forgetting method based on a large language model, comprising the following steps:
[0006] Based on the entity disambiguation dataset and the forgetting dataset, after constructing the comparative learning samples, data preprocessing was performed to remove irrelevant information. The vocabulary of the LLaMA3 model was used to segment and encode the text to generate the dataset.
[0007] Based on the LLaMA3 model, improvements are made by adding projection layers and contrastive learning modules. Based on the improved LLaMA3 model, model training is performed according to the dataset to build a large language model. The constructed large language model has the capabilities of entity disambiguation and forgetting.
[0008] Preferably, in the process of constructing contrastive learning samples, the contrastive learning samples include positive samples and negative samples, wherein the positive samples include the text of the target entity, which is encoded using the LLaMA3 model to obtain its representation; the negative samples include two types, type one is the text containing entities that need to be forgotten, which is the object whose similarity the model should reduce, and type two is the text containing other entities that are easily confused with the target entity.
[0009] Preferably, in the process of improving the LLaMA3 model, a projection layer is added after the output layer of the LLaMA3 model to map the high-dimensional text representation to the contrastive learning space, and integrate the modules required for contrastive learning, including contrastive loss calculation and positive and negative sample matching.
[0010] Preferably, during the model training process, for each sample, the cosine similarity is used to calculate the similarity between its feature representation and the positive sample and the negative sample, and the contrast loss is used to measure the effect of the model in distinguishing between positive and negative samples.
[0011] Preferably, in the process of using contrast loss to measure the effect of the model in distinguishing positive and negative samples, the loss function InfoNCE loss of contrastive learning is used to train the model to distinguish positive and negative samples. The formula is as follows:
[0012]
[0013] Among them, z i represents sample i, represents the positive sample of sample i, represents the negative sample of sample i, including the entity to be forgotten, sim(.) represents the similarity function, and τ represents the temperature hyperparameter.
[0014] Preferably, in the process of constructing a large language model, a forgetting weighted loss function is constructed for negative samples of entities that need to be forgotten, which is used to increase the weight of negative samples of entities that need to be forgotten in the loss function, thereby prompting the model to reduce the similarity to these entities.
[0015] Preferably, during the model training process, the model parameters are updated by back propagation based on the loss value.
[0016] The present invention also discloses an entity disambiguation and forgetting system based on a large language model, which is used to implement the above-mentioned entity disambiguation and forgetting method based on a large language model, including:
[0017] The data processing module is used to construct comparative learning samples based on the entity disambiguation dataset and the forgetting dataset, perform data preprocessing to remove irrelevant information, and use the vocabulary of the LLaMA3 model to segment and encode the text to generate a dataset;
[0018] The large language model construction module is used to improve the LLaMA3 model by adding a projection layer and a contrastive learning module, and to train the model based on the improved LLaMA3 model and the data set to build a large language model, so that the constructed large language model has the ability of entity disambiguation and forgetting.
[0019] The present invention discloses the following technical effects:
[0020] (1) The present invention uses the LLaMA3 model as its basic architecture, which fundamentally solves the many defects exposed by previous classic machine learning models (SVM, decision tree), makes the model more efficient in entity disambiguation and forgetting tasks, and improves the accuracy and robustness of the model.
[0021] (2) The present invention adds a projection layer after the output layer of the LLaMA3 model to map the high-dimensional text representation to the contrastive learning space. After that, it is further trained through the contrastive learning module to learn to distinguish different entity representations, enhance the entity disambiguation capability, and improve the model's ability to distinguish different entities.
[0022] (3) The LLaMA3 model in the invention has been pre-trained on a large-scale corpus and has rich language knowledge and context understanding capabilities. During the training process, negative samples of contrastive learning are used to construct the model, which reduces the sensitivity of the model to the entities that need to be forgotten, thereby achieving the purpose of forgetting. By combining the LLaMA3 large language model with contrastive learning, its powerful language understanding and representation capabilities can be fully utilized, the accuracy of entity disambiguation can be improved, and effective forgetting of specific entity information can be achieved. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0024] Figure 1 It is a schematic diagram of the overall framework of the entity disambiguation and forgetting method based on a large language model described in the present invention;
[0025] Figure 2 A schematic diagram of a framework for data preparation according to the present invention;
[0026] Figure 3 It is a schematic diagram of the framework of the LLaMA3 model described in the present invention;
[0027] Figure 4 It is a schematic diagram of the framework of the comparison module described in the present invention. DETAILED DESCRIPTION
[0028] To make the purpose, technical scheme and advantages of the embodiments of the present application clearer, the technical scheme in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all of the embodiments. The components of the embodiments of the present application generally described and shown in the drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the application claimed for protection, but merely represents the selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without making creative work belong to the scope of protection of the present application.
[0029] like Figure 1-4 As shown, the present invention provides an entity disambiguation and forgetting method based on a large language model, the method comprising the following steps:
[0030] S1: Data preparation: Figure 2As shown, the given text data is preprocessed, including determining the entity set, constructing sample data, and annotating the data to finally obtain the standardized text;
[0031] Step S1: Data preparation, specifically including the following sub-steps:
[0032] S11: Determine the entity set
[0033] The target entity set is the entity set that the model needs to correctly identify and disambiguate; the forgotten entity set is the entity set that the model needs to forget, such as outdated, sensitive or private entities;
[0034] S12: Constructing sample data
[0035] It includes positive samples and negative samples. Positive samples are texts containing target entities, which are encoded using the LLaMA3 model to obtain their representations. There are two types of negative samples. Type 1 is texts containing entities that need to be forgotten, which are objects whose similarity should be reduced by the model. Type 2 is texts containing other entities that are easily confused with the target entity, which enhances the model's ability to distinguish.
[0036] S13: Data Annotation
[0037] Use named entity recognition (NER) tools (such as spaCy, BioBERT, ScispaCy) to identify and annotate entities in the text to ensure that the entity information of each sample is accurate. Finally, remove redundant characters, unify the case, and form standardized text.
[0038] S2: Model design: Figure 1 As shown, the standardized text is input into the LLaMA3 model, and after the projection layer and contrastive learning module, the total loss value is output to guide the optimization of model parameters and minimize the contrastive loss and forgetting penalty;
[0039] First, the LLaMA model's tokenizer is used to decompose the text into words or subword units and convert the text into integer representations. Then, the text is passed through the attention layer and encoding layer in sequence, and then through the LLaMA model's multi-layer Transformer encoder to generate context-related feature representations.
[0040] Step S2: Model design, specifically including the following sub-steps:
[0041] S21: LLaMA3 model feature extraction
[0042] Load the LLaMA3 pre-trained model and use its powerful text encoding capabilities to obtain the vector representation of the text from the middle layer or the last layer of the model. The specific steps are as follows:
[0043] Use the LLaMA3 model's tokenizer to decompose the text into words or subword units, encode the tokenization results into input IDs, and finally generate an attention mask, which is converted into an integer representation. Prepare an input dictionary containing input IDs and attention masks, where the input ID marks the corresponding ID sequence, and the attention mask marks the position of the actual content. Map the input ID to an embedding vector to represent the semantic information of the word. The input ID passes through the embedding layer to obtain the initial word vector representation, that is, the embedding representation. Then, the position information is added to the embedding representation, and then passes through the attention layer and the encoding layer in sequence, and through the multi-layer Transformer encoder of the LLaMA model to generate context-related feature representations.
[0044] S22: Add contrastive learning module
[0045] Add a projection layer after the output layer of the LLaMA3 model to map the high-dimensional text representation to the contrastive learning space and integrate the modules required for contrastive learning, including contrastive loss calculation, positive and negative sample matching, etc. The specific steps are as follows:
[0046] Taking the sequence-level feature representation generated in S21 as input, a linear layer (projection layer) is used to map the high-dimensional vector to a low-dimensional representation space, and a nonlinear activation function (such as ReLU) is added to enhance the representation ability to generate a low-dimensional feature representation, that is, a numerical vector suitable for contrastive learning;
[0047] like Figure 4 As shown in the figure, low-dimensional feature representation is used as input, and sample pairs are constructed after the contrastive learning module, including positive sample pairs and negative sample pairs. Then, cosine similarity is used to calculate the similarity. For positive sample pairs, the similarity of positive sample pairs is shortened to enhance the model's recognition ability of the same entity. For negative sample pairs, the similarity of negative sample pairs is extended to enhance the model's ability to distinguish different entities, especially the entities that need to be forgotten. Finally, the loss value and forgetting loss are calculated, and the total loss value is finally obtained to guide the optimization of model parameters and minimize the contrast loss and forgetting penalty.
[0048] S3: Constructing loss function: Multiple functions are used for calculation in S22, including: contrast loss function, forgetting loss function and loss adjustment of forgetting mechanism.
[0049] Step S3 constructs the loss function, the details are as follows:
[0050] S31: Contrastive loss function
[0051] The InfoNCE loss function of contrastive learning is used to train the model to distinguish positive and negative samples. The formula is as follows:
[0052]
[0053] Among them, z i represents sample i, represents the positive sample of sample i, represents the negative sample of sample i, including the entity to be forgotten, sim(.) represents the similarity function, and τ represents the temperature hyperparameter;
[0054] S32: Forgetting loss function
[0055] Specifically targeting negative samples of entities to be forgotten, the model is prompted to reduce its sensitivity to these entities. The formula is as follows:
[0056]
[0057] Among them, w k represents the weight of the negative sample of the entity to be forgotten, Negative sample feature vector representing the entity to be forgotten;
[0058] S33: Loss adjustment of forgetting mechanism
[0059] For negative samples of entities that need to be forgotten, increase their weight in the loss function to encourage the model to reduce the similarity of these entities. The weighted loss formula is as follows:
[0060]
[0061]
[0062] Among them, λ is the forgetting weight coefficient, W k is the weight of the negative sample of the entity to be forgotten, It represents the contrast loss for the entity that needs to be forgotten, and λ represents the forgetting weight coefficient, which controls the strength of forgetting.
[0063] The output of the character-level network, that is, the vector representation of each sentence in the text, is used as input. The output and context vector (that is, the final hidden state) of each time step are obtained through a bidirectional recurrent neural network. The hidden layer dimension is a model hyperparameter and needs to be adjusted according to the specific data set and training process.
[0064] S4: Model training: Based on the loss value, the model parameters are updated through back propagation, so that the model can better pull positive samples closer and push negative samples away. This process is repeated until the model converges. After training, the model has a stronger entity disambiguation ability and achieves forgetting of specified entities.
[0065] Step S4: Model training, specifically including the following sub-steps:
[0066] S41: Training Strategy
[0067] Use appropriate batch size to improve training efficiency and stability, adopt learning rate decay or adaptive adjustment strategy to avoid oscillation during training, and use methods such as Dropout and weight decay to prevent overfitting;
[0068] S42: Backpropagation and Optimization
[0069] Use the optimization algorithm (AdamW) to update the model parameters and minimize the loss function;
[0070] S43: Model Evaluation
[0071] Including entity disambiguation performance evaluation, forgetting effect evaluation and model generalization ability evaluation.
[0072] The present invention applies the LLaMA3 model to the tasks of entity disambiguation and forgetting, and combines the large-scale language model with the entity disambiguation and forgetting method based on contrastive learning, which can give full play to the powerful representation ability of LLMs, improve the accuracy of entity disambiguation, and achieve effective forgetting of specific entity information. The model: 1. First, the initial text is processed and features are extracted using the LLaMA3 model; 2. A projection layer is added after the output layer of the LLaMA3 model to map the high-dimensional text representation to the contrastive learning space and generate a low-dimensional feature representation, that is, a numerical vector suitable for contrastive learning; 3. With the low-dimensional feature representation as input, sample pairs are constructed through the contrastive learning module, including positive sample pairs and negative sample pairs, and then the cosine similarity is used to calculate the similarity. For the positive sample pairs, the similarity of the positive sample pairs is brought closer to enhance the model's recognition ability of the same entity. For the negative sample pairs, the similarity of the negative sample pairs is pushed further away to improve the model's ability to distinguish different entities, especially the entities that need to be forgotten. Finally, the loss value and the forgetting loss are calculated to obtain the total loss value; 4. Finally, based on the loss value, the model parameters are updated through back propagation so that the model can better bring the positive samples closer and push the negative samples further away. This process is repeated until the model converges. After training, the model has a stronger entity disambiguation ability and achieves forgetting of the specified entity.
[0073] The present invention adopts self-supervised learning, does not rely on manually designed features, and can automatically learn the distinguishing features of entities from data; applies contrastive learning to construct positive and negative sample pairs, can dynamically learn new entity features, adapt to new contexts, and use the powerful representation ability of large-scale language models to flexibly respond to long-tail entities and new knowledge updating needs, especially in fine-tuning and precise control of large models, greatly improving efficiency and expanding the limitations of traditional methods.
[0074] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0075] In the description of the present invention, it should be understood that the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features. In the description of the present invention, the meaning of "plurality" is two or more, unless otherwise clearly and specifically defined.
[0076] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention is also intended to include these modifications and variations.
Claims
1. A method for entity disambiguation and forgetting based on a large language model, characterized in that: The following steps are involved: Based on the entity disambiguation dataset and the forgetting dataset, after constructing the comparative learning samples, data preprocessing was performed to remove irrelevant information. The vocabulary of the LLaMA3 model was used to segment and encode the text to generate the dataset. Based on the LLaMA3 model, improvements are made by adding a projection layer and a contrastive learning module, and based on the improved LLaMA3 model, model training is performed according to the data set to construct a large language model, so that the constructed large language model has the capabilities of entity disambiguation and forgetting.
2. The entity disambiguation and forgetting method based on a large language model according to claim 1, characterized in that: In the process of constructing contrastive learning samples, the contrastive learning samples include positive samples and negative samples, wherein the positive samples include the text of the target entity, which is encoded using the LLaMA3 model to obtain its representation; the negative samples include two types, type one is the text containing the entity that needs to be forgotten, as the object whose similarity the model should reduce, and type two is the text containing other entities that are easily confused with the target entity.
3. The entity disambiguation and forgetting method based on a large language model according to claim 2, characterized in that: In the process of improving the LLaMA3 model, a projection layer is added after the output layer of the LLaMA3 model to map the high-dimensional text representation to the contrastive learning space and integrate the modules required for contrastive learning, including contrastive loss calculation and positive and negative sample matching.
4. The entity disambiguation and forgetting method based on a large language model according to claim 3, characterized in that: During the model training process, for each sample, the cosine similarity is used to calculate the similarity of its feature representation with the positive and negative samples, and the contrast loss is used to measure the effectiveness of the model in distinguishing positive and negative samples.
5. The entity disambiguation and forgetting method based on a large language model according to claim 4, characterized in that: In the process of using contrast loss to measure the effect of the model in distinguishing positive and negative samples, the contrastive learning loss function InfoNCE loss is used to train the model to distinguish positive and negative samples. The formula is as follows: Among them, z i represents sample i, represents the positive sample of sample i, represents the negative sample of sample i, including the entity to be forgotten, sim(.) represents the similarity function, and τ represents the temperature hyperparameter.
6. The entity disambiguation and forgetting method based on a large language model according to claim 5, characterized in that: In the process of building a large language model, a forgetting weighted loss function is constructed for the negative samples of entities that need to be forgotten. It is used to increase the weight of the negative samples of entities that need to be forgotten in the loss function, so as to prompt the model to reduce the similarity of these entities.
7. The entity disambiguation and forgetting method based on a large language model according to claim 6, characterized in that: During the model training process, the model parameters are updated through back propagation based on the loss value.
8. An entity disambiguation and forgetting system based on a large language model, used to implement an entity disambiguation and forgetting method based on a large language model as described in any one of claims 1 to 7, characterized in that: include: The data processing module is used to construct comparative learning samples based on the entity disambiguation dataset and the forgetting dataset, perform data preprocessing to remove irrelevant information, and use the vocabulary of the LLaMA3 model to segment and encode the text to generate a dataset; A large language model construction module is used to improve the LLaMA3 model by adding a projection layer and a contrastive learning module, and to perform model training based on the improved LLaMA3 model and the data set to construct a large language model, so that the constructed large language model has the capabilities of entity disambiguation and forgetting.
Citation Information
Patent Citations
Short text entity disambiguation method based on multi-task learning
CN115081445A
Retrieval model training method and device, knowledge question-answering method and device, equipment and medium
CN117290782A
Data enhancement method and device based on multi-modal language alignment
CN119153017A
Cited By
Content detection model training and fine tuning method, content detection method, device, equipment, medium and product
CN120950979A