Large model triple extraction method based on entity large model knowledge injection
By injecting the output of the entity recognition model into the large model and adopting a special training strategy, the hallucination problems and insufficient detailed query capabilities of the large model in long text triple extraction are solved, achieving higher extraction accuracy and domain adaptability.
Patent Information
- Application Number
- CN202411892736.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-20
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2044-12-20
AI Technical Summary
In the prior art, large models have hallucinations problems when extracting long text triplets, and their details query capabilities for longer texts are weak, making it difficult to adapt to different fields and text types.
By injecting the output of the entity recognition model into the large model, the large model's ability in long text processing and detailed query is enhanced, and the generalization ability of the model is improved through specially designed data and three-stage training strategies, reducing the illusion of entity extraction.
It improves the accuracy of triple extraction, enhances the adaptability of the model in different fields and text types, and provides an effective technical means for the construction of knowledge graphs.
Smart Images

Figure CN120067338A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of information extraction, and more specifically, to a method for extracting triples from large models based on knowledge injection of entity large models. Background Art
[0002] In the construction of knowledge graphs, triples usually refer to a combination consisting of two entities and a relationship, expressed as (entity 1, relationship, entity 2). This structure can describe various semantic relationships between entities, such as "Einstein" - "was born in" - "Germany". Triples are the basic units for storing and expressing knowledge in knowledge graphs. Traditional methods for triple extraction using deep learning identify entities in the input text by using transformer series models or variants, and then predict the relationship based on the entity positions. This method relies on a large amount of labeled data and has limited generalization ability, making it difficult to adapt to new domains. With the rise of large models, the method of using large models for triple extraction effectively utilizes pre-trained knowledge, reduces the dependence on labeled data, and improves generalization. However, the method of using large models for triple extraction has the following difficulties and challenges:
[0003] 1. Improving the ability to detect entities in long texts: It is necessary to solve the hallucination problem that occurs when large models extract triples from long texts, that is, incorrectly extracting non-existent entities or missing the recognition of entities, and incorrectly generating non-existent entity relationships. This requires enhancing the entity detection ability of large models through knowledge injection from external entity models. One of the most critical problems faced by common comment filtering and analysis techniques when processing social media texts is the lack of context understanding. Social media comments usually contain rich emotional colors and complex contextual information, as well as phenomena such as ellipsis, typos, and creative spellings that often appear in comments, which all increase the difficulty of text understanding. The techniques in the background art often rely on standardized text inputs and clear grammatical structures, and the non-standard nature of social media texts makes these techniques difficult to adapt. Therefore, the lack of context understanding is a major challenge faced by the above techniques when processing social media comments.
[0004] 2. Integrating entity recognition models: Integrating entity recognition models into large models with different strategies to improve the performance of triple extraction tasks is a technical challenge. It is necessary to find effective methods to enable large models to utilize external models to improve the overall performance.
[0005] 3. Models with strong domain adaptability: Traditional entity recognition models rely on domain data. Therefore, it is difficult to design a model that can adapt to different domains and reduce error propagation. Many analysis techniques may perform well on specific types of data or within specific domains, but their performance may decline sharply when faced with new or diverse data. For example, a model trained on specific product reviews may have difficulty adapting to the analysis of reviews for other products. The diversity and complexity of social media reviews require analysis techniques to have good generalization ability and be able to adapt to texts in different languages, topics, and contexts. However, many existing techniques often focus on specific types of texts or specific domains and lack sufficient generalization ability. This limits the application of these techniques in different scenarios and also increases the difficulty of technology deployment and maintenance. At the same time, considering the imperfection of entity models, it is necessary to design appropriate training tasks to improve the accuracy and generalization ability of the model through adversarial training.
[0006] 4. Text diversity and data annotation: The diversity and complexity of text data, as well as different text types, make the triple extraction task challenging. In addition, creating a high-quality annotated dataset is a time-consuming and laborious task. Especially for texts in specific domains or languages, the lack of sufficient scale and quality of training data will limit the performance and generalization ability of the model. Moreover, the rapid change of social media texts also means that the annotated data may quickly become outdated and need to be continuously updated and maintained. The quality problem of annotated data will also affect the accuracy of the analysis results. If the annotated data is biased or incorrect, then the trained model will also inherit these biases, resulting in inaccurate analysis results. In addition, for comments on emerging topics or special contexts, it may be difficult to find sufficient annotated data for effective training. Therefore, the dependence on high-quality annotated data is another major challenge faced by existing technologies. Domain adaptability is also a challenge in triple extraction. The model may perform poorly in unseen domains, so it is necessary to solve the technical problem of effectively extracting triples from text data to construct a knowledge graph. Summary of the Invention
[0007] The present invention provides a large model triple extraction method based on entity large model knowledge injection, which solves the technical problems of hallucinations in existing large models and weak detail query ability for long texts.
[0008] To solve the above technical problems, the technical solution of the present invention is as follows:
[0009] The present invention provides a large model triple extraction method based on entity large model knowledge injection, including the following steps:
[0010] Obtain the input text and task prompt words of the triple to be extracted;
[0011] Input the input text into a pre-trained large entity recognition model to obtain the entity information of the input text, where the entity information includes the entities and entity positions in the input text;
[0012] Input the input text, task prompt words, and entity information into a pre-trained large triple extraction model to obtain the triple extraction result of the input text.
[0013] The present invention splits the triple extraction task, performs entity recognition in a large entity recognition model, and then provides the entity position information to the large triple extraction model. With prior knowledge, the large triple extraction model can pay more attention to judging entity types and relationships between entities, while retaining a certain ability to distinguish entity information. The subsequent extraction accuracy will be higher, reducing hallucinations, and solving the technical problems of hallucinations in large models and weak detail query capabilities for relatively long texts in the prior art.
[0014] Further, for the pre-trained large entity recognition model and the pre-trained large triple extraction model, the training process includes:
[0015] Construct a first training set, where the first training set includes input text, triple extraction prompt words, and entity information;
[0016] Use the first training set to train the large triple extraction model to obtain a pre-trained large triple extraction model;
[0017] Construct a second training set, where the second training set includes input text and entity extraction prompt words;
[0018] Use the second training set to train the entity recognition model to obtain a preliminarily trained entity recognition model;
[0019] Connect the output of the preliminarily trained entity recognition model to the input of the pre-trained large triple extraction model, and freeze the parameters of the pre-trained large triple extraction model;
[0020] Construct a third training set, where the third training set includes input text, triple extraction prompt words, and entity information, and the data in the third training set is different from that in the first training set;
[0021] Use the third training set to train the connection of the output of the preliminarily trained entity recognition model and the pre-trained large triple extraction model to obtain a pre-trained entity recognition model.
[0022] Further, the following processing is performed on the entity information in the first training set:
[0023] Retain 50% of the entity information;
[0024] Replace 30% of the entity information with errors, including the replacement of entity location information and the replacement of entities;
[0025] Completely empty 20% of the entity information.
[0026] Furthermore, train the triple extraction large model using the first training set to obtain a pre-trained triple extraction large model, including:
[0027] Introduce the first LoRa module into the preset first large model and add a text embedding layer to obtain a triple extraction large model;
[0028] Directly input the input text and triple extraction prompt words in the first training set into the backbone network of the preset large model, and convert the entity information in the first training set into a text embedding representation through the text embedding layer and then input it into the backbone network of the preset large model for training. Among them, when training, freeze the parameters of the backbone network of the preset large model and only fine-tune the parameters of the text embedding layer and the LoRa module;
[0029] During the training process, first use the first loss function to calculate the error between the predicted value and the true value under the current parameters of the triple extraction large model, and then calculate the gradient of the first loss function with respect to each parameter through the backpropagation algorithm;
[0030] According to the calculated gradient and the preset learning rate, use the gradient descent or its variant optimization algorithm to update the parameters of the triple extraction large model until the first loss function of the triple extraction large model is less than the preset value or reaches the preset number of iterations to obtain a pre-trained triple extraction large model.
[0031] Furthermore, the first loss function includes:
[0032] Text generation cross-entropy loss L triple_generation :
[0033]
[0034] In the formula, N represents the total number of samples, y ij represents the true probability that the triple extraction large model generates the j-th word at the i-th position, and P ij represents the predicted probability that the triple extraction large model generates the j-th word at the i-th position;
[0035] Entity location output loss L triple_position :
[0036]
[0037] In the formula, m represents the number of entities, Denotes the true position of the first character of the i-th entity, Denotes the true position of the last character of the i-th entity, Denotes the predicted position of the first character of the i-th entity, Denotes the predicted position of the last character of the i-th entity, w 1 and w 2 Denotes the weight parameter;
[0038] Entity number loss L triple_quantity :
[0039] L triple_quantity = log(1 + |P - T|)
[0040] Where P is the predicted number of entities and T is the true number of entities;
[0041] The first loss function L triplr_total is:
[0042]
[0043] Where λ 1 、λ 2 and λ 3 are hyperparameters.
[0044] Furthermore, using the second training set to train the entity recognition large model, a preliminarily trained entity recognition large model is obtained, including:
[0045] Adopting a preset second large model as the entity recognition large model;
[0046] During the training process, all parameters of the model will be updated during the fine-tuning process. First, use the second loss function to calculate the error between the predicted value and the true value under the current model parameters, and then calculate the gradient of the second loss function with respect to each parameter through the backpropagation algorithm.
[0047] Parameter update: According to the calculated gradient and the preset learning rate, use the gradient descent or its variant optimization algorithm to update the parameters of the model until the second loss function of the entity recognition large model is less than the preset value or reaches the preset number of iterations, and a pre-trained entity recognition large model is obtained.
[0048] Furthermore, the number of parameters of the preset second large model is less than the number of parameters of the preset first large model.
[0049] Furthermore, the second loss function L ner_generation , includes:
[0050]
[0051] Further, training the output of the preliminarily trained large entity recognition model after connection with the pre-trained large triple extraction model using the third training set to obtain a pre-trained large entity recognition model, including:
[0052] Introduce a second LoRa module into the preliminarily trained large entity recognition model, and only update the parameters of the second LoRa module during the training process;
[0053] Process the input text in the third training set through the preliminarily trained large entity recognition model to generate entity recognition results;
[0054] Use the generated entity recognition results, the input text in the third training set, and the triple extraction prompt words in the third training set as the input of the pre-trained large triple extraction model to obtain the output of the pre-trained large triple extraction model;
[0055] Calculate the third loss function according to the output of the pre-trained large triple extraction model;
[0056] Update the parameters of the second LoRa module according to the gradient of the third loss function.
[0057] Further, the third loss function includes:
[0058]
[0059] In the formula, L ner_quantity represents the generation loss.
[0060] Compared with the prior art, the beneficial effects of the technical solution of the present invention are:
[0061] The present invention proposes a method for enhancing the triple extraction performance of a large model through entity model knowledge injection. By injecting the output of the entity recognition model into the large model, the present invention enhances the ability of the large model in long text processing and detail query. At the same time, through specially designed data and three-stage training strategies, the generalization ability of the model is improved, and the hallucination of entity extraction is reduced. This method not only improves the accuracy of triple extraction, but also enhances the adaptability of the model in different fields and text types, providing an effective technical means for the construction of knowledge graphs. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] Figure 1 It is a schematic flowchart of a method for large model triple extraction based on entity large model knowledge injection provided by an embodiment of the present invention;
[0063] Figure 2 It is a schematic diagram of training a large triple extraction model using the first training set provided by an embodiment of the present invention;
[0064] Figure 3 Schematic diagram of using the third training set to train the output of the preliminarily trained large entity recognition model after connection and the pre-trained large triple extraction model provided by the embodiment of the present invention;
[0065] Figure 4 Schematic diagram of the enhanced large model triple extraction process provided by the embodiment of the present invention. Detailed implementation manners
[0066] The drawings are only for illustrative purposes and should not be construed as a limitation of this patent;
[0067] To better illustrate this embodiment, some components in the drawings are omitted, enlarged or reduced, and do not represent the size of the actual product;
[0068] For those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.
[0069] The technical solutions of the present invention will be further described below with reference to the drawings and embodiments.
[0070] Embodiment 1
[0071] The present invention provides a large model triple extraction method based on entity large model knowledge injection, as Figure 1 shown, including the following steps:
[0072] Obtain the input text of the triple to be extracted and the task prompt word;
[0073] Input the input text into the pre-trained large entity recognition model to obtain the entity information of the input text, where the entity information includes the entities and entity positions of the input text;
[0074] Input the input text, task prompt word and entity information into the pre-trained large triple extraction model to obtain the triple extraction result of the input text.
[0075] The present invention splits the triple extraction task, performs entity recognition in the large entity recognition model, and then provides the entity position information to the large triple extraction model. With the prior knowledge, the large triple extraction model can pay more attention to judging entity types and the relationships between entities, while retaining a certain ability to distinguish entity information. The subsequent extraction accuracy will be higher, reducing hallucinations and solving the technical problems of hallucinations in large models and weak detail query ability for long texts in the prior art.
[0076] Embodiment 2
[0077] On the basis of Embodiment 1, this embodiment further discloses the following content:
[0078] The training process of the pre-trained large entity recognition model and the pre-trained large triple extraction model includes:
[0079] Construct a first training set, which includes input text, triple extraction prompt words, and entity information;
[0080] Use the first training set to train the large triple extraction model to obtain a pre-trained large triple extraction model;
[0081] Construct a second training set, which includes input text and entity extraction prompt words;
[0082] Use the second training set to train the large entity recognition model to obtain a preliminarily trained large entity recognition model;
[0083] Connect the output of the preliminarily trained large entity recognition model with the input of the pre-trained large triple extraction model, and freeze the parameters of the pre-trained large triple extraction model;
[0084] Construct a third training set, which includes input text, triple extraction prompt words, and entity information, and the data in the third training set is different from the data in the first training set;
[0085] Use the third training set to train the connection between the output of the preliminarily trained large entity recognition model and the pre-trained large triple extraction model to obtain a pre-trained large entity recognition model.
[0086] In a specific embodiment, the entire training process includes three stages:
[0087] The training objective of the first stage is to enable the large triple extraction model to learn to use the text index information output by the large entity recognition model for triple extraction. After the large entity recognition model is trained, it can output the position and text segment of the entity in the text, but the large triple extraction model cannot directly use this information. Therefore, the model architecture needs to be transformed and corresponding tasks need to be set for strengthening;
[0088] The training objective of the second stage is to train a large entity recognition model for outputting the position and text segment of the entity in the text for subsequent entity recognition tasks;
[0089] The training objective of the third stage is to assemble the two large models, freeze the parameters of the large triple extraction model, only fine-tune the parameters of the large entity recognition model, and use the capabilities of the large-parameter large model to enhance the capabilities of the parameters of the large entity recognition model, further reducing the hallucination problem in the process of generating small models.
[0090] Embodiment 3
[0091] On the basis of Embodiment 2, the following content is further disclosed in this embodiment:
[0092] Further, the entity information in the first training set is processed as follows:
[0093] Retain 50% of the entity information and retain the correct entity position information during annotation;
[0094] Perform error replacement on 30% of the entity information, including the replacement of entity position information and entities. For example:
[0095] Replacement of position information:
[0096] Input: "end_index":16,"start_index":1,"text":"Cartier"
[0097] Replace with: "end_index":18,"start_index":4,"text":"Cartier"
[0098] It is the replacement of entity content, replaced with data not in the text:
[0099] Input: "end_index":16,"start_index":1,"text":"Diya"
[0100] Completely empty 20% of the entity information.
[0101] In this embodiment, because for the entity recognition large model, even a trained entity recognition large model cannot guarantee 100% accuracy in the position and text segment of the output entity in the text. To avoid the triple extraction large model being overly dependent on entity position information and improve the generalization and fault tolerance of the model, this embodiment transforms the triple extraction training set with multi-domain manual annotations, modifies and replaces the entity position information in an appropriate proportion. Through the above data setting method, when there are errors in the input position information or generated entities, the model has a certain degree of fault tolerance, and to a certain extent, reduces the 100% dependence of the model on entity position information.
[0102] Further, use the first training set to train the triple extraction large model to obtain a pre-trained triple extraction large model, including:
[0103] Introduce the first LoRa module into the preset first large model and add a text embedding layer to obtain the triple extraction large model;
[0104] Input text and triple extraction prompt words in the first training set are directly input into the backbone network of a preset large model, and entity information in the first training set is converted into a text embedding representation through the text embedding layer and then input into the backbone network of the preset large model for training. During training, the parameters of the backbone network of the preset large model are frozen, and only the parameters of the text embedding layer and the LoRa module are fine-tuned;
[0105] During the training process, first use the first loss function to calculate the error between the predicted value and the true value under the current triple extraction large model parameters, and then calculate the gradient of the first loss function with respect to each parameter through the backpropagation algorithm;
[0106] According to the calculated gradient and the preset learning rate, use the gradient descent or its variant optimization algorithm to update the parameters of the triple extraction large model until the first loss function of the triple extraction large model is less than the preset value or reaches the preset number of iterations, and obtain the pre-trained triple extraction large model.
[0107] In a specific embodiment, during the training of the first stage, the model is as Figure 2 shown, and the model is combined according to the model architecture shown in Figure 2 . Denote the large model with large parameters for triple extraction as LLM_triple. Based on LLM_triple (taking llama3 - 70b as an example), introduce the LoRa (Low-Rank Adaptation) module for fine-tuning;
[0108] Figure 2 The function of the text embedding layer in
[0109] For each piece of training data, the json string at the input 2 entity positions is converted into an embedding representation that the model can understand through the embedding layer. This step converts the index information of the entity in the text into a numerical form.
[0110] In this embodiment, the format of the first training set is as follows:
[0111] The first training set includes input text, triple extraction prompt words, and entity information:
[0112] Input text: "The original distance from Andy Lau is only one Cartier. Hahaha #This operation is stable# Andy Lau# Watch";
[0113] Entity information: "[[{"end_index":16,"start_index":13,"text":"Cartier"},{"end_index":32,"start_index":29,"text":"Andy Lau"},{"end_index":36,"start_index":34,"text":"watch"}]]";
[0114] Prompt:
[0115] ## Role
[0116] You are an expert in analyzing social media knowledge graphs. You are very good at extracting and identifying specified entities from text and judging entity relationships. Please strictly extract the entities in the text according to the task requirements and judge the relationships between the entities:
[0117] ## Entity Definition
[0118] 1. <Person Name>: A person name is a name used to identify and distinguish individuals. In different cultures and languages, a person name may include a surname (family name), a given name, a middle name, a nickname, a title, etc. A person name is usually used to refer to a specific person, which can be a historical figure, a public figure, a fictional character, or any individual.
[0119] 2. <Organization Name>: Refers to a collective composed of individuals, which can be a company, a government agency, a non-profit organization, a school, etc. ...
[0121] ## Relationship Definition:
[0122] 1. <Parent>: Commonly appears in the text as "father / mother of...", "as the parent of...", "one of the parents of...", etc.
[0123] 2. <Child>: Commonly appears in the text as "son / daughter of...", "as the child of...", "one of the descendants of...", etc. ...
[0125] Annotation Principle:
[0126] 1. **Entity Extraction**: Carefully read the input text and extract all entities that meet the definition according to the "Entity Definition".
[0127] 2. **Entity Relationship Classification**: Carefully analyze the context where each entity is located, extract entity pairs that meet the "Entity Relationship Definition", and output the relationship type between the entities.
[0128] 3. **Format Requirements**: The triple should include a subject (person, organization, or object), a relation, and an object (person, place, or object).
[0129] ## Article to be Extracted
[0130] {input_text}
[0131] Output Format:
[0132] Please output strictly according to the following format. If there is no entity, the entity result should be an empty list. If there is no entity relation, the triple result should be an empty list. The specific answer format is as follows:
[0133] Entity Result: [{"end_index": end index, "start_index": start index, "text": "Andy Lau"},... {"end_index": end index, "start_index": start index, "text": "Cartier"}]
[0134] Triple Result: [{"ent1": "Entity 1", "relation": "Relation 1", "ent2": "Entity 2"},... {"ent1": "Entity 3", "relation": "Relation 2", "ent2": "Entity 4"}]
[0135] Note: {input_text} will be filled with the input text. The ellipsis in the definition represents other incomplete definitions. This embodiment only shows an example.
[0136] For the above input, the large model outputs the following string:
[0137] Entities: [{"end_index": 16, "start_index": 13, "text": "Cartier"}, {"end_index": 32, "start_index": 29, "text": "Andy Lau"}, {"end_index": 36, "start_index": 34, "text": "watch"}]
[0138] Entity Relations: [{"ent1": "watch", "relation": "brand", "ent2": "Cartier"}, {"ent1": "Cartier", "relation": "spokesperson", "ent2": "Andy Lau"}]
[0139] In a specific embodiment, a generation task method is used to complete the triple extraction task, that is, input text + triple extraction prompt words;
[0140] The input text and the triple extraction prompt words are constructed into triple prompt words and directly input into the LLM_triple backbone network, so that LLM_triple can receive both the original text features and entity position features at the same time. Then, the model is fine-tuned using the LoRa of the training set of triples.
[0141] During training, the parameters of the LLM_tripled backbone network are frozen (that is, the parameters of the model are not adjusted during the training process to avoid losing the capabilities learned by the model during pre-training), and only the parameters of the text embedding layer and the large model LoRa layer are fine-tuned during the fine-tuning training.
[0142] Design of the loss function:
[0143] Text generation cross-entropy loss L triple_generation : Focus on the difference between the text generated by the model and the target text. For batch data, the cross-entropy loss is usually the average of the losses of each sample. N is the total number of data in each batch, and y ij represents the true probability that the triple extraction large model generates the j-th word at the i-th position, and P ij represents the predicted probability that the triple extraction large model generates the j-th word at the i-th position:
[0144]
[0145] In the formula, N represents the total number of samples, and y ij represents the true probability that the triple extraction large model generates the j-th word at the i-th position, and P ij represents the predicted probability that the triple extraction large model generates the j-th word at the i-th position;
[0146] Entity Position Output Loss L triple_position : This embodiment requires a loss function to measure the accuracy of the entity position index. This can be achieved by calculating the difference between the entity start and end indices predicted by the model and the true indices:
[0147]
[0148] In the formula, m represents the number of entities, represents the true position of the first character of the i-th entity, represents the true position of the last character of the i-th entity, represents the predicted position of the first character of the i-th entity, represents the predicted position of the last character of the i-th entity, w 1 and w 2Represents a weight parameter, a weight parameter used to balance the importance of the start position and the end position, which can be adjusted according to the actual requirements of the production environment. In this embodiment, w 1 = 0.6, w 2 = 0.4, because determining the start index of an entity is more important than determining the end index of an entity. This total loss function will accumulate the position losses of all entities, thereby providing a measure of the overall position prediction performance for the model;
[0149] Entity quantity loss L triple_quantity : Calculate the error by taking the logarithm of the ratio of the predicted value to the true value, which is applicable to the case where both the predicted value and the true value are positive numbers:
[0150] L triple_quantity = log(1 + |P - T|)
[0151] In the formula, P is the predicted entity quantity, and T is the true entity quantity;
[0152] The total loss function combines the above loss functions to form a total loss function for optimizing entity recognition, position prediction, and entity quantity alignment simultaneously. The first loss function L triple_total can be defined as::
[0153]
[0154] In the formula, λ 1 、λ 2 and λ 3 are hyperparameters, hyperparameters used to balance the importance of different loss terms. In this embodiment, the values are λ 1 = 0.5, λ 2 = 0.3, λ 3 = 0.2, and the specific values can be appropriately adjusted according to the production task. This loss function can be used for optimization during the training process. By minimizing this loss function, the model will pay more attention to the position index of the entity and the entity quantity in addition to the accuracy of text generation.
[0155] Furthermore, training the entity recognition large model using the second training set to obtain a preliminarily trained entity recognition large model, including:
[0156] Adopting a preset second large model as the entity recognition large model;
[0157] During the training process, all parameters of the model will be updated during the fine-tuning process. First, use the second loss function to calculate the error between the predicted value and the true value under the current model parameters, and then calculate the gradient of the second loss function with respect to each parameter through the backpropagation algorithm.
[0158] Parameter Update: According to the calculated gradient and the preset learning rate, use gradient descent or its variant optimization algorithm to update the parameters of the model until the second loss function of the entity recognition large model is less than the preset value or reaches the preset number of iterations, and obtain the pre-trained entity recognition large model.
[0159] In a specific embodiment, the second training set includes input text and entity extraction prompt words:
[0160] Input text: "The only thing between me and Andy Lau is just a Cartier watch. Hahaha #This move is steady# Andy Lau# Watch"
[0161] Entity extraction prompt words:
[0162] ## Role
[0163] You are an expert in the field of social media speech analysis. You are very good at reading social media texts and extracting useful information from them. Next, I will give you a piece of input text. Please extract the entities related to the beauty industry according to the following requirements.
[0164] Task Explanation:
[0165] Entity recognition, namely named entity recognition, requires extracting entities of specified dimensions from the original text;
[0166] Dimension: That is, the category of the entity. An entity can belong to multiple categories at the same time;
[0167] There are N types of dimension categories, which are: brand, category, product name...
[0168] Dimension Definition:
[0169] 1. <Brand>: Brands of beauty (cosmetics, skin care, beauty instruments, perfumes, supporting tools), daily chemicals (household chemicals), etc., including Chinese and English names, such as: Yves Saint Laurent, MAC;
[0170] 2. <Category>: Categories of beauty (cosmetics, skin care, daily chemicals (household chemicals)), etc., such as: lipstick, facial mask, hair removal device, etc.; ...
[0172] Annotation Principle:
[0173] 1. **Precise Extraction**: When there are no entities / phrases that meet the given dimensions in the article, do not extract any other entities;
[0174] 2. **Complete Extraction**: Entities of certain dimensions may be a phrase rather than a simple noun; ...
[0176] Input Text
[0177] {input_text}
[0178] Output format:
[0179] Please output strictly according to the following format. If there is no entity, the entity result will be an empty list. The specific answer format is as follows:
[0180] Entity result: [{"end_index": end index, "start_index": start index, "text": "entity", "type": "entity type"},...{"end_index": end index, "start_index": start index, "text": "entity", "type": "entity type"}].
[0181] The corresponding output of the large model:
[0182] [{"end_index": 6, "start_index": 3, "text": "Andy Lau", "type": "person's name"}, {"end_index": 16, "start_index": 13, "text": "Cartier", "type": "brand"}, {"end_index": 36, "start_index": 34, "text": "watch", "type": "category"}]
[0183] In the above output of the large model, "text" represents the entity, "start_index" represents the position of the first character of the entity in the sentence, "end_index" is the index of the last character of the entity in the sentence, and "type" represents the type of the entity.
[0184] Furthermore, the number of parameters of the preset second large model is less than that of the preset first large model.
[0185] In a specific embodiment, a multi-domain entity recognition training set is used in advance to fine-tune a large model with small parameters, denoted as LLM_ner. The model can be selected from large models with parameters below 10b such as gemma-2b, minicpm-4b, qwen2.5-1.5b, etc. (b represents tens of billions of parameters). The purpose of selecting a large model with small parameters is that the inference speed of the model is proportional to the number of parameters of the model. Because there are requirements for speed in the production environment, and the large model with small parameters only needs to perform an entity recognition task and does not require a model with such complex parameters.
[0186] Fine-tuning strategy selection: Adopt the SFT (Full-parameter fine-tuning) strategy, where all parameters of the model will be updated during fine-tuning. The loss function used is the same as the previous triple training task:
[0187]
[0188] Furthermore, use the third training set to train the output of the preliminarily trained large entity recognition model after connection and the pre-trained large triple extraction model to obtain a pre-trained large entity recognition model, including:
[0189] Introduce a second LoRa module into the preliminarily trained large entity recognition model, and only update the parameters of the second LoRa module during training;
[0190] Process the input text in the third training set through the preliminarily trained large entity recognition model to generate entity recognition results;
[0191] Use the generated entity recognition results, the input text in the third training set, and the triple extraction prompt words in the third training set as the input of the pre-trained large triple extraction model to obtain the output of the pre-trained large triple extraction model;
[0192] Calculate the third loss function according to the output of the pre-trained large triple extraction model;
[0193] Update the parameters of the second LoRa module according to the gradient of the third loss function.
[0194] In a specific embodiment, combine the models according to the Figure 3 shown model architecture. The overall model architecture is divided into three parts: a large entity recognition model, a text embedding layer, and a large triple extraction model. At this time, the parts that need to freeze parameters include the fine-tuned large model LLM_triple backbone parameters, as well as the parameters of the text embedding layer and the LLM_ner backbone parameters, and the LoRa module of LLM_triple. The parts that need parameter update include the parameters of the LoRa layer for LLM_ner.
[0195] Reuse the loss of the previous triple training + the generation loss of the large entity recognition model LLM_ner:
[0196]
[0197] In the formula, L ner_quantity represents the generation loss.
[0198] After completing the above steps, the entity recognition model and the triple extraction model will work better together to improve the performance of the overall system.
[0199] After being trained with data from multiple fields, the inference of subsequent new data, such as Figure 4 shown, each piece of data will first pass through LLM_ner to obtain an entity index string, then pass through the text embedding layer to convert it into entity position information that can be understood by the large model, and then be combined with the task prompt words and input into the triple extraction large model to obtain the triple result. The specific process is as follows. Since the triple extraction model can read entity index information, the entity recognition model can be replaced with any entity recognition model in other fields, which has a certain generality.
[0200] The same or similar reference numerals correspond to the same or similar components;
[0201] The terms used to describe the positional relationship in the drawings are for illustrative purposes only and should not be construed as a limitation of this patent;
[0202] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, and are not intended to limit the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all implementation manners here. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the claims of the present invention.
Claims
1. A method for extracting large model triples based on entity large model knowledge injection, characterized in that: The following steps are involved: Obtain the input text and task prompt words of the triples to be extracted; Input the input text into a pre-trained entity recognition model to obtain entity information of the input text, wherein the entity information includes the entity and entity position of the input text; The input text, task prompt words and entity information are input into a pre-trained triple extraction large model to obtain a triple extraction result of the input text.
2. The method for extracting large model triples based on entity large model knowledge injection according to claim 1 is characterized in that: The training process of the pre-trained entity recognition large model and the pre-trained triple extraction large model includes: Constructing a first training set, wherein the first training set includes input text, triple extraction prompt words and entity information; Using the first training set to train a large triple extraction model to obtain a pre-trained large triple extraction model; Constructing a second training set, wherein the second training set includes input text and entity extraction prompt words; Using the second training set to train the entity recognition large model to obtain a preliminarily trained entity recognition large model; Connecting the output of the preliminarily trained entity recognition large model to the input of the pre-trained triple extraction large model, and freezing the parameters of the pre-trained triple extraction large model; Constructing a third training set, wherein the third training set includes input text, triple extraction prompt words and entity information, and the data of the third training set is different from the data of the first training set; The output of the preliminarily trained entity recognition large model after the third training set training connection is used with the pre-trained triple extraction large model to obtain a pre-trained entity recognition large model.
3. The method for extracting large model triples based on entity large model knowledge injection according to claim 2 is characterized in that: The entity information of the first training set is processed as follows: Keep 50% of the entity information; 30% of the entity information is incorrectly replaced, including the replacement of entity location information and the replacement of entities; Leave 20% of the entity information completely blank.
4. The method for extracting large model triples based on entity large model knowledge injection according to claim 3 is characterized in that: The first training set is used to train a large triple extraction model to obtain a pre-trained large triple extraction model, including: Introduce the first LoRa module into the preset first large model, and add a text embedding layer to obtain a triple extraction large model; The input text and triple extraction prompt words in the first training set are directly input into the backbone network of the preset large model, and the entity information in the first training set is converted into a text embedding representation through the text embedding layer and then input into the backbone network of the preset large model for training, wherein the parameters of the backbone network of the preset large model are frozen during training, and only the parameters of the text embedding layer and the LoRa module are fine-tuned; During the training process, the first loss function is first used to calculate the error between the predicted value and the true value under the parameters of the current triple extraction large model, and then the gradient of the first loss function with respect to each parameter is calculated through the back propagation algorithm; According to the calculated gradient and the preset learning rate, the parameters of the triple extraction large model are updated using the gradient descent or its variant optimization algorithm until the first loss function of the triple extraction large model is less than the preset value or reaches the preset number of iterations to obtain the pre-trained triple extraction large model.
5. The method for extracting large model triples based on entity large model knowledge injection according to claim 4 is characterized in that: The first loss function includes: Cross entropy loss L for text generation triple_generation : In the formula, N represents the total number of samples, y ij represents the true probability of the triple extraction model generating the jth word at the ith position, P ij represents the predicted probability of the triple extraction model generating the jth word at the ith position; Entity position output loss L triple_position : In the formula, m represents the number of entities, represents the actual position of the first word of the i-th entity, Indicates the actual position of the last word of the i-th entity, represents the predicted position of the first word of the i-th entity, represents the predicted position of the last word of the i-th entity, w1 and w2 represent weight parameters; Entity loss L triple_quantity : L triple_quantity =log(1+|P-T|) Where P is the predicted number of entities and T is the actual number of entities; The first loss function L triple_total for: Where λ1, λ2 and λ3 are hyperparameters.
6. The method for extracting large model triples based on entity large model knowledge injection according to claim 4 is characterized in that: The entity recognition large model is trained using the second training set to obtain a preliminarily trained entity recognition large model, including: The preset second largest model is used as the entity recognition large model; During the training process, all parameters of the model will be updated in the fine-tuning process. First, the second loss function is used to calculate the error between the predicted value and the true value under the current model parameters, and then the gradient of the second loss function with respect to each parameter is calculated through the back-propagation algorithm. Parameter update: According to the calculated gradient and the preset learning rate, the model parameters are updated using the gradient descent or its variant optimization algorithm until the loss function of the entity recognition model is less than the preset value or reaches the preset number of iterations, thus obtaining the pre-trained entity recognition model.
7. The method for extracting large model triples based on entity large model knowledge injection according to claim 6 is characterized in that: The parameter quantity of the preset second large model is smaller than the parameter quantity of the preset first large model.
8. The method for extracting large model triples based on entity large model knowledge injection according to claim 6 is characterized in that: The second loss function L ner_generation ,include:
9. The method for extracting large model triples based on entity large model knowledge injection according to any one of claims 6 to 8, characterized in that: The output of the preliminarily trained entity recognition large model after the third training set training connection and the pre-trained triple extraction large model are used to obtain a pre-trained entity recognition large model, including: Introduce a second LoRa module into the preliminarily trained entity recognition model, and only update the parameters of the second LoRa module during the training process; Process the input text in the third training set through the preliminarily trained entity recognition large model to generate entity recognition results; Using the generated entity recognition result, the input text in the third training set, and the triple extraction prompt word in the third training set as inputs of the pre-trained triple extraction large model to obtain the output of the pre-trained triple extraction large model; Calculate the third loss function based on the output of the pre-trained triple extraction large model; Update the parameters of the second LoRa module according to the third loss function gradient.
10. The method for extracting large model triples based on entity large model knowledge injection according to claim 9, characterized in that: The third loss function includes: Where, L ner_quantity Represents generation loss.
Citation Information
Patent Citations
End-to-end speech recognition model compression method based on cross distillation
CN116072107A
Knowledge graph construction method based on IE-Triple
CN117252258A
BIM knowledge graph construction method and system based on label confusion learning
CN117473102A
Entity relation joint extraction method and device, equipment and medium
CN117634463A
Ship domain knowledge graph construction method and device based on large model, and medium
CN117828102A
Cited By
Poem entity extraction model training method, poem entity extraction method and poem entity extraction equipment
CN120277219A