Knowledge graph construction method based on large language model

By pre-training and fine-tuning the large language model, combining the target industry data, dynamically update the knowledge graph, the problem of insufficient understanding of the large language model in specific fields is solved, and efficient and accurate knowledge graph construction and maintenance is achieved to adapt to the rapidly changing knowledge environment.

CN120372016AInactive Publication Date: 2025-07-25BEIJING HUAYUAN TECH CO LTD

Patent Information

Application Number
CN202410118810.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-29
Publication Date
2025-07-25
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The limited understanding of the large language model in local, specific domain knowledge or specific word usage situations results in high cost of building and updating knowledge graphs and lagging information, making it difficult to adapt to rapidly changing knowledge environments.

Method used

By pre-training and fine-tuning the large language model, combining the text data of the target industry, setting up guidance models and analysis rules, dynamically updating the knowledge graph, realizing the extraction and storage of entity relationships, using lightweight deep convolution to capture local feature information, using similarity calculations for entity fusion, and dynamically maintaining the timeliness of the knowledge graph.

Benefits of technology

It improves the semantic understanding and relationship extraction capabilities of the knowledge graph, enhances the adaptability and universality of the model, reduces the construction and maintenance costs, and maintains the accuracy and timeliness of the knowledge graph.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120372016A_ABST
    Figure CN120372016A_ABST
Patent Text Reader

Abstract

The invention discloses a knowledge graph construction method based on a large language model, and relates to the field of natural language processing and deep learning, and the method comprises the steps: obtaining large-scale text data of a target industry, evaluating the calculation resources needed by the large language model, and pre-training the large language model; performing fine adjustment on the large language model by using small-batch and high-quality target industry text data; evaluating the accuracy of the large language model; determining entity relationship extraction requirements, setting a guide model, and extracting entity categories, relationship categories and attribute information; analyzing an extraction result output by the large language model, converting the extraction result into a triple structure, performing entity fusion on extracted entity categories, storing triple data, and generating a knowledge graph of the target industry; and dynamically updating and maintaining the large language model and the knowledge graph. The method has high universality, can adapt to multiple industries, and can keep the timeliness and accuracy of the knowledge graph.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical fields of natural language processing and deep learning, and particularly relates to a method for constructing a knowledge graph based on a large language model. Background Art

[0002] Large language models can learn knowledge from large-scale corpora and achieve state-of-the-art performance in various natural language processing tasks such as entity recognition and relation extraction. Large language models can consider the context information on both the left and right sides of a word simultaneously, thus understanding the context and meaning of the word more comprehensively and having better capabilities in understanding context information. However, the understanding of large language models for local, specific domain knowledge or specific word usage scenarios may be relatively limited.

[0003] Knowledge graphs represent knowledge in a structured manner and play an important role in practical applications such as question answering and recommendation. However, it is difficult for knowledge graphs to include all possible entities, relations, and attributes, and this incompleteness easily leads to the lack of important information, restricting the utility of knowledge graphs in various applications. The Chinese invention patent with the publication number CN117033608A discloses a knowledge graph generative question answering method and system based on a large language model, including: constructing fine-tuning training data for the large language model, where the training data includes prompt statements, question sets, and answer sets; among them, the prompt statements include prompt templates and instance data; fine-tuning the large language model based on LoRA; providing a question answering knowledge base for the large language model fine-tuned by LoRA through a subgraph retrieval strategy; using the large language model fine-tuned by LoRA as a question answering inference model, inputting the question text into the question answering inference model, and the question answering inference model generates question answers based on the provided question answering knowledge base. Although this technical solution constructs the graph information and the question together as model prompt statements to generate question answers, thereby ensuring that the answers are more accurate and traceable. However, fine-tuning using LoRA requires a large amount of computing resources and time, and this optimization method will increase the complexity and time cost of training. At the same time, the knowledge graph cannot reflect the information changes in the real world in a timely manner, resulting in the lag of information update in the knowledge graph and no longer being accurate and effective. The construction and maintenance of the knowledge graph usually require a large amount of human resources to manually organize, annotate, and update a large amount of knowledge, with high costs.

[0004] In summary, although large language models have broad generality, their in-depth understanding and adaptability for specific domains are limited, and it is difficult to capture local, specific words, and specific domain information; accurately extracting entities and their relationships in complex texts is crucial for constructing high-quality knowledge graphs, but it remains a challenge. Summary of the Invention

[0005] This application aims to solve at least one of the technical problems in the related art to some extent. For this purpose, one object of this application is to propose a method, system, electronic device, and readable storage medium for constructing a knowledge graph based on a large language model. By making full use of advanced natural language processing technologies, a more intelligent and efficient solution for knowledge management and retrieval is provided. Through continuous learning and updating, this application can adapt to the ever-changing knowledge environment and provide strong support for applications in various industries.

[0006] In the first aspect disclosed by this application, a method for constructing a knowledge graph based on a large language model is provided. The method includes:

[0007] Obtain a large amount of text data in the target industry, clean and label the large amount of text data, evaluate the required computing resources according to the size and parameter settings of the large language model, and use the large amount of text data to pre-train the large language model;

[0008] Load the weights of the pre-trained large language model, and use a small amount of text data in the target industry to fine-tune the large language model;

[0009] Evaluate the accuracy of the large language model;

[0010] Determine the requirements for entity relationship extraction, set up a guiding model, define rules for extracting entity categories, relationship categories, and attribute information from the large amount of text data, and extract entity categories, relationship categories, and attribute information according to the rules;

[0011] Formulate parsing rules, parse the extraction results output by the large language model, convert the extraction results into a triple structure, perform entity fusion on the extracted entity categories, store the triple data, and generate a knowledge graph of the target industry;

[0012] Dynamically update and maintain the large language model and the knowledge graph.

[0013] The large language model includes:

[0014] The large language model is designed based on the Transformer architecture and is pre-trained in an unsupervised learning manner in the large amount of text data in the target industry to learn the general representation of the language. The overall structure of the large language model includes: input features, multi-head attention, lightweight depth convolution, residual connection, feed-forward neural network, lightweight depth convolution, residual connection, and output. The lightweight depth convolution captures local, specific word, and specific domain information.

[0015] The step of loading the weights of the pre-trained large language model and using a small amount of text data in the target industry to fine-tune the large language model includes:

[0016] Define the problems adapted to the target industry, set clear goals, create an annotation dataset related to the target industry, ensure coverage of various contexts within the target industry, and divide the dataset into a training set, a validation set, and a test set;

[0017] Load the weights of the pre-trained large language model;

[0018] Fine-tune the large language model using a small-scale text data of the target industry;

[0019] Define specific tasks adapted to the target industry, add an adaptation layer to the large language model, match the output of the target industry tasks, and adjust the learning rate for the fine-tuning task;

[0020] Solidify the parameters of the fine-tuned large language model and the intermediate parameters during the fine-tuning process, retain the feature information learned by the large language model, and save the network structure of the large language model and the dimension sizes of the input and output.

[0021] The steps for evaluating the accuracy of the large language model include:

[0022] Evaluate the performance of the large language model using the validation set, and monitor the key indicators of Accuracy, Precision, Recall, and F1 score of the large language model;

[0023] According to the validation results, perform hyperparameter tuning. The hyperparameters include the learning rate, batch size, and fine-tuning layer;

[0024] Record the parameter values of the adjusted parameters and the corresponding metrics;

[0025] Use the test set to verify the performance of the large language model in the real scenario.

[0026] The requirements for determining entity relationship extraction include:

[0027] The entity categories, relationship categories, and attribute information to be extracted. The relationship category is a single relationship, a multiple relationship, or the relationship between entities. Clearly define each entity category and relationship category, and establish a clear entity-relationship-attribute model.

[0028] The steps for setting up the guiding model and defining the rules for extracting entity categories, relationship categories, and attribute information from large-scale text data include:

[0029] Clarify the tasks to be completed by the large language model;

[0030] Determine the extraction rules. Entities must be in the same sentence and there must be specific relationship words between entities;

[0031] Define the extraction scope of entities and relationships;

[0032] Specify the output format of the large language model.

[0033] The steps of formulating parsing rules, parsing the extraction results output by the large language model, converting the extraction results into a triple structure, performing entity fusion on the extracted entity categories, storing the triple data, and generating a knowledge graph for the target industry include:

[0034] Define the triple structure, including the types of entities and relationships;

[0035] Formulate parsing rules, parse the extraction results output by the large language model according to the parsing rules, and convert the extraction results into a triple structure;

[0036] Perform entity fusion on the extracted entity categories;

[0037] Store the triple data and import it into the Neo4j database;

[0038] Generate a knowledge graph for the target industry.

[0039] The entity fusion of the extracted entity categories includes:

[0040] Perform entity fusion using similarity calculation. Entities with similarity higher than the specified threshold are fused in terms of relationships and attributes; analyze the extracted relationships, use the trained large language model to infer the existing relationships for the unknown triples in the test set, and the model outputs the probability or confidence score of the existence of the relationship, indicating the degree of certainty of the large language model for each prediction; the similarity calculation method is as follows:

[0041]

[0042] Among them, Slimilarity represents similarity, q represents the query, c represents the content, w represents the word in q, z k represents the k-th strategy, k is a positive integer p(wz k ) is the probability of the word appearing under the condition of the strategy in z k , and p(z k c) is the probability of the strategy z k appearing under the condition of c.

[0043] The steps of dynamically updating and maintaining the large language model and the knowledge graph include:

[0044] Regularly introduce new knowledge graph data;

[0045] Update the large language model and the knowledge graph through online learning to maintain the consistency and timeliness of the knowledge graph with the actual knowledge;

[0046] Regularly monitor the health status of the knowledge graph;

[0047] Clean up outdated or incorrect information.

[0048] The second aspect disclosed in this application provides a knowledge graph construction system based on a large language model, and the system includes:

[0049] A data processing and pre-training module, which is used to obtain large-scale text data of the target industry, clean and annotate the large-scale text data, evaluate the required computing resources according to the size and parameter settings of the large language model, and pre-train the large language model using the large-scale text data;

[0050] A large language model fine-tuning and target industry adaptation module, which is used to load the pre-trained large language model weights and fine-tune the large language model using small-scale text data of the target industry;

[0051] An accuracy evaluation module for the large language model, which is used to evaluate the accuracy of the large language model;

[0052] An entity recognition and relationship extraction module, which is used to determine the requirements for entity relationship extraction, set a guiding model, define rules for extracting entity categories, relationship categories, and attribute information from large-scale text data, and extract entity categories, relationship categories, and attribute information according to the rules;

[0053] A knowledge graph construction module, which is used to formulate parsing rules, parse the extraction results output by the large language model, convert the extraction results into a triple structure, perform entity fusion on the extracted entity categories, store the triple data, and generate a knowledge graph of the target industry;

[0054] A dynamic update and maintenance module, which is used to dynamically update and maintain the large language model and the knowledge graph.

[0055] The third aspect disclosed in this application provides an electronic device, which includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps in a knowledge graph construction method based on a large language model.

[0056] The fourth aspect disclosed in this application provides a readable storage medium, which stores a computer program. The computer program is suitable for being loaded by a processor to execute the steps in the knowledge graph construction method based on a large language model.

[0057] Compared with the prior art, for a knowledge graph construction method based on a large language model proposed in this application, the advantages of this application are:

[0058] 1. This application has strong semantic understanding and relation extraction capabilities. By pre-training on large-scale text data, the large language model can learn rich language representations and semantic information, so it has strong semantic understanding and relation extraction capabilities.

[0059] 2. This application has strong generality and adaptability. The large language model covers a large amount of general language knowledge during pre-training, so it has strong generality and can adapt to multiple fields and tasks; it can be fine-tuned on data from different fields to make the model more adaptable to the context and entity relationships of a specific field, thus improving the adaptability of the model.

[0060] 3. This application can be iteratively updated and continuously learned by continuously introducing new data and online learning, so as to maintain the timeliness and accuracy of the knowledge graph. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] Figure 1 is a schematic flowchart of a method for constructing a knowledge graph based on a large language model provided by this application;

[0062] Figure 2 is a schematic flowchart of the knowledge graph construction process in the method for constructing a knowledge graph based on a large language model provided by this application;

[0063] Figure 3 is a schematic diagram of the overall architecture of the large language model provided by this application;

[0064] Figure 4 is a schematic diagram of a system for constructing a knowledge graph based on a large language model provided by this application;

[0065] Figure 5 is a schematic diagram of the structure of an electronic device provided by this application;

[0066] Figure 6 is a schematic diagram of the structure of a computer-readable storage medium provided by this application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0067] To better understand this application, more detailed descriptions of various aspects of this application will be made with reference to the accompanying drawings. It should be understood that these detailed descriptions are only descriptions of the exemplary embodiments of this application and do not limit the scope of this application in any way. Throughout the specification, the same reference numerals refer to the same elements. The expression "and / or" includes any and all combinations of one or more of the associated listed items.

[0068] As used herein, terms such as "substantially", "about" and similar terms are used as terms indicating approximation, rather than terms indicating degree, and are intended to account for the inherent deviations in measured or calculated values that would be recognized by a person of ordinary skill in the art. Additionally, in this application, the order of description of each step process does not necessarily represent the order in which these processes occur in actual operation, unless otherwise clearly specified or derivable from the context.

[0069] It should also be understood that expressions such as "comprising", "including", "having", "containing" and / or "including" are open-ended rather than closed-ended expressions in this specification, which means that there are the stated features, elements and / or components, but do not exclude the existence of one or more other features, elements, components and / or their combinations. In addition, when an expression such as "at least one of..." appears after a list of listed features, it modifies the entire list of features, rather than just a single element in the list. In addition, when describing the embodiments of this application, the use of "may" means "one or more embodiments of this application". And the term "exemplary" is intended to refer to an example or illustration.

[0070] Unless otherwise defined, all terms used herein (including engineering terms and scientific and technical terms) have the same meaning as the ordinary understanding of a person of ordinary skill in the art to which this application pertains. It should also be understood that unless clearly stated in this application, words defined in a common dictionary should be interpreted as having a meaning consistent with their meaning in the context of the relevant art, and should not be interpreted in an idealized or overly formal sense.

[0071] It should be noted that, without conflict, the embodiments in this application and the features in the embodiments may be combined with each other. The following will describe this application in detail with reference to the drawings and in combination with the embodiments.

[0072] Embodiment 1

[0073] Figure 1 This application provides a schematic flow diagram of a method for constructing a knowledge graph based on a large language model. As Figure 1 shown, a method for constructing a knowledge graph based on a large language model includes:

[0074] S101: Obtain a large amount of text data in the target industry, clean and label the large amount of text data, evaluate the required computing resources according to the size and parameter settings of the large language model, and pre-train the large language model using the large amount of text data.

[0075] When training large language models, it is necessary to prepare training data, computing resources, and network structures for the target industry; the training data needs to be large in quantity and high in quality, and an industry-specific vocabulary should be added to the training data to make the large language model more sensitive to the proprietary names of the target industry in subsequent knowledge graph extraction tasks; the evaluation method of computing resources needs to be measured according to the size and parameter settings of the large language model.

[0076] (1) Obtain large-scale text data for the target industry: The large-scale text data includes structured and unstructured data; open-source data can be used for training data construction, which has the advantages of high data quality and clear industry division, and data suitable for the target industry can be selectively integrated; add industry-related proprietary vocabulary, including industry terms, entity names, etc., to ensure that the model has a better understanding of the context of a specific industry; add multi-turn dialogue data to improve the multi-turn reasoning ability of the large language model; conduct data quality evaluation to ensure the accuracy, integrity, and diversity of the data; use statistical tools to understand the distribution, characteristics, and defects of the data.

[0077] (2) Clean and annotate the large-scale text data: Perform preprocessing operations such as cleaning, deduplication, and annotation on the large-scale text data to improve the quality and usability of the data.

[0078] (3) Evaluate the required computing resources according to the size and parameter settings of the large language model: Large language models require a large amount of computing resources for training, including GPU or TPU acceleration hardware devices; if large-scale data or large models need to be processed, consider adopting a distributed training strategy and using multiple computers or parallel computing resources to accelerate the training process.

[0079] (4) Pre-train the large language model: Pre-train the large language model on the large-scale text data. Specifically, add lightweight deep convolution to the large language model to capture local feature information of the large-scale text data of the target industry to obtain general language expression ability.

[0080] Such as Figure 3As shown, the large language model is pre-trained in an unsupervised manner on large-scale text data in a specific industry based on the Transformer architecture to learn the general representation of language. Compared with the Transformer, the large language model achieves bidirectional context understanding. During the training process, the large language model can consider the context information in both the left and right directions simultaneously, thus better capturing the semantic relationships between words and sentences. The overall structure of the large language model includes: input, multi-head attention, lightweight depth convolution, residual network, feed-forward neural network, lightweight depth convolution, residual network, and output. The lightweight depth convolution is used to capture local feature information and save computing resources. The input includes the addition of word representations, segment representations, and position representations of large-scale text data in a specific industry. The output hidden representation is used to identify entities in the text. The identified entities can be nodes for constructing a knowledge graph. Using the representation output by the large language model, the relationships between entities can be extracted from the text. The hidden representation output by the large language model contains context information, semantic information, local, specific word, and specific domain information. Understanding the relevance between entities, semantic relationships, and local information of words in the text helps to construct a knowledge graph more accurately.

[0081] S102: Load the weights of the pre-trained large language model, and fine-tune the large language model using small batches of high-quality text data in the target industry.

[0082] Use small batches of high-quality target industry data to train the large language model, so that the large language model learns the morphological and syntactic features in the target industry, enabling better understanding ability in subsequent extraction tasks, thereby improving the accuracy. While fine-tuning the model, clearly define the entity categories, relationship categories, and attribute information to be extracted in the knowledge graph.

[0083] (1) Define the problems adapted to the target industry and set clear goals: To perform supervised fine-tuning, create an annotated dataset related to the target industry to ensure coverage of various contexts within the target industry. Divide the dataset into a training set, a validation set, and a test set for model training, tuning, and evaluation.

[0084] (2) Load the weights of the pre-trained large language model and freeze some layers if necessary to retain general features.

[0085] (3) Define the specific tasks adapted to the target industry, add an adaptation layer to the large language model to match the output of the target industry tasks: For the fine-tuning task, adjust the learning rate, use a smaller learning rate to ensure stable convergence, select an appropriate loss function according to the nature of the task, and use cross-entropy loss in entity recognition. Use the fine-tuning dataset to train the model, apply the backpropagation algorithm, and select a suitable optimizer.

[0086] (4) Large language model preservation: Solidify the parameters of the fine-tuned large language model and the intermediate parameters during the fine-tuning process, retain the feature information learned by the large language model, save the network structure of the large language model and the dimension sizes of the input and output, and ensure normal prediction during reproduction.

[0087] S103: Evaluate the accuracy of the large language model.

[0088] Restore the network structure of the large language model, perform an overall reproduction of the large language model, load the validation data and use the large language model for prediction, evaluate the performance of the large language model, and adjust the parameters of the large language model based on the test results.

[0089] (1) Use the validation set to evaluate the model performance, detect key metrics such as precision, recall, and F1 score; according to the validation results, perform hyperparameter tuning, and hyperparameters include learning rate, batch size, and selection of fine-tuning layers, etc. Each time the parameters are adjusted, record the parameter values and the corresponding metrics.

[0090] The use of the validation set to evaluate the performance of the large language model, monitor the key metrics of Accuracy, Precision, Recall, and F1 score of the large language model, including:

[0091] Accuracy: Accuracy is the accuracy, an evaluation metric for classification tasks, indicating the proportion of samples correctly classified by the model in the total number of samples. The formula is:

[0092]

[0093] Among them, TP represents the true positive, the number of samples correctly predicted as the positive class by the large language model; TN represents the true negative, the number of samples correctly predicted as the negative class by the large language model; FP represents the false positive, the number of samples incorrectly predicted as the positive class by the large language model; FN represents the false negative, the number of samples incorrectly predicted as the negative class by the large language model.

[0094] Precision: Precision is the precision, indicating the proportion of samples actually being the positive class among the samples classified as the positive class. The formula is:

[0095]

[0096] Among them, TP represents the true positive, the number of samples correctly predicted as the positive class by the large language model; FP represents the false positive, the number of samples incorrectly predicted as the positive class by the large language model.

[0097] Recall: Recall is the recall rate, indicating the proportion of samples actually being the positive class that are correctly classified as the positive class. The formula is:

[0098]

[0099] Among them, TP represents true positive examples, which are the number of samples correctly predicted as positive classes by the large language model; FN represents false negative examples, which are the number of samples wrongly predicted as negative classes by the large language model.

[0100] F1 score: Considering both precision and recall and balancing the relationship between them, the formula is:

[0101]

[0102] Among them, Precision is precision and Recall is recall.

[0103] (2) Use an independent test set to verify the performance of the model in a real scenario to ensure that the large language model has good generalization ability on unseen data.

[0104] S104: Determine the requirements for entity relationship extraction, set up a guiding model, define the rules for extracting entity categories, relationship categories, and attribute information from large-scale text data, and extract entity categories, relationship categories, and attribute information according to the rules.

[0105] Set up a guiding template for entity relationship extraction. The guiding template needs to include key information such as the category, quantity, extraction rules, extraction requirements, and scope of the extraction elements to ensure that the large language model can accurately understand the user's intention.

[0106] (1) Determine the goals of the entity relationship extraction task: the entity categories, relationship categories, and attribute information to be extracted, determine the specific requirements for extraction, whether it is a single relationship, multiple relationships, or the relationships between entities; clearly define each entity category and relationship category, and establish a clear entity-relationship-attribute model;

[0107] (2) Set up a guiding template: Clearly define the task that the large language model needs to complete. The task description should not be ambiguous or contradictory. Define the rules for extracting entities, relationships, and attributes from text data, including keywords and context information; specify the extraction requirements, that entities must be in the same sentence and there must be specific relationship words between entities; define the extraction scope of entities and relationships, and perform extraction within a paragraph, document, or other specific scope;

[0108] (3) Deploy the trained large language model to actual applications: Monitor the performance of the large language model, and continuously iterate and improve the guiding template and the model according to the feedback in actual use to adapt to the changes in the field and user needs.

[0109] S105: Develop parsing rules to parse the extraction results output by the large language model, convert the extraction results into a triple structure, perform entity fusion on the extracted entity categories, store the triple data, and generate a knowledge graph for the target industry.

[0110] As Figure 2 shown, parse the results extracted by the large language model according to the rules and convert them into a graph-based data structure. The graph-based data structure is a triple structure. Select an appropriate data import method for data storage according to the number of triple data. The graph database uses Neo4j. During the data storage process, entities need to be fused and relationships completed to ensure the integrity of the knowledge graph.

[0111] (1) Define the triple structure, including the types of entities and relationships: The triple data structure includes a subject entity, an object entity, and a relationship.

[0112] (2) Develop rules: Develop parsing rules to parse the extraction results output by the large language model according to the parsing rules, convert the extraction results into a triple structure, and develop entity fusion rules.

[0113] (3) According to the entity fusion rules, fuse the extracted entities: Use similarity calculation for entity fusion. Entities with similarity higher than the specified threshold are fused in terms of relationships and attributes. Analyze the extracted relationships and use the trained large language model to infer the existing relationships for the unknown triples in the test set. The model outputs the probability or confidence score of the existence of the relationship, indicating the degree of certainty of the large language model for each prediction. The similarity calculation formula is:

[0114]

[0115] where Slimilarity represents similarity, q represents the query, c represents the content, w represents the word in q, z k represents the kth strategy, k is a positive integer, p(wz k ) is the probability of the word appearing under the condition of the strategy z k , and p(z k c) is the probability of the strategy z k appearing under the condition of c.

[0116] (4) Store the triple data and import it into the Neo4j database for use: Select an appropriate import method according to the size of the data volume; for small-scale data, directly import it using the Cypher language of Neo4j; for large-scale data, use the batch import tool or graph database migration tool of Neo4j; use the selected import method to import the data into the Neo4j database; ensure that the imported data format meets the requirements of Neo4j, and check whether the nodes, relationships, and attributes in the knowledge graph are correctly defined;

[0117] (5) Generate the knowledge graph of the target industry.

[0118] S106: Dynamically update and maintain the large language model and the knowledge graph.

[0119] Regularly introduce new knowledge graph data, and update the large language model and the knowledge graph through online learning to maintain their consistency and timeliness with actual knowledge; set up a monitoring mechanism to regularly check the health status of the knowledge graph and the addition of entity relationships; timely clean up the redundant information in the knowledge graph. When the structure of the knowledge graph changes, it is necessary to select the data to be retained and the data to be deleted to maintain the accuracy and integrity of the knowledge graph.

[0120] (1) Regularly monitor the changes or updates of the data source, and collect the newly introduced knowledge graph data regularly through API integration;

[0121] (2) The new knowledge graph data contains information related to the existing entities in the existing knowledge graph. Identify and fuse these new entities, use entity linking technology to identify different expressions of the same entity, and merge them into one entity; update the relationships between the existing entities, and complete the possibly missing relationships; the new data reveals new associations between the existing entities, and integrate these relationships into the knowledge graph.

[0122] Embodiment 2

[0123] Figure 4 is a schematic diagram of a knowledge graph construction system based on a large language model provided by an embodiment of the present application. As Figure 4 shown, a knowledge graph construction system based on a large language model, the system includes:

[0124] A data processing and pre-training module, configured to obtain a large amount of text data of the target industry, clean and annotate the large amount of text data, evaluate the required computing resources according to the size and parameters of the large language model, and pre-train the large language model using the large amount of text data;

[0125] A large language model fine-tuning and target industry adaptation module, configured to load the pre-trained large language model weights and fine-tune the large language model using a small amount of text data of the target industry;

[0126] An accuracy evaluation module for a large language model, which is used to evaluate the accuracy of the large language model;

[0127] An entity recognition and relationship extraction module, which is used to determine the requirements for entity relationship extraction, set a guiding model, define rules for extracting entity categories, relationship categories, and attribute information from large-scale text data, and extract entity categories, relationship categories, and attribute information according to the rules;

[0128] A knowledge graph construction module, which is used to formulate parsing rules, parse the extraction results output by the large language model, convert the extraction results into a triple structure, perform entity fusion on the extracted entity categories, store the triple data, and generate a knowledge graph for the target industry;

[0129] A dynamic update and maintenance module, which is used to dynamically update and maintain the large language model and the knowledge graph.

[0130] Embodiment 3

[0131] Figure 5 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. As Figure 5 shown, according to another aspect of the present application, an electronic device 500 is also provided. The electronic device 500 may include one or more processors and one or more memories. Among them, computer-readable code is stored in the memory, and when the computer-readable code is run by one or more processors, it can execute a knowledge graph construction method based on a large language model.

[0132] The method or system according to the embodiment of the present application can also be implemented by means of Figure 5 the architecture of the electronic device shown. As Figure 5As shown, the electronic device 500 may include a bus 501, one or more CPUs 502, a read-only memory (ROM) 503, a random access memory (RAM) 504, a communication port 505 connected to a network, an input / output component 506, a hard disk 507, etc. The storage device in the electronic device 500, such as the ROM 503 or the hard disk 507, may store a method for constructing a knowledge graph based on a large language model provided in this application. A method for constructing a knowledge graph based on a large language model may, for example, include: obtaining a large amount of text data in a target industry, cleaning and annotating the large amount of text data, evaluating the required computing resources according to the size and parameter settings of the large language model, and pre-training the large language model using the large amount of text data; loading the pre-trained large language model weights, and fine-tuning the large language model using a small amount of text data in the target industry; evaluating the accuracy of the large language model; determining the requirements for entity relationship extraction, setting up a guiding model, defining rules for extracting entity categories, relationship categories, and attribute information from the large amount of text data, and extracting entity categories, relationship categories, and attribute information according to the rules; formulating parsing rules, parsing the extraction results output by the large language model, converting the extraction results into a triple structure, performing entity fusion on the extracted entity categories, storing the triple data, and generating a knowledge graph for the target industry; dynamically updating and maintaining the large language model and the knowledge graph. Further, the electronic device 500 may also include a user interface 508. Of course, Figure 5 the architecture shown is only exemplary, and when implementing different devices, one or more components in the electronic device shown may be omitted according to actual needs. Figure 5

[0133] Embodiment 4

[0134] Figure 6 is a schematic diagram of the structure of a computer-readable storage medium provided by an embodiment of this application. As Figure 6 shown, it is a computer-readable storage medium 600 according to an embodiment of this application. Computer-readable instructions are stored on the computer-readable storage medium 600. When the computer-readable instructions are run by a processor, a method for constructing a knowledge graph based on a large language model according to an embodiment of this application described with reference to the above drawings can be executed. The storage medium 600 includes, but is not limited to, for example, volatile memory and / or non-volatile memory. Volatile memory may, for example, include random access memory (RAM) and cache memory, etc. Non-volatile memory may, for example, include read-only memory (ROM), hard disk, flash memory, etc.

[0135] It should be understood that the methods, apparatuses, and devices of the present application can be implemented in many ways. For example, the methods, apparatuses, and devices of the present application can be implemented by software, hardware, firmware, or any combination of software, hardware, and firmware. The above order of steps for the method is only for illustration, and the steps of the method of the present application are not limited to the specific order described above, unless otherwise specifically stated. In addition, in some embodiments, the present application can also be implemented as a program recorded in a recording medium, and these programs include machine-readable instructions for implementing the method according to the present application. Therefore, the present application also covers a recording medium storing a program for executing the method according to the present application.

[0136] In addition, parts of the above technical solutions provided in the embodiments of the present application that are consistent with the implementation principles of the corresponding technical solutions in the prior art are not described in detail to avoid unnecessary repetition.

[0137] As described above, the specific embodiments have further elaborated on the purpose, technical solutions, and beneficial effects of the present invention. It should be understood that the above description is only the specific embodiments of the present invention and is not used to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for constructing a knowledge graph based on a large language model, characterized in that, Including the following steps: Obtain a large amount of text data in the target industry, clean and annotate the large amount of text data, evaluate the required computing resources according to the size and parameter settings of the large language model, and use the large amount of text data to pre-train the large language model; Load the pre-trained large language model weights, and use a small amount of text data in the target industry to fine-tune the large language model; Evaluate the accuracy of the large language model; Determine the requirements for entity relationship extraction, set up a guiding model, define the rules for extracting entity categories, relationship categories, and attribute information from the large amount of text data, and extract entity categories, relationship categories, and attribute information according to the rules; Formulate parsing rules, parse the extraction results output by the large language model, convert the extraction results into a triple structure, perform entity fusion on the extracted entity categories, store the triple data, and generate a knowledge graph for the target industry; Dynamically update and maintain the large language model and the knowledge graph.

2. The method for constructing a knowledge graph based on a large language model according to claim 1, wherein The step of evaluating the accuracy of the large language model includes: Use the validation set to evaluate the performance of the large language model, and monitor the key indicators of Accuracy, Precision, Recall, and F1 score of the large language model; According to the validation results, perform hyperparameter tuning. The hyperparameters include the learning rate, batch size, and fine-tuning layer; Record the parameter values of the adjusted parameters and the corresponding indicators; Use the test set to verify the performance of the large language model in the real scenario.

3. The method for constructing a knowledge graph based on a large language model according to claim 2, wherein, The step of using the validation set to evaluate the performance of the large language model and monitoring the key indicators of Accuracy, Precision, Recall, and F1 score of the large language model includes: Accuracy: Accuracy is the accuracy, an evaluation indicator for classification tasks, indicating the proportion of samples correctly classified by the model in the total number of samples. The formula is: Among them, TP represents the true positive example, the number of samples correctly predicted as the positive class by the large language model; TN represents the true negative example, the number of samples correctly predicted as the negative class by the large language model; FP represents the false positive example, the number of samples wrongly predicted as the positive class by the large language model; FN represents the false negative example, the number of samples wrongly predicted as the negative class by the large language model; Precision: Precision is the precision, indicating the proportion of samples actually being the positive class among the samples classified as the positive class. The formula is: Among them, TP represents the true positive example, the number of samples correctly predicted as the positive class by the large language model; FP represents the false positive example, the number of samples wrongly predicted as the positive class by the large language model; Recall: Recall is the recall rate, indicating the proportion of samples actually being the positive class that are correctly classified as the positive class. The formula is: Among them, TP represents the true positive example, the number of samples correctly predicted as the positive class by the large language model; FN represents the false negative example, the number of samples wrongly predicted as the negative class by the large language model; F1 score: Considering both precision and recall comprehensively and balancing the relationship between the two. The formula is: Among them, Precision is the precision and Recall is the recall rate.

4. The method for constructing a knowledge graph based on a large language model according to claim 1, wherein The step of setting up the guiding model and defining the rules for extracting entity categories, relationship categories, and attribute information from the large amount of text data includes: Clarify the tasks to be completed by the large language model; Determine the extraction rules, where entities must be in the same sentence and there must be specific relational words between entities; Define the extraction scope of entities and relationships; Specify the output format of the large language model.

5. The method for constructing a knowledge graph based on a large language model according to claim 1, wherein, The steps of formulating parsing rules, parsing the extraction results output by the large language model, converting the extraction results into a triple structure, performing entity fusion on the extracted entity categories, storing the triple data, and generating a knowledge graph for the target industry include: Define the triple structure, including the types of entities and relationships; Formulate parsing rules, parse the extraction results output by the large language model according to the parsing rules, and convert the extraction results into a triple structure; Perform entity fusion on the extracted entity categories; Store the triple data and import it into the Neo4j database; Generate a knowledge graph for the target industry.

6. The method for constructing a knowledge graph based on a large language model according to claim 5, wherein The entity fusion of the extracted entity categories includes: Perform entity fusion using similarity calculation. Entities with similarity higher than the specified threshold are fused in terms of relationships and attributes; analyze the extracted relationships, use the trained large language model to infer the existing relationships for unknown triples in the test set, and the model outputs the probability or confidence score of the existence of the relationship, indicating the certainty of the large language model for each prediction; the similarity calculation formula is: Among them, Slimilarity represents similarity, q represents the query, c represents the content, w represents the word in q, and z k represents the k-th strategy, where k is a positive integer, and p(wz k ) is the probability of the word appearing under the condition of the strategy of z k , and p(z k c) is the probability of the strategy z k appearing under the condition of c.

7. The method for constructing a knowledge graph based on a large language model according to claim 1, wherein The steps of dynamically updating and maintaining the large language model and the knowledge graph include: Regularly introduce new knowledge graph data; Update the large language model and the knowledge graph through online learning to maintain the consistency and timeliness of the knowledge graph with actual knowledge; Regularly monitor the health status of the knowledge graph; Clean up outdated or incorrect information.

8. A knowledge graph construction system based on a large language model, characterized in that, The system includes: A data processing and pre-training module for obtaining large-scale text data of the target industry, cleaning and annotating the large-scale text data, evaluating the required computing resources according to the size and parameters of the large language model, and pre-training the large language model using the large-scale text data; A large language model fine-tuning and target industry adaptation module for loading the pre-trained large language model weights and fine-tuning the large language model using small-scale text data of the target industry; An accuracy evaluation module for the large language model to evaluate the accuracy of the large language model; An entity recognition and relationship extraction module for determining the requirements for entity relationship extraction, setting a guiding model, defining the rules for extracting entity categories, relationship categories, and attribute information from large-scale text data, and extracting entity categories, relationship categories, and attribute information according to the rules; A knowledge graph construction module for formulating parsing rules, parsing the extraction results output by the large language model, converting the extraction results into a triple structure, performing entity fusion on the extracted entity categories, storing the triple data, and generating a knowledge graph for the target industry; A dynamic update and maintenance module for dynamically updating and maintaining the large language model and the knowledge graph.

9. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it realizes the steps in a method for constructing a knowledge graph based on a large language model as described in any one of claims 1-7.

10. A readable storage medium, characterized in that, The readable storage medium stores a computer program, and the computer program is suitable for being loaded by a processor to execute a method for constructing a knowledge graph based on a large language model as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Knowledge graph generation type question answering method and system based on large language model

    CN117033608A

  • Chinese triple extraction method based on BERT model

    CN113901820A

  • Power grid dispatching entity relation joint extraction method and system

    CN116523042A

  • Knowledge graph generation method and device, storage medium and electronic equipment

    CN116955646A

  • Knowledge graph construction method based on IE-Triple

    CN117252258A

Cited By

  • Knowledge base processing method and system for power field

    CN120745783A

  • Knowledge graph construction method, device and system based on large model

    CN121436141A

  • Knowledge graph dynamic construction method and system for advanced planning and scheduling

    CN121684471A