A method, system and terminal device for generating a knowledge graph
Through the combination of clustering and GraphRAG large language model, knowledge graphs in the field of innovation and entrepreneurship are generated, the problem of insufficient data labeling is solved, and knowledge graph completion is achieved through the screening of LDA theme models, improving the accuracy and coverage of the graph.
Patent Information
- Application Number
- CN202510436265.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-04-09
AI Technical Summary
In the field of innovation and entrepreneurship, the existing technology relies on manual annotation data sets, resulting in few annotation data sets and uneven data quality, making it difficult to generate accurate knowledge graphs.
By uploading the target text data for clustering, a task sample of GraphRAG prompt words is constructed, and entities and relationships are extracted using the GraphRAG large language model to generate a knowledge graph. In addition, by uploading triplet data from the external knowledge base, using the LDA topic model to filter triplet data of strong related topics, and adding it to the parquet file of GraphRAG large language model to achieve knowledge graph completion.
Reduce the dependence on manually labeled data sets, improve the accuracy of the generation of knowledge graphs, and support the completion of knowledge graphs, enhancing the coverage and accuracy of the graphs.
Smart Images

Figure CN119940513B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of knowledge graph generation, and particularly relates to a knowledge graph generation method, system and terminal device. Background Art
[0002] In the field of innovation and entrepreneurship, knowledge graphs play an important role. They help to quickly and accurately match services suitable for startup teams, solve problems such as the difficulty for startup teams to find appropriate services and resources, information asymmetry between entrepreneurs and service providers, and uneven service quality, and then support the intelligent decision-making of innovation and entrepreneurship services and the optimal allocation of resources, improving the efficiency and quality of the innovation and entrepreneurship ecosystem.
[0003] Currently, when generating knowledge graphs, it is common to use deep learning models to extract knowledge (entity recognition and relationships) for generation. However, when deep learning models extract knowledge, they often rely on manually labeled data sets, and have high requirements for the quality of the data sets. However, in the field of innovation and entrepreneurship, the labeled data sets are relatively few and the data quality is uneven. The method of using deep learning models to extract knowledge and generate graphs is difficult to generate, which is not conducive to ensuring the accuracy of the generated graphs. Summary of the Invention
[0004] The present invention provides a knowledge graph generation method, system and terminal device to at least solve the problem of difficult knowledge graph generation caused by few labeled data sets and uneven data quality.
[0005] In a first aspect, the present invention provides a knowledge graph generation method, which includes:
[0006] Upload target text data, where the target text data includes multiple texts; the target text data is data collected in advance for generating a knowledge graph;
[0007] Cluster the target text data to obtain several clusters;
[0008] Extract one text from each of the obtained clusters to obtain a set of target texts;
[0009] Use each target text to construct a task example of a GraphRAG prompt word to obtain a set of target task examples;
[0010] Submit the obtained target task examples to the GraphRAG large language model;
[0011] Use the GraphRAG large language model to extract entities and relationships in the target text data according to the target task examples, and generate a knowledge graph.
[0012] In one embodiment, the method further includes:
[0013] Upload an external knowledge base; the external knowledge base is a knowledge graph collected in the corresponding field;
[0014] Extract the triple data of the external knowledge base;
[0015] Input each triple data of the extracted external knowledge base into a pre-trained LDA topic model for calculation, and correspondingly output the probabilities of each triple data of the external knowledge base under the strongly related topics and each weakly related topic; among them, the strongly related topics and each weakly related topic are the most relevant topic and other topics in the output topics determined based on the topic-word distribution matrix generated when the LDA topic model is trained.
[0016] For each triple data of the external knowledge base, respectively judge whether its probability under the strongly related topic is greater than its probability under each of the weakly related topics:
[0017] If so, add the triple data to the parquet file of the GraphRAG large language model; if not, discard the triple data.
[0018] The parquet file includes the create_final_text_units.parquet data file, the create_final_entities.parquet entity file, and the create_final_relationships.parquet file.
[0019] In one embodiment, adding the triple data to the parquet file of the GraphRAG large language model includes:
[0020] Step 1, generate and store text units, specifically:
[0021] Reverse generate text content from the target triple data, denoted as the original text content; the target triple data is the triple data to be added to the parquet file of the GraphRAG large language model.
[0022] Calculate the MD5 hash value of the original text content and use it as the first hash ID;
[0023] Obtain the text length of the original text content;
[0024] Create a text unit containing the first hash ID, the original text content, and the text length;
[0025] Append the text unit to the create_final_text_units.parquet data file and save it;
[0026] Step 2, establish the association between entities and text units, specifically as follows:
[0027] Calculate the MD5 hash value of the text unit to obtain the unit ID;
[0028] Calculate the MD5 hash value of the head entity of the target triple data to obtain the entity hash ID of the head entity;
[0029] Calculate the MD5 hash value of the tail entity of the target triple data to obtain the entity hash ID of the tail entity;
[0030] Traverse the head entity and the tail entity of the target triple data;
[0031] For the currently traversed entity, perform the following steps M respectively:
[0032] Step M: Locate the record in the "name" field of the create_final_entities.parquet entity file whose name is the name of the entity of the target triple data, and add "[entity hash ID, unit ID]" corresponding to the entity under the "text_unit_ids" field of the record; if the entity of the target triple data cannot be located, add a new record in the create_final_entities.parquet entity file, then add the name of the entity to the "name" field of the newly added record, and add "[entity hash ID, unit ID]" corresponding to the entity to the "text_unit_ids" field of the newly added record;
[0033] Step 3, relationship data construction, specifically as follows:
[0034] Read the relationship data in the create_final_relationships.parquet file of the GraphRAG large language model;
[0035] Create a relationship record in the storage format of the create_final_relationships.parquet of the GraphRAG large language model for the target triple data;
[0036] Append the relationship record to the create_final_relationships.parquet file and save it.
[0037] In one embodiment, for the most relevant topic and other topics in the output topic determined based on the topic-word distribution matrix generated when the LDA topic model training is completed, the implementation method includes:
[0038] Based on the above-mentioned topic-lexical distribution matrix, top-n word segments with the highest probabilities under each output topic of the LDA topic model are obtained correspondingly; n is an empirical value;
[0039] Based on the top-n word segments with the highest probabilities under each output topic, the most relevant topic among each output topic of the LDA topic model is determined as the strong relevant topic, and other topics among each output topic of the LDA topic model are determined as weak relevant topics.
[0040] In one embodiment, based on the top-n word segments with the highest probabilities under each output topic, the most relevant topic among each output topic of the LDA topic model is determined as the strong relevant topic, and the specific implementation method includes:
[0041] For each topic among each output topic, calculate the sum of the probabilities of its top-n word segments;
[0042] For each output topic, take the output topic with the largest calculated sum of probabilities as the strong relevant topic.
[0043] In one embodiment, a training method of the LDA topic model includes:
[0044] Prepare a training set;
[0045] Construct an LDA topic model and set model parameters to obtain an initial LDA topic model; setting model parameters includes setting the number of LDA topics n_components, document-topic prior parameter doc_topic_prior, and topic-word prior parameter topic_word_prior, and also includes setting a random seed random_state, where n_components = 2, doc_topic_prior = 0.001, topic_word_prior = 1.0, and random_state = 0;
[0046] Use the training set to train the initial LDA topic model to obtain a trained LDA topic model.
[0047] In one embodiment, the method further includes:
[0048] Visualize the generated knowledge graph.
[0049] In a second aspect, the present invention provides a knowledge graph generation system, and the system includes:
[0050] A first uploading unit, configured to upload target text data, where the target text data includes multiple texts; the target text data is pre-collected data for generating a knowledge graph;
[0051] A clustering unit for clustering the target text data to obtain a number of clusters;
[0052] A text extraction unit for extracting a text from each cluster obtained by the clustering unit to obtain a set of target texts;
[0053] An example construction unit for constructing a task example of a GraphRAG prompt word for each target text obtained by the text extraction unit to obtain a set of target task examples;
[0054] An example submission unit for submitting each target task example obtained by the example construction unit to the GraphRAG large language model;
[0055] An entity relationship extraction unit for using the GraphRAG large language model to extract entities and relationships in the target text data according to the target task examples and generate a knowledge graph.
[0056] In one embodiment, the system further includes:
[0057] A second upload unit for uploading an external knowledge base; the external knowledge base is any knowledge graph collected in the corresponding field;
[0058] A data extraction unit for extracting triple data of the external knowledge base;
[0059] A probability calculation unit for inputting each triple data of the extracted external knowledge base into a pre-trained LDA topic model for operation, and correspondingly outputting the probability of each triple data of the external knowledge base under the strongly related topic and each weakly related topic; wherein, the strongly related topic and each weakly related topic are the most relevant topic and other topics in the output topics determined based on the topic-lexical distribution matrix generated when the LDA topic model is trained;
[0060] A data processing unit for respectively determining whether the probability of each triple data of the external knowledge base under the strongly related topic is greater than the probability under each of the weakly related topics: if so, adding the triple data to the parquet file of the GraphRAG large language model; if not, discarding the triple data.
[0061] In a third aspect, the present invention provides a terminal device, including:
[0062] A processor;
[0063] A memory for storing execution instructions of the processor;
[0064] Wherein, the processor is configured to execute the methods described in the above aspects.
[0065] As can be seen from the above technical solutions, the present invention has the following advantages:
[0066] (1) The present invention clusters the target text data to obtain several clusters, and extracts one text from each of the obtained clusters to obtain a set of target texts. Then, task examples of GraphRAG prompts are constructed using each target text. After that, the constructed task examples are submitted to the GraphRAG large language model. Then, using this GraphRAG large language model, according to the submitted task examples (target task examples), entities and relationships in the above target text data are extracted to generate a knowledge graph, avoiding the dependence on manually labeled data sets, and to a certain extent helping to reduce the requirements for the quality of data sets. At the same time, it can also enable the present invention to submit diverse examples to GraphRAG, which in turn helps to improve the accuracy of GraphRAG in completing tasks, thereby improving the accuracy of the knowledge graph generated by the present invention.
[0067] (2) The present invention can also add triple data from an external knowledge base to the parquet file of the GraphRAG large language model, which is convenient for users to complete the knowledge graph. Description of the Drawings
[0068] In order to more clearly illustrate the technical solutions of the present invention, the accompanying drawings required for description will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0069] Figure 1 is a schematic flowchart of the method according to an embodiment of the present invention.
[0070] Figure 2 is a schematic block diagram of the system according to an embodiment of the present invention.
[0071] Figure 3 is a schematic block diagram of the terminal device according to an embodiment of the present invention. Detailed Embodiments
[0072] GraphRAG large language model: That is, the Graph-based Retrieval-Augmented Generation large language model, which is a model that combines knowledge graph and retrieval-augmented generation (RAG) technology.
[0073] The knowledge graph generation method provided by the present invention incorporates clustering, and constructs task examples of GraphRAG prompts based on the clustering results and submits them to the GraphRAG large language model for subsequent use of the GraphRAG large language model to generate a knowledge graph. In addition, the present invention also supports uploading an external knowledge base for graph completion.
[0074] The following will describe in detail the specific implementation steps of the knowledge graph generation method. For illustration rather than limitation, specific details such as specific system structures and technologies are proposed to thoroughly understand the embodiments of the present invention. However, those skilled in the art should clearly understand that the present invention can also be implemented in other embodiments without these specific details.
[0075] It should be understood that when used in the specification of the present invention, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations. The terms "comprising", "including", "having", and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.
[0076] Statements such as "in one embodiment" or "in some embodiments" described in the present invention mean that the specific features, structures, or characteristics described in the embodiment are included in one or more embodiments of the present invention. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments" that appear in different parts of the present invention do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways.
[0077] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.
[0078] Please refer to Figure 1 , which shows the knowledge graph generation method provided by an embodiment of the present invention. The method includes the following steps (S101 to S106).
[0079] Step S101: Upload target text data.
[0080] The target text data is data collected in advance for generating a knowledge graph.
[0081] Taking the field of innovation and entrepreneurship as an example, the target text data is data for generating a knowledge graph in the field of innovation and entrepreneurship that has been pre-collected.
[0082] Understandably, the above target text data includes multiple texts.
[0083] Understandably, those skilled in the art can collect relevant text data according to actual needs. Exemplarily, when collecting the target text data, the categories of the collected data can be divided first, and then the data can be collected according to these categories. For example, taking the innovation and entrepreneurship knowledge graph in the above field of innovation and entrepreneurship as an example, data can be collected from ten aspects: start-up enterprise information, innovation and entrepreneurship activities, innovation and entrepreneurship consultations, innovation technology patents, business incubators, business services, business courses, market dynamics, investment and financing, and expert talents. Specifically, Mita AI Search can be used to search for data sources in the above ten aspects to determine from which websites (such as Chuangyebang, IT Juzi, etc.) to crawl data. Then, the Octopus Collector can be used to set the request headers, paging rules, etc. to collect website data. After that, the collected website data is preprocessed. For example, data cleaning is performed first, and then data desensitization. Data cleaning may include, but is not limited to: removing duplicate and irrelevant (such as irrelevant to the knowledge graph to be generated) data. Data desensitization: processing sensitive fields in the data.
[0084] Understandably, the target text data is data after preprocessing, and the preprocessing includes data cleaning and data desensitization.
[0085] Step S102: Cluster the target text data to obtain several clusters.
[0086] Optionally, the k-means clustering algorithm is used to cluster the target text data.
[0087] In specific implementation, the silhouette coefficient and the sum of squared errors within the cluster can be observed by adjusting the k value to explore the influence of different k values on the clustering effect.
[0088] In this embodiment, k = 10 is used as the number of clustering clusters. Thus, the k-means clustering algorithm is used to cluster the target text data, and 10 clusters can be obtained.
[0089] Step S103: Extract one text from each of the obtained clusters to obtain a set of target texts.
[0090] In this embodiment, one text is randomly extracted from each of the obtained 10 clusters to obtain 10 target texts.
[0091] In addition, it should be noted that, in order to increase the representativeness of the extracted documents, in the specific implementation, the centroid nearest neighbor method and / or central tendency indicators can also be used to determine the most representative text in each type of cluster, and then extract it as the text extracted from the corresponding cluster.
[0092] Step S104: Use each target text to construct a task example of the GraphRAG prompt, and obtain a set of target task examples.
[0093] The GraphRAG prompt is the prompt of the GraphRAG large language model.
[0094] In the specific implementation, the task example includes the target text, entity recognition and attribute extraction, and also includes relationship construction. Among them, entity recognition and attribute extraction are used to specify the entities (hereinafter referred to as "target entities") to be recognized from the target text in the example, and relationship construction is used to specify which relationships between the target entities are to be constructed.
[0095] Taking the generation of the above-mentioned innovation and entrepreneurship knowledge graph as an example, the target entities can include: technology, product, service, platform, management, culture, economy, research, consulting, investment, founder, CEO, incubator, AI, startup, survey, user, financing, innovation, learning, decision making, a total of 21 types of entity categories. The relationship construction can be designed by those skilled in the art according to actual needs.
[0096] Step S105: Submit the obtained target task examples to the GraphRAG large language model.
[0097] This step S105 can provide representative task examples for the large model, which helps to increase the diversity and coverage of the examples in the GraphRAG large language model, and then deepen the model's understanding of the task framework and the expected output style.
[0098] Submitting diverse examples to GraphRAG in this step also helps to improve the efficiency of GraphRAG in completing tasks, and then helps to improve the efficiency of generating the knowledge graph of the present invention.
[0099] Step S106: Use the GraphRAG large language model to extract entities and relationships from the target text data according to the target task example, and generate a knowledge graph.
[0100] In specific implementation, a prompt is constructed through the GraphRAG method, and a structured instruction is written, requiring the large model to complete the task of extracting entities and relationships from the input text, and specifying the output format according to the target task example. Finally, the GraphRAG large language model is used to identify the entities and corresponding relationships in the input text, and the identified entities and relationships are used to generate a graph and store it in the parquet file of the model.
[0101] In some embodiments of the present invention, the above method further includes the following steps S201 to S204.
[0102] Step S201: Upload an external knowledge base.
[0103] The external knowledge base is a knowledge graph collected in the corresponding field. That is, the external knowledge base is another graph related to the graph generated in step S106.
[0104] Step S202: Extract the triple data of the external knowledge base.
[0105] Step S203: Input each triple data of the extracted external knowledge base into a pre-trained LDA topic model for calculation, and correspondingly output the probability of each triple data of the external knowledge base under the strongly related topic and each weakly related topic.
[0106] In this embodiment, the strongly related topic and each weakly related topic are the most relevant topic and other topics in the output topics determined based on the topic-lexical distribution matrix generated when the LDA topic model is trained.
[0107] Step S204: For each triple data of the external knowledge base, respectively judge whether the probability under the strongly related topic is greater than the probability under each of the weakly related topics:
[0108] If so, add the triple data to the parquet file of the GraphRAG large language model; if not, discard the triple data.
[0109] In this embodiment, adding the triple data to the target knowledge graph means storing the triple data in the parquet file of the GraphRAG large language model.
[0110] It can be understood that the above external knowledge base can be any other graph related to the graph generated in step S106.
[0111] The design and use of the above steps S201 to S204 help the user to complete the graph complementation, thus helping the user to construct a more comprehensive graph of knowledge. In addition, it also helps the user to complete the graph complementation subsequently, thereby overcoming the shortcoming that GraphRAG does not support incremental updates, realizing the dynamic addition of knowledge triples in the graph, and thus realizing the update of the content of the knowledge graph.
[0112] During subsequent use, when new knowledge data needs to be added to the graph, even if only one triple or a few triples of data need to be added, these triples of data to be added can be directly uploaded and added in the form of an external knowledge base, which is convenient to use and convenient for timely updating the content of the knowledge graph.
[0113] The parquet file includes the create_final_text_units.parquet data file, the create_final_entities.parquet entity file, and the create_final_relationships.parquet file.
[0114] The create_final_text_units.parquet data file, the create_final_entities.parquet entity file, and the create_final_relationships.parquet file are the default parquet files for GraphRAG to store the generated knowledge graph.
[0115] In some embodiments of the present invention, adding triple data to the parquet file of the GraphRAG large language model includes:
[0116] Step 1, generate and store text units, specifically:
[0117] Reverse generate text content from the target triple data, denoted as the original text content; the target triple data is the triple data to be added to the parquet file of the GraphRAG large language model;
[0118] Calculate the MD5 hash value of the original text content and use it as the first hash ID;
[0119] Obtain the text length of the original text content;
[0120] Create a text unit containing the first hash ID, the original text content, and the text length (create it according to the default format in the create_final_text_units.parquet data file of GraphRAG);
[0121] Append the said text unit to the create_final_text_units.parquet data file of the GraphRAG large language model and save it;
[0122] Step 2: Establish the association between entities and text units, specifically as follows:
[0123] Calculate the MD5 hash value of the said text unit to obtain the unit ID;
[0124] Calculate the MD5 hash value of the head entity of the target triple data to obtain the entity hash ID of the head entity;
[0125] Calculate the MD5 hash value of the tail entity of the target triple data to obtain the entity hash ID of the tail entity;
[0126] Traverse the head entity and the tail entity of the target triple data;
[0127] For the entity currently traversed, perform the following steps M respectively:
[0128] Step M: Locate the record in the "name" field of the create_final_entities.parquet entity file of the GraphRAG large language model with the name of this entity in the target triple data, and add "[entity hash ID, unit ID]" corresponding to this entity under the "text_unit_ids" field of the record; if the entity in the target triple data cannot be located, add a new record in the create_final_entities.parquet entity file, then add the name of this entity to the "name" field of the newly added record, and add "[entity hash ID, unit ID]" corresponding to this entity to the "text_unit_ids" field of the newly added record;
[0129] Step 3: Relationship data construction, specifically as follows:
[0130] Read the relationship data in the create_final_relationships.parquet file of the GraphRAG large language model;
[0131] Create a relationship record in the storage format of the create_final_relationships.parquet of the GraphRAG large language model for the target triple data;
[0132] Append the said relationship record to the create_final_relationships.parquet file and save it.
[0133] Optionally, the method for reversely generating text content from the target triple data may be: concatenating the elements of the target triple data in sequence to form a text content, with commas separating the elements.
[0134] When currently traversing to the head entity of the target triple data, the specific implementation of step M is:
[0135] Locate the record in the "name" field of the create_final_entities.parquet entity file of the GraphRAG large language model whose name is the name of the head entity of the target triple data, and add the "[entity hash ID, unit ID]" corresponding to the head entity under the "text_unit_ids" field of the record;
[0136] If the head entity cannot be located, a new record is added to the create_final_entities.parquet entity file, and then the name of the head entity is added to the "name" field of the newly added record, and the "[entity hash ID, unit ID]" corresponding to the head entity is added to the "text_unit_ids" field of the newly added record.
[0137] It can be understood that the entity hash ID of the head entity is denoted as the second hash ID, the entity hash ID of the tail entity is denoted as the third hash ID, the "[entity hash ID, unit ID]" corresponding to the head entity is "[second hash ID, unit ID]", and the "[entity hash ID, unit ID]" corresponding to the tail entity is "[third hash ID, unit ID]".
[0138] When currently traversing to the tail entity of the target triple data, step M includes:
[0139] Locate the record in the "name" field of the create_final_entities.parquet entity file of the GraphRAG large language model whose name is the name of the tail entity of the target triple data, and add the "[entity hash ID, unit ID]" corresponding to the tail entity under the "text_unit_ids" field of the record;
[0140] If the tail entity cannot be located, a new record is added to the create_final_entities.parquet entity file, and then the name of the tail entity is added to the "name" field of the newly added record, and the "[entity hash ID, unit ID]" corresponding to the tail entity is added to the "text_unit_ids" field of the newly added record.
[0141] Understandably, the entity hash ID of the head entity is denoted as the second hash ID, and the entity hash ID of the tail entity is denoted as the third hash ID. The "[entity hash ID, unit ID]" corresponding to the head entity is "[second hash ID, unit ID]", and the "[entity hash ID, unit ID]" corresponding to the tail entity is "[third hash ID, unit ID]".
[0142] In some embodiments of the present invention, for the most relevant topic and other topics in the output topics determined based on the topic-word distribution matrix generated when the LDA topic model training is completed, the implementation method includes:
[0143] Based on the topic-word distribution matrix, correspondingly obtain the top n word segments with the highest probabilities under each output topic of the LDA topic model; n is an empirical value;
[0144] Based on the top n word segments with the highest probabilities under each output topic, determine that the most relevant topic in each output topic of the LDA topic model is the strongly relevant topic, and determine that all other topics in each output topic of the LDA topic model are weakly relevant topics.
[0145] Exemplarily, based on the topic-word distribution matrix, correspondingly obtaining the top n word segments with the highest probabilities under each output topic of the LDA topic model, the implementation method is:
[0146] For each topic involved in the topic-word distribution matrix, extract all the word segments under it and their corresponding probabilities;
[0147] For each topic involved in the topic-word distribution matrix, sort all the extracted word segments in descending order according to the extracted probabilities;
[0148] For each topic involved in the topic-word distribution matrix, select the top n word segments with the highest probabilities.
[0149] n is a positive integer, and its specific value can be set by those skilled in the art according to experience.
[0150] In this embodiment, n = 5.
[0151] Optionally, based on the top n word segments with the highest probabilities under each output topic, determining that the most relevant topic in each output topic of the LDA topic model is the strongly relevant topic, the specific implementation method includes:
[0152] For each topic in each output topic, calculate the sum of the probabilities of the top n word segments with the highest probabilities under it;
[0153] For each output topic, take the output topic with the largest calculated sum of probabilities as the strongly relevant topic.
[0154] Optionally, the training method of the LDA topic model includes:
[0155] Step 1: Prepare the training set;
[0156] Step 2: Construct an LDA topic model and set the model parameters to obtain an initial LDA topic model; setting the model parameters includes setting the number of LDA topics n_components, the document-topic prior parameter doc_topic_prior, and the topic-word prior parameter topic_word_prior, and also includes setting the random seed random_state, where n_components = 2, doc_topic_prior = 0.001, topic_word_prior = 1.0, and random_state = 0;
[0157] Step 3: Use the training set to train the initial LDA topic model to obtain a trained LDA topic model.
[0158] The above Step 1 includes:
[0159] Prepare the text data to obtain a text data set, which is the training set.
[0160] Preparing the text data means collecting text data related to the knowledge graph to be generated within the target domain, where the target domain is the domain of the knowledge graph to be generated.
[0161] The above Step 2 specifically includes:
[0162] Create an LDA topic model and perform parameter setting for the created LDA topic model to obtain an initial LDA topic model.
[0163] The parameter setting includes setting the number of LDA topics, prior parameters, and random seed.
[0164] Create an LDA topic model: Instantiate the LatentDirichletAllocation class.
[0165] The prior parameters include the document-topic distribution α and the topic-word distribution η.
[0166] The prior parameters are represented by doc_topic_prior and topic_word_prior. The random seed is represented by random_state. The number of topics is represented by n_components. The document-topic distribution α is represented by doc_topic_prior. The topic-word distribution η is represented by topic_word_prior.
[0167] In this embodiment, the number of LDA topics n_components is set to 2, the prior parameters doc_topic_prior = 0.001 and topic_word_prior = 1.0 are set, and the random seed random_state = 0 is fixed to ensure the reproducibility of the experiment.
[0168] Step 3 specifically includes the following steps 1101 to 1104.
[0169] Step 1101: Segment the text data.
[0170] The word segmentation in Step 1101 includes resource loading, text cleaning, and word segmentation.
[0171] Resource loading includes stop word list loading and custom dictionary loading.
[0172] The user can prepare the stop word list and custom dictionary in advance and store them locally.
[0173] Stop word list loading: Read the stop words from the local file stop word list and construct a stop word set.
[0174] Custom dictionary loading: Load the custom dictionary from the local to ensure the correct segmentation of specific words or professional terms required by the user.
[0175] Optionally, the user can load the custom dictionary through the jieba.load_userdict() function of the jieba word segmentation library.
[0176] Text cleaning and word segmentation:
[0177] Read the texts in the text dataset one by one to generate a document list lines;
[0178] Use the jieba.lcut() function of the jieba word segmentation library to perform accurate mode word segmentation on the texts in the document list lines to obtain a first word list;
[0179] According to the stop word set, remove the stop words in the first word list to obtain a filtered word list filtered_words.
[0180] After loading the custom dictionary using jieba.load_userdict(), use the jieba.lcut() function to perform word segmentation on the text. Jieba will perform word segmentation according to the words in the custom dictionary, thus obtaining more accurate results.
[0181] Understandably, in the document list `lines`, each line represents a text. In the first word list, each line represents the word segmentation result of a text. In the filtered word list `filtered_words`, each line represents the filtered word segmentation result of a text's word segmentation result.
[0182] Step 1102: Concatenate each line of word segmentation in the word list `filtered_words` into a string separated by spaces (i.e., there is a space between each word segmentation), generating the `tokenized_docs` list.
[0183] In the `tokenized_docs` list, each sub-list corresponds to the word segmentation of a text.
[0184] Step 1103: Convert the `tokenized_docs` list into a sparse matrix `X`.
[0185] In the sparse matrix `X`, the rows represent documents, the columns represent terms, and the values in the matrix represent the frequency of terms in the documents.
[0186] Step 1104: Based on the sparse matrix `X`, train the initial LDA topic model to obtain a trained LDA topic model.
[0187] The trained LDA topic model can output a document-topic distribution matrix.
[0188] In the document-topic distribution matrix, each row represents a text, each column represents a topic, and the matrix elements represent the probability distribution of the text on each topic.
[0189] In the topic-term distribution matrix, each row represents a topic, each column represents a term (i.e., word segmentation), and the matrix elements represent the probability distribution of the term on each topic.
[0190] The method of this embodiment can retain the strongly relevant triples in the external knowledge base through the trained LDA topic model, fuse them into the model, and achieve knowledge completion of the knowledge graph.
[0191] In some embodiments of the present invention, the method further includes: visualizing the generated knowledge graph.
[0192] Visualizing the generated knowledge graph is more intuitive.
[0193] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0194] Figure 2A knowledge graph generation system provided by the present invention, the system includes:
[0195] A first upload unit 201 for uploading target text data, where the target text data includes multiple texts; the target text data is pre-collected data for generating a knowledge graph;
[0196] A clustering unit 202 for clustering the target text data to obtain several clusters;
[0197] A text extraction unit 203 for respectively extracting a text from each cluster obtained by the clustering unit to obtain a set of target texts;
[0198] An example construction unit 204 for constructing a task example of a GraphRAG prompt word for each target text obtained by the text extraction unit to obtain a set of target task examples;
[0199] An example submission unit 205 for submitting each target task example obtained by the example construction unit to the GraphRAG large language model;
[0200] An entity relationship extraction unit 206 for using the GraphRAG large language model and based on the target task examples, extracting entities and relationships in the target text data to generate a knowledge graph.
[0201] In some embodiments of the present invention, the system further includes:
[0202] A second upload unit for uploading an external knowledge base; the external knowledge base is any knowledge graph collected in the corresponding field;
[0203] A data extraction unit for extracting triple data of the external knowledge base;
[0204] A probability calculation unit for inputting each triple data of the extracted external knowledge base into a pre-trained LDA topic model for operation, and correspondingly outputting the probability of each triple data of the external knowledge base under a strongly related topic and each weakly related topic; wherein, the strongly related topic and each weakly related topic are the most relevant topic and other topics in the output topics determined based on the topic-lexical distribution matrix generated when the LDA topic model is trained;
[0205] A data processing unit for respectively determining, for each triple data of the external knowledge base, whether its probability under the strongly related topic is greater than its probability under each of the weakly related topics: if so, adding the triple data to the parquet file of the GraphRAG large language model; if not, discarding the triple data.
[0206] In some embodiments of the present invention, the system further includes:
[0207] A graph visualization unit for visualizing the generated knowledge graph.
[0208] In some embodiments of the present invention, the graph visualization unit includes a first visualization module, a second visualization module, and a third visualization module.
[0209] In use, when the user selects different visualization modules, the graphs in the corresponding visualization forms can be obtained.
[0210] Optionally, the first visualization module is used to read the entities and relationships of the generated knowledge graph into a DataFrame using pandas, and then export it in the excel format.
[0211] Optionally, the second visualization module is used to visualize the generated knowledge graph using plotly.
[0212] In specific implementation, the nodes in the network graph can be laid out based on the Fruchterman-Reingold force-directed algorithm through the spring_layout method of networkx; the attributes of the nodes and edges, including positions, colors, texts, etc., can be set through the Scatter of plotly; the graph layout can be displayed through the Figure method.
[0213] Optionally, the third visualization module is used to visualize the generated knowledge graph using Neo4j.
[0214] By constructing the obtained graph into a graph database using Neo4j, the relevant nodes of the graph can be effectively stored and managed. In particular, the query speed for nodes (i.e., entities) and relationships is fast, and the nodes can be moved with strong interactivity.
[0215] For the embodiments of the knowledge graph generation system in this embodiment, this system and the knowledge graph generation methods in the above embodiments belong to the same inventive concept. For the details not described in detail in the embodiments of the knowledge graph generation system, reference can be made to the embodiments of the knowledge graph generation method in the previous text.
[0216] Figure 3 It is a schematic structural diagram of a terminal 300 provided by an embodiment of the present invention, and this terminal 300 can be used to execute the method provided by the embodiment of the present invention.
[0217] Among them, the terminal 300 may include: a processor 310, a memory 320, and a communication unit 330. These components communicate via one or more buses. Those skilled in the art can understand that the structure of the server shown in the figure does not constitute a limitation on the present invention. It can be a bus structure, a star structure, and may also include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.
[0218] Among them, the memory 320 can be used to store the execution instructions of the processor 310. The memory 320 can be implemented by any type of volatile or non-volatile storage terminal or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk. When the execution instructions in the memory 320 are executed by the processor 310, the terminal 300 can execute some or all of the steps in the above method embodiments.
[0219] The processor 310 is the control center of the storage terminal, connecting various parts of the entire electronic terminal through various interfaces and lines. By running or executing software programs and / or modules stored in the memory 320, and calling the data stored in the memory, it executes various functions of the electronic terminal and / or processes data. The processor may be composed of an integrated circuit (IC). For example, it may be composed of a single packaged IC, or may be composed of multiple packaged ICs with the same or different functions connected together. For example, the processor 310 may only include a central processing unit (CPU). In the embodiment of the present invention, the CPU may be a single operation core or may include multiple operation cores.
[0220] The communication unit 330 is used to establish a communication channel, so that the storage terminal can communicate with other terminals. It receives user data sent by other terminals or sends user data to other terminals.
[0221] The technical effects that can be achieved in this embodiment can be referred to the description above, and will not be elaborated here.
[0222] For the same and similar parts among the various embodiments in this specification, reference can be made to each other.
[0223] The foregoing description of the disclosed embodiments enables those skilled in the art to practice or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Thus, the present invention is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A knowledge graph generation method, characterized in that the method include: Upload target text data, wherein the target text data includes multiple texts; the target text data is pre-collected data for generating a knowledge graph; Cluster the target text data to obtain several clusters; Extract a text from each obtained cluster to obtain a set of target texts; Use each target text to construct a task example of GraphRAG prompt words to obtain a set of target task examples; Submit the obtained target task examples to the GraphRAG large language model; Utilizing the GraphRAG large language model, based on the target task examples, extracting entities and relationships in the target text data to generate a knowledge graph; The method also includes: Upload an external knowledge base; the external knowledge base is a knowledge graph collected in the corresponding field; Extracting triple data from the external knowledge base; Input each triple data of the extracted external knowledge base into the pre-trained LDA topic model for operation, and output the probability of each triple data of the external knowledge base under the strongly related topics and each weakly related topic; wherein the strongly related topics and each weakly related topics are the most related topics and other topics in the output topics determined based on the topic-vocabulary distribution matrix generated when the LDA topic model training is completed; For each triple data of the external knowledge base, determine whether its probability under the strongly related topic is greater than its probability under each of the weakly related topics: If so, the triple data is added to the parquet file of the GraphRAG large language model; if not, the triple data is discarded.
2. The knowledge graph generation method according to claim 1, characterized in that: The parquet files include create_final_text_units.parquet data files, create_final_entities.parquet entity files and create_final_relationships.parquet files; Add triple data to the parquet file of the GraphRAG large language model, including: Step 1: Generate and store text units, specifically: Reversely generate text content from the target triple data, which is recorded as the original text content; the target triple data is the triple data to be added to the parquet file of the GraphRAG large language model; Calculate the MD5 hash value of the original text content and use it as the first hash ID; Get the text length of the original text content; Creating a text unit including the first hash ID, the original text content, and the text length; Append the text unit to the create_final_text_units.parquet data file and save it; Step 2: Establish the association between entity and text unit, specifically: Calculate the MD5 hash value of the text unit to obtain a unit ID; Calculate the MD5 hash value of the first entity of the target triple data to obtain the entity hash ID of the first entity; Calculate the MD5 hash value of the tail entity of the target triple data to obtain the entity hash ID of the tail entity; Traverse the first entity and the last entity of the target triple data; For the currently traversed entity, perform the following steps M respectively: Step M, locate the record whose "name" field in the create_final_entities.parquet entity file is the name of the entity of the target triple data, and add "[entity hash ID, unit ID]" corresponding to the entity under the "text_unit_ids" field of the record; if the entity of the target triple data cannot be located, add a new record in the create_final_entities.parquet entity file, then add the name of the entity to the "name" field of the newly added record, and add "[entity hash ID, unit ID]" corresponding to the entity to the "text_unit_ids" field of the newly added record; Step 3: Relational data construction, specifically: Read the relationship data in the create_final_relationships.parquet file of the GraphRAG large language model; According to the storage format of create_final_relationships.parquet described in the GraphRAG large language model, create relationship records in this storage format for the target triple data; Append the relationship records to the create_final_relationships.parquet file and save it.
3. The knowledge graph generation method according to claim 1, characterized in that: The most relevant topics and other topics among the output topics determined based on the topic-vocabulary distribution matrix generated when the LDA topic model training is completed, and the implementation method includes: Based on the topic-vocabulary distribution matrix, the top n word segments with probability ranking under each output topic of the LDA topic model are obtained; n is an empirical value; Based on the top n word segments with probability ranking under each output topic, the most relevant topic in each output topic of the LDA topic model is determined to be the strongly relevant topic, and the other topics in each output topic of the LDA topic model are determined to be weakly relevant topics.
4. The knowledge graph generation method according to claim 3, characterized in that: Based on the top n word segments with probability ranking under each output topic, the most relevant topic among the output topics of the LDA topic model is determined as the strongly relevant topic. The specific implementation method includes: For each topic in each output topic, calculate the sum of the probabilities of the top n word segments; For each output topic, the output topic with the largest sum of calculated probabilities is taken as the strongly related topic.
5. The knowledge graph generation method according to claim 1, characterized in that: The training method of the LDA topic model includes: Prepare training set; Construct an LDA topic model and set model parameters to obtain an initial LDA topic model; setting model parameters includes setting the number of LDA topics n_components, document-topic prior parameter doc_topic_prior and topic-word prior parameter topic_word_prior, and also setting the random seed random_state, where n_components=2, doc_topic_prior=0.001, topic_word_prior=1.0, random_state=0; Use the training set to train the initial LDA topic model to obtain the trained LDA topic model.
6. The knowledge graph generation method according to claim 1, characterized in that: The method also includes: Visualize the generated knowledge graph.
7. A knowledge graph generation system, characterized in that: The system includes: A first uploading unit is used to upload target text data, wherein the target text data includes a plurality of texts; the target text data is pre-collected data for generating a knowledge graph; A clustering unit, used for clustering the target text data to obtain a plurality of clusters; A text extraction unit is used to extract a text from each cluster obtained by the clustering unit to obtain a set of target texts; A sample construction unit is used to construct a task sample of GraphRAG prompt words for each target text obtained by the text extraction unit to obtain a set of target task samples; The sample submission unit is used to submit the target task samples obtained by the sample construction unit to the GraphRAG large language model; An entity relationship extraction unit, configured to extract entities and relationships in the target text data and generate a knowledge graph based on the target task examples using the GraphRAG large language model; The system also includes: A second uploading unit is used to upload an external knowledge base; the external knowledge base is any knowledge graph collected in the corresponding field; A data extraction unit, used for extracting triple data from the external knowledge base; A probability calculation unit is used to input each triple data of the extracted external knowledge base into a pre-trained LDA topic model for calculation, and correspondingly output the probability of each triple data of the external knowledge base under a strongly related topic and each weakly related topic; wherein the strongly related topic and each weakly related topic are the most related topic and other topics in the output topics determined based on the topic-vocabulary distribution matrix generated when the LDA topic model training is completed; The data processing unit is used to determine, for each triple data of the external knowledge base, whether its probability under the strongly related topic is greater than its probability under each weakly related topic: if so, add the triple data to the parquet file of the GraphRAG large language model; if not, discard the triple data.
8. A terminal device, characterized in that: include: processor; A memory for storing execution instructions of the processor; The processor is configured to execute the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Knowledge graph construction method and system based on large language model
CN117150050A
Nuclear power DCS operation and maintenance knowledge graph construction system and method based on large language model
CN118607632A