Knowledge graph generation method and system and terminal equipment

Through the combination of clustering and GraphRAG large language model, the problem of insufficient data labeling in the knowledge graph generation in the field of innovation and entrepreneurship is solved, and the high-accurate knowledge graph generation is achieved, and intelligent decision-making and resource optimization configuration of knowledge graphs are supported.

CN119940513AActive Publication Date: 2025-05-06SHANDONG INSPUR INNOVATION & ENTREPRENEURSHIP TECH CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510436265.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-05-06
Estimated Expiration
2045-04-09

AI Technical Summary

Technical Problem

In the field of innovation and entrepreneurship, the existing technology relies on manual annotation data sets, resulting in few annotation data sets and uneven data quality, making it difficult to generate accurate knowledge graphs.

Method used

By uploading the target text data for clustering, a task sample of GraphRAG prompt words is constructed, and entities and relationships are extracted using the GraphRAG large language model to generate a knowledge graph. In addition, by uploading triple data from the external knowledge base, using the LDA topic model to filter the data of strong related topics, and adding it to the parquet file of the GraphRAG large language model to achieve the completion of the knowledge graph.

Benefits of technology

Reliance on manual labeled data sets is reduced, the accuracy of knowledge graph generation is improved, the sample diversity and coverage of GraphRAG large language model is enhanced, and the intelligent decision-making and resource optimization configuration of knowledge graphs are supported.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940513A_ABST
    Figure CN119940513A_ABST
Patent Text Reader

Abstract

The invention provides a knowledge graph generation method and system and terminal equipment, and belongs to the technical field of knowledge graph generation. The method comprises the steps that target text data is uploaded, and the target text data comprises a plurality of texts; the target text data is pre-collected data used for generating a knowledge graph; clustering the target text data to obtain a plurality of class clusters; extracting a text from each obtained class cluster to obtain a group of target texts; constructing a task sample of a GraphRAG cue word by utilizing each target text to obtain a group of target task samples; the obtained target task examples are submitted to a GraphRAG large language model; and by utilizing the GraphRAG large language model, according to the target task sample, extracting entities and relationships in the target text data, and generating a knowledge graph. The method is used for solving the problem that the knowledge graph is difficult to generate due to few annotated data sets and uneven data quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of knowledge graph generation, and specifically relates to a knowledge graph generation method, system and terminal device. Background Art

[0002] In the field of innovation and entrepreneurship, knowledge graphs play an important role. They help to quickly and accurately match services suitable for entrepreneurial teams, solve problems such as entrepreneurial teams having difficulty finding suitable services and resources, information asymmetry between entrepreneurs and service providers, and uneven service quality. They then support intelligent decision-making and optimal resource allocation for innovative and entrepreneurial services, and improve the efficiency and quality of the innovative and entrepreneurial ecosystem.

[0003] At present, deep learning models are commonly used to extract knowledge (entity recognition and relationships) for knowledge graph generation. However, deep learning models often need to rely on manually annotated data sets to extract knowledge, which requires high data set quality. However, there are relatively few annotated data sets in the field of innovation and entrepreneurship, and the data quality varies. The method of using deep learning models to extract knowledge to generate graphs is difficult to generate, which is not conducive to ensuring the accuracy of the generated graphs. Summary of the invention

[0004] The present invention provides a knowledge graph generation method, system and terminal device to at least solve the problem of difficulty in generating knowledge graphs caused by small number of annotated data sets and uneven data quality.

[0005] In a first aspect, the present invention provides a method for generating a knowledge graph, the method comprising: Upload target text data, wherein the target text data includes multiple texts; the target text data is pre-collected data for generating a knowledge graph; Cluster the target text data to obtain several clusters; Extract a text from each obtained cluster to obtain a set of target texts; Use each target text to construct a task example of GraphRAG prompt words to obtain a set of target task examples; Submit the obtained target task examples to the GraphRAG large language model; Utilizing the GraphRAG large language model, based on the target task examples, entities and relationships in the target text data are extracted to generate a knowledge graph.

[0006] In one embodiment, the method further comprises: Upload an external knowledge base; the external knowledge base is a knowledge graph collected in the corresponding field; Extracting triple data from the external knowledge base; Input each triple data of the extracted external knowledge base into the pre-trained LDA topic model for operation, and output the probability of each triple data of the external knowledge base under the strongly related topics and each weakly related topic; wherein the strongly related topics and each weakly related topics are the most related topics and other topics in the output topics determined based on the topic-vocabulary distribution matrix generated when the LDA topic model training is completed; For each triple data of the external knowledge base, determine whether its probability under the strongly related topic is greater than its probability under each of the weakly related topics: If so, the triple data is added to the parquet file of the GraphRAG large language model; if not, the triple data is discarded.

[0007] The parquet files include a create_final_text_units.parquet data file, a create_final_entities.parquet entity file, and a create_final_relationships.parquet file.

[0008] In one embodiment, adding triple data to the parquet file of the GraphRAG large language model includes: Step 1: Generate and store text units, specifically: Reversely generate text content from the target triple data, which is recorded as the original text content; the target triple data is the triple data to be added to the parquet file of the GraphRAG large language model; Calculate the MD5 hash value of the original text content and use it as the first hash ID; Get the text length of the original text content; Creating a text unit including the first hash ID, the original text content, and the text length; Append the text unit to the create_final_text_units.parquet data file and save it; Step 2: Establish the association between entity and text unit, specifically: Calculate the MD5 hash value of the text unit to obtain a unit ID; Calculate the MD5 hash value of the first entity of the target triple data to obtain the entity hash ID of the first entity; Calculate the MD5 hash value of the tail entity of the target triple data to obtain the entity hash ID of the tail entity; Traverse the first entity and the last entity of the target triple data; For the currently traversed entity, perform the following steps M respectively: Step M, locate the record whose "name" field in the create_final_entities.parquet entity file is the name of the entity of the target triple data, and add "[entity hash ID, unit ID]" corresponding to the entity under the "text_unit_ids" field of the record; if the entity of the target triple data cannot be located, add a new record in the create_final_entities.parquet entity file, then add the name of the entity to the "name" field of the newly added record, and add "[entity hash ID, unit ID]" corresponding to the entity to the "text_unit_ids" field of the newly added record; Step 3: Relational data construction, specifically: Read the relationship data in the create_final_relationships.parquet file of the GraphRAG large language model; According to the storage format of create_final_relationships.parquet described in the GraphRAG large language model, create relationship records in this storage format for the target triple data; Append the relationship records to the create_final_relationships.parquet file and save it.

[0009] In one embodiment, the most relevant topics and other topics among the output topics are determined based on the topic-vocabulary distribution matrix generated when the LDA topic model training is completed, and the implementation method includes: Based on the topic-vocabulary distribution matrix, the top n word segments with probability ranking under each output topic of the LDA topic model are obtained; n is an empirical value; Based on the top n word segments with probability ranking under each output topic, the most relevant topic in each output topic of the LDA topic model is determined to be the strongly relevant topic, and the other topics in each output topic of the LDA topic model are determined to be weakly relevant topics.

[0010] In one embodiment, based on the top n word segments with probability ranking under each output topic, the most relevant topic among the output topics of the LDA topic model is determined as the strongly relevant topic. The specific implementation method includes: For each topic in each output topic, calculate the sum of the probabilities of the top n word segments; For each output topic, the output topic with the largest sum of calculated probabilities is taken as the strongly related topic.

[0011] In one embodiment, a method for training an LDA topic model includes: Prepare training set; Construct an LDA topic model and set model parameters to obtain an initial LDA topic model; setting model parameters includes setting the number of LDA topics n_components, document-topic prior parameter doc_topic_prior and topic-word prior parameter topic_word_prior, and also setting the random seed random_state, where n_components=2, doc_topic_prior=0.001, topic_word_prior=1.0, random_state=0; Use the training set to train the initial LDA topic model to obtain the trained LDA topic model.

[0012] In one embodiment, the method further comprises: Visualize the generated knowledge graph.

[0013] In a second aspect, the present invention provides a knowledge graph generation system, the system comprising: A first uploading unit is used to upload target text data, wherein the target text data includes a plurality of texts; the target text data is pre-collected data for generating a knowledge graph; A clustering unit, used for clustering the target text data to obtain a plurality of clusters; A text extraction unit is used to extract a text from each cluster obtained by the clustering unit to obtain a set of target texts; A sample construction unit is used to construct a task sample of GraphRAG prompt words for each target text obtained by the text extraction unit to obtain a set of target task samples; The sample submission unit is used to submit the target task samples obtained by the sample construction unit to the GraphRAG large language model; An entity relationship extraction unit is used to use the GraphRAG large language model to extract entities and relationships in the target text data according to the target task examples to generate a knowledge graph.

[0014] In one embodiment, the system further comprises: A second uploading unit is used to upload an external knowledge base; the external knowledge base is any knowledge graph collected in the corresponding field; A data extraction unit, used for extracting triple data from the external knowledge base; A probability calculation unit is used to input each triple data of the extracted external knowledge base into a pre-trained LDA topic model for calculation, and correspondingly output the probability of each triple data of the external knowledge base under a strongly related topic and each weakly related topic; wherein the strongly related topic and each weakly related topic are the most related topic and other topics in the output topics determined based on the topic-vocabulary distribution matrix generated when the LDA topic model training is completed; The data processing unit is used to determine, for each triple data of the external knowledge base, whether its probability under the strongly related topic is greater than its probability under each weakly related topic: if so, add the triple data to the parquet file of the GraphRAG large language model; if not, discard the triple data.

[0015] In a third aspect, the present invention provides a terminal device, comprising: processor; A memory for storing execution instructions of the processor; The processor is configured to execute the methods described in the above aspects.

[0016] It can be seen from the above technical solutions that the present invention has the following advantages: (1) The present invention clusters the target text data to obtain a number of clusters, and extracts a text from each of the obtained clusters to obtain a group of target texts, and then uses each target text to construct a task example of a GraphRAG prompt word, and then submits each constructed task example to the GraphRAG large language model, and then uses the GraphRAG large language model to extract the entities and relationships in the above target text data to generate a knowledge graph based on the submitted task examples and target task examples, thereby avoiding dependence on manually annotated data sets, and to a certain extent helping to reduce the requirements for data set quality. At the same time, it can also enable the present invention to submit diverse examples to GraphRAG, which in turn helps to improve the accuracy of GraphRAG in completing tasks, thereby improving the accuracy of the knowledge graph generated by the present invention.

[0017] (2) The present invention can also add triple data from an external knowledge base to the parquet file of the GraphRAG large language model, which makes it convenient for users to complete the knowledge graph. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solution of the present invention, the accompanying drawings required for use in the description will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.

[0019] Figure 1 is a schematic flow chart of a method according to an embodiment of the present invention.

[0020] Figure 2 is a schematic block diagram of a system according to an embodiment of the present invention.

[0021] Figure 3 is a schematic block diagram of a terminal device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0022] GraphRAG large language model: Graph-based Retrieval-Augmented Generation large language model, a model that combines knowledge graph (Knowledge Graph) and retrieval-augmented generation (RAG) technology.

[0023] The knowledge graph generation method provided by the present invention incorporates clustering, and constructs a task sample of GraphRAG prompt words based on the clustering results and submits it to the GraphRAG large language model for subsequent use of the GraphRAG large language model to generate a knowledge graph. In addition, the present invention also supports uploading external knowledge bases for graph completion.

[0024] The specific execution steps of the knowledge graph generation method will be described in detail below. For the purpose of illustration rather than limitation, specific details such as specific system structures and technologies are proposed to provide a thorough understanding of the embodiments of the present invention. However, it should be clear to those skilled in the art that the present invention can also be implemented in other embodiments without these specific details.

[0025] It should be understood that when used in the present specification, the term "comprising" indicates the presence of the described features, integral bodies, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integral bodies, steps, operations, elements, components and / or their collections. The terms "comprising", "including", "having" and their variations all mean "including but not limited to", unless otherwise specifically emphasized.

[0026] The phrases such as "one embodiment" or "some embodiments" described in the present invention mean that the specific features, structures or characteristics described in the embodiment are included in one or more embodiments of the present invention. Therefore, the phrases such as "in one embodiment", "in some embodiments", "in some other embodiments", "in some other embodiments" etc. that appear in different places in the present invention do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways.

[0027] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0028] Please refer to Figure 1 , which shows a knowledge graph generation method provided by an embodiment of the present invention, which includes the following steps (S101 to S106).

[0029] Step S101: Upload target text data.

[0030] The target text data is pre-collected data used to generate the knowledge graph.

[0031] Taking the field of innovation and entrepreneurship as an example, the target text data is pre-collected data in the field of innovation and entrepreneurship used to generate a knowledge graph.

[0032] It can be understood that the above target text data includes multiple texts.

[0033] It is understandable that those skilled in the art can collect relevant text data according to actual needs. Exemplarily, when collecting target text data, the collected data can be divided into categories first, and then data can be collected according to these categories. For example, taking the generation of the above-mentioned innovation and entrepreneurship knowledge graph in the field of innovation and entrepreneurship as an example, data can be collected from ten aspects: start-up enterprise information, innovation and entrepreneurship activities, innovation and entrepreneurship information, innovation technology patents, entrepreneurship incubators, entrepreneurship services, entrepreneurship courses, market dynamics, investment and financing, and expert talents. Specifically, the Secret Tower AI search can be used to search the data sources of the above ten aspects to determine which websites (such as Entrepreneurship Helper, IT Orange, etc.) to crawl data from. After that, the Octopus Collector can be used to set the request header, page turning rules, etc. to collect website data. After that, the collected website data is pre-processed, such as data cleaning first, and then data desensitization. Data cleaning may include but is not limited to: removing duplicate and irrelevant data (such as irrelevant to the knowledge graph to be generated). Data desensitization: Processing sensitive fields in the data.

[0034] It can be understood that the target text data is preprocessed data, and the preprocessing includes data cleaning and data desensitization.

[0035] Step S102: clustering the target text data to obtain a number of clusters.

[0036] Optionally, the target text data is clustered using a k-means clustering algorithm.

[0037] In specific implementation, the silhouette coefficient and intra-cluster square error can be observed by adjusting the k value to explore the impact of different k values ​​on the clustering effect.

[0038] In this embodiment, k=10 is used as the number of clusters, and thus 10 clusters can be obtained by clustering the target text data using the k-means clustering algorithm.

[0039] Step S103: extract a text from each obtained cluster to obtain a group of target texts.

[0040] In this embodiment, one text is randomly extracted from each of the 10 obtained clusters to obtain 10 target texts.

[0041] It should also be noted that in order to increase the representativeness of the extracted files, in the specific implementation, the centroid nearest neighbor method and / or the central tendency index can also be used to determine the most representative text in each cluster, and then extract it as the text extracted from the corresponding cluster.

[0042] Step S104: construct a task example of GraphRAG prompt words using each target text to obtain a set of target task examples.

[0043] GraphRAG prompt words, that is, the prompt words (prompt) of the GraphRAG large language model.

[0044] In specific implementation, the task examples include target text, entity recognition and attribute extraction, and relationship construction, where entity recognition and attribute extraction are used to specify the entities to be recognized from the target text in the sample (hereinafter referred to as "target entities"), and relationship construction is used to specify the relationships between the target entities to be constructed.

[0045] Taking the generation of the above-mentioned innovation and entrepreneurship knowledge graph as an example, the target entities may include: technology, product, service, platform, management, culture, economy, research, consulting, investment, founder, CEO, incubator, AI, startup, survey, user, financing, innovation, learning, decision making, a total of 21 entity categories. The relationship construction can be designed by technicians in this field according to actual needs.

[0046] Step S105: Submit the obtained target task examples to the GraphRAG large language model.

[0047] This step S105 can provide representative task examples for the large model, which helps to increase the diversity and coverage of examples in the GraphRAG large language model, and then deepen the model's understanding of the task framework and expected output style.

[0048] This step submits diverse samples to GraphRAG, which also helps to improve the efficiency of GraphRAG in completing tasks, and then helps to improve the efficiency of the present invention in generating knowledge graphs.

[0049] Step S106: Utilizing the GraphRAG large language model, based on the target task examples, extract entities and relationships in the target text data to generate a knowledge graph.

[0050] In the specific implementation, the prompt is constructed through the GraphRAG method, and structured instructions are written to require the large model to complete the task of extracting entities and relationships from the input text. The output format is specified based on the target task examples, and finally the GraphRAG large language model is used to identify the entities and corresponding relationships in the input text, and a graph is generated using the identified entities and relationships and stored in the model's parquet file.

[0051] In some embodiments of the present invention, the above method further includes the following steps S201 to S204.

[0052] Step S201: Upload external knowledge base.

[0053] The external knowledge base is a knowledge graph collected in the corresponding field, that is, the external knowledge base is other graphs related to the graph generated in step S106.

[0054] Step S202: extracting triple data from the external knowledge base.

[0055] Step S203: input each triple data of the extracted external knowledge base into the pre-trained LDA topic model for operation, and output the probability of each triple data of the external knowledge base under the strongly related topic and each weakly related topic accordingly.

[0056] In this embodiment, the strongly related topics and the weakly related topics are the most related topics and other topics in the output topics determined based on the topic-vocabulary distribution matrix generated when the LDA topic model training is completed.

[0057] Step S204: for each triple data of the external knowledge base, determine whether its probability under the strongly related topic is greater than its probability under each of the weakly related topics: If so, the triple data is added to the parquet file of the GraphRAG large language model; if not, the triple data is discarded.

[0058] In this embodiment, the triple data is added to the target knowledge graph, that is, the triple data is stored in the parquet file of the GraphRAG large language model.

[0059] It can be understood that the external knowledge base mentioned above can be any other graph related to the graph generated in step S106.

[0060] The design and use of the above steps S201 to S204 are helpful to help users complete the graph, thereby helping users build a more comprehensive knowledge graph. In addition, it also helps users to complete the graph in the future, thereby overcoming the shortcoming that GraphRAG does not support incremental updates, and dynamically adding knowledge triples to the graph, thereby updating the content of the knowledge graph.

[0061] During subsequent use, when new knowledge data needs to be added to the graph, even if there is only one triple or several triples that need to be added, these triples that need to be added can be directly uploaded and added in the form of an external knowledge base, which is easy to use and facilitates timely updating of the knowledge graph content.

[0062] The parquet files include a create_final_text_units.parquet data file, a create_final_entities.parquet entity file, and a create_final_relationships.parquet file.

[0063] The create_final_text_units.parquet data file, create_final_entities.parquet entity file, and create_final_relationships.parquet file are the default parquet files that GraphRAG uses to store the knowledge graph it generates.

[0064] In some embodiments of the present invention, adding triple data to the parquet file of the GraphRAG large language model includes: Step 1: Generate and store text units, specifically: Reversely generate text content from the target triple data, which is recorded as the original text content; the target triple data is the triple data to be added to the parquet file of the GraphRAG large language model; Calculate the MD5 hash value of the original text content and use it as the first hash ID; Get the text length of the original text content; Create a text unit including the first hash ID, the original text content, and the text length (created according to the default format in the create_final_text_units.parquet data file of GraphRAG); Append the text unit to the create_final_text_units.parquet data file of the GraphRAG large language model and save it; Step 2: Establish the association between entity and text unit, specifically: Calculate the MD5 hash value of the text unit to obtain a unit ID; Calculate the MD5 hash value of the first entity of the target triple data to obtain the entity hash ID of the first entity; Calculate the MD5 hash value of the tail entity of the target triple data to obtain the entity hash ID of the tail entity; Traverse the first entity and the last entity of the target triple data; For the currently traversed entity, perform the following steps M respectively: Step M, locate the record whose "name" field in the create_final_entities.parquet entity file of the GraphRAG large language model is the name of the entity of the target triple data, and add the "[entity hash ID, unit ID]" corresponding to the entity under the "text_unit_ids" field of the record; if the entity of the target triple data cannot be located, add a new record in the create_final_entities.parquet entity file, then add the name of the entity to the "name" field of the newly added record, and add the "[entity hash ID, unit ID]" corresponding to the entity to the "text_unit_ids" field of the newly added record; Step 3: Relational data construction, specifically: Read the relationship data in the create_final_relationships.parquet file of the GraphRAG large language model; According to the storage format of create_final_relationships.parquet described in the GraphRAG large language model, create relationship records in this storage format for the target triple data; Append the relationship records to the create_final_relationships.parquet file and save it.

[0065] Optionally, the method of reversely generating text content from the target triple data may be: sequentially concatenating the elements of the target triple data to form a text content, with the elements separated by commas.

[0066] When the first entity of the target triple data is traversed, step M is specifically implemented as follows: Locate the record whose "name" field in the create_final_entities.parquet entity file of the GraphRAG large language model is the name of the first entity of the target triple data, and add "[entity hash ID, unit ID]" corresponding to the first entity under the "text_unit_ids" field of the record; If the first entity cannot be located, a new record is added to the create_final_entities.parquet entity file, and then the name of the first entity is added to the "name" field of the newly added record, and the "[entity hash ID, unit ID]" corresponding to the first entity is added to the "text_unit_ids" field of the newly added record.

[0067] It can be understood that the entity hash ID of the first entity is recorded as the second hash ID, and the entity hash ID of the last entity is recorded as the third hash ID. The "[entity hash ID, unit ID]" corresponding to the first entity is "[second hash ID, unit ID]", and the "[entity hash ID, unit ID]" corresponding to the last entity is "[third hash ID, unit ID]".

[0068] When the tail entity of the target triple data is currently traversed, step M includes: Locate the record whose "name" field in the create_final_entities.parquet entity file of the GraphRAG large language model is the name of the tail entity of the target triple data, and add "[entity hash ID, unit ID]" corresponding to the tail entity under the "text_unit_ids" field of the record; If the tail entity cannot be located, a new record is added to the create_final_entities.parquet entity file, and then the name of the tail entity is added to the "name" field of the newly added record, and the "[entity hash ID, unit ID]" corresponding to the tail entity is added to the "text_unit_ids" field of the newly added record.

[0069] It can be understood that the entity hash ID of the first entity is recorded as the second hash ID, and the entity hash ID of the last entity is recorded as the third hash ID. The "[entity hash ID, unit ID]" corresponding to the first entity is "[second hash ID, unit ID]", and the "[entity hash ID, unit ID]" corresponding to the last entity is "[third hash ID, unit ID]".

[0070] In some embodiments of the present invention, the most relevant topics and other topics in the output topics are determined based on the topic-vocabulary distribution matrix generated when the LDA topic model training is completed, and the implementation method includes: Based on the topic-vocabulary distribution matrix, the top n word segments with probability ranking under each output topic of the LDA topic model are obtained; n is an empirical value; Based on the top n word segments with probability ranking under each output topic, the most relevant topic in each output topic of the LDA topic model is determined to be the strongly relevant topic, and the other topics in each output topic of the LDA topic model are determined to be weakly relevant topics.

[0071] Exemplarily, based on the topic-vocabulary distribution matrix, the top n word segments with probability ranking under each output topic of the LDA topic model are obtained, and the implementation method is: For each topic involved in the topic-vocabulary distribution matrix, extract all the sub-words and their corresponding probabilities; For each topic involved in the topic-vocabulary distribution matrix, all extracted word segments are sorted in descending order according to the probability of extraction; For each topic involved in the topic-vocabulary distribution matrix, select the top n word segments with the highest probability ranking.

[0072] n is a positive integer, and its specific value can be set by those skilled in the art based on experience.

[0073] In this embodiment, n=5.

[0074] Optionally, based on the top n word segments with probability ranking under each output topic, determining the most relevant topic among each output topic of the LDA topic model as the strongly relevant topic, the specific implementation method includes: For each topic in each output topic, calculate the sum of the probabilities of the top n word segments; For each output topic, the output topic with the largest sum of calculated probabilities is taken as the strongly related topic.

[0075] Optionally, the training method of the LDA topic model includes: Step 1: Prepare the training set; Step 2: Build an LDA topic model and set model parameters to obtain an initial LDA topic model; setting model parameters includes setting the number of LDA topics n_components, document-topic prior parameter doc_topic_prior and topic-word prior parameter topic_word_prior, and also setting the random seed random_state, where n_components=2, doc_topic_prior=0.001, topic_word_prior=1.0, random_state=0; Step three: Use the training set to train the initial LDA topic model to obtain the trained LDA topic model.

[0076] The above step 1 includes: Prepare text data to obtain a text dataset, which is a training set.

[0077] Prepare text data, that is, collect text data related to the knowledge graph to be generated in the target domain, where the target domain is the domain of the knowledge graph to be generated.

[0078] The above step 2 specifically includes: An LDA topic model is created, and parameters of the created LDA topic model are set to obtain an initial LDA topic model.

[0079] The parameter setting includes setting the number of LDA topics, prior parameters and random seeds.

[0080] Create an LDA topic model: Instantiate the LatentDirichletAllocation class.

[0081] The prior parameters include the document-topic distribution α and the topic-word distribution η.

[0082] Prior parameters are denoted by doc_topic_prior and topic_word_prior. Random seed is denoted by random_state. Number of topics is denoted by n_components, document-topic distribution α is denoted by doc_topic_prior, and topic-word distribution η is denoted by topic_word_prior.

[0083] In this embodiment, the number of LDA topics n_components is set to 2, the prior parameters doc_topic_prior is set to 0.001 and topic_word_prior is set to 1.0, and the random seed random_state is fixed to 0 to ensure the reproducibility of the experiment.

[0084] The above step three specifically includes the following steps 1101 to 1104.

[0085] Step 1101, segment the text data.

[0086] The word segmentation described in step 1101 includes resource loading, text cleaning and word segmentation.

[0087] Resource loading includes stop word list loading and custom dictionary loading.

[0088] Users can prepare stop word lists and custom dictionaries in advance and store them locally.

[0089] Stop word list loading: Read stop words from the local file stop word list and build a stop word set.

[0090] Custom dictionary loading: Load custom dictionaries from local machines to ensure correct segmentation of specific words or professional terms required by users.

[0091] Optionally, users can load custom dictionaries through the jieba.load_userdict() function of the jieba word segmentation library.

[0092] Text cleaning and word segmentation: Read the text in the text dataset one by one and generate a document list lines; Use the jieba.lcut() function of the jieba word segmentation library to perform precise word segmentation on the text in the document list lines to obtain the first word list; According to the stop word set, the stop words in the first word list are removed to obtain a filtered word list filtered_words.

[0093] After using jieba.load_userdict() to load the custom dictionary, use the jieba.lcut() function to segment the text. Jieba will segment the text based on the words in the custom dictionary to get more accurate results.

[0094] It can be understood that in the document list lines, each line represents a text. In the first word list, each line represents the word segmentation result of a text. In the filtered word list filtered_words, each line represents the filtered word segmentation result of a text.

[0095] Step 1102 , concatenate each line of words in the word list filtered_words into strings separated by spaces (ie, each word is separated by a space), and generate a tokenized_docs list.

[0096] In the tokenized_docs list, each sublist corresponds to a word segmentation of the text.

[0097] Step 1103, convert the tokenized_docs list into a sparse matrix X.

[0098] In the sparse matrix X, rows represent documents, columns represent terms, and the values ​​in the matrix represent the frequency of terms in the documents.

[0099] Step 1104: Based on the sparse matrix X, the initial LDA topic model is trained to obtain a trained LDA topic model.

[0100] The trained LDA topic model can output the document-topic distribution matrix.

[0101] In the document-topic distribution matrix, each row represents a text, each column represents a topic, and the matrix elements represent the probability distribution of the text on each topic.

[0102] In the topic-word distribution matrix, each row represents a topic, each column represents a word (i.e., a segmentation word), and the matrix elements represent the probability distribution of the word on each topic.

[0103] The method of this embodiment can retain strongly related triples in the external knowledge base through the trained LDA topic model, integrate them into the model, and realize the knowledge completion of the graph.

[0104] In some embodiments of the present invention, the method further includes: visualizing the generated knowledge graph.

[0105] Visualizing the generated knowledge graph is more intuitive.

[0106] It should be understood that the order of execution of the steps in the above embodiment does not necessarily mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of the present invention.

[0107] Figure 2 A knowledge graph generation system provided by the present invention includes: A first uploading unit 201 is used to upload target text data, wherein the target text data includes a plurality of texts; the target text data is pre-collected data for generating a knowledge graph; A clustering unit 202, used to cluster the target text data to obtain a plurality of clusters; The text extraction unit 203 is used to extract a text from each cluster obtained by the clustering unit to obtain a group of target texts; A sample construction unit 204 is used to construct a task sample of GraphRAG prompt words for each target text obtained by the text extraction unit to obtain a set of target task samples; A sample submission unit 205 is used to submit each target task sample obtained by the sample construction unit to the GraphRAG large language model; The entity relationship extraction unit 206 is used to use the GraphRAG large language model to extract entities and relationships in the target text data according to the target task examples to generate a knowledge graph.

[0108] In some embodiments of the present invention, the system further comprises: A second uploading unit is used to upload an external knowledge base; the external knowledge base is any knowledge graph collected in the corresponding field; A data extraction unit, used for extracting triple data from the external knowledge base; A probability calculation unit is used to input each triple data of the extracted external knowledge base into a pre-trained LDA topic model for calculation, and correspondingly output the probability of each triple data of the external knowledge base under a strongly related topic and each weakly related topic; wherein the strongly related topic and each weakly related topic are the most related topic and other topics in the output topics determined based on the topic-vocabulary distribution matrix generated when the LDA topic model training is completed; The data processing unit is used to determine, for each triple data of the external knowledge base, whether its probability under the strongly related topic is greater than its probability under each weakly related topic: if so, add the triple data to the parquet file of the GraphRAG large language model; if not, discard the triple data.

[0109] In some embodiments of the present invention, the system further comprises: The graph visualization unit is used to visualize the generated knowledge graph.

[0110] In some embodiments of the present invention, the graph visualization unit includes a first visualization module, a second visualization module and a third visualization module.

[0111] When using it, the user selects different visualization modules to obtain the map in the corresponding visualization form.

[0112] Optionally, the first visualization module is used to read the entities and relationships of the generated knowledge graph into a DataFrame using pandas, and then export it to an Excel format.

[0113] Optionally, the second visualization module is used to visualize the generated knowledge graph using plotly.

[0114] In specific implementation, the nodes in the network graph can be laid out through the spring_layout method of networkx based on the Fruchterman-Reingold force-directed algorithm; the properties of nodes and edges can be set through the Scatter of plotly, including position, color, text, etc.; the graph layout can be displayed through the Figure method.

[0115] Optionally, the third visualization module is used to visualize the generated knowledge graph using Neo4j.

[0116] By building the resulting graph into a graph database using Neo4j, the relevant nodes of the graph can be effectively stored and managed. In particular, the query speed for nodes (i.e. entities) and relationships is fast, the nodes can be moved, and the interactivity is strong.

[0117] The embodiment of the knowledge graph generation system of this embodiment belongs to the same inventive concept as the knowledge graph generation method of the above-mentioned embodiments. For details not described in detail in the embodiment of the knowledge graph generation system, please refer to the embodiment of the knowledge graph generation method in the previous text.

[0118] Figure 3 The present invention provides a schematic diagram of the structure of a terminal 300 according to an embodiment of the present invention. The terminal 300 may be used to execute the method according to an embodiment of the present invention.

[0119] The terminal 300 may include: a processor 310, a memory 320 and a communication unit 330. These components communicate via one or more buses. Those skilled in the art will appreciate that the server structure shown in the figure does not limit the present invention, and it may be a bus structure or a star structure, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.

[0120] The memory 320 can be used to store the execution instructions of the processor 310, and the memory 320 can be implemented by any type of volatile or non-volatile storage terminal or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk. When the execution instructions in the memory 320 are executed by the processor 310, the terminal 300 can perform some or all of the steps in the following method embodiments.

[0121] The processor 310 is the control center of the storage terminal, and uses various interfaces and lines to connect various parts of the entire electronic terminal. It runs or executes software programs and / or modules stored in the memory 320, and calls data stored in the memory to perform various functions of the electronic terminal and / or process data. The processor can be composed of an integrated circuit (IC), for example, it can be composed of a single packaged IC, or it can be composed of a plurality of packaged ICs with the same or different functions. For example, the processor 310 can include only a central processing unit (CPU). In the embodiment of the present invention, the CPU can be a single computing core or multiple computing cores.

[0122] The communication unit 330 is used to establish a communication channel so that the storage terminal can communicate with other terminals, receive user data sent by other terminals or send user data to other terminals.

[0123] The technical effects that can be achieved by this embodiment can be found in the above description and will not be repeated here.

[0124] The same and similar parts between the various embodiments in this specification can be referenced to each other.

[0125] The above description of the disclosed embodiments enables one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A knowledge graph generation method, characterized in that the method include: Upload target text data, wherein the target text data includes multiple texts; the target text data is pre-collected data for generating a knowledge graph; Cluster the target text data to obtain several clusters; Extract a text from each obtained cluster to obtain a set of target texts; Use each target text to construct a task example of GraphRAG prompt words to obtain a set of target task examples; Submit the obtained target task examples to the GraphRAG large language model; Utilizing the GraphRAG large language model, based on the target task examples, entities and relationships in the target text data are extracted to generate a knowledge graph.

2. The knowledge graph generation method according to claim 1, characterized in that: The method also includes: Upload an external knowledge base; the external knowledge base is a knowledge graph collected in the corresponding field; Extracting triple data from the external knowledge base; Input each triple data of the extracted external knowledge base into the pre-trained LDA topic model for operation, and output the probability of each triple data of the external knowledge base under the strongly related topics and each weakly related topic; wherein the strongly related topics and each weakly related topics are the most related topics and other topics in the output topics determined based on the topic-vocabulary distribution matrix generated when the LDA topic model training is completed; For each triple data of the external knowledge base, determine whether its probability under the strongly related topic is greater than its probability under each of the weakly related topics: If so, the triple data is added to the parquet file of the GraphRAG large language model; if not, the triple data is discarded.

3. The knowledge graph generation method according to claim 2, characterized in that: The parquet files include create_final_text_units.parquet data files, create_final_entities.parquet entity files and create_final_relationships.parquet files; Add triple data to the parquet file of the GraphRAG large language model, including: Step 1: Generate and store text units, specifically: Reversely generate text content from the target triple data, which is recorded as the original text content; the target triple data is the triple data to be added to the parquet file of the GraphRAG large language model; Calculate the MD5 hash value of the original text content and use it as the first hash ID; Get the text length of the original text content; Creating a text unit including the first hash ID, the original text content, and the text length; Append the text unit to the create_final_text_units.parquet data file and save it; Step 2: Establish the association between entity and text unit, specifically: Calculate the MD5 hash value of the text unit to obtain a unit ID; Calculate the MD5 hash value of the first entity of the target triple data to obtain the entity hash ID of the first entity; Calculate the MD5 hash value of the tail entity of the target triple data to obtain the entity hash ID of the tail entity; Traverse the first entity and the last entity of the target triple data; For the currently traversed entity, perform the following steps M respectively: Step M, locate the record whose "name" field in the create_final_entities.parquet entity file is the name of the entity of the target triple data, and add "[entity hash ID, unit ID]" corresponding to the entity under the "text_unit_ids" field of the record; if the entity of the target triple data cannot be located, add a new record in the create_final_entities.parquet entity file, then add the name of the entity to the "name" field of the newly added record, and add "[entity hash ID, unit ID]" corresponding to the entity to the "text_unit_ids" field of the newly added record; Step 3: Relational data construction, specifically: Read the relationship data in the create_final_relationships.parquet file of the GraphRAG large language model; According to the storage format of create_final_relationships.parquet described in the GraphRAG large language model, create relationship records in this storage format for the target triple data; Append the relationship records to the create_final_relationships.parquet file and save it.

4. The knowledge graph generation method according to claim 2, characterized in that: The most relevant topics and other topics among the output topics determined based on the topic-vocabulary distribution matrix generated when the LDA topic model training is completed, and the implementation method includes: Based on the topic-vocabulary distribution matrix, the top n word segments with probability ranking under each output topic of the LDA topic model are obtained; n is an empirical value; Based on the top n word segments with probability ranking under each output topic, the most relevant topic in each output topic of the LDA topic model is determined to be the strongly relevant topic, and the other topics in each output topic of the LDA topic model are determined to be weakly relevant topics.

5. The knowledge graph generation method according to claim 4, characterized in that: Based on the top n word segments with probability ranking under each output topic, the most relevant topic among the output topics of the LDA topic model is determined as the strongly relevant topic. The specific implementation method includes: For each topic in each output topic, calculate the sum of the probabilities of the top n word segments; For each output topic, the output topic with the largest sum of calculated probabilities is taken as the strongly related topic.

6. The knowledge graph generation method according to claim 1, characterized in that: The training method of the LDA topic model includes: Prepare training set; Construct an LDA topic model and set model parameters to obtain an initial LDA topic model; setting model parameters includes setting the number of LDA topics n_components, document-topic prior parameter doc_topic_prior and topic-word prior parameter topic_word_prior, and also setting the random seed random_state, where n_components=2, doc_topic_prior=0.001, topic_word_prior=1.0, random_state=0; Use the training set to train the initial LDA topic model to obtain the trained LDA topic model.

7. The knowledge graph generation method according to claim 1, characterized in that: The method also includes: Visualize the generated knowledge graph.

8. A knowledge graph generation system, characterized in that: The system includes: A first uploading unit is used to upload target text data, wherein the target text data includes a plurality of texts; the target text data is pre-collected data for generating a knowledge graph; A clustering unit, used for clustering the target text data to obtain a plurality of clusters; A text extraction unit is used to extract a text from each cluster obtained by the clustering unit to obtain a set of target texts; A sample construction unit is used to construct a task sample of GraphRAG prompt words for each target text obtained by the text extraction unit to obtain a set of target task samples; The sample submission unit is used to submit the target task samples obtained by the sample construction unit to the GraphRAG large language model; An entity relationship extraction unit is used to use the GraphRAG large language model to extract entities and relationships in the target text data according to the target task examples to generate a knowledge graph.

9. The knowledge graph generation system according to claim 8, characterized in that: The system also includes: A second uploading unit is used to upload an external knowledge base; the external knowledge base is any knowledge graph collected in the corresponding field; A data extraction unit, used for extracting triple data from the external knowledge base; A probability calculation unit is used to input each triple data of the extracted external knowledge base into a pre-trained LDA topic model for calculation, and correspondingly output the probability of each triple data of the external knowledge base under a strongly related topic and each weakly related topic; wherein the strongly related topic and each weakly related topic are the most related topic and other topics in the output topics determined based on the topic-vocabulary distribution matrix generated when the LDA topic model training is completed; The data processing unit is used to determine, for each triple data of the external knowledge base, whether its probability under the strongly related topic is greater than its probability under each weakly related topic: if so, add the triple data to the parquet file of the GraphRAG large language model; if not, discard the triple data.

10. A terminal device, characterized in that: include: processor; A memory for storing execution instructions of the processor; The processor is configured to execute the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Construction method and device of power field knowledge graph and electronic equipment

    CN115587190A

  • Theme-based knowledge base construction method, storage medium and electronic equipment

    CN116467460A

  • Knowledge graph construction method and system based on large language model

    CN117150050A

  • Nuclear power DCS operation and maintenance knowledge graph construction system and method based on large language model

    CN118607632A

  • Knowledge graph construction method and device for target culture resources and electronic equipment

    CN119539044A