Knowledge graph construction method and device and medium

Through screening and large language model processing methods, a high-quality knowledge graph is constructed, which solves the problems of low efficiency and unstable construction of quality in the existing technology, and achieves more efficient and accurate knowledge graph construction, which is especially suitable for vertical fields.

CN120104608APending Publication Date: 2025-06-06ZHEJIANG LAB

Patent Information

Application Number
CN202510595178.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-09
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The existing knowledge graph construction methods are inefficient and unstable in quality, especially in vertical fields, with low coverage and practicality, making it difficult to support complex domain knowledge requirements.

Method used

By collecting pending texts, filtering out target texts that meet preset conditions for the correlation with the specified domain, using the pre-constructed target large language model to build initial triplets, and obtaining high-quality target triplets through filtering and verification, and finally building a knowledge graph.

Benefits of technology

It improves the quality and accuracy of the knowledge graph, reduces the impact of noisy data, improves construction efficiency, and ensures the domain specificity and practicality of the knowledge graph.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120104608A_ABST
    Figure CN120104608A_ABST
Patent Text Reader

Abstract

The invention discloses a knowledge graph construction method and apparatus, and a medium. The method comprises the steps of collecting a to-be-processed text for constructing a knowledge graph in a specified field; screening out a target text from the to-be-processed text, wherein the correlation between the target text and the specified field meets a preset condition; processing the target text through a pre-constructed target large language model to construct an initial triple; filtering the initial triple to obtain a target triple; and constructing the knowledge graph through the target triple. Therefore, through strict correlation screening conditions, texts irrelevant to the specified field are removed, the influence of noise data on knowledge graph construction is reduced, and the quality and accuracy of the knowledge graph are improved. On the basis of the target texts, the knowledge triples are efficiently and accurately extracted from a large number of target texts by means of the powerful language understanding and generating capacity of the large language model, the labor cost overhead of manual extraction is reduced, and the knowledge graph construction efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of knowledge graph construction, and in particular to a method, device and medium for constructing a knowledge graph. Background Art

[0002] Knowledge Graph (KG), as a structured knowledge representation, displays entities, attributes and their relationships in a graph structure. It plays a vital role in the field of natural language processing (NLP) and provides a solid foundation for machines to understand semantics and reasoning. It is therefore widely used in question-answering systems, recommendation systems, information retrieval and other tasks.

[0003] The current knowledge graph can be constructed based on pre-set rules, but this method requires domain experts to spend a lot of time and energy on manual verification, resulting in low construction efficiency and uneven quality of knowledge graphs. Another feasible way is to construct it through a statistical model, which can effectively improve the construction efficiency. However, the construction of a statistical model can only extract one relationship from a single sentence, which will lose a lot of factual information in long texts and affect the quality of the knowledge graph. In addition, the current knowledge graph has low coverage and practicality in vertical fields such as geology and medical health, and it is difficult to support complex domain knowledge needs.

[0004] It can be seen that how to efficiently construct high-quality knowledge graphs to meet the needs of knowledge graphs in different fields is an urgent problem that needs to be solved by technical personnel in this field. Summary of the invention

[0005] In view of this, one aspect of the present application provides a method for constructing a knowledge graph, the method comprising: Collect the text to be processed for building the knowledge graph in the specified field; Filtering out target texts whose relevance to the specified field meets preset conditions from the texts to be processed; Processing the target text through a pre-built target large language model to construct initial triples; After filtering the initial triples, a target triple is obtained; The knowledge graph is constructed through the target triples.

[0006] Optionally, the step of selecting target texts whose relevance to the designated field meets preset conditions from the texts to be processed includes: Get pre-built relevance word projects; Based on the relevance prompt word project, the relevance score between the text to be processed and the specified field is determined through the target large language model; wherein the larger the relevance score, the stronger the relevance between the text to be processed and the specified field; The text whose relevance score is greater than a score threshold is used as the target text.

[0007] Optionally, after filtering the initial triples, obtaining target triples includes: Get a pre-built verification word project; Based on the verification prompt word project, determining whether the target text supports the initial triple through the target large language model; If supported, retain the initial triplet; If not supported, a supporting recheck signal is obtained; and based on the supporting recheck signal, it is determined whether to retain the initial triplet.

[0008] Optionally, after filtering the initial triples, obtaining target triples includes: Get pre-built data cleaning prompt word projects; Based on the data cleaning prompt word project, the initial triples are subjected to semantic deduplication processing and semantic normalization processing through the target large language model to obtain the target triples.

[0009] Optionally, after filtering the initial triples, obtaining target triples includes: Get pre-built conflict word projects; Based on the conflict prompt word project, determining the triple pairs having logical conflicts in the initial triples through the target large language model; Obtaining a conflict recheck signal of the triplet pair; If the conflict recheck signal indicates that the triplet pair conflicts, then the triplet pair is eliminated; If the conflict secondary check signal indicates that the triplet pair does not conflict, retaining the triplet pair; An initial triplet after conflict resolution based on the conflict recheck signal is used as the target triplet.

[0010] Optionally, the processing of the target text by using a pre-built target large language model to construct an initial triple includes: Get pre-built target cue word projects; Based on the target prompt word project, entity fields and relationship fields are extracted from the target text through the target large language model; and the initial triples are constructed according to the entity fields and the relationship fields.

[0011] Optionally, constructing the knowledge graph through the target triples includes: The target triples are constructed into the knowledge graph by a target graph structure tool; The knowledge graph is stored in a target knowledge graph library.

[0012] Another aspect of the present application provides a knowledge graph construction device, the device comprising: A module for collecting texts to be processed, used to collect texts to be processed for building a knowledge graph in a specified field; A target text screening module is used to screen out target texts whose relevance to the specified field meets preset conditions from the texts to be processed; A triple construction module, used for processing the target text through a pre-constructed target large language model to construct initial triples; A triplet filtering module, used for filtering the initial triplet to obtain a target triplet; A knowledge graph construction module is used to construct the knowledge graph through the target triples.

[0013] Another aspect of the present application provides a device for constructing a knowledge graph, including a memory and a processor, wherein the memory stores a computer program that can be executed on the processor, and when the processor executes the program, the steps of the method for constructing the knowledge graph are implemented.

[0014] Another aspect of the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method for constructing the knowledge graph.

[0015] The present application provides a method, device and medium for constructing a knowledge graph, which has the following beneficial effects: through strict relevance screening conditions, texts irrelevant to the specified field are removed, the impact of noise data on the construction of the knowledge graph is reduced, and the quality and accuracy of the knowledge graph is improved. And the screened target text is closer to the specified field, ensuring that the constructed knowledge graph can provide accurate structural knowledge for tasks in the specified field. On the basis of the target text, with the help of the powerful language understanding and generation capabilities of the large language model, knowledge triples can be efficiently and accurately extracted from a large amount of target text, reducing the labor cost overhead of manual extraction and improving the efficiency of knowledge graph construction. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1A schematic diagram of a process for constructing a knowledge graph provided in an embodiment of the present application; Figure 2 A schematic diagram of the principle of a method for constructing a knowledge graph provided in an embodiment of the present application; Figure 3 A schematic diagram of a flow chart of a method for constructing a knowledge graph provided in another embodiment of the present application; Figure 4 A schematic diagram of the structure of a knowledge graph provided in an embodiment of the present application; Figure 5 A schematic diagram of the structure of a knowledge graph construction device provided in an embodiment of the present application; Figure 6 A schematic diagram of the structure of a knowledge graph construction device provided in another embodiment of the present application.

[0017] The accompanying drawings are marked as follows: 50 is a text collection module to be processed, 51 is a target text screening module, 52 is a triple construction module, 53 is a triple filtering module, 54 is a knowledge graph construction module, 60 is a memory, 61 is a processor, 62 is a display screen, 63 is an input and output interface, 64 is a communication interface, 65 is a power supply, 66 is a communication bus, 601 is a computer program, 602 is an operating system, and 603 is data. DETAILED DESCRIPTION

[0018] The terms used in this application are for the purpose of describing specific embodiments only and are not intended to limit this application. The singular forms of "a", "said" and "the" used in this application and the appended claims are also intended to include plural forms unless the context clearly indicates other meanings. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more associated listed items.

[0019] It should be understood that although the terms first, second, third, etc. may be used in the present application to describe various information, these information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".

[0020] Figure 1 A schematic diagram of a method for constructing a knowledge graph provided in an embodiment of the present application is shown in FIG. Figure 1 As shown, the method includes: S10: Collect the text to be processed for building the knowledge graph in the specified field; In order to ensure the comprehensiveness of the knowledge in the knowledge graph and provide more comprehensive data information support for subsequent practical business, in a specific embodiment, a large amount of text in a specified field is collected as the text to be processed for constructing the knowledge graph.

[0021] In an optional embodiment, crawler technology can be used to collect web page content related to a specified field from the Internet. For example, in the medical field, the text to be processed can be crawled from professional medical websites, online medical journals, medical health forums, etc. In an optional embodiment, when crawling the text to be processed, it can be crawled by setting keywords (for example, "cardiovascular disease", "drug therapy", etc.), or by setting crawling rules (for example, limiting crawling to specific sections of medical websites), so as to obtain a large amount of text related to the specified field.

[0022] Of course, the text to be processed in the specified field can also be collected from public data sets, specific resources in the specified field, literature and social media, etc. This application does not limit the way and means of obtaining the text to be processed. It should be noted that in a specific embodiment, in order to ensure the comprehensiveness of the knowledge graph, in an optional embodiment, data information of different levels and angles is obtained from different channels to ensure the diversity of the text to be processed.

[0023] Therefore, by collecting a large amount of text to be processed, a large amount of data information related to the specified field can be obtained, which provides rich data support for the construction of the knowledge graph and helps to build a comprehensive knowledge graph.

[0024] S11: Filter out target texts whose relevance to the specified field meets preset conditions from the texts to be processed; Figure 2 A schematic diagram of the principle of a method for constructing a knowledge graph provided in an embodiment of the present application is shown in FIG. Figure 2 As shown, in a specific embodiment, after obtaining the text to be processed, in order to improve the quality of subsequent knowledge graph construction, the text to be processed is screened so as to eliminate texts that are not highly relevant to the specified field from a large number of texts to be processed.

[0025] Specifically, in an optional embodiment, a high-quality keyword library of a specified field can be pre-built. Based on the keyword library, the correlation between the text to be processed and the specified field is analyzed, and the text whose correlation meets the expectation is screened out as the target text. When screening, it can be determined based on the proportion of keywords in the keyword library appearing in the text, and the text whose proportion is greater than the proportion threshold is selected. Among them, the higher the proportion, the higher the correlation between the text to be processed and the specified field.

[0026] In another optional embodiment, word embedding technology (e.g., Word2Vec, GloVe, etc.) can be used to convert the vocabulary in the text to be processed into a vector representation. And calculate the similarity between the text to be processed and the keyword vector of the specified field. For example, the vector of the keyword "deep learning" is compared with the vector of each vocabulary in the text to be processed, and the overall semantic similarity of the text is calculated. Finally, the text with a similarity greater than the similarity threshold is selected as the target text, wherein the greater the similarity, the higher the correlation between the text to be processed and the specified field.

[0027] It should be noted that the purpose of step S11 is mainly to select texts that are highly consistent with the specified field from a large number of texts to be processed, so as to ensure the quality of the knowledge graph constructed subsequently. In fact, this application does not limit the method for selecting the target text.

[0028] S12: Process the target text through a pre-built target large language model to construct initial triples; Further, such as Figure 2 As shown, in order to quickly construct an initial triple based on the target text, in an optional embodiment, the target text is inferred through a pre-built target large language model to output an initial triple with a structure of (entity 1, relationship, entity 2).

[0029] It should be noted that in a specific embodiment, in order to ensure the quality of the construction of the initial triples, the target large language model can select a vertical large language model of a specified field. For example, if the specified field is medical health, a model such as Bert-Med trained with medical knowledge can be selected. Thus, based on the accurate recognition of professional terms in vertical fields, the construction accuracy of the initial triples is improved.

[0030] Of course, if there is no vertical large language model in the specified field, you can choose a general large language model as the target large language model, such as the GPT series large model and the Bert large model. In order to improve the accuracy of the construction of the initial triples, you can obtain the training data set of the specified field to fine-tune the target large language model, thereby ensuring the performance of the target large language model in the specified field.

[0031] S13: After filtering the initial triples, the target triples are obtained; S14: Construct a knowledge graph through target triples.

[0032] like Figure 2As shown, in order to further improve the quality of the knowledge graph, in an optional embodiment, after obtaining the initial triples, the initial triples are filtered. Specifically, one of the initial triple pairs with the same semantics is removed. Delete the semantically incomplete triples and the logically different triples. Finally, the obtained target triples are converted into a unified standard format. At this point, after obtaining high-quality target triples, the knowledge graph is constructed through the target triples.

[0033] In an optional embodiment, the cleaned and corrected target triple set can be constructed into a knowledge graph through the NetworkX tool, and the constructed knowledge graph can be stored using the Neo4j database to support efficient query.

[0034] In an optional embodiment, taking the legal field as an example, relevant provisions of the labor law are selected for knowledge graph construction, and compared with the knowledge graph construction method based on SpaCy. Among them, SpaCy is an open source natural language processing (NLP) library. Table 1 is a comparison table of different knowledge graph construction methods.

[0035] Table 1 Comparison of different knowledge graph construction methods

[0036] As shown in Table 1, in an optional embodiment, the knowledge graph corresponding to the text of the relevant provisions of the Labor Law is constructed by a knowledge graph construction method based on SpaCy, and by a knowledge graph construction method provided in this application (i.e., a construction method based on the LLMs large model). Obviously, the knowledge graph construction method based on the large language model (LLMs) performs well and is significantly better than the Spacy-based method. The LLMs method has significant advantages in terms of entity coverage, relationship diversity, relationship accuracy, and recall, and can build a higher quality and more comprehensive knowledge graph.

[0037] It should be noted that in Table 1, the number of relationships refers to the relationship between any two entities, and the same relationship may exist between the statistical numbers. Relationship diversity refers to the number of all relationships that are not the same. It can be seen that the knowledge graph construction method provided in this application can construct a high-quality and comprehensive knowledge graph for vertical fields.

[0038] Therefore, the knowledge graph construction method provided in the embodiment of the present application removes texts irrelevant to the specified field through strict relevance screening conditions, reduces the impact of noise data on the knowledge graph construction, and improves the quality and accuracy of the knowledge graph. And the screened target text is closer to the specified field, ensuring that the constructed knowledge graph can provide accurate structural knowledge for tasks in the specified field. On the basis of the target text, with the help of the powerful language understanding and generation capabilities of the large language model, knowledge triples can be efficiently and accurately extracted from a large amount of target text, reducing the labor cost overhead of manual extraction and improving the efficiency of knowledge graph construction.

[0039] In an optional embodiment, the target texts whose relevance to the specified field meets the preset conditions are screened out from the texts to be processed, including: Get pre-built relevance word projects; Based on the relevance cue word engineering, the relevance score between the text to be processed and the specified field is determined through the target large language model; the larger the relevance score, the stronger the relevance between the text to be processed and the specified field; The texts with relevance scores greater than the score threshold are taken as target texts.

[0040] In order to efficiently and accurately select target texts that are highly relevant to a specified field from a large amount of text to be processed, as an optional embodiment, a relevance prompt project that can guide the target large language model to perform text screening can be pre-built. The target large language model can include but is not limited to the BERT large language model and the RoBERTa large language model.

[0041] In an optional embodiment, keywords of the specified field are extracted. Further, based on the relevance prompt project, keywords of the specified field and the text to be processed are embedded into the target large language model, thereby outputting a relevance score between the text to be processed and the specified field.

[0042] It should be noted that, in an optional embodiment, when calculating the relevance score, the similarity between the text to be processed and the keyword can be calculated, and the relevance score can be assigned based on the similarity value. In a specific embodiment, the higher the similarity, the higher the corresponding relevance score, which represents the higher relevance between the text to be processed and the specified field. It is worth noting that when calculating the similarity, the cosine similarity algorithm can be used, which is not limited in this application.

[0043] Furthermore, after the target large language model outputs the relevance score of each text to be processed, a scoring threshold is set according to the actual business needs, and the target text that meets the actual business needs is selected based on the scoring threshold. Specifically, in fields with high accuracy requirements, such as the medical field, a higher scoring threshold can be set to select high-quality data as much as possible. Of course, in an optional embodiment, the scoring threshold can be dynamically adjusted according to the quality feedback of the preliminary screening results to achieve the best screening effect.

[0044] Therefore, the knowledge graph construction method provided in the embodiment of the present application can accurately screen out target texts that are highly relevant to the specified field and reduce misjudgment and noise data through the combination of relevance prompt word engineering and large language models. Using the screened target text as the basic data for knowledge graph construction can improve the accuracy and reliability of the knowledge graph, so that it can better reflect the knowledge structure in the field. The knowledge graph constructed based on the target text can provide high-quality data information support for applications in specific fields such as medical health.

[0045] In an optional embodiment, after filtering the initial triples, a target triple is obtained, including: Get a pre-built verification word project; Based on the verification prompt word project, the target large language model is used to determine whether the target text supports the initial triples; If supported, keep the initial triple; If not supported, obtain a supporting re-verification signal; and determine whether to retain the initial triplet based on the supporting re-verification signal.

[0046] It is understandable that when constructing the initial triples through the target large language model, due to the existence of the large language model, that is, the content output by the large language model may be inconsistent with objective facts, logically contradictory, irrelevant to the input, or contain fictitious details. In order to avoid the occurrence of the above technical problems and ensure the quality of the triples used to construct the knowledge graph, in an optional embodiment, after obtaining the initial triples, the initial triples are re-verified to ensure the factuality and logic of the initial triples.

[0047] Specifically, a verification prompt project is pre-constructed to guide the target large language model to re-verify the initial triples. Furthermore, based on the verification prompt project, both the target text and the initial triples are input into the target large language model, so as to determine whether each initial triple can be supported by the corresponding target text through reasoning of the target large language model.

[0048] In an optional embodiment, when the output result of the target large language model is support, the initial triple represented can be supported by the corresponding target text, that is, the quality of the initial triple meets the requirements and can be retained. Of course, if the output result of the target large language model is unsupported, the initial triple generated by the representation cannot be supported by the corresponding target text, that is, it may not be true, and the initial triple needs to be further verified.

[0049] In an optional embodiment, the initial triplet with the output result of unsupported can be manually rechecked. When the supportive recheck signal output after manual check is characterized as supported, the corresponding initial triplet can be included at this time, otherwise the corresponding initial triplet is removed.

[0050] Therefore, the knowledge graph construction method provided in the embodiment of the present application ensures, through a strict re-verification process, that only triples that are explicitly supported by the target text can enter the knowledge graph, thereby reducing the introduction of erroneous information and improving the reliability of the knowledge in the knowledge graph.

[0051] In an optional embodiment, after filtering the initial triples, a target triple is obtained, including: Get pre-built data cleaning prompt word projects; Based on the data cleaning prompt word project, the target triples are obtained after semantic deduplication and semantic normalization processing of the initial triples through the target large language model.

[0052] On the basis of the above embodiments, in order to further improve the quality of the target triples and thus improve the quality of the knowledge graph, as an optional embodiment, after re-verifying the initial triples, the initial triples are further subjected to semantic deduplication and semantic normalization.

[0053] Specifically, a data cleaning prompt project for guiding the target large language model to perform data cleaning is pre-built, and based on the data cleaning prompt project, the initial triples that have passed re-verification are input into the target large language model for processing.

[0054] Semantic deduplication refers to identifying and removing duplicate or highly similar content through semantic similarity. The core of this process is to determine whether the fields in the initial triples are repeated at the semantic level, rather than just the exact consistency of the text. In an optional embodiment, the similarity between the initial triples can be calculated using the target large language model, and the initial triple pairs with a similarity greater than a threshold are selected.

[0055] In a specific embodiment, the greater the similarity, the more semantically similar the two initial triples are. When the similarity is greater than a threshold, the two initial triples can be regarded as the same semantic triples. At this time, any one of the two initial triples can be eliminated and only one of them can be retained.

[0056] Semantic normalization refers to unifying different expressions of the same concept in texts from different sources into standard terms. For example, "myocardial infarction" and "heart attack" are unified into "acute myocardial infarction". At the same time, triplets of different expressions are converted into standardized forms.

[0057] Therefore, the knowledge graph construction method provided in the embodiment of the present application reduces redundant information and inconsistent expressions in the knowledge graph through semantic deduplication and normalization processing, and improves the overall quality of the knowledge graph.

[0058] Figure 3 A schematic flow chart of a method for constructing a knowledge graph provided in another embodiment of the present application, wherein after filtering the initial triples, the target triples are obtained, including: S30: Obtain a pre-built conflict prompt word project; S31: Based on the conflict prompt word engineering, the target large language model is used to determine the triple pairs with logical conflicts in the initial triples; On the basis of the above embodiments, in order to further improve the quality of the target triples, in an optional embodiment, a conflict prompt project is pre-constructed to guide the target large language model to perform logical conflict detection on the initial triples.

[0059] Furthermore, based on the conflict prompt project, the initial triples after re-verification, semantic deduplication and normalization are used as the input of the target large language model, so that the target large language model can perform a logical check on the input initial triples.

[0060] In a specific embodiment, the entities in the initial triples are aligned to check whether there are cases involving the same entity but inconsistent relationships or attributes. For example, (Apple, headquarters, Beijing) and (Apple, headquarters, California) conflict. In addition, the relationships in the initial triples can be compared to check whether there are logically contradictory relationships. For example, (A, greater than, B) and (B, greater than, A) conflict.

[0061] S32: Obtaining a conflict recheck signal of the triplet pair; S33: if the conflict recheck signal indicates a triplet pair conflict, then the triplet pair is removed; S34: if the conflict secondary check signal indicates that the triple pair does not conflict, retain the triple pair; S35: taking the initial triplet after conflict resolution based on the conflict recheck signal as the target triplet.

[0062] like Figure 3 As shown, in a specific embodiment, when the target large language model determines that there are logically conflicting triple pairs, in order to improve the detection accuracy, the triple pairs are further rechecked for conflicts. Specifically, the recheck can be performed manually, and a conflict recheck signal after the manual check is obtained.

[0063] When the conflict recheck signal indicates that the triple pair conflicts, the triple pair can be eliminated. When the conflict secondary check signal indicates that the triple pair does not conflict, the corresponding triple pair can be retained. Furthermore, after all triples are conflict-resolved, the retained triples are used as target triples that can be used to construct the knowledge graph.

[0064] It should be noted that, in an optional embodiment, the data cleaning prompt project and the conflict prompt project can be combined, that is, semantic deduplication, normalization and logical conflict detection can be performed on the initial triples at the same time.

[0065] Therefore, the knowledge graph construction method provided in the embodiment of the present application ensures that the knowledge in the knowledge graph is logically free of contradictions through a strict conflict detection and resolution process, thereby improving the credibility and reliability of the knowledge graph.

[0066] In an optional embodiment, the target text is processed by a pre-built target large language model to construct an initial triple, including: Get pre-built target cue word projects; Based on the target prompt word project, the target large language model is used to extract entity fields and relationship fields from the target text; and initial triples are constructed based on the entity fields and relationship fields.

[0067] In a specific embodiment, in order to achieve efficient and accurate construction of the knowledge graph, it is crucial to accurately extract the entity fields and relationship fields in the target text. In an optional embodiment, a target prompt project is pre-built to guide the large language model to accurately identify the key information in the text and convert it into a structured triple form.

[0068] Furthermore, based on the target prompt project, the target text is input into the target large language model for processing. Specifically, the target text is segmented into paragraphs or sentences so that the model can process sentence by sentence or paragraph by paragraph. Furthermore, the target prompt words are used to guide the large language model to extract entities from the text. For example, the prompt word may be "extract all entities related to [specified field] from this text". At the same time, the target prompt words are used to guide the large language model to extract the relationship between entities from the text. For example, the prompt word may be "determine the relationship between [entity 1] and [entity 2] in the text". Furthermore, the initial triples are constructed based on the extracted entities and relationships.

[0069] Therefore, the knowledge graph construction method provided in the embodiment of the present application can efficiently extract structured knowledge from a large amount of text through the combination of target prompt word engineering and large language model, greatly reducing the workload of manual knowledge extraction and improving the efficiency of knowledge graph construction. In addition, standardized entity expression and accurate relationship extraction ensure the quality of the initial triples, providing a reliable data foundation for subsequent knowledge graph construction and application.

[0070] Figure 4 A schematic diagram of the structure of a knowledge graph provided in an embodiment of the present application. In an optional embodiment, a knowledge graph is constructed based on a target triple, including: Through the target graph structure tool, the target triples are constructed into a knowledge graph; The knowledge graph is stored through the target knowledge graph library.

[0071] In a specific embodiment, the target triples can be constructed into a knowledge graph through a target graph structure tool, wherein the target graph structure tool may include but is not limited to NetworkX, Apache Jena, and OrientDB. The tool for constructing the knowledge graph can be selected according to actual business needs, and this application does not limit this.

[0072] like Figure 4 As shown, in a specific embodiment, the circles in the constructed knowledge graph represent entities, and the directed lines represent the relationships and directions between entities. For example, Figure 4 There is a relationship X1 between entity B and entity A shown. For example, entity B is China, relationship X1 is capital, entity A is Beijing, that is, the corresponding target unit group is (China, capital, Beijing).

[0073] Furthermore, in order to support the storage and fast query of large-scale data, in an optional embodiment, the constructed knowledge graph is stored in a target knowledge graph library, wherein the target knowledge graph library includes but is not limited to Neo4j and BlazeGraph.

[0074] Therefore, through the target knowledge graph library, the knowledge graph can be efficiently stored and managed to support the storage and rapid query of large-scale data, and conveniently support the continuous updating and expansion of the knowledge graph to adapt to the dynamic changes of domain knowledge.

[0075] In order to make the technical personnel in this field more clear about the construction method of the knowledge graph provided by this application, the following explanation is given using "metformin" in the medical and health field as an example.

[0076] In a specific embodiment, a collection of medical and health related texts is collected, including public data sets, domain-specific resources, web crawler data, etc. At the same time, a pre-built relevance prompt project is obtained, and the collected text to be processed is output to the target large language model for relevance scoring. In the embodiment of the present application, the scoring threshold is set to 0.8.

[0077] For example, the text to be processed and the associated prompt project are as follows: <You are a text screening expert, responsible for screening out the content that is highly relevant to metformin from the given text to be processed. Please judge whether the text belongs to metformin based on the content and give a relevance score (the score ranges from 0 to 1, with 1 being the highest relevance).

[0078] The pending text 1 is: "Metformin is a first-line drug for the treatment of type 2 diabetes. It works by reducing glucose production in the liver and improving insulin sensitivity. However, it is contraindicated in patients with severe renal impairment due to the risk of lactic acidosis"; Pending text 2 is: "Metformin is also used to control blood sugar levels in people with type 2 diabetes. It is generally safe but should be avoided in people with severe kidney disease. Some studies suggest that metformin may reduce the risk of cardiovascular disease, although this has not been fully confirmed" > In a specific embodiment, the correlation score between the target large language model outputting the text to be processed 1 and "metformin" is 0.98, and the correlation score between the text to be processed 2 and "metformin" is 0.96, both greater than 0.8. Therefore, both the text to be processed 1 and the text to be processed 2 are retained, and the text to be processed 1 and the text to be processed 2 are used as the target texts, that is, the target text 1 and the target text 2.

[0079] Furthermore, based on the target Prompt project, initial triples are constructed for the text to be processed 1 and the text to be processed 2 through a large language model.

[0080] For example, the target Prompt project is as follows: <You are a knowledge extraction expert. Please extract entities and their relations from the following text. The output should be in the form of (entity 1, relation, entity 2).

[0081] Target text 1 is: "Metformin is a first-line drug for the treatment of type 2 diabetes. It works by reducing glucose production in the liver and improving insulin sensitivity. However, it is contraindicated in patients with severe renal impairment due to the risk of lactic acidosis"; Target text 2 is: "Metformin is also used to control blood sugar levels in people with type 2 diabetes. It is generally safe, but should be avoided in people with severe kidney disease. Some studies suggest that metformin may reduce the risk of cardiovascular disease, although this has not been fully proven" > Correspondingly, the initial triples output by the target large language model for the target text 1 include: 1. (metformin, treatment, type 2 diabetes); 2. (metformin, reduce, glucose production in the liver); 3. (metformin, improve, insulin sensitivity); 4. (metformin, contraindications, severe renal impairment); 5. (metformin, risk, lactic acidosis).

[0082] The initial triples output for target text 2 include: 1. (metformin, management, blood sugar levels); 2. (metformin, contraindicated in, severe kidney disease); 3. (metformin, reduces, the risk of cardiovascular disease).

[0083] Furthermore, based on the verification prompt project, the initial triples obtained from the target text 1 and the target text 2 are re-verified through the target large language model.

[0084] For example, the verification prompt project is: <You are an expert in factual and logical verification. Please verify whether the given target text supports the following triples. If it does, output "supported"; if it does not, output "not supported".

[0085] Target text 1 is: "Metformin is a first-line drug for the treatment of type 2 diabetes. It works by reducing glucose production in the liver and improving insulin sensitivity. However, it is contraindicated in patients with severe renal impairment due to the risk of lactic acidosis"; Target text 2 is: "Metformin is also used to control blood sugar levels in people with type 2 diabetes. It is generally safe, but should be avoided in people with severe kidney disease. Some studies suggest that metformin may reduce the risk of cardiovascular disease, although this has not been fully proven" > Correspondingly, the support results of the initial triples corresponding to the target text 1 output by the target large language model are: 1. (metformin, treatment, type 2 diabetes) is supported; 2. (metformin, reduce, glucose production in the liver) is supported; 3. (metformin, improve, insulin sensitivity) is supported; 4. (metformin, contraindications, severe renal impairment) is supported; 5. (metformin, risk, lactic acidosis) is supported.

[0086] The support results for the initial triples corresponding to target text 2 are: 1. (Metformin, management, blood sugar level) is supported; 2. (Metformin, contraindicated in, severe kidney disease) is supported; 3. (Metformin, reduces, the risk of cardiovascular disease) is not supported.

[0087] Thus, the unsupported initial triad (metformin, reduced, risk of cardiovascular disease) was eliminated.

[0088] Furthermore, in an optional embodiment, the data cleaning prompt project and the conflict prompt project are combined to perform semantic deduplication, normalization, and logical conflict detection on the initial triples.

[0089] For example, the merged Prompt project is: <You are an expert in knowledge graph optimization and logical conflict resolution. Please process the following triples related to entity [entity field]. Your task is: 1. Remove duplicate triplets.

[0090] 2. Convert all triples into a unified standard form.

[0091] 3. Detect and resolve logical conflicts between triplets.

[0092] The triads to be addressed include: 1. (metformin, treatment, type 2 diabetes); 2. (metformin, reduction, glucose production in the liver); 3. (metformin, improvement, insulin sensitivity); 4. (metformin, contraindications, severe renal impairment); 5. (metformin, risk, lactic acidosis); 6. (metformin, management, blood glucose levels); 7. (metformin, contraindicated in, severe renal disease). > In a specific embodiment, the target large language model outputs the following results: 1. (metformin, treatment, type 2 diabetes) is retained; 2. (metformin, reduction, glucose production in the liver) is retained; 3. (metformin, improvement, insulin sensitivity) is retained; 4. (metformin, contraindications, severe renal impairment) is retained; 5. (metformin, risk, lactic acidosis) is retained; 6. (metformin, management, blood sugar level) is retained; 7. (metformin, contraindicated in, severe kidney disease) is to be verified. Among them, the triple (metformin, management, blood sugar level) and the triple (metformin, contraindicated in, severe kidney disease) are conflicting and need to be further verified.

[0093] After manual determination, the triplet (metformin, management, blood glucose level) and the triplet (metformin, contraindicated in, severe renal disease) were different mechanisms and did not conflict, so both could be retained.

[0094] In the above embodiments, the method for constructing a knowledge graph is described in detail. The present application also provides an embodiment corresponding to a device for constructing a knowledge graph.

[0095] Figure 5 A schematic diagram of the structure of a knowledge graph construction device provided in an embodiment of the present application is shown in FIG. Figure 5 As shown, the device comprises: A to-be-processed text collection module 50 is used to collect to-be-processed texts for constructing a knowledge graph in a specified field; The target text screening module 51 is used to screen out target texts whose relevance to the specified field meets preset conditions from the texts to be processed; A triple construction module 52, for processing the target text through a pre-constructed target large language model to construct an initial triple; A triplet filtering module 53, used to filter the initial triplet to obtain the target triplet; The knowledge graph construction module 54 is used to construct a knowledge graph through target triples.

[0096] In addition, the knowledge graph construction device provided in the embodiment of the present application also includes: A prompt word project acquisition module is used to acquire a pre-built correlation prompt word project; A relevance score determination module is used to determine the relevance score between the text to be processed and the specified field through the target large language model based on the relevance prompt word engineering; wherein the larger the relevance score, the stronger the relevance between the text to be processed and the specified field; The target text determination module is used to take the text whose relevance score is greater than the score threshold as the target text.

[0097] The prompt word project acquisition module is also used to obtain a pre-built verification prompt word project; The support processing module is used to determine whether the target text supports the initial triplet based on the verification prompt word project and through the target large language model; if supported, the initial triplet is retained; if not supported, a support re-verification signal is obtained; and whether to retain the initial triplet is determined based on the support re-verification signal.

[0098] The prompt word project acquisition module is also used to obtain a pre-built data cleaning prompt word project; The data cleaning module is used to obtain the target triplet after semantic deduplication and semantic normalization processing of the initial triplet through the target large language model based on the data cleaning prompt word project.

[0099] The prompt word project acquisition module is also used to acquire the pre-built conflict prompt word project; The conflict resolution module is used to determine the triple pairs with logical conflicts in the initial triples based on the conflict prompt word project and through the target large language model; obtain the conflict re-verification signal of the triple pair; if the conflict re-verification signal indicates that the triple pair conflicts, the triple pair is eliminated; if the conflict secondary verification signal indicates that the triple pair does not conflict, the triple pair is retained; and the initial triple after conflict resolution based on the conflict re-verification signal is used as the target triple.

[0100] The prompt word project acquisition module is also used to acquire a pre-built target prompt word project; The triple construction module is also used to extract entity fields and relationship fields from the target text through the target large language model based on the target prompt word project; and to construct initial triples based on the entity fields and relationship fields.

[0101] The knowledge graph construction module is also used to construct the target triples into a knowledge graph through the target graph structure tool; The storage module is used to store the knowledge graph through the target knowledge graph library.

[0102] Figure 6 A schematic diagram of a knowledge graph construction device provided in another embodiment of the present application is shown in FIG. Figure 6 As shown, the knowledge graph construction device includes: a memory 60, for storing computer programs; Processor 61 is used to implement the steps of the method for constructing a knowledge graph as mentioned in the above embodiment when executing a computer program.

[0103] The knowledge graph construction device provided in this embodiment may include but is not limited to a tablet computer, a laptop computer, or a desktop computer.

[0104] Among them, the processor 61 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 61 can be implemented in at least one hardware form of a digital signal processor (Digital Signal Processor, referred to as DSP), a field programmable gate array (Field-Programmable Gate Array, referred to as FPGA), and a programmable logic array (Programmable Logic Array, referred to as PLA). The processor 61 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also known as a central processing unit (Central Processing Unit, referred to as CPU); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 61 may be integrated with a graphics processing unit (Graphics Processing Unit, referred to as GPU), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 61 may also include an artificial intelligence (Artificial Intelligence, referred to as AI) processor, which is used to process computing operations related to machine learning.

[0105] The memory 60 may include one or more computer-readable storage media, which may be non-transitory. The memory 60 may also include a high-speed random access memory, and a non-volatile memory, such as one or more disk storage devices, flash memory storage devices. In this embodiment, the memory 60 is at least used to store the following computer program 601, wherein, after the computer program is loaded and executed by the processor 61, it can implement the relevant steps of the method for constructing the knowledge graph disclosed in any of the aforementioned embodiments. In addition, the resources stored in the memory 60 may also include an operating system 602 and data 603, etc., and the storage method may be temporary storage or permanent storage. Among them, the operating system 602 may include Windows, Unix, Linux, etc. Data 603 may include, but is not limited to, relevant data involved in the method for constructing the knowledge graph, etc.

[0106] In some embodiments, the knowledge graph construction device may also include a display screen 62, an input and output interface 63, a communication interface 64, a power supply 65 and a communication bus 66.

[0107] Those skilled in the art will understand that Figure 6 The structure shown in does not constitute a limitation on the device for constructing a knowledge graph, and may include more or fewer components than shown in the figure.

[0108] The knowledge graph construction device provided in the embodiment of the present application includes a memory and a processor. When the processor executes the program stored in the memory, it can implement the knowledge graph construction method in the above embodiment.

[0109] It should be noted that, although the operations are depicted in a specific order in the accompanying drawings, this should not be understood as requiring these operations to be performed in the specific order shown or to be performed sequentially, or requiring all illustrated operations to be performed to achieve the desired results. In some cases, multitasking and parallel processing may be advantageous. In addition, the separation of various system modules and components in the above-described embodiments should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product, or packaged into multiple software products.

Claims

1. A method for constructing a knowledge graph, characterized in that: The method comprises: Collect the text to be processed for building the knowledge graph in the specified field; Filtering out target texts whose relevance to the specified field meets preset conditions from the texts to be processed; Processing the target text through a pre-built target large language model to construct initial triples; After filtering the initial triples, a target triple is obtained; The knowledge graph is constructed through the target triples.

2. The method for constructing a knowledge graph according to claim 1, wherein: The step of selecting target texts whose relevance to the specified field meets preset conditions from the texts to be processed includes: Get pre-built relevance word projects; Based on the relevance prompt word project, the relevance score between the text to be processed and the specified field is determined through the target large language model; wherein the larger the relevance score, the stronger the relevance between the text to be processed and the specified field; The text whose relevance score is greater than a score threshold is used as the target text.

3. The method for constructing a knowledge graph according to claim 1, wherein: After filtering the initial triples, the target triples are obtained, including: Get a pre-built verification word project; Based on the verification prompt word project, determining whether the target text supports the initial triple through the target large language model; If supported, retain the initial triplet; If not supported, a supporting recheck signal is obtained; and based on the supporting recheck signal, it is determined whether to retain the initial triplet.

4. The method for constructing a knowledge graph according to claim 1, wherein: After filtering the initial triples, the target triples are obtained, including: Get pre-built data cleaning prompt word projects; Based on the data cleaning prompt word project, the initial triples are subjected to semantic deduplication processing and semantic normalization processing through the target large language model to obtain the target triples.

5. The method for constructing a knowledge graph according to any one of claims 1 to 4, characterized in that: After filtering the initial triples, the target triples are obtained, including: Get pre-built conflict word projects; Based on the conflict prompt word project, determining the triple pairs having logical conflicts in the initial triples through the target large language model; Obtaining a conflict recheck signal of the triplet pair; If the conflict recheck signal indicates that the triplet pair conflicts, then the triplet pair is eliminated; If the conflict secondary check signal indicates that the triplet pair does not conflict, retaining the triplet pair; An initial triplet after conflict resolution based on the conflict recheck signal is used as the target triplet.

6. The method for constructing a knowledge graph according to claim 1, wherein: The target text is processed by a pre-built target large language model to construct an initial triple, including: Get pre-built target cue word projects; Based on the target prompt word project, entity fields and relationship fields are extracted from the target text through the target large language model; and the initial triples are constructed according to the entity fields and the relationship fields.

7. The method for constructing a knowledge graph according to claim 1, wherein: The step of constructing the knowledge graph through the target triples includes: The target triples are constructed into the knowledge graph by a target graph structure tool; The knowledge graph is stored in a target knowledge graph library.

8. A device for constructing a knowledge graph, characterized in that: The device comprises: A module for collecting texts to be processed, used to collect texts to be processed for building a knowledge graph in a specified field; A target text screening module is used to screen out target texts whose relevance to the specified field meets preset conditions from the texts to be processed; A triple construction module, used for processing the target text through a pre-constructed target large language model to construct initial triples; A triplet filtering module, used for filtering the initial triplet to obtain a target triplet; A knowledge graph construction module is used to construct the knowledge graph through the target triples.

9. A device for constructing a knowledge graph, comprising a memory and a processor, wherein the memory stores a computer program that can be run on the processor, characterized in that: When the processor executes the program, the steps of the method for constructing a knowledge graph described in any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method for constructing a knowledge graph described in any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Knowledge retrieval method and device

    CN119202137A

  • Electric power knowledge graph construction method based on large language model

    CN119357408A

  • Knowledge graph open domain construction and RAG question and answer method and device and storage medium

    CN119416882A

Cited By

  • Knowledge graph updating method, electronic equipment and storage medium

    CN120744139A

  • LLM model-based text analysis method and device, medium and equipment

    CN120929610A

  • A text analysis method, apparatus, medium, and device based on an LLM model

    CN120929610B

  • Knowledge graph construction method and device, electronic equipment and storage medium

    CN121009197A

  • A knowledge graph construction method and device, electronic equipment and storage medium

    CN121009197B