QA-driven large-model hierarchical knowledge graph parameterization construction method

By employing a QA-driven, hierarchical knowledge graph parameterization construction method, and utilizing pre-trained language models to generate structured question-and-answer data, this method performs information granularity decomposition and atomic node construction. This solves the problems of low efficiency and redundancy in existing knowledge graph construction technologies, and achieves efficient and refined knowledge extraction and graph representation.

CN120911583APending Publication Date: 2025-11-07GUANGDONG TECSUN SCIENCE & TECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511085613.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-04
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

In existing technologies, knowledge graph construction methods rely on manual annotation, which is costly, inefficient, and has poor scalability. Furthermore, they fail to effectively decompose unstructured text, lack a unified node organization logic, result in chaotic graph hierarchies, numerous redundant nodes, and imperfect user interaction adjustments, making it impossible to achieve flexible editing and dynamic updates.

Method used

By employing a QA-driven, hierarchical knowledge graph parameterization construction method, we utilize pre-trained language models to generate structured question-and-answer data, perform information granularity decomposition, and combine atomic node construction to improve the precision of knowledge extraction and the representation of the knowledge graph.

Benefits of technology

It enables automatic extraction of knowledge units from unstructured documents, reduces the cost of manual intervention, constructs a knowledge graph with a clear structure and complete logic, avoids redundancy and duplication, and improves the precision of knowledge extraction and the expressive power of the graph.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120911583A_ABST
    Figure CN120911583A_ABST
Patent Text Reader

Abstract

According to the QA-driven large-model hierarchical knowledge graph parameterization construction method provided by the invention, the unstructured document is segmented, and the structured question and answer pairs are generated in combination with the language model, so that automatic extraction from the original text to the knowledge unit is realized, and the manual participation cost is reduced; then, semantic affiliation and hierarchical relations are established between the extracted question content and entity categories, atomic nodes and intermediate nodes, systematic hierarchical organization of the knowledge graph is achieved, the structure is clear, and logic is complete. Moreover, a semantic clustering algorithm is adopted to classify bottom nodes, and intermediate nodes are automatically constructed on the premise of meeting graph structure constraint conditions, so that redundancy and repetition are avoided; and finally, through information granularity decomposition driven by question and answer pairs and in combination with atomic node construction, the fine degree of knowledge extraction and the expressive power of graph representation are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of knowledge graph, and particularly relates to a method for parameterized construction of a QA-driven large model hierarchical knowledge graph. BACKGROUND

[0002] As a structured semantic network, knowledge graph has been widely applied in intelligent question answering, recommendation system, search engine and other scenarios. Traditional knowledge graph construction methods mostly rely on manual annotation, rule extraction or domain expert maintenance, which have problems such as high cost, low efficiency and poor scalability. With the development of pre-trained language models, the semantic understanding and question answering generation capabilities based on large models have been significantly improved, providing new possibilities for realizing high-quality and low-cost automatic knowledge extraction.

[0003] In the prior art, some methods attempt to combine question answering generation for knowledge extraction, but still face many challenges: for example, unstructured text cannot be effectively disassembled and segmented, there is a lack of unified node organization logic, atomic information cannot be effectively attributed to semantic categories, and the graph hierarchy is chaotic with many redundant nodes, affecting the overall usability and structural clarity. In addition, the interactive adjustment support for users is also not perfect, and the flexible editing and dynamic updating of graph content cannot be realized.

[0004] Therefore, the problems in the prior art need to be solved. SUMMARY

[0005] The present application provides a method for parameterized construction of a QA-driven large model hierarchical knowledge graph, which solves the defects in the prior art by using question and answer pairs to drive information granularity decomposition and combining atomic node construction to significantly improve the precision of knowledge extraction and the expressiveness of graph representation.

[0006] The present application provides a method for parameterized construction of a QA-driven large model hierarchical knowledge graph, comprising: The user inputted document to be processed is divided according to a preset word threshold to obtain multiple paragraphs, and each paragraph is bound with a unified document identifier; The paragraphs with the same document identifier are processed in parallel, and a pre-trained language model is called to generate a structured question and answer data list for each paragraph; The question content is extracted from the question and answer data list, and the entity category recognition is performed on the question content to obtain multiple entity types to form the top nodes of the knowledge graph; According to the entity category, specific information is extracted from the question content to generate atomic nodes, and each atomic node is attributed to the corresponding entity category to form the bottom nodes of the knowledge graph; The bottom nodes are semantically clustered to form multiple levels of intermediate nodes; According to the semantic hierarchical relationship of the top-level node, the multi-level intermediate nodes and the bottom-level nodes, a knowledge graph corresponding to the to-be-processed document is constructed.

[0007] According to the method for parameterized construction of a QA-driven large model hierarchical knowledge graph provided in the application, before the step of processing each paragraph with the same document identifier in parallel and calling a pre-trained language model to generate a structured question and answer data list for each paragraph, the method further comprises the following steps: According to the mapping relationship between the preset number of words and the number of question and answer pairs, the number of question and answer pairs required to be generated for each paragraph is estimated according to the text length of each paragraph, so as to guide the parallel configuration and resource allocation of the subsequent question and answer generation process.

[0008] According to the method for parameterized construction of a QA-driven large model hierarchical knowledge graph provided in the application, after the step of extracting question content from the question and answer data list and performing entity category identification on the question content to obtain a plurality of entity types, the method further comprises the following steps: For each identified entity type, a corresponding abstract name and semantic definition are provided.

[0009] According to the method for parameterized construction of a QA-driven large model hierarchical knowledge graph provided in the application, after the step of constructing the knowledge graph corresponding to the to-be-processed document according to the semantic hierarchical relationship of the top-level node, the multi-level intermediate nodes and the bottom-level nodes, the method further comprises the following steps: An input knowledge graph adjustment instruction of a user is acquired, and the knowledge graph adjustment instruction is used to indicate an editing operation on a node or an edge relationship in the knowledge graph. According to the knowledge graph adjustment instruction, an entity name of a node, a node level or an edge relationship with other nodes in the knowledge graph is added, modified or deleted. The data structure of the knowledge graph is updated.

[0010] According to the method for parameterized construction of a QA-driven large model hierarchical knowledge graph provided in the application, the step of performing semantic clustering on the bottom-level nodes to form multi-level intermediate nodes specifically comprises the following steps: A constraint condition is constructed, and the constraint condition comprises a graph hierarchical depth limit, a node naming length limit and a number limit of nodes at each level. Under the premise of meeting the constraint condition, a clustering algorithm is used to perform hierarchical clustering from bottom to top on the bottom-level nodes, to generate intermediate nodes with increasing semantic abstraction degrees, and to establish an upper and lower hierarchical relationship between nodes.

[0011] According to the method for parameterized construction of a QA-driven large model hierarchical knowledge graph provided by the application, the step of extracting specific information from the question content according to the entity category to generate atomic nodes and attributing each atomic node to a corresponding entity category to form the bottom layer nodes of the knowledge graph comprises the following steps: According to the entity category, the question content is subjected to semantic analysis to extract specific information with clear semantics to generate atomic nodes. According to the semantic features and semantic hierarchical relationships of each atomic node, the entity category to which the atomic node is attributed is determined.

[0012] According to the method for parameterized construction of a QA-driven large model hierarchical knowledge graph provided by the application, the step of constructing the knowledge graph corresponding to the to-be-processed document according to the semantic hierarchical relationships of the top layer nodes, the multi-level intermediate nodes and the bottom layer nodes comprises the following steps: The bottom layer nodes are classified into corresponding intermediate nodes according to semantic features. The intermediate nodes are classified into corresponding entity category nodes according to semantic features.

[0013] The application further provides a device for parameterized construction of a QA-driven large model hierarchical knowledge graph, comprising: A paragraph segmentation module is configured to segment a to-be-processed document input by a user into multiple paragraphs according to a preset word threshold, and each paragraph is bound to a unified document identifier. A question and answer data module is configured to process each paragraph with the same document identifier in parallel, and call a pre-trained language model to generate a structured question and answer data list for each paragraph. A top layer node module is configured to extract question content from the question and answer data list, identify the entity category of the question content, obtain multiple entity types, and form the top layer nodes of the knowledge graph. A bottom layer node module is configured to extract specific information from the question content according to the entity category to generate atomic nodes, and attribute each atomic node to a corresponding entity category to form the bottom layer nodes of the knowledge graph. An intermediate node module is configured to perform semantic clustering on the bottom layer nodes to form multi-level intermediate nodes. A graph construction module is configured to construct the knowledge graph corresponding to the to-be-processed document according to the semantic hierarchical relationships of the top layer nodes, the multi-level intermediate nodes and the bottom layer nodes.

[0014] The application further provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method for parameterized construction of a hierarchical knowledge graph of a large model driven by QA according to any of the above when executing the program.

[0015] The application further provides a non-transitory computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the method for parameterized construction of a hierarchical knowledge graph of a large model driven by QA according to any of the above.

[0016] The application further provides a computer program product comprising a computer program, wherein the computer program, when executed by a processor, implements the method for parameterized construction of a hierarchical knowledge graph of a large model driven by QA according to any of the above.

[0017] The method for parameterized construction of a hierarchical knowledge graph of a large model driven by QA provided by the application realizes automatic extraction from original text to knowledge units by segmenting unstructured documents and generating structured question-answer pairs in combination with a language model, thereby reducing the cost of manual participation; then, semantic attribution and hierarchical relationships are established between the extracted question content and entity categories, atomic nodes and intermediate nodes, thereby realizing systematic hierarchical organization of the knowledge graph, and the structure is clear and the logic is complete. Moreover, the bottom nodes are classified by using a semantic clustering algorithm, and the intermediate nodes are automatically constructed under the premise of meeting the constraint conditions of the graph structure, thereby avoiding redundancy and repetition; finally, the information granularity decomposition driven by the question-answer pairs is combined with the atomic node construction, thereby significantly improving the fine degree of knowledge extraction and the expressiveness of graph representation. BRIEF DESCRIPTION OF DRAWINGS

[0018] In order to more clearly illustrate the technical solutions in the application or the prior art, the following will briefly introduce the drawings needed in the embodiments or the prior art description. Obviously, the drawings in the following description are some embodiments of the application, and for those skilled in the art, other drawings can also be obtained without creative labor.

[0019] Figure 1 is a flowchart of the method for parameterized construction of a hierarchical knowledge graph of a large model driven by QA provided by the application; Figure 2 is a schematic diagram of entity category extraction provided by the application; Figure 3 is a schematic diagram of an atomic node provided by the application; Figure 4 is a schematic diagram of knowledge graph construction provided by the application; Figure 5It is a structure schematic view of the device for parameterized construction of a QA-driven large model hierarchical knowledge graph provided by the application. Figure 6 It is a structure schematic view of an electronic device provided by the application. DETAILED DESCRIPTION

[0020] To make the objectives, technical solutions and advantages of the present application clearer, the technical solutions in the present application will be described clearly and completely below with reference to the drawings in the present application. Obviously, the described embodiments are some of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the protection scope of the present application.

[0021] To solve the problems in the prior art, the present application provides a method for parameterized construction of a QA-driven large model hierarchical knowledge graph, so as to significantly improve the fine degree of knowledge extraction and the expression of graph representation by information granularity decomposition driven by question and answer pairs and combined with atomic node construction. The method for parameterized construction of a QA-driven large model hierarchical knowledge graph will be described as follows, as shown in Figure 1 The method includes but is not limited to the following steps: Step 110: The user-inputted document to be processed is divided according to a preset word threshold to obtain multiple paragraphs, and each paragraph is bound with a unified document identifier.

[0022] In step 110, after receiving the user-inputted document to be processed, the text content is first automatically divided according to the preset word threshold. The division logic can be set based on the number of words, syntactic boundaries or semantic integrity, etc. Each divided text is taken as an independent processing unit, and is uniformly bound with the document identifier (Document ID) corresponding to the document to realize subsequent aggregation and traceability operations.

[0023] For example, if it is set that each paragraph does not exceed 500 words, and the original document contains 2000 words, then the original document is divided into 4 paragraphs, which are marked as DocA_P1, DocA_P2, …, DocA_P4 respectively, and all the paragraphs belong to the document identifier DocA.

[0024] Step 120: The paragraphs with the same document identifier are processed in parallel, and a pre-trained language model is called to generate a structured question and answer data list for each paragraph.

[0025] In step 120, the paragraphs bound to the same document identifier are processed in parallel. Each paragraph is sent to a pre-trained language model (such as GPT, T5, ERNIE, etc.) for question and answer pair generation. The generated structured question and answer data list includes questions, answers, context paragraphs, and related meta-information, forming a semantic summary representation at the paragraph level.

[0026] Preferably, the DeepSeek-V3 large model inference can be used to generate potential question and answer pairs, and output standardized data structures. Because DeepSeek-V3 has advantages in efficient long text processing, complex reasoning capability, and leading Chinese knowledge in processing text understanding tasks, it can meet the text understanding needs in multiple fields. It helps to extract more key question and answer pairs. After generation, users can view the generated QA and check relevance, accuracy, and fluency. If the generated QA pair is not satisfactory, it can be regenerated. The specific prompt words are as follows: Task: You are a question and answer generation assistant. Your task is to generate question and answer pairs based on the knowledge mentioned in the user input text. Please generate question and answer pairs according to the following steps: 1. Understand and summarize the theme of this long text 2. What key information and concepts are involved in this long text? 3. Decompose and reorganize multiple information and concepts 4. Generate question and answer pairs based on these key information and concepts.

[0027] Step 130, extract question content from the question and answer data list, and perform entity category recognition on the question content to obtain multiple entity types to form the top nodes of the knowledge graph.

[0028] In step 130, all question content is extracted from the above structured question and answer data list, and entity category recognition is performed on each question sentence. Entity name recognition model (NER) or multi-classification model can be used to identify the semantic categories involved in the question stem, such as "product function", "technical principle", "application scenario", etc.

[0029] The recognized entity categories are used as the top nodes of the knowledge graph. Each type of question will be assigned to a corresponding entity category to construct the semantic backbone structure of the knowledge graph.

[0030] Step 140, according to the entity category, extract specific information from the question content to generate atomic nodes, and assign each atomic node to the corresponding entity category to form the bottom nodes of the knowledge graph.

[0031] In step 140, on the basis of obtaining the entity category, the semantic analysis is performed on each question and answer pair to extract specific factual information from the question and the answer, and an atomic knowledge node (i.e., a bottom node) is constructed. The node usually has a unique semantic direction and a fine granularity, such as “image recognition”, “ResNet algorithm”, “supporting multi-classification”, and the like.

[0032] Each atomic node is attributed to the corresponding top node according to its semantic category, thereby establishing a one-to-many attribution relationship of “entity category → atomic information” and forming an initial bottom structure of the graph.

[0033] In step 150, the bottom nodes are subjected to semantic clustering to form multi-level intermediate nodes.

[0034] In step 150, after the bottom nodes are constructed, the atomic nodes are subjected to semantic clustering processing. The node content is subjected to vectorization representation by embedding models (such as BERT Embedding, S-BERT), and hierarchical clustering, K-Means or a self-defined semantic aggregation strategy is adopted to perform clustering under the premise of satisfying the following constraint conditions: a maximum graph level depth limit (such as no more than 5 layers); a node naming length limit (such as no more than 15 characters); a control of the number of nodes at each level (such as a lower limit of no less than 3 atomic nodes under each node).

[0035] After clustering, a plurality of intermediate semantic nodes are generated, and hierarchical connections are established with the upper and lower nodes to realize structure organization with an incremental semantic abstraction degree.

[0036] In step 160, according to the semantic hierarchical relationship of the top nodes, the multi-level intermediate nodes and the bottom nodes, a knowledge graph corresponding to the to-be-processed document is constructed.

[0037] In step 160, a complete knowledge graph is constructed in combination with the semantic hierarchical relationship among the top nodes, the intermediate nodes and the bottom atomic nodes. The construction process is from bottom to top, the bottom atomic nodes are aggregated to the intermediate nodes, the intermediate nodes are further attributed to the top entity categories, and a multi-level semantic graph structure is formed.

[0038] Each edge relationship represents semantic connections such as belonging, containing or subordination, and the whole graph structure can be derived into a graph database (such as Neo4j) or a knowledge file (such as RDF, OWL) for persistent storage and calling.

[0039] The graph not only has a clear semantic hierarchy, but also facilitates subsequent query, visualization analysis and intelligent reasoning and the like applications.

[0040] As a further optional embodiment, before the step of processing each paragraph with the same document identifier in parallel, and calling the pre-trained language model to generate a structured question and answer data list for each paragraph, the method further comprises: Based on the text length of each paragraph, a number of question and answer pairs required to be generated for each paragraph is estimated according to a preset mapping relationship between the number of words and the number of question and answer pairs, so as to guide the parallel configuration and resource allocation of the subsequent question and answer generation process.

[0041] In this embodiment, in order to improve the resource utilization efficiency and processing performance of the question and answer generation process, before processing each paragraph with the same document identifier in parallel and calling the pre-trained language model to generate a structured question and answer data list, the system first analyzes the text length of each paragraph.

[0042] Specifically, the system presets a set of mapping relationship between the number of words and the number of question and answer pairs, for example: 1 effective question and answer pair is expected to be generated for every 50 words, or a more complex nonlinear mapping function is established according to the actual training data experience. After reading the actual number of words of each paragraph, the system estimates the number of question and answer pairs that should be generated for the paragraph according to the mapping relationship.

[0043] Then, the system configures the generation task of the pre-trained language model in parallel according to the estimated number of question and answer pairs of each paragraph, for example, by using thread pool or distributed task scheduling system, dynamically allocating computing resources (such as GPU core or CPU thread), and preferentially allocating more resources to the paragraphs with more expected number of question and answer pairs, so as to improve the overall generation efficiency and load balancing.

[0044] This embodiment can dynamically adjust the generation strategy according to the complexity of the paragraph content, improve the quality and system throughput of the question and answer generation, and is especially suitable for application scenarios of processing large documents or batch documents.

[0045] As a further optional embodiment, after the step of extracting question content from the question and answer data list, and performing entity category identification on the question content to obtain a plurality of entity types, the method further comprises: For each identified entity type, a corresponding abstract name and semantic definition are provided.

[0046] For example, Figure 2As shown, in this embodiment, for each identified entity type, a corresponding abstract name and semantic definition are provided. Specifically, the system first matches the identified entity categories based on a predefined entity type mapping table, with each entity category corresponding to an abstract name, such as abstracting "university name", "company name", and other entities as "organization", abstracting "date of birth", "registration time", and other entities as "time point", and abstracting "product name", "service name", and other entities as "object identifier", etc. Subsequently, the system calls the semantic analysis module to provide corresponding semantic definitions for each abstract name, such as: "organization" refers to a social organization unit with legal entity attributes or operational functions; "time point" refers to a timestamp type of information that marks the occurrence of an event or records the time of a certain behavior, etc.

[0047] The above abstract names and their semantic definitions will be stored in the entity type ontology library of the knowledge graph, for subsequent unified structure modeling, entity alignment, reasoning, and display of the graph, ensuring semantic consistency and scalability of multi-source heterogeneous data during graph construction.

[0048] As a further optional embodiment, after the step of constructing the knowledge graph corresponding to the to-be-processed document according to the semantic hierarchical relationship of the top-level node, the multi-level intermediate node, and the bottom-level node, the system further includes: obtaining a knowledge graph adjustment instruction input by a user, the knowledge graph adjustment instruction being used to indicate an editing operation on a node or edge relationship in the knowledge graph; performing an adding, modifying, or deleting operation on an entity name of a node, a node level, or an edge relationship with other nodes in the knowledge graph according to the knowledge graph adjustment instruction; updating the data structure of the knowledge graph.

[0049] In this embodiment, after constructing the knowledge graph corresponding to the to-be-processed document according to the semantic hierarchical relationship of the top-level node, the multi-level intermediate node, and the bottom-level node, the system also supports interactive editing and adjustment of the knowledge graph, including the following steps: First, the system provides a graphical user interface (GUI) for displaying the constructed knowledge graph, which includes multiple nodes and edges representing the semantic relationship between the nodes, and the user can perform selection operations on the nodes or edges through the interface.

[0050] When the user performs operations such as clicking, right-clicking, or dragging in the interface, the system parses a knowledge graph adjustment instruction, which includes: operation type (such as adding, modifying, or deleting), target object (node or edge), and corresponding parameters (such as new name, new level, relationship direction, etc.).

[0051] Specifically, the adjustment of the node can include: Modification of entity name: the user selects a node, enters a new name in the pop-up edit box, and the system automatically replaces the entity tag of the node with the new value; Adjustment of node level: when the user wants to promote an intermediate node to the top node or sink to the bottom node, the target level can be selected by dragging or menu, and the system reorganizes its position in the graph according to the semantic relationship; Modification of edge relationship: the user can select two nodes and specify a new semantic relationship, and the system updates the type of the edge between the two nodes or adds / deletes the edge.

[0052] After the above editing operations are completed, the system will update the internal data structure of the knowledge graph to ensure that the node level, entity information and edge relationship are consistent with the user's instructions. At the same time, the system also supports the buffer recording of editing instructions such as "undo / redo" to ensure the flexibility and safety of user operations.

[0053] The updated knowledge graph data structure will be automatically synchronized and saved to the backend database for subsequent calling, display or reasoning.

[0054] As a further optional embodiment, the step of semantically clustering the bottom nodes to form multiple levels of intermediate nodes specifically includes: Constructing constraint conditions, including graph level depth limit, node naming length limit and number of nodes per level limit; Under the premise of meeting the constraint conditions, the bottom nodes are hierarchically clustered from bottom to top using a clustering algorithm to generate intermediate nodes with increasing semantic abstraction and establish the superior-inferior hierarchical relationship between nodes.

[0055] In this embodiment, first, constraint conditions are constructed to guide the clustering process, including but not limited to: Graph level depth limit: set the constructed knowledge graph to not exceed the preset level number, for example, not more than 4 levels, to ensure the controllability and readability of the graph structure; Node naming length limit: limit the character length of the node name, for example, not more than 20 characters, to avoid the inconvenience of information display caused by long node names; Limit the number of nodes per level: set an upper limit to the number of nodes per level, for example, each level contains a maximum of 50 nodes, to prevent overcrowding of information or low retrieval efficiency caused by too many nodes in a layer.

[0056] Then, under the premise of meeting the above constraint conditions, the bottom nodes (i.e. the original nodes obtained by entity recognition and fact extraction) are hierarchically clustered from bottom to top using a semantic clustering algorithm. The specific process is as follows: Semantic representation generation: generate semantic vector representation for each bottom node, which can be generated by large models or embedding models (such as BERT, Word2Vec, etc.), capturing the semantic information of the node; Initial clustering: all bottom nodes are regarded as initial clusters, the semantic similarity between nodes is calculated, and clustering algorithm (such as K-means, hierarchical clustering, DBSCAN, etc.) is used to classify similar nodes into a class; Intermediate node generation: extract representative keywords or phrases as the name of intermediate node for each class of nodes, forming semantic abstraction level; Recursive clustering: continue to perform semantic clustering on the intermediate nodes generated in the previous step, and generate higher level intermediate nodes layer by layer until the constraint of graph level depth is met; Hierarchical relationship establishment: record the inclusion relationship between each layer of intermediate nodes and their subordinate nodes during the clustering process, and construct the complete superior and inferior hierarchical structure.

[0057] This embodiment helps to automatically construct a knowledge graph with clear semantic hierarchical structure, improves knowledge organization and retrieval efficiency, and provides support for subsequent knowledge query and reasoning.

[0058] As a further optional embodiment, the step of extracting specific information from the question content according to the entity category to generate atomic nodes, and attributing each atomic node to the corresponding entity category to constitute the bottom nodes of the knowledge graph, specifically includes: According to the entity category, the semantic analysis of the question content is carried out, and the specific information with clear semantics is extracted to generate atomic nodes; According to the semantic features and semantic hierarchical relationship of each atomic node, the entity category to which the atomic node belongs is determined.

[0059] As Figure 3 shown, in this embodiment, according to the entity category (such as person, time, place, institution, product, etc.) identified in the foregoing steps, the semantic understanding model pre-trained is combined to perform semantic analysis on each question content. The semantic analysis includes processing such as word segmentation, part-of-speech tagging, named entity recognition, and dependency syntax analysis, so as to locate the key words or phrases with clear semantic boundaries, which are the candidate atomic nodes.

[0060] Then, the system determines the specific meaning expressed by the candidate atomic nodes in a specific semantic environment according to the context relationship of the candidate atomic nodes in the sentence, and marks them as atomic nodes with independent semantics. For example, in the question "Who is the founder of Huawei?", "Huawei" can be identified as an entity of the "organization" category, "founder" can be used as a role relationship, and "Ren Zhengfei" can be extracted as an atomic node in the answer and attributed to the "person" category in subsequent question and answer matching.

[0061] Further, the system determines the entity category to which the atomic node belongs according to the hierarchical relationship of the atomic node in the knowledge structure, such as the membership relationship, type-instance relationship, superior-inferior relationship, etc. For example, "Ren Zhengfei" as an instance node of the "person" category will be attributed to the upper entity category "person"; if there is an intermediate category "entrepreneur", a three-level structure can be constructed: "person" -> "entrepreneur" -> "Ren Zhengfei".

[0062] Finally, all atomic nodes will establish a membership relationship with their corresponding entity categories in the form of a graph structure, forming the bottom nodes of the knowledge graph and laying the foundation for the extraction and connection of relationships between entities in subsequent tasks. This process not only enhances the semantic clarity of the nodes in the knowledge graph, but also improves the practicality and scalability of the graph in multi-level queries, semantic reasoning, and other tasks.

[0063] As a further optional embodiment, the step of constructing the knowledge graph corresponding to the to-be-processed document according to the semantic hierarchical relationship of the top-level nodes, the multi-level intermediate nodes, and the bottom-level nodes specifically includes: classifying the bottom-level nodes into corresponding intermediate nodes according to semantic features; classifying the intermediate nodes into corresponding entity category nodes according to semantic features.

[0064] As shown in Figure 4 In this embodiment, semantic analysis is performed on each bottom-level node to extract its key words, context relationship, syntax structure, and other semantic features. Then, according to the pre-set semantic classification rules or the semantic similarity calculation results generated by the large language model, the most matched intermediate node of each bottom-level node is determined, and the bottom-level node is classified into the corresponding intermediate node, thereby forming a mapping relationship from the bottom-level to the intermediate node.

[0065] Then, the system also performs semantic feature analysis on all intermediate nodes, and further classifies them into corresponding entity category nodes (i.e., top-level nodes) by judging their semantic categories or matching degrees with entity categories. The entity category nodes can include but are not limited to pre-defined general categories such as "person", "organization", "time", "place", "attribute", "action", and "relationship".

[0066] Through the above classification process, the semantic mapping of the bottom layer node → the intermediate node → the entity category node is realized, and then the knowledge graph with a clear semantic hierarchical structure is constructed, which is beneficial to subsequent entity aggregation, relationship reasoning and semantic retrieval and other application operations in the graph.

[0067] The device for parameterized construction of a QA-driven large model hierarchical knowledge graph provided by the application is described below, as shown in Figure 5 The device for parameterized construction of a QA-driven large model hierarchical knowledge graph described below can be correspondingly referred to the method for parameterized construction of a QA-driven large model hierarchical knowledge graph described above.

[0068] A device for parameterized construction of a QA-driven large model hierarchical knowledge graph comprises: A paragraph segmentation module 510 is configured to segment a user input document to be processed according to a preset word threshold to obtain multiple paragraphs, and each paragraph is bound to a unified document identifier. A question and answer data module 520 is configured to process each paragraph with the same document identifier in parallel, and call a pre-trained language model to generate a structured question and answer data list for each paragraph. A top node module 530 is configured to extract question content from the question and answer data list, and perform entity category recognition on the question content to obtain multiple entity types to form top nodes of a knowledge graph. A bottom node module 540 is configured to extract specific information from the question content according to the entity category to generate atomic nodes, and attribute each atomic node to a corresponding entity category to form bottom nodes of the knowledge graph. An intermediate node module 550 is configured to perform semantic clustering on the bottom nodes to form multiple levels of intermediate nodes. A graph construction module 560 is configured to construct a knowledge graph corresponding to the document to be processed according to the semantic hierarchical relationship among the top nodes, the multiple levels of intermediate nodes and the bottom nodes.

[0069] Figure 6 An entity structure diagram of an electronic device is shown in Figure 6 The electronic device can include a processor 610, a communications interface 620, a memory 630 and a communications bus 640, wherein the processor 610, the communications interface 620 and the memory 630 complete mutual communication through the communications bus 640. The processor 610 can call a logical instruction in the memory 630 to execute a method for parameterized construction of a QA-driven large model hierarchical knowledge graph, which comprises: splitting the user-inputted to-be-processed document according to a preset word threshold to obtain a plurality of paragraphs, each paragraph being bound with a unified document identifier; performing parallel processing on the paragraphs with the same document identifier, and calling a pre-trained language model to generate a structured question and answer data list for each paragraph; extracting question content from the question and answer data list, and performing entity category identification on the question content to obtain a plurality of entity types to form top nodes of a knowledge graph; extracting specific information from the question content according to the entity categories to generate atomic nodes, and attributing each atomic node to a corresponding entity category to constitute bottom nodes of the knowledge graph; performing semantic clustering on the bottom nodes to form a plurality of intermediate nodes; constructing a knowledge graph corresponding to the to-be-processed document according to a semantic hierarchical relationship among the top nodes, the plurality of intermediate nodes and the bottom nodes.

[0070] In addition, the logic instructions in the memory 630 described above can be implemented in the form of a software function unit and sold or used as an independent product, which can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.

[0071] On the other hand, the present application also provides a computer program product, which comprises a computer program, the computer program can be stored on a non-transitory computer readable storage medium, and the computer program can be executed by a processor to enable a computer to execute the method of parameterized construction of a QA-driven large model hierarchical knowledge graph provided by the above-mentioned methods, which comprises: splitting the user-inputted to-be-processed document according to a preset word threshold to obtain a plurality of paragraphs, each paragraph being bound with a unified document identifier; performing parallel processing on the paragraphs with the same document identifier, and calling a pre-trained language model to generate a structured question and answer data list for each paragraph; extracting question content from the question and answer data list, and performing entity category recognition on the question content to obtain a plurality of entity types to form top-level nodes of the knowledge graph; extracting specific information from the question content according to the entity category to generate atomic nodes, and attributing each atomic node to a corresponding entity category to constitute bottom-level nodes of the knowledge graph; performing semantic clustering on the bottom-level nodes to form a plurality of intermediate nodes; constructing the knowledge graph corresponding to the to-be-processed document according to the semantic hierarchical relationship among the top-level nodes, the plurality of intermediate nodes and the bottom-level nodes.

[0072] In another aspect, the present application also provides a non-transitory computer readable storage medium having a computer program stored thereon, the computer program being executed by a processor to implement the method of parameterized construction of a large model of a knowledge graph driven by QA provided by the above method, the method comprising: splitting the to-be-processed document input by a user according to a preset word threshold to obtain a plurality of paragraphs, each paragraph being bound to a unified document identifier; performing parallel processing on each paragraph having the same document identifier, and calling a pre-trained language model to generate a structured question and answer data list for each paragraph; extracting question content from the question and answer data list, and performing entity category recognition on the question content to obtain a plurality of entity types to form top-level nodes of the knowledge graph; extracting specific information from the question content according to the entity category to generate atomic nodes, and attributing each atomic node to a corresponding entity category to constitute bottom-level nodes of the knowledge graph; performing semantic clustering on the bottom-level nodes to form a plurality of intermediate nodes; constructing the knowledge graph corresponding to the to-be-processed document according to the semantic hierarchical relationship among the top-level nodes, the plurality of intermediate nodes and the bottom-level nodes.

[0073] The device embodiments described above are only schematic, wherein the units shown as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, i.e., can be located in one place, or can be distributed on a plurality of network units. Part or all of the modules can be selected according to actual needs to achieve the purpose of the present embodiment scheme. Those skilled in the art can understand and implement without creative labor.

[0074] Those skilled in the art can clearly understand the technical solutions of the various embodiments from the above description of the embodiments, and the various embodiments can be implemented by means of software with the necessary general hardware platforms, and of course, can also be implemented by hardware. Based on such understanding, the above technical solutions, essentially or in other words, the part of the prior art that makes a contribution, can be embodied in the form of a software product, which can be stored in a computer readable storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, and the like, and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0075] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, rather than limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for some technical features therein; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for QA-driven hierarchical knowledge graph parameterization construction of large models, characterized in that, The method comprises the following steps: The user inputted document to be processed is divided according to a preset word threshold to obtain multiple paragraphs, and each paragraph is bound to a unified document identifier; Parallel processing is performed on the paragraphs with the same document identifier, and a pre-trained language model is called to generate structured question and answer data list for each paragraph; The question content is extracted from the question and answer data list, and the entity category recognition is performed on the question content to obtain multiple entity types to form the top nodes of the knowledge graph; According to the entity category, specific information is extracted from the question content to generate atomic nodes, and each atomic node is attributed to the corresponding entity category to constitute the bottom nodes of the knowledge graph; Semantic clustering is performed on the bottom nodes to form multiple levels of intermediate nodes; According to the semantic hierarchical relationship among the top nodes, the multiple levels of intermediate nodes and the bottom nodes, the knowledge graph corresponding to the document to be processed is constructed.

2. The method of claim 1, wherein, Before the step of parallel processing the paragraphs with the same document identifier and calling the pre-trained language model to generate the structured question and answer data list for each paragraph, the method further comprises the following steps: According to the text length of each paragraph, the number of question and answer pairs required to be generated for each paragraph is estimated according to a preset mapping relationship between the number of words and the number of question and answer pairs, so as to guide the parallel configuration and resource allocation in the subsequent question and answer generation process.

3. The method of claim 1, wherein, After the step of extracting the question content from the question and answer data list and performing the entity category recognition on the question content to obtain multiple entity types, the method further comprises the following steps: For each recognized entity type, a corresponding abstract name and semantic definition are provided.

4. The method of claim 1, wherein, After the step of constructing the knowledge graph corresponding to the document to be processed according to the semantic hierarchical relationship among the top nodes, the multiple levels of intermediate nodes and the bottom nodes, the method further comprises the following steps: An adjustment instruction of the knowledge graph inputted by a user is acquired, and the adjustment instruction of the knowledge graph is used to indicate an editing operation on the nodes or edge relationship in the knowledge graph; According to the adjustment instruction of the knowledge graph, an entity name of a node, a node level or an edge relationship with other nodes in the knowledge graph is added, modified or deleted; The data structure of the knowledge graph is updated.

5. The method of claim 1, wherein, The step of performing semantic clustering on the bottom nodes to form multiple levels of intermediate nodes specifically comprises the following steps: A constraint condition is constructed, and the constraint condition comprises a graph hierarchical depth limit, a node naming length limit and a number limit of nodes at each level; On the premise of meeting the constraint condition, a clustering algorithm is used to perform hierarchical clustering from bottom to top on the bottom nodes to generate intermediate nodes with increasing semantic abstraction degree, and an upper and lower hierarchical relationship among the nodes is established.

6. The method of claim 1, wherein, The step of extracting specific information from the question content according to the entity category to generate atomic nodes and attributing each atomic node to the corresponding entity category to constitute the bottom nodes of the knowledge graph specifically comprises the following steps: According to the entity category, semantic analysis is performed on the question content to extract specific information with clear semantics to generate atomic nodes; According to the semantic features and semantic hierarchical relationship of each atomic node, the entity category to which the atomic node is attributed is determined.

7. The method of claim 1, wherein, The step of constructing the knowledge graph corresponding to the to-be-processed document according to the semantic hierarchical relationship of the top-level node, the multi-level intermediate node, and the bottom-level node specifically includes the following steps. The bottom-level nodes are classified into corresponding intermediate nodes according to semantic features; The intermediate nodes are classified into corresponding entity category nodes according to semantic features.

8. An apparatus for QA-driven hierarchical knowledge graph parameterization construction of a large model, characterized in that, The method comprises the following steps: A paragraph segmentation module is configured to segment a to-be-processed document input by a user into multiple paragraphs according to a preset word threshold, each paragraph being bound to a unified document identifier; A question and answer data module is configured to process each paragraph with the same document identifier in parallel, and to call a pre-trained language model to generate a structured question and answer data list for each paragraph; A top-level node module is configured to extract question content from the question and answer data list, and to perform entity category recognition on the question content to obtain multiple entity types, so as to form top-level nodes of a knowledge graph; A bottom-level node module is configured to extract specific information from the question content according to the entity categories, to generate atomic nodes, and to attribute each atomic node to a corresponding entity category, so as to form bottom-level nodes of the knowledge graph; An intermediate node module is configured to perform semantic clustering on the bottom-level nodes to form multi-level intermediate nodes; A graph construction module is configured to construct a knowledge graph corresponding to the to-be-processed document according to the semantic hierarchical relationship of the top-level node, the multi-level intermediate node, and the bottom-level node.

9. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor executes the program to implement the method for parameterized construction of a large model hierarchical knowledge graph driven by QA according to any one of claims 1 to 7. 10.A non-transitory computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the method for parameterized construction of a large model hierarchical knowledge graph driven by QA according to any one of claims 1 to 7.

Citation Information

Cited By

  • Method and device for generating contract knowledge graph

    CN121119095A