A hierarchical perception-based knowledge graph construction method
Patent Information
- Application Number
- CN202610926360.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-25
- Publication Date
- 2026-09-25
AI Technical Summary
当面对包含章、节、小节等多层级结构的教材或技术手册等长篇复杂文档时,这种平面化的分割与局部抽取方式会产生突出问题:平面化切分将原本语义关联的知识点打散至不同片段,导致跨段落、跨章节的全局上下文信息丢失,关键实体和关系因缺乏上下文支撑而无法被正确识别;局部抽取仅利用片段内部有限信息生成实体表示,缺乏上级章节摘要的高阶语义补充,导致实体特征区分度不足,同时抽取出的关系仅反映片段内平面关联,无法体现文档固有的层级语义,最终形成的图谱拓扑结构扁平化;此外,由于缺乏文档物理层级和语义层级的双重约束,现有的全局关系发现机制容易将分布于不同章节、无直接关联的实体错误连接,产生大量虚假的跨章节关系,严重降低知识图谱的准确性和可靠性
[0035](1)通过将文档原生层级结构显式建模为文档层级树并作为统一先验,避免了传统扁平化切分造成的上下文割裂,保留了跨段落、跨章节的全局语义关联,显著提升了知识抽取的完整性。
Smart Images

Figure CN122817474A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of natural language processing and knowledge engineering, specifically to a hierarchical knowledge graph construction method, which is particularly suitable for the automated knowledge structuring and knowledge graph generation of long and complex documents such as textbooks, technical manuals, and standards. Background Technology
[0002] Knowledge graphs, which organize entities, attributes, and semantic relationships in a structured form, have become an important infrastructure for tasks such as intelligent retrieval, question answering systems, educational resource modeling, and decision support. With the rapid development of large-scale pre-trained language models, the knowledge graph construction paradigm has gradually shifted from the traditional approach based on rule templates, feature engineering, and supervised learning to a generative construction paradigm based on large language models.
[0003] However, the aforementioned mainstream methods are primarily designed for processing short texts or flat text blocks. Their core idea is to segment the input text according to a fixed length or simple semantic boundaries, and then extract entities and relationships independently from each segment. When faced with long and complex documents such as textbooks or technical manuals with multi-level structures including chapters, sections, and subsections, this flat segmentation and local extraction approach presents significant problems: flat segmentation scatters semantically related knowledge points across different segments, resulting in the loss of global contextual information across paragraphs and chapters, and key entities and relationships cannot be correctly identified due to a lack of contextual support; local extraction only utilizes limited information within a segment to generate entity representations, lacking higher-order semantic supplementation from higher-level chapter summaries, leading to insufficient entity feature discriminability. Furthermore, the extracted relationships only reflect flat relationships within the segment, failing to reflect the inherent hierarchical semantics of the document, ultimately resulting in a flattened graph topology; in addition, due to the lack of dual constraints from the document's physical and semantic levels, existing global relationship discovery mechanisms are prone to incorrectly connecting entities distributed across different chapters without direct connections, generating a large number of false cross-chapter relationships, severely reducing the accuracy and reliability of the knowledge graph.
[0004] To address the aforementioned issues, some existing technical solutions attempt to incorporate document structure information into the knowledge graph construction process. For example, some solutions utilize document format, layout, and multi-source content for knowledge integration; some achieve graph organization from local to global by constructing entity graphs and community summaries; and some introduce retrieval enhancement mechanisms to expand contextual coverage. However, all of these methods treat document structure information merely as auxiliary metadata in the segmentation, retrieval, or candidate expansion stages, failing to systematically integrate it into core construction stages such as semantic representation enhancement, topological regularization, and global relation denoising. Specifically, existing solutions lack mechanisms to model chapter hierarchies as unified structural priors, fail to utilize ancestor summaries to enrich low-order entity representations, do not construct relation decay functions based on hierarchical distance to suppress cross-hierarchical noisy connections, and lack global relation edge self-pruning strategies based on structural consistency.
[0005] In summary, there is currently a lack of a comprehensive knowledge graph construction method that can use the original hierarchical structure of documents as a unified prior and integrate it into knowledge extraction, semantic enhancement, topological folding, and global relationship optimization. Summary of the Invention
[0006] To address the core shortcomings of traditional knowledge graph construction methods, which primarily target short texts or flat text blocks and employ fixed-length segmentation leading to the loss of global contextual information across paragraphs, sparse entity semantics, flattened graph topology, and the generation of numerous false cross-chapter relationships due to the lack of dual constraints of document physical and semantic hierarchy, this invention proposes a hierarchy-aware knowledge graph construction method. This method innovatively integrates the document hierarchy structure as a unified prior into the entire knowledge graph construction process. It constructs an explicit graph skeleton through hierarchical summarization and local extraction, and achieves semantic refinement and structural reorganization by combining multi-granular semantic enhancement and role-driven ontology folding. Furthermore, it performs global entity disambiguation, relation scoring, and adaptive pruning based on a hierarchy-aware distance decay function and structural consistency score, thereby constructing a knowledge graph that possesses both semantic expressiveness and hierarchical structural consistency.
[0007] To achieve the above objectives, the present invention adopts the following technical solution:
[0008] A hierarchical knowledge graph construction method includes the following steps:
[0009] 1) Obtain the document to be processed, identify the hierarchical headings and the corresponding body text in the document, and build a document hierarchy tree containing structural nodes and parent-child containment relationships based on the hierarchical level of the headings.
[0010] 2) Perform bottom-up hierarchical summary generation on each node in the document hierarchy tree to obtain the summary text of each level node; take each leaf node as a unit, concatenate the title, summary text and body text of the leaf node and input them into the preset extraction interface, extract entities through the preset extraction interface and record the entity name, extract the semantic relationship between entities within the same leaf node and record the relationship type, form the original relationship edge set, and establish the correspondence between each entity and its leaf node.
[0011] 3) For each entity extracted in step 2), find all ancestor nodes of the leaf node to which the entity belongs in the document hierarchy tree, obtain the summary text of each ancestor node, sort and concatenate the summary text of each ancestor node according to the hierarchical distance, and concatenate it with the name of the entity and the summary text of the leaf node to which it belongs. Input the summaries together into the semantic refinement interface, and generate an enhanced semantic description of the entity through the semantic refinement interface. Input the enhanced semantic description into the semantic vectorization model to generate the feature vector of the entity. Perform role recognition according to the local topic scope, divide the entity into core entities and subordinate entities, and convert the relationship edge connecting the core entity and the subordinate entity into a hierarchical subordinate edge from the core entity to the subordinate entity.
[0012] 4) For each entity feature vector generated in step 3), calculate the similarity between any two entity feature vectors, sort them according to similarity, and extract high-confidence candidate entity pairs based on similarity jump points. Input the high-confidence candidate entity pairs into the synonym determination interface, and perform semantic equivalence judgment through the synonym determination interface. Merge entities judged as synonyms into the same entity node. For any two merged entity nodes, determine their respective leaf nodes according to the entity-leaf node correspondence established in step 2). Find the nearest common ancestor node of the two leaf nodes in the document hierarchy tree, calculate the sum of the path distances from the two leaf nodes to the nearest common ancestor node as the hierarchical distance, calculate the hierarchical decay coefficient based on the hierarchical distance, and perform a weighted sum of the hierarchical decay coefficient, the feature vector similarity of the two entity nodes, and the neighborhood overlap of the two entity nodes to obtain the relationship score. Filter the original relationship edges according to the relationship score, and retain those with scores higher than a preset threshold as candidate relationship edges.
[0013] 5) For each candidate relation edge selected in step 4), the names of the entity nodes at both ends of the candidate relation edge, the enhanced semantic description, and the names of the neighboring entities of each entity node in the knowledge graph are concatenated and input into the rationality review interface. The rationality review interface is used to evaluate the semantic rationality and generate a rationality score for the candidate relation edge. If the score is lower than a preset threshold, the candidate relation edge is deleted. For the retained relation edges, the structural consistency score between the entities at both ends of the relation edge is calculated. Relation edges with a structural consistency score lower than a preset threshold are deleted. The output is a knowledge graph composed of entity nodes, retained relation edges, hierarchical subordinate edges, and anchor edges from entities to leaf nodes.
[0014] In step 1), when identifying the hierarchical headings and their corresponding body text in the document, the hierarchical level of the heading is determined according to at least one of the following features: heading indentation, font size, font style, and numbering format. The hierarchical level includes at least three levels: chapter, section, and subsection. The text content between each heading and the next heading of the same level is determined as the body text corresponding to that heading.
[0015] When constructing a document hierarchy tree, let the hierarchical structure of the original document be represented as a hierarchy tree T = (VT, ET), where VT represents the set of structural nodes in the document and ET represents the set of parent-child inclusion relationships between hierarchy nodes; let V(1) be the set of nodes at level 1, where the value of 1 corresponds to levels such as books, chapters, sections, and subsections; if node u ∈ V(1) is the parent node of node v ∈ V(1+1), then establish the parent-child relationship edge (u, v) ∈ ET; by traversing from top to bottom, determine the parent node of each non-root node, and finally construct a complete document hierarchy tree.
[0016] In step 2), when performing bottom-up hierarchical summary generation for each node in the document hierarchy tree, for leaf nodes, the corresponding text of the leaf node is input into a preset extraction interface, and the summary text of the leaf node is generated through the preset extraction interface; for non-leaf nodes, the summary texts of each child node of the non-leaf node are concatenated and input into the preset extraction interface, and the summary text of the non-leaf node is generated through the preset extraction interface.
[0017] The summary generation is performed in a post-order traversal manner, that is, all child nodes are processed first and then the parent node is processed, ensuring that when each node generates a summary, the summaries of all its child nodes have been generated. The entire summary generation process starts from the leaf node at the bottom level of the document hierarchy tree and goes up layer by layer, eventually generating summary text for all level nodes including the root node.
[0018] When extracting entities and relations on a per-leaf-node basis, the title, summary text, and body text of the leaf node are concatenated in sequence to form a structured input text. The structured input text is then submitted to a preset extraction interface, which outputs the set of entities and the set of relation edges within the scope of the leaf node. Each entity records its name, and each relation edge records its head entity, tail entity, and relation type.
[0019] When establishing the correspondence between each entity and its leaf node, for each entity e extracted from the leaf node p, an anchor edge is created pointing from the leaf node p to the entity e, and recorded as (p, e): all anchor edges constitute the mapping set Φ from the entity to the document hierarchy tree node; the mapping set Φ is used to determine the source location of the entity and calculate the hierarchical distance between entities.
[0020] In step 3), when sorting and concatenating the summary texts of each ancestor node according to hierarchical distance, for entity e, let its leaf node be p0, the parent node of p0 be p1, the parent node of p1 be p2, and so on upwards until the root node; obtain the summary texts corresponding to each ancestor node p1, p2, ..., pk; concatenate them in order from closest to furthest from the leaf node where the entity is located, so that the ancestor node summaries with closer hierarchical distances are arranged at the front, thereby giving higher attention weight to the direct superior context that is most closely related to the semantics of the entity.
[0021] When generating an enhanced semantic description of an entity, the entity name, the summary text of its leaf node, and the summary text of its ancestor nodes at all levels, concatenated in sequence, are input into the semantic refinement interface. A semantic refinement instruction is set, and the semantic refinement interface outputs the enhanced semantic description of the entity after fusing local context and multi-granularity hierarchical context. This enhanced semantic description contains the entity's direct information within the local scope and its semantic positioning information in the global document hierarchy.
[0022] When the enhanced semantic description is input into the semantic vectorization model to generate the feature vector of the entity, the enhanced semantic description is input into the semantic encoder based on the pre-trained language model. After forward inference calculation by the model, the dense feature vector of the entity is output. The feature vector is used for similarity calculation, entity disambiguation and relation scoring in subsequent steps.
[0023] When determining roles based on local topic scope, entities within the local topic scope are input into a preset extraction interface for role determination, and the entities are divided into core entities and subordinate entities. When a relationship edge points from a subordinate entity to a core entity, the relationship edge is reversed and converted into a hierarchical subordinate edge pointing from a core entity to a subordinate entity. When a relationship edge already points from a core entity to a subordinate entity, it is directly marked as a hierarchical subordinate edge to construct a local hierarchical classification structure. When the entities at both ends of a relationship edge are core entities or both are subordinate entities, the edge is retained as a planar semantic relationship edge without being collapsed.
[0024] In step 4), when performing global entity disambiguation based on entity feature vectors, for each entity, the similarity of its feature vectors with all other entities is calculated, and the similarity is sorted from high to low. The similarity is converted into a distance metric to obtain an ascending distance sequence. The difference between adjacent distances is calculated, and the position with the largest difference is taken as the cutoff position. The candidate entity pairs before the cutoff position are retained as high-confidence candidate entity pairs.
[0025] When sorting by similarity and truncating high-confidence candidate entity pairs based on similarity jump points, for entity ei, calculate the feature vector similarity between it and all other entities; sort by similarity from high to low to form a similarity sequence; calculate the difference between adjacent similarities, take the position with the largest difference as the truncation position, and retain the candidate entity pairs before the truncation position as high-confidence candidate entity pairs.
[0026] When performing semantic equivalence judgment on high-confidence candidate entity pairs input to the synonym judgment interface, the entity names and their enhanced semantic descriptions in the candidate entity pairs are concatenated and input into the synonym judgment interface. A synonym judgment instruction is set, and the synonym judgment interface outputs a binary classification judgment result on whether the entities are synonyms. Entities judged as synonyms are merged into equivalence classes through a disjoint-set data structure, and all entities in the same equivalence class are merged into the same global entity node.
[0027] When searching for the nearest common ancestor (LCA) node of two leaf nodes in the document hierarchy tree, for any two entity nodes, their respective leaf nodes are determined according to the mapping set Φ. When searching for the LCA node of two leaf nodes in the document hierarchy tree, the sum of the path distances from the two leaf nodes to the LCA node is calculated as the hierarchical distance. The smaller the hierarchical distance, the closer the two entities are in the document hierarchy structure and the closer their semantic relationship.
[0028] The hierarchical attenuation coefficient is calculated based on the hierarchical distance using the following formula: HADD(u,v)=exp(-λ·d(u,v)) Where HADD(u, v) is the hierarchical decay coefficient between entity u and entity v, d(u, v) is the sum of the path distances from the leaf nodes of entity u and entity v to their nearest common ancestor node, and λ is a preset decay constant.
[0029] The relationship score is obtained by weighted summing of the hierarchical attenuation coefficient, the similarity of the feature vectors of the two entity nodes, and the neighborhood overlap of the two entity nodes, using the following formula: Score(u,v)=α·cos(z u ,z v )+β·AA(u,v)+γ·HADD(u,v) Among them, cos(Z) u , z v ) represents the cosine similarity of entity vectors, AA(u,v) represents the neighborhood overlap based on the Adamic-Adar index, HADD(u,v) represents the hierarchical decay coefficient, and α, β, and γ are preset weight coefficients.
[0030] When filtering candidate relationship edges based on relationship scores, all candidate entity pairs are sorted in descending order according to the calculated relationship scores. Candidate relationship edges with scores higher than a preset threshold are retained, or candidate relationship edges with higher rankings are retained according to a preset ratio, and then proceed to the subsequent review stage.
[0031] In step 5), when the entity node name, enhanced semantic description, and neighbor entity names are concatenated and input into the rationality review interface for semantic rationality evaluation, the entity node names, enhanced semantic descriptions, and neighbor entity name lists of each entity node in the knowledge graph at both ends of the candidate relationship edge are concatenated into a review prompt text and input into the rationality review interface. The rationality review interface outputs the rationality score of the candidate relationship edge; if the score is lower than the preset rationality threshold, the candidate relationship edge is deleted.
[0032] The structural consistency score is calculated using the following formula: SCS(u,v)=w1·cos(z u ,z v )+w2·NS(u,v)+w3·HADD(u,v) in, N(u) represents the set of neighboring entities of entity u, and w1, w2, and w3 are preset weight coefficients; relationships with structural consistency scores lower than the preset structural threshold are deleted.
[0033] The final output knowledge graph consists of entity nodes, preserved semantic relationship edges, hierarchical subordinate edges, and anchor edges from entities to leaf nodes. Each entity node in the knowledge graph is associated with a corresponding position in the document hierarchy tree through anchor edges, achieving explicit alignment between the knowledge graph and the document hierarchy structure.
[0034] The present invention achieves the following beneficial effects through the above technical solution:
[0035] (1) By explicitly modeling the native hierarchical structure of the document as a document hierarchy tree and using it as a unified prior, the contextual fragmentation caused by traditional flat segmentation is avoided, and the global semantic association across paragraphs and chapters is preserved, which significantly improves the completeness of knowledge extraction.
[0036] (2) By using bottom-up hierarchical summarization and multi-granularity ancestor context enhancement, the entity semantic sparsity problem caused by local extraction is effectively alleviated, and the entity representation's distinguishability and semantic richness are improved.
[0037] (3) By reconstructing the local planar relational network through role-driven ontology folding, the taxonomic features of the graph are enhanced, effectively overcoming the defects of topological flattening.
[0038] (4) By introducing a hierarchical perceived distance decay function and a structural consistency score, hierarchical proximity constraints and posterior adaptive pruning are applied to cross-chapter entity relationships, which suppresses the generation of false relationships and significantly improves the reliability of relationship modeling.
[0039] In summary, by using the document hierarchy as a unified prior throughout the entire process of knowledge extraction, semantic enhancement, topological folding, and global relation optimization, this invention can construct a high-quality knowledge graph that combines semantic expressiveness with hierarchical structure consistency. It is suitable for the knowledge organization of long and complex documents such as textbooks and technical manuals. Attached Figure Description
[0040] Figure 1 This is a flowchart illustrating the hierarchical awareness-based knowledge graph construction method provided in this embodiment of the invention.
[0041] Figure 2 This is a schematic diagram of the overall framework for constructing a knowledge graph based on hierarchical perception, provided in an embodiment of the present invention.
[0042] Figure 3 This is a schematic diagram illustrating the mapping relationship between a document hierarchy tree and a knowledge graph entity provided in an embodiment of the present invention.
[0043] Figure 4 This is a visual comparison diagram of the results generated by different knowledge graph construction methods. Detailed Implementation
[0044] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.
[0045] This embodiment uses a classic third edition textbook in the field of artificial intelligence as the document to be processed. A knowledge graph is constructed from the entire textbook, comprising 12 chapters and approximately 300,000 words, to illustrate the specific implementation process of this invention. The overall workflow of this invention is as follows: Figure 1 As shown, the five core stages are: document hierarchy tree construction, hierarchical summarization and local extraction, semantic enhancement and ontology folding, entity disambiguation and relation scoring, and LLM review and structural pruning. Figure 1 In this invention, "document hierarchy tree construction" corresponds to step 1), "hierarchical summarization and local extraction" corresponds to step 2), "semantic enhancement and ontology folding" corresponds to step 3), "entity disambiguation and relation scoring" corresponds to step 4), and "LLM review and structural pruning" corresponds to step 5. The overall framework of this invention is as follows: Figure 2 As shown, the input document passes through the structure-aware layer, the semantic construction layer, and the optimization output layer in sequence. The structure-aware layer is responsible for hierarchical tree reconstruction and anchoring, the semantic construction layer is responsible for entity enhancement and relation extraction, and the optimization output layer is responsible for global disambiguation and structural pruning, ultimately generating a knowledge graph that is explicitly aligned with the document's hierarchical structure.
[0046] In step 1), the PDF document of the teaching material to be processed is first obtained. The hierarchical headings in the document are identified, and their hierarchical levels are determined according to font size (first-level headings are bold, size 3; second-level headings are bold, size 4; third-level headings are bold, size 4), and numbering format (e.g., "Chapter 3", "3.1", "3.1.1"), etc. Specifically, "Chapter 3 Machine Learning" and "Chapter 5 Deep Learning" are marked as L1 level (chapter level), "3.1 Supervised Learning" and "3.2 Unsupervised Learning" are marked as L2 level (section level), and "3.1.1 Linear Regression" and "3.1.2 Logistic Regression" are marked as L3 level (subsection level). The text content between each heading and the next heading of the same level is determined as the corresponding body text.
[0047] Establishing parent-child relationships from top to bottom: If node u∈V(1) is the parent node of node v∈V(1+1), then establish a parent-child relationship edge (u,v)∈ET, and finally construct a complete document hierarchy tree T. The document hierarchy tree structure constructed in this embodiment is as follows: Figure 3As shown in the left half, the root node is "Books", and its child nodes are "Chapter 3" and "Chapter 5"; the child nodes of "Chapter 3" are "3.1" and "3.2"; the child nodes of "3.1" are "3.1.1" and "3.1.2", and so on. This embodiment contains a total of 1 root node, 12 L1 level nodes, 47 L2 level nodes, and 186 L3 level (leaf) nodes.
[0048] In step 2), a post-order traversal is performed on the document hierarchy tree to generate a hierarchical summary, and local entity and relation extraction is performed for each leaf node.
[0049] For a leaf node p, such as section "3.1.1", the text corresponding to this leaf node is truncated to 512 tokens and input into a preset extraction interface. In this embodiment, the preset extraction interface uses the DeepSeek-V3.2 large language model, with an inference temperature set to 0.1. The input prompt is: "Please summarize the core technical content of the following text in no more than 50 words", and the output is the summary s(p) of this leaf node. For a non-leaf node u, such as section "3.1", the summaries of all its child nodes, namely "3.1.1" and "3.1.2", are concatenated in order, with a total length not exceeding 1024 tokens. This is then input into the DeepSeek-V3.2 model again, with the prompt: "Combining the summaries of the following sub-sections, generate the core overview of this section", and the output is the summary s(u) of this non-leaf node. For example, the summary of "section 3.1" is: "This section mainly introduces linear regression and logistic regression algorithms in supervised learning, covering loss function construction and gradient descent optimization strategies."
[0050] Using each leaf node p as the smallest extraction unit, the title, summary text, and body text of that leaf node are concatenated in sequence to form structured input text. A preset extraction interface is called, which uses the DeepSeek-V3.2 model and is configured with a few-shot prompt word template. The output is constrained to be in JSON format. Entities are extracted and their names are recorded. Simultaneously, semantic relationships between entities within the same leaf node are extracted and their relationship types are recorded.
[0051] For example, in section "3.1.1", the entities "linear regression", "loss function (MSE)", and "gradient descent" are extracted; the relations extracted are "linear regression has a loss function" and "gradient descent optimizes linear regression". The correspondence between each entity and its corresponding leaf node, i.e., anchor edges, is established. For entity e ("linear regression"), its corresponding leaf node p ("3.1.1"), an anchor edge (p, e) is created and recorded in the mapping set Φ. For example... Figure 3As shown in the upper right half, entities E1 and E2 are vertically associated with nodes at the section level through anchor edges, i.e., the edges represented by vertical dashed lines in the figure, thus establishing an explicit correspondence between knowledge graph entities and the physical location of documents. Figure 3 The lower right corner further illustrates the result of role-driven ontology folding: within the local theme of a section, the core entity is located at the center, and subordinate entities point to the core entity through hierarchical subordinate edges, i.e., the edges represented by dashed lines with arrows in the figure, thus constructing a core-subordinate classification hierarchy structure, reconstructing the originally flat planar relationship network into a hierarchical topology with taxonomic features.
[0052] In step 3), for each entity extracted in step 2), multi-granularity ancestor context enhancement is performed, and ontology folding is performed based on role recognition.
[0053] For entity e, i.e., "linear regression", its leaf node p0, i.e., "3.1.1", is located in the document hierarchy tree. The parent node p1, i.e., "3.1", and its summary "Supervised Learning Core Algorithm..." are retrieved, along with the grandparent node p2, i.e., "Chapter 3", and its summary "Machine Learning Overview...". The nodes are then concatenated according to their hierarchical distance, from closest to furthest, i.e., in the order of p1 and p2. The entity name "linear regression", its leaf node summary, i.e., "...linear regression principle...", and the concatenated ancestor summary are input into the semantic refinement interface. In this embodiment, the semantic refinement interface uses the DeepSeek-V3.2 model, and the refinement instruction is set as: "Generating an enhanced semantic description of this entity by combining the background information of the parent chapter, requiring it to include its domain, core definition, and function." The output enhanced semantic description is: "Linear regression is a classic algorithm in supervised learning, belonging to the Machine Learning chapter. It is mainly used to predict continuous numerical values by minimizing the mean squared error loss function and fitting a linear relationship using gradient descent."
[0054] The enhanced semantic description described above is input into the semantic vectorization model to generate the feature vector of the entity. In this embodiment, the semantic vectorization model uses the all-MiniLM-L6-v2 pre-trained model, and the output is a 384-dimensional dense feature vector z. e .
[0055] Within the local topic scope, specifically section "3.1", entity input is processed through a pre-defined extraction interface for role determination. Based on the determination results, "Linear Regression" and "Logistic Regression" are classified as core entities (Core), while "Loss Function" and "Gradient Descent" are classified as subordinate entities (Non-core). For relationship edges connecting core and subordinate entities, if a subordinate entity points to a core entity (e.g., "Loss Function belongs to Linear Regression"), the relationship edge is reversed and converted into a hierarchical subordinate edge pointing from a core entity to a subordinate entity (e.g., "Linear Regression includes Loss Function"). If both ends of the relationship edge are core entities (e.g., "Linear Regression vs. Logistic Regression"), it is retained as a planar semantic relationship edge. Through this folding operation, a framework is constructed as follows: Figure 3 The “core-subordinate” classification hierarchy is shown on the lower right.
[0056] In step 4), global entity disambiguation is performed on each entity feature vector generated in step 3), and relationship scores are calculated based on hierarchical perception distance decay.
[0057] Based on the 187 entity feature vectors generated in step 3), a nearest neighbor retrieval index is built using the FAISS library. For each entity ei, the cosine similarity of its feature vectors with all other entities is calculated. Similarity sequences are formed by sorting them from high to low. The difference between adjacent similarities is calculated, and the position with the largest difference is taken as the cutoff position. Candidate entity pairs before the cutoff position are retained as high-confidence candidate entity pairs. For example, for entity SGD, there are candidate entities Stochastic Gradient Descent and Stochastic Gradient Descent. The high-confidence candidate entity pairs are input into the synonym determination interface. In this embodiment, DeepSeek-V3.2 is used. The names and enhanced semantic descriptions of both entities are input, and the instruction is set to determine whether they are the same entity, with an answer of yes or no. After determination, SGD and Stochastic Gradient Descent are synonyms and are merged into the same entity node through a union lookup set; while Linear Regression and Logistic Regression, although semantically similar, are not synonyms and are not merged.
[0058] For any two merged entity nodes u and v, their respective leaf nodes are determined according to the mapping set Φ. In this embodiment, entity u, i.e., "linear regression", is anchored to leaf node "3.1.1", and entity v, i.e., "overfitting", is anchored to leaf node "3.2.2". The nearest common ancestor (LCA) node is found in the document hierarchy tree, which is "Chapter 3", i.e., level L1. The sum of the path distances d(u,v) from the two leaf nodes to this LCA node is calculated: the leaf node '3.1.1' to u needs to trace back 2 levels (3.1 and Chapter 3 in sequence) to its LCA node, so the path distance is 2. The distance from v to the LCA is 2, so d(u,v) = 4. Using a preset attenuation constant λ = 0.3, the hierarchical attenuation coefficient HADD(u,v) = exp(-0.3 × 4) ≈ 0.301 is calculated according to the formula.
[0059] Calculate the cosine similarity of entity vectors cos(z) u , z v The Adamic-Adar index has a neighborhood overlap of AA(u,v) = 0.65, and the neighborhood overlap is AA(u,v) = 0.45. In this embodiment, the Adamic-Adar index is normalized to the [0,1] interval by Min-Max before participating in the weighted calculation. Preset weight coefficients α = 0.4, β = 0.3, and γ = 0.3 are set, and the relationship score is calculated according to the formula: Score(u,v)=0.4×0.65+0.3×0.45+0.3×0.301=0.4853. In this embodiment, the preset threshold for the relationship score is 0.45. Since 0.4853 is greater than 0.45, the candidate relationship edge, namely "linear regression and overfitting correlation", is retained and enters the candidate relationship edge set.
[0060] In step 5), for each candidate relation edge selected in step 4), semantic rationality is first checked, and then structural consistency self-pruning is performed.
[0061] The names of the entity nodes at both ends of the candidate relationship edge, the enhanced semantic description, and the names of their respective neighboring entities in the knowledge graph are concatenated and input into the rationality review interface. This embodiment uses the DeepSeek-V3.2 model, with the instruction set as: "Please determine whether the following entity pairs are reasonably connected in the knowledge graph, and give a score between 0 and 1." For example, input entity A is "linear regression," whose neighbors are "loss function, gradient descent," and entity B is "overfitting," whose neighbors are "regularization, validation set." The model outputs a rationality score of 0.2, which is lower than the preset rationality threshold of 0.6. Therefore, this relationship is determined to be a cross-chapter noise edge and is deleted.
[0062] For each retained relation edge, a structural consistency score (SCS) is calculated. Preset weight coefficients are set: w1 = 0.5, w2 = 0.2, and w3 = 0.3. For a given retained relation edge, its neighborhood support is calculated according to the definition of NS(u,v) in the structural consistency score formula, resulting in NS(u,v) = 0.3. Combining this with a cosine similarity of 0.6 and a HADD coefficient of 0.5, the SCS is calculated as: SCS = 0.5 × 0.6 + 0.2 × 0.3 + 0.3 × 0.5 = 0.51. In this embodiment, the preset structural threshold is 0.5. Since 0.51 is greater than 0.5, this relation edge is ultimately retained.
[0063] Finally, the output is a knowledge graph composed of entity nodes, semantically preserving edges, hierarchical subordinate edges, and anchor edges from entities to leaf nodes. Each entity node in this graph is associated with a corresponding position in the document hierarchy tree through anchor edges, achieving explicit alignment between the knowledge graph and the document hierarchy structure.
[0064] To verify the effectiveness of the present invention, this embodiment constructs a knowledge graph across 12 chapters throughout the book and sets up comparative and ablation experiments.
[0065] AutoKG (a general extractive graph construction method based on large language models), EDC (an extract-definition-normalization decoupling method), and GraphRAG (a graph RAG method based on entity graphs and community summaries) were selected as baselines for comparison. Due to the lack of complete publicly labeled data in long textbooks, this embodiment adopts a two-level evaluation scheme: In the truth evaluation subset, computer science experts were invited to manually annotate the first three chapters, totaling 500 entities and 1200 relations, to construct a reference graph Ggt; in the LLM-assisted evaluation of the entire book, the GPT-4 model was used for blind scoring across the entire book, based on two dimensions: entity semantic quality (ES), with scores ranging from 0 to 1, and relation rationality (RS), with scores ranging from 0 to 10.
[0066] In this embodiment, DeepSeek-V3.2 is used as the basic generative model for knowledge extraction and optimization, with a temperature set to 0.1. The all-MiniLM-L6-v2 model is used for entity embedding, and the FAISS index is used for nearest neighbor retrieval. In the relationship scoring stage, the hierarchical decay constant λ is set to 0.3, and the structural consistency self-pruning threshold is set to 0.6.
[0067] On manually annotated subsets, precision (P), recall (R), and F1 scores were calculated based on mapped entities, and the mapping edge connectivity (MEC) was introduced to evaluate the ability of the graph structure to preserve the document hierarchy. The results are shown in Table 1.
[0068] Table 1 Truth evaluation results on manually labeled subsets
[0069] As shown in Table 1, the F1 score of this invention reaches 0.7617, which is 0.1264 higher than the second-best GraphRAG, indicating that this invention significantly outperforms existing methods in the joint extraction accuracy of entities and relations. In particular, the MEC score of this invention reaches 0.6010, which is much higher than EDC's 0.3745 and GraphRAG's 0.3724, proving that this invention can effectively maintain the original hierarchical topology of the document and significantly suppress false relationships across levels, i.e., relational illusions, through hierarchical-aware distance decay and structural consistency pruning.
[0070] The results of the LLM auxiliary evaluation across the entire book are shown in Table 2.
[0071] Table 2. Results of LLM-related assessments across the entire book.
[0072] As shown in Table 2, the entity semantic quality (ES) of this invention is as high as 0.9754, which is 0.2056 higher than that of AutoKG, proving that the entity features have extremely high semantic discriminability and richness through multi-granularity ancestor context enhancement; the relation rationality (RS) of this invention reaches 7.3759, which is 0.3509 higher than that of GraphRAG, proving that the relation edges selected based on hierarchical perception scoring and self-pruning mechanism have stronger logical rationality and reliability.
[0073] To verify the effectiveness of each core component of this invention, an ablation experiment was conducted in this embodiment. Using a model without the aforementioned three modules as a baseline, three modules—Multi-granularity Context Enhancement (MGCE), Role-Driven Folding (Folding), and Hierarchical Awareness Distance Attenuation (HADD)—were gradually introduced, and their impact on the ES and RS indices was observed. The results are shown in Table 3.
[0074] Table 3 Ablation Experiment Results
[0075] As shown in Table 3, both ES and RS showed an upward trend after gradually introducing various modules based on the baseline model. Specifically, the introduction of MGCE significantly improved ES from 0.7333 to 0.9050, verifying the significant mitigation effect of multi-granularity ancestor summarization on the entity semantic sparsity problem; the introduction of Folding improved RS from 7.0615 to 7.1809, verifying the optimization effect of role-driven folding on the local topology; and finally, the introduction of HADD further improved RS to 7.3759, proving that the hierarchical perceptual distance decay function can effectively suppress cross-chapter noisy connections and is a key element in improving the reliability of relation modeling.
[0076] Visualization results of knowledge graphs generated by different methods, such as Figure 4 As shown in the figure, the top left is AutoKG, the top right is EDC, the bottom left is GraphRAG, and the bottom right is the present invention, HAKG. From Figure 4 The significant differences in the morphology of the spectral structures produced by each method can be clearly seen:
[0077] Figure 4 The top left graph generated by AutoKG has an extremely loose distribution of nodes, with messy and disordered connections between entities. It lacks a clear central hub and hierarchical structure, and a large number of distant entities across chapters are incorrectly associated, indicating that it has produced a serious illusion of relationships due to the lack of hierarchical priors. Figure 4The top right corner, EDC, is an improvement over AutoKG, with some relief from entity redundancy issues. However, the overall topology of the graph is still planar, with core and subordinate concepts mixed together, failing to reflect the inherent chapter-section-subsection hierarchical logic of the document. Figure 4 The lower left side, GraphRAG, demonstrates a strong ability to integrate global information. The central region of its graph shows a high-density cluster of local communities, indicating that it can capture some semantic relationships across paragraphs through community summaries. However, there are still a large number of isolated subgraphs and sparse branches on the periphery of the graph, and there is a lack of clear hierarchical bridging edges between communities, indicating that it is difficult to organize local communities into a unified hierarchical knowledge system. Figure 4 The graph generated by HAKG in the lower right corner of this invention exhibits the clearest hierarchical organization: the top layer clearly identifies several large core node clusters corresponding to "chapter" level topics; within each core cluster, secondary nodes corresponding to "sections" are radially distributed; the outermost "section" level subordinate entities are tightly arranged around their corresponding core entities. Planar semantic relationship edges and hierarchical subordinate edges are distinguished by different line types, with hierarchical subordinate edges represented by directed dashed lines, clearly showing the hierarchical constraint direction of "core entities pointing to subordinate entities." The entire graph's relationship edges are compact and orderly, with almost no long-distance noise edges across levels. This visualization result fully demonstrates that this invention, by using the document hierarchical structure as a unified prior throughout the entire process, can construct a high-quality knowledge graph with rich semantic expression and a consistent hierarchical structure.
[0078] Those skilled in the art will understand that the preset extraction interface, semantic refinement interface, synonym determination interface, and rationality verification interface in the above embodiments are not limited to implementations based on DeepSeek-V3.2. In alternative embodiments, the above interfaces can be implemented using open-source large language models such as OpenAI GPT-4 API, LLaMA-3-70B-Instruct local deployment, or Tongyi Qianwen, requiring only adjustments to the corresponding prompt word templates to achieve the same technical effect. In another alternative embodiment, entity and relation extraction can also be achieved by combining a named entity recognition model based on BERT-BiLSTM-CRF with a relation extraction model based on dependency parsing. All the above alternative solutions fall within the protection scope of this invention.
[0079] In summary, this embodiment successfully constructs a high-quality knowledge graph that combines semantic expressiveness with hierarchical structure consistency by using the document hierarchy as a unified prior throughout the entire process of knowledge extraction, semantic enhancement, topological folding, and global relation optimization. Comparative experiments, ablation experiments, and visualization analysis all verify the significant progress of this invention compared to existing technologies, confirming its practical value and industrial application prospects in the knowledge organization of long and complex documents such as textbooks and technical manuals.
Claims
1. A method for constructing a knowledge graph based on hierarchical awareness, characterized in that, Includes the following steps: 1) Obtain the document to be processed, identify the hierarchical headings and the corresponding body text in the document, and build a document hierarchy tree containing structural nodes and parent-child inclusion relationships based on the hierarchical level of the headings; 2) Perform bottom-up hierarchical summary generation on each node in the document hierarchy tree to obtain the summary text of each level node; take each leaf node as a unit, concatenate the title, summary text and body text of the leaf node and input them into the preset extraction interface, extract entities through the preset extraction interface and record the entity name, extract the semantic relationship between entities within the same leaf node and record the relationship type, form the original relationship edge set, and establish the correspondence between each entity and its leaf node; 3) For each entity extracted in step 2), find all ancestor nodes of the leaf node to which the entity belongs in the document hierarchy tree, obtain the summary text of each ancestor node, sort and concatenate the summary text of each ancestor node according to the hierarchical distance, and concatenate it with the name of the entity and the summary text of the leaf node to which it belongs. Input the summaries together into the semantic refinement interface, and generate an enhanced semantic description of the entity through the semantic refinement interface; input the enhanced semantic description into the semantic vectorization model to generate the feature vector of the entity; perform role recognition according to the local topic scope, divide the entity into core entities and subordinate entities, and convert the relationship edge connecting the core entity and the subordinate entity into a hierarchical subordinate edge from the core entity to the subordinate entity. 4) For each entity feature vector generated in step 3), calculate the similarity between any two entity feature vectors, sort them according to similarity, and extract high-confidence candidate entity pairs based on similarity jump points. Input the high-confidence candidate entity pairs into the synonym determination interface, and perform semantic equivalence judgment through the synonym determination interface. Merge entities judged as synonyms into the same entity node. For any two merged entity nodes, determine their respective leaf nodes according to the entity-leaf node correspondence established in step 2). Find the nearest common ancestor node of the two leaf nodes in the document hierarchy tree, calculate the sum of the path distances from the two leaf nodes to the nearest common ancestor node as the hierarchical distance, calculate the hierarchical decay coefficient based on the hierarchical distance, and perform a weighted sum of the hierarchical decay coefficient, the feature vector similarity of the two entity nodes, and the neighborhood overlap of the two entity nodes to obtain the relationship score. Filter the original relationship edges according to the relationship score, and retain those with scores higher than a preset threshold as candidate relationship edges. 5) For each candidate relation edge selected in step 4), the names of the entity nodes at both ends of the candidate relation edge, the enhanced semantic description, and the names of the neighboring entities of each entity node in the knowledge graph are concatenated and input into the rationality review interface. The rationality review interface is used to evaluate the semantic rationality and generate a rationality score for the candidate relation edge. If the score is lower than a preset threshold, the candidate relation edge is deleted. For the retained relation edges, the structural consistency score between the entities at both ends of the relation edge is calculated. Relation edges with a structural consistency score lower than a preset threshold are deleted. The output is a knowledge graph composed of entity nodes, retained relation edges, hierarchical subordinate edges, and anchor edges from entities to leaf nodes.
2. The method according to claim 1, characterized in that, Step 1) of identifying hierarchical headings and corresponding body text in the document specifically includes: determining the hierarchical level of the heading according to at least one of the following features: heading indentation, font size, font style, and numbering format. The hierarchical level includes at least three levels: chapter, section, and subsection. The text content between each heading and the next heading of the same level is determined as the body text corresponding to that heading.
3. The method according to claim 1, characterized in that, Step 2) involves generating a bottom-up hierarchical summary for each node in the document hierarchy tree. Specifically, this includes: for leaf nodes, inputting the corresponding text into a preset extraction interface to generate the summary text of the leaf node; for non-leaf nodes, concatenating the summary texts of each child node of the non-leaf node and inputting them into a preset extraction interface to generate the summary text of the non-leaf node.
4. The method according to claim 1, characterized in that, Step 2) establishes the correspondence between each entity and its leaf node, specifically including: for each entity e extracted from the leaf node p, create an anchor edge from the leaf node p to the entity e, recorded as (p, e); all anchor edges constitute a mapping set from entities to document hierarchy tree nodes; the mapping set is used to determine the source location of entities and calculate the hierarchical distance between entities.
5. The method according to claim 1, characterized in that, In step 3), the summary texts of each ancestor node are sorted and concatenated according to the hierarchical distance. Specifically, this includes: for entity e, let its leaf node be p0, the parent node of p0 be p1, the parent node of p1 be p2, and so on upwards until the root node; obtain the summary texts corresponding to each ancestor node p1, p2, ..., pk; and concatenate them in order of distance from the leaf node where the entity is located, so that the ancestor node summaries with closer hierarchical distances are arranged at the beginning.
6. The method according to claim 1, characterized in that, In step 3), role recognition is performed according to the local topic scope, and entities are divided into core entities and subordinate entities. For the relationship edge connecting the core entity and the subordinate entity, the relationship edge is converted into a hierarchical subordinate edge pointing from the core entity to the subordinate entity. Specifically, this includes: inputting entities within the local topic scope into a preset extraction interface for role determination, and dividing the entities into core entities and subordinate entities; when the relationship edge points from the subordinate entity to the core entity, the relationship edge is converted in reverse to a hierarchical subordinate edge pointing from the core entity to the subordinate entity; when the relationship edge has already pointed from the core entity to the subordinate entity, it is directly marked as a hierarchical subordinate edge to construct a local hierarchical classification structure.
7. The method according to claim 1, characterized in that, Step 4) involves sorting by similarity and selecting high-confidence candidate entity pairs based on similarity jump points. Specifically, this includes: for entity ei, calculating the feature vector similarity between it and all other entities; sorting by similarity from high to low to form a similarity sequence; calculating the difference between adjacent similarities, taking the position with the largest difference as the cutoff position, and retaining the candidate entity pairs before the cutoff position as high-confidence candidate entity pairs.
8. The method according to claim 1, characterized in that, In step 4), the hierarchical attenuation coefficient is calculated based on the hierarchical distance, specifically using the following formula: HADD(u,v)=exp(-λ·d(u,v)) Where HADD(u, v) is the hierarchical decay coefficient between entity u and entity v, d(u, v) is the sum of the path distances from the leaf nodes of entity u and entity v to their nearest common ancestor node, and λ is a preset decay constant.
9. The method according to claim 1, characterized in that, In step 4), the relationship score is obtained by weighted summation of the hierarchical attenuation coefficient, the feature vector similarity between the two entity nodes, and the neighborhood overlap between the two entity nodes, specifically using the following formula: Score(u,v)=α·cos(z u ,z v )+β·AA(u,v)+γ·HADD(u,v) Among them, cos(z) u ,z v ) represents the cosine similarity of entity vectors, AA(u,v) represents the neighborhood overlap based on the Adamic-Adar index, HADD(u,v) represents the hierarchical decay coefficient, and α, β, and γ are preset weight coefficients.
10. The method according to claim 1, characterized in that, In step 5), the structural consistency score is calculated using the following formula: SCS(u,v)=w1·cos(z u ,z v )+w2·NS(u,v)+w3·HADD(u,v) in, N(u) represents the set of neighboring entities of entity u, and w1, w2, and w3 are preset weight coefficients.