A document semantic DOM network construction method and system

By constructing a DOM network for PDF documents using a large language model, the problem of existing technologies being unable to understand deep semantic logical relationships is solved, enabling efficient semantic querying and intelligent document processing capabilities, and improving the level of document intelligence.

CN120611104BActive Publication Date: 2025-11-25SHANGHAI YILIAN INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511088141.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-05
Publication Date
2025-11-25
Estimated Expiration
2045-08-05

AI Technical Summary

Technical Problem

Existing technologies cannot effectively understand the deep semantic logic relationships of PDF documents, resulting in insufficient intelligent retrieval and analysis capabilities, and poor adaptability to different templates and fields.

Method used

We employ a large language model to perform semantic analysis on PDF documents, construct a DOM network with hierarchical and cross-hierarchical semantic relationships, and establish a cross-hierarchical semantic network by identifying node types and semantic roles.

Benefits of technology

It achieves deep semantic understanding of PDF documents, supports complex semantic queries, improves the accuracy and efficiency of information retrieval, and can be used in knowledge graphs and intelligent question answering systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120611104B_ABST
    Figure CN120611104B_ABST
Patent Text Reader

Abstract

The application discloses a document semantic DOM network construction method and system, belongs to the technical field of document processing and natural language processing, and the PDF document semantic DOM network construction method executes the following steps through a computer: a large language model is used to perform semantic role and function type analysis on information units extracted from a PDF document, and semantic nodes are generated; a preliminary hierarchical tree structure is constructed in combination with layout features and semantic content, and an optimization algorithm is used to adjust the tree; relationship confidence is calculated based on the judgment of the large language model and a preset weighting formula, so that a semantic link across levels is established, and finally a semantic DOM network is formed; and the output semantic network can reveal deep argumentation logic of a document, supports advanced semantic queries, and provides a high-quality structured data basis for downstream knowledge graph construction and an intelligent question answering system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of document processing and natural language processing, and in particular to a method and system for constructing a document semantic DOM network. Background Technology

[0002] PDF, as a cross-platform electronic document format that maintains consistent content formatting, has been widely used in many fields such as scientific papers, legal contracts, and technical manuals. However, the PDF format was designed with a greater emphasis on visual consistency than on machine-readable structured semantics, which poses a significant challenge to automated and intelligent document processing.

[0003] To address this issue, several structured processing methods have been proposed in existing technologies, one representative technique being based on layout analysis and heuristic rules. These methods typically infer the hierarchical structure of a document by recognizing layout features such as font size, bolding, indentation, and position. For example, it might identify the largest, centered text as a first-level heading and paragraphs with bullet points or numbering as list items. However, these methods have a fundamental flaw:

[0004] First, it completely lacks true semantic understanding capabilities. It can recognize that "this is a paragraph," but it cannot understand whether the paragraph is "presenting an argument," "providing evidence," or "defining a concept." This approach often results in a tree of physical structure rather than a network of semantic logic, hindering deep intelligent retrieval and analysis. Second, due to its reliance on fixed parent-child hierarchical relationships, it cannot identify and express complex semantic connections across different chapters. For example, an experimental result in Chapter 5 of a document might support a core argument presented in Chapter 2; such long-distance, non-hierarchical logical relationships are completely uncaptured by layout- and rule-based methods. Furthermore, the rule base of such methods requires manual maintenance, has poor generalization ability, and is extremely unadaptable to documents with different templates and domains. It fails to effectively utilize the deep reasoning capabilities of large language models to solve the aforementioned semantic problems. Therefore, a semantic processing method and system for PDF documents based on large language models is urgently needed. Summary of the Invention

[0005] The purpose of this invention is to address the shortcomings of existing technologies by proposing a document semantic DOM network construction method and system.

[0006] To achieve the above objectives, the present invention adopts the following technical solution:

[0007] A method for constructing a document semantic DOM network includes the following steps:

[0008] Step S101: Traverse the layout-block-tree of the PDF document to generate an initial semantic information stream containing multiple information units;

[0009] Step S102: Use a large language model to perform semantic analysis on the information units in the semantic information flow to generate semantic nodes containing node types and semantic roles;

[0010] Step S103: Based on the layout features and semantic content between the semantic nodes, construct a semantic DOM tree with hierarchical relationships;

[0011] Step S104: Based on the semantic nodes, analyze and establish cross-level semantic relationships between nodes to form the semantic DOM network.

[0012] Furthermore, the step of performing semantic analysis using a large language model in step S102 specifically includes:

[0013] Based on a preset prompt word template that includes a structured output format definition, a request containing the content of the information unit to be analyzed is generated.

[0014] The request is sent to the large language model;

[0015] The system receives the analysis results returned by the large language model, which contain preset structured data (such as JSON format), and extracts the node type and semantic role from the analysis results.

[0016] Furthermore, the step of constructing a hierarchical semantic DOM tree in step S103 specifically includes:

[0017] Based on the layout features of the PDF document, including font size, indentation, or numbering, the semantic nodes are initially classified into hierarchical levels.

[0018] The large language model is used to confirm or correct the results of the preliminary hierarchical division at the semantic level in order to determine the final parent-child node relationship.

[0019] Furthermore, the method also includes optimizing the semantic DOM tree after its construction, the optimization steps including at least one of the following:

[0020] Based on a preset semantic similarity threshold, adjacent semantic nodes of the same type are merged.

[0021] Alternatively, an isolated semantic node whose parent is the root node can be remounted to a new parent node based on its semantic relevance to other nodes in the document.

[0022] Furthermore, the step of analyzing and establishing cross-level semantic relationships in step S104 specifically includes:

[0023] The large language model is used to determine the semantic relationship type between at least two semantic nodes and to generate a preliminary relationship confidence score.

[0024] Furthermore, step S104 also includes:

[0025] Based on the preliminary relationship confidence level, and combined with the physical distance information and node type combination information between the at least two semantic nodes, the final relationship confidence level is calculated using a preset weighted formula.

[0026] Furthermore, the data structure of the semantic node includes: a unique node identifier, a node type, node content, a semantic role, and a list for storing semantic relationships with other nodes.

[0027] The present invention also provides a document semantic DOM network construction system for implementing the above method.

[0028] A document semantic DOM network construction system, comprising:

[0029] The information flow generation module is configured to traverse the layout-block-tree of the PDF document to generate an initial semantic information flow containing multiple information units;

[0030] The semantic analysis module is configured to perform semantic analysis on information units in the semantic information stream using a large language model to generate semantic nodes containing node types and semantic roles.

[0031] The DOM tree construction module is configured to construct a hierarchical semantic DOM tree based on the layout features and semantic content between the semantic nodes.

[0032] The semantic network forming module is configured to analyze and establish cross-level semantic relationships between the semantic nodes to form the semantic DOM network.

[0033] Furthermore, the semantic analysis module is further configured as follows:

[0034] Based on a preset prompt word template that includes a structured output format definition, a request containing the content of the information unit to be analyzed is generated.

[0035] The request is sent to the large language model;

[0036] It also receives the analysis results returned by the large language model, which contain preset structured data (such as JSON format), in order to extract the node type and semantic role.

[0037] The present invention also provides a computer-readable storage medium for implementing the above-described method and system.

[0038] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0039] By recognizing the physical structure of a document (such as headings and paragraphs) through a large language model, and identifying the logical function (such as "argument" and "evidence") and semantic role (such as "posing a question" and "data support") of each content unit, the understanding of the document is no longer limited to "what it is", but goes deeper into "what it is doing", thus improving the understanding of the document.

[0040] By establishing cross-level semantic relationships (such as "support" and "refute"), the tree structure of the document is upgraded to a network structure, which reveals the hidden and complex argumentation within the document. For example, it can clearly link the experimental data in Chapter 5 with the core arguments in Chapter 2, solving a fundamental problem that existing technologies cannot solve.

[0041] Based on the semantic network constructed above, it can support complex semantic queries far exceeding traditional keyword retrieval. Users are no longer limited to finding "paragraphs containing a certain word," but can make advanced retrieval requests based on logical relationships, such as "find all evidence refuting 'argument A'" or "list all definitions and application examples of 'concept B'," thereby improving the accuracy and efficiency of information acquisition. In addition, by outputting the semantic DOM network, it can not only directly improve the level of intelligence in document processing, but also serve as a high-quality data source, seamlessly connecting to more advanced downstream AI applications such as automatic knowledge graph construction, intelligent question answering systems, and automated report generation, thus having higher practical value and application prospects. Attached Figure Description

[0042] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used together with the embodiments of the invention to explain the invention and do not constitute a limitation thereof.

[0043] Figure 1 This is a flowchart illustrating the document semantic DOM network construction method proposed in this invention;

[0044] Figure 2 This is a structural block diagram of the document semantic DOM network construction system proposed in this invention;

[0045] Figure 3 This is a schematic diagram illustrating the DOM tree construction and optimization process in an embodiment of the present invention;

[0046] Figure 4 This is a schematic diagram of semantic nodes and semantic network structure in an embodiment of the present invention. Detailed Implementation

[0047] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.

[0048] In the description of this invention, it should be understood that the terms "upper", "lower", "front", "rear", "left", "right", "top", "bottom", "inner", "outer", etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this invention.

[0049] This invention provides a method and system for constructing a document semantic DOM network.

[0050] like Figure 1-4 As shown, a method for constructing a document semantic DOM network specifically includes the following steps:

[0051] Step S101: Generate a semantic information flow based on the layout block tree of the PDF document;

[0052] Step S102: Perform semantic analysis on the information flow using a large language model;

[0053] Step S103: Construct and optimize a semantic DOM tree with hierarchical relationships;

[0054] Step S104: Analyze and establish cross-level semantic relationships to form the final semantic network.

[0055] As attached Figure 2 As shown, a document semantic DOM network construction system for implementing the above method includes: an information flow generation module, a semantic analysis module, a DOM tree construction module, and a semantic network formation module. The functions of each module will be explained in detail in conjunction with the following process.

[0056] In the technical solution of this invention:

[0057] Data structures are semantic nodes, representing a unit in a document that has independent semantic meaning, where:

[0058] The specific implementation definition of semantic nodes is as follows:

[0059] self.id: A globally unique identifier for a node, such as a UUID.

[0060] `self.type`: The type of the node, whose value is an enumeration of node types (NodeType). This enumeration type can include: UNKNOWN, SECTION, PARAGRAPH, CONCEPT, ARGUMENT, EVIDENCE, REFERENCE, DEFINITION, TABLE, FORMULA, etc.

[0061] self.content: The original text content of the node.

[0062] self.role: The semantic role of the node, describing its function in the context, such as "asking a question" or "providing background information".

[0063] self.children: A list used to store the IDs of its child nodes, forming a tree-like hierarchical relationship.

[0064] self.relations: A list used to store semantic relationships with other nodes.

[0065] self.attributes: A dictionary used to store other attributes of the node, such as font, position, etc.

[0066] self.context: A dictionary used to store the context information of this node, such as information about its immediate predecessor node.

[0067] Semantic relations are used to define non-hierarchical relationships between nodes, and their specific implementation can be defined as follows:

[0068] self.source_id: The ID of the source node.

[0069] self.target_id: The ID of the target node.

[0070] self.type: The type of the relation, whose value is an enumeration of relation types (RelationType), such as SUPPORTS, CONTRASTS, EXAMPLE_OF, etc.

[0071] self.confidence: Describes the confidence level of the relationship, and is a floating-point number.

[0072] self.attributes: Other attributes used to store relationships.

[0073] The various modules and steps in the embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0074] Information Flow Generation Module: This module preprocesses the input PDF document, transforming it from a physical structure into a linear semantic information flow containing preliminary contextual relationships. The specific implementation steps are as follows:

[0075] First, the system receives the PDF document and parses it into a layout-block-tree using existing PDF parsing libraries (such as pdfplumber, PyMuPDF, etc.). Each node in the tree represents a physical block in the document, such as a text block, image block, or table block, and retains its physical attributes such as position, size, and font.

[0076] The information flow generation module performs a preorder traversal of the layout block tree. During the traversal, for each node, the following operations are performed:

[0077] Extract node information: Extract the text content, font, font size, coordinates, and other block_info information of the current node;

[0078] Analyze context information: Analyze the position and logical relationship between the current node and its parent and sibling nodes to form preliminary context information;

[0079] Generate a message unit: Combine block_info and context into a message unit and add it to a message flow list;

[0080] Finally, after the traversal is complete, the information flow generation module outputs an ordered list containing all information units, i.e., the semantic information flow, and passes it to the semantic analysis module.

[0081] Semantic Analysis Module: The semantic analysis module is used to perform deep semantic understanding of information streams using a large language model (LLM). The semantic analysis module receives a list of information streams and processes each information unit in it to create semantic nodes containing rich semantic information.

[0082] In a preferred embodiment of the present invention, the semantic analysis module is further subdivided into the following internal units:

[0083] Prompt word generation unit: Responsible for dynamically and automatically generating structured requests, i.e. prompt words, based on preset templates and the context of the current information unit.

[0084] Large Language Model Interaction Unit: Responsible for managing API communication with the large language model, including sending generated requests, handling network latency and retries, and receiving returned analysis results.

[0085] Semantic parsing and node creation unit: responsible for parsing the structured data (such as JSON) returned by the large language model, and instantiating SemanticNode and SemanticRelation objects based on the parsed information.

[0086] The semantic analysis module performs the following steps for each information unit:

[0087] a) Construct Prompt: Generate a request based on our pre-defined prompt template used to identify node types and semantic roles. A specific example of this template is as follows: # Role

[0088] You are a top-notch document structure analysis expert.

[0089] # Task

[0090] Analyze the given "current text block" and combine it with its "context information" to output its most likely node type and semantic role in JSON format.

[0091] # JSON output format

[0092] {

[0093] "type": "NodeType",

[0094] "role": "Describes the role of this text block in the text".

[0095] "reasoning": Briefly explain the reasons for your judgment.

[0096] }

[0097] # Contextual Information

[0098] Previous node type: {{previous_node_type}}

[0099] # Current text block

[0100] {{current_block_content}}

[0101] (b) Send the generated request to the large language model (e.g., via API call to GPT-4, Wenxin Yiyan, etc.), and the LLM will return a JSON-formatted string containing the analysis results;

[0102] (c) Parse the JSON string returned by LLM, extract information such as type and role, and combine it with the original text content to create a SemanticNode instance.

[0103] (d) After generating a certain number of semantic nodes, the semantic analysis module will also call the LLM to analyze the relationships between nodes, especially non-hierarchical semantic relationships (such as SUPPORTS, CONTRASTS, etc.). The process is similar to that described above, using specific prompt word templates, the LLM is asked to determine the relationship type between two given nodes and provide an initial confidence level. This relationship information is created as SemanticRelation instances and stored in the relations list of the corresponding nodes.

[0104] DOM Tree Building Module: The DOM tree building module is used to organize an unordered list of semantic nodes into a semantic DOM tree with a strict hierarchical relationship. The specific steps of implementing the DOM tree building module are as follows:

[0105] The DOM tree construction module employs a hybrid strategy of "layout feature preprocessing + LLM semantic confirmation" to determine parent-child relationships. First, it performs preliminary hierarchical division based on node layout features (e.g., nodes with font sizes greater than 16pt and bold tend to be higher-level nodes). Then, the preliminary results are passed to LLM for semantic confirmation or correction (for example, LLM can determine that one text block is a "summary" of another text block based on content, thus establishing their parent-child relationship). The specific process is as follows:

[0106] Constructing a tree structure: Based on the determined parent-child relationships, organize all SemanticNodes into a tree, with the children list of each node storing the IDs of all its child nodes;

[0107] Optimize the tree structure: After the initial tree construction, the optimize_tree algorithm is executed for optimization. The specific optimization strategies include:

[0108] Semantic node merging: Traverse the tree. If two adjacent PARAGRAPH type sibling nodes are found, calculate the semantic similarity of their contents. If the similarity is higher than a preset threshold (e.g., 0.95), merge the two nodes into one.

[0109] Orphaned node remounting: Find the EVIDENCE type node whose parent node is the root node, calculate its semantic similarity with all ARGUMENT type nodes, and remount it as a child node under the most similar ARGUMENT node.

[0110] Semantic Network Formation Module: The semantic network formation module is the final step. It is responsible for forming a network-like knowledge structure based on the already established DOM tree and utilizing the previously stored cross-level relationships.

[0111] This module traverses each SemanticNode in the DOM tree and reads its list of relations. For each SemanticRelation instance in the list, it constructs a directed edge from the source node to the target node in the graph. The type of the edge is defined by the relation's type, and the weight of the edge can be defined by confidence.

[0112] Preferably, a weighted formula is used to calculate the final relationship confidence level. This formula integrates the initial judgment of LLM, the physical distance information between nodes, and prior knowledge of node type combinations. A specific formula example is shown below:

[0113]

[0114] It is the preliminary relation confidence returned by the large language model in step S102, with a value range of [0, 1]. This is the normalized physical distance between the source and target nodes. For example, you can calculate the difference in the number of lines between two nodes in a document and normalize it to the interval [0, 1]. Generally, the closer the physical distance, the greater the likelihood of a stronger semantic relationship. This item will contribute a higher score; This is a prior knowledge reward value based on node type combinations. It is a pre-defined, discrete reward score. For example, when the relation type being judged is "SUPPORTS", if the source node type is "EVIDENCE" and the target node type is "ARGUMENT", this is a classic argument combination, and a reward score can be set accordingly. =1; For other less typical combinations, you can set... =0.

[0115] , , These are preset weighting coefficients, and they satisfy... + + =1, the above weights are used to balance the impact of different factors on the final confidence level;

[0116] It should be further explained that the weighting coefficient , , It was determined through experiments and optimization on a labeled validation dataset, with the aim of finding an optimal balance point that yields the best calculated confidence level. It best matches the true relationship of human annotation. For example, in a typical scientific paper dataset, the judgment of LLM can be found ( ) usually dominates, but physical distance ( ) and type combination ( As an effective auxiliary feature, it can significantly correct the judgment of LLM in certain fuzzy scenarios, such as weight allocation. =0.7, =0.2, =0.1, to perform data-driven parameter selection.

[0117] At this point, the PDF document semantic DOM network is complete, and can be used for subsequent downstream tasks such as deep semantic retrieval, knowledge question answering, or knowledge graph construction.

[0118] Example 1

[0119] Process a typical PDF of a research paper in the field of computer science, which usually includes chapters such as title, authors, abstract, introduction, related work, methodology, experiments, conclusion, and references.

[0120] Input: A PDF research paper in the field of computer science.

[0121] The processing steps are as follows:

[0122] Step S101: The system uses the information flow generation module to parse the paper into an ordered information flow containing various text blocks and figure blocks.

[0123] Step S102: The semantic analysis module processes the information flow. For example, using a large language model, it identifies a paragraph in the "Introduction" section as NodeType.ARGUMENT (argument), with the role of "presenting the core problem to be solved in this paper." At the same time, it identifies a chart in the "Experiments" section and the text description below it as NodeType.EVIDENCE (evidence), with the role of "demonstrating the performance advantages of the method in this paper."

[0124] Step S103: The DOM tree construction module constructs a hierarchical DOM tree with the paper title as the root node and each chapter as a child node, based on the paper's chapter number (e.g., 1, 1.1, 2), font size, and other layout features, combined with the semantic confirmation of LLM.

[0125] Step S104: The core task of the semantic network formation module is to establish cross-level relationships. The module analyzes the relationship between the EVIDENCE node in the "Experiments" section and the ARGUMENT node in the "Introduction" section, calculating the final confidence level by calling the LLM and combining it with our proposed weighted formula. It can establish a SUPPORTS relationship from the EVIDENCE node to the ARGUMENT node with high confidence.

[0126] The resulting semantic DOM network can reproduce the chapter structure of the paper and reveal the core argument chain of the paper through semantic links such as SUPPORTS: that is, which experimental result supports which initial argument, so that researchers can quickly understand the core contribution and argument logic of the paper.

[0127] Example 2:

[0128] In this embodiment, a typical API (Application Programming Interface) technical document PDF is processed, which usually includes module introduction, function definition, parameter list, code example and return value description.

[0129] Input: A PDF document containing API technical information.

[0130] The processing steps are as follows:

[0131] Step S101: The system parses the API document into an information stream;

[0132] Step S102: The semantic analysis module uses LLM to identify semantic units unique to the technical document, such as:

[0133] A function signature of NodeType.FUNCTION_DEFINITION is identified, the table below it is NodeType.PARAMETER_DESCRIPTION, and a code block is NodeType.CODE_EXAMPLE;

[0134] Step S103: The DOM tree building module organizes each function definition and its related parameters, code examples, etc., into an independent subtree, forming a clear modular hierarchical structure.

[0135] Step S104: The semantic network forming module focuses on establishing references and dependencies between different semantic units. For example, the system will establish an EXAMPLE_OF (is an example of...) relationship from the CODE_EXAMPLE node to the FUNCTION_DEFINITION node it demonstrates.

[0136] The resulting semantic DOM network acts like an interactive API map, allowing developers to browse the API structure hierarchically and perform efficient queries through semantic links, such as "find all code examples that use the 'create_user' function," thereby improving the usability of technical documentation and development efficiency.

[0137] To better understand the technical solution of this application, the following comparative experiments will further illustrate the point.

[0138] Experimental environment: A standard deep learning server equipped with an NVIDIA A100 GPU was used on the Ubuntu operating system, and the experiment was conducted using Python and the PyTorch framework.

[0139] Dataset: A dataset of 500 PDF documents from different fields (including scientific papers, legal contracts, and technical manuals) was constructed. All documents were meticulously annotated by humans, including the semantic role of each text block (such as argument, evidence) and the logical relationships between blocks (such as support, rebuttal), serving as the "Gold Standard" for evaluation (Ground Truth).

[0140] Comparative example:

[0141] Baseline method: This is a representative method based on layout analysis and heuristic rules, which does not use a large language model for semantic recognition.

[0142] This invention employs all the technical solutions described in this invention.

[0143] Evaluation indicators:

[0144] Role Accuracy: The percentage of correctly identified semantic roles (arguments, evidence, etc.) out of the total.

[0145] Relation F1 Score: A comprehensive evaluation of the accuracy of relation detection. ) and recall rate ( The performance index (SPI) is the core standard for measuring the performance of relation detection. Its calculation method is as follows:

[0146]

[0147]

[0148]

[0149] In the relationship detection task:

[0150] True Positives are genuine examples, which are semantic relationships correctly identified by the method of this invention and that actually exist in the manually annotated "gold standard".

[0151] False Positives are semantic relationships that are identified by the method of this invention but do not actually exist in the "gold standard".

[0152] False Negatives are false negatives, which are semantic relationships that exist in the "gold standard" but that the method of this invention has failed to identify.

[0153] The specific results are shown in Table 1.

[0154]

[0155] Table 1

[0156] As can be seen from the above structure, the method of the present invention has a significant advantage over the baseline method in key semantic understanding metrics. The present invention innovatively introduces a large language model for deep semantic analysis and network construction, which is not a simple technical improvement, but a significant improvement in technical effectiveness, and can solve the problem that existing technologies cannot effectively understand the deep logical relationships of documents.

[0157] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A method for constructing a document semantic DOM network, characterized in that, Includes the following steps: Step S101: Traverse the layout block tree of the PDF document to generate an initial semantic information stream containing multiple information units; Step S102: Use a large language model to perform semantic analysis on the information units in the semantic information flow to generate semantic nodes that represent the document's argumentation structure. The semantic nodes are specifically classified into at least one of argument-type nodes, evidence-type nodes, definition-type nodes, or concept-type nodes, and include semantic roles. The semantic roles further specify the function of the information units in the argumentation logic as at least one of proposing an argument, providing supporting evidence, or explaining a concept definition. Step S103: Based on the layout features and semantic content between semantic nodes, construct a hierarchical semantic DOM tree. This step specifically includes: Based on the layout features of the PDF document, a preliminary hierarchical division of semantic nodes is performed; The large language model is used to confirm or correct the results of the initial hierarchical division at the semantic level in order to determine the final parent-child node relationship. Step S104: Based on semantic node analysis, establish cross-level semantic relationships between nodes to form a semantic DOM network. The cross-level semantic relationships are used to characterize the argumentation logic between nodes. Specifically, the argumentation logic includes at least one of support, refutation, or example relationships. This step specifically includes: Use a large language model to determine the semantic relationship type between at least two semantic nodes and generate preliminary relationship confidence. Based on the preliminary relationship confidence, and combined with the physical distance information between at least two semantic nodes and the node type combination information, the final relationship confidence is calculated using a preset weighted formula.

2. The document semantic DOM network construction method according to claim 1, characterized in that, The step of semantic analysis using a large language model in step S102 specifically includes: Based on a preset prompt word template that includes a structured output format definition, a request containing the content of the information unit to be analyzed is generated. The request is sent to the large language model; The system receives the analysis results returned by the large language model, which contain preset structured data, and extracts the node type and semantic role from the analysis results.

3. The document semantic DOM network construction method according to claim 2, characterized in that, The method further includes optimizing the semantic DOM tree after its construction, and the optimization steps include at least one of the following: Based on a preset semantic similarity threshold, adjacent semantic nodes of the same type are merged. Alternatively, an isolated semantic node whose parent is the root node can be remounted to a new parent node based on its semantic relevance to other nodes in the document.

4. The document semantic DOM network construction method according to claim 1, characterized in that, The data structure of a semantic node includes: a unique node identifier, a node type, node content, a semantic role, and a list for storing semantic relationships with other nodes.

5. A document semantic DOM network construction system, used to implement the document semantic DOM network construction method as described in any one of claims 1 to 4, characterized in that, include: The information flow generation module is configured to traverse the layout block tree of the PDF document to generate an initial semantic information flow containing multiple information units. Specifically, this module is configured as follows: Perform a preorder traversal of the layout block tree, and extract the text content and layout features of the current node in the layout block tree during the traversal. It also analyzes the position and logical relationship between the current node and its parent and sibling nodes to form information representing the initial context relationship; The text content, layout features, and information representing the initial contextual relationships are combined into an information unit, and all information units are aggregated to form an initial semantic information flow. The semantic analysis module is configured to perform semantic analysis on information units in the semantic information flow using a large language model to generate semantic nodes that represent the document's argumentation structure. These semantic nodes are specifically classified as at least one of argument-type nodes, evidence-type nodes, definition-type nodes, or concept-type nodes, and include semantic roles. The semantic roles further specify the function of the information unit in the argumentation logic as at least one of presenting an argument, providing supporting evidence, or elucidating a concept definition. The DOM tree construction module is configured to perform preliminary hierarchical division of semantic nodes based on the layout features of the PDF document, and to use a large language model to confirm or correct the results of the preliminary hierarchical division at the semantic level, so as to construct a semantic DOM tree with hierarchical relationships. The semantic network forming module is configured to analyze and establish cross-level semantic relationships between semantic nodes to represent the argumentation logic, thereby forming a semantic DOM network. The cross-level semantic relationships are used to represent the argumentation logic between nodes. The argumentation logic specifically includes at least one of support, refutation, or example relationships. The module is specifically configured to: use a large language model to determine the semantic relationship type between at least two semantic nodes and generate a preliminary relationship confidence score; and calculate the final relationship confidence score using a preset weighted formula based on the preliminary relationship confidence score, the physical distance information between at least two semantic nodes, and the node type combination information.

6. The document semantic DOM network construction system as described in claim 5, characterized in that, The semantic analysis module is further configured as follows: Based on a preset prompt word template that includes a structured output format definition, a request containing the content of the information unit to be analyzed is generated. The request is sent to the large language model; It also receives the analysis results returned by the large language model, which contain preset structured data, in order to extract the node type and semantic role.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the document semantic DOM network construction method as described in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Knowledge base construction method based on large language model and intelligent question and answer method

    CN120278257A

  • Text information structured recovery method and system based on large language model and application

    CN120409443A