Global context analysis method and system for complex document
By performing layout analysis, picture text semantic analysis and global knowledge graph construction on complex documents, combined with knowledge graph embedding and graph attention network model, the problem of difficulty in digging deep into document semantic structure and knowledge correlation in existing technologies is solved, and efficient utilization of document information and improvement of value is achieved.
Patent Information
- Application Number
- CN202510486491.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-04-18
AI Technical Summary
The prior art is difficult to deeply explore the inherent semantic structure and knowledge relationships of documents, which limits the effective use of information.
Provide a global context analysis method for complex documents. It generates a catalog through layout analysis, recognizes picture text information and performs semantic analysis, builds a global knowledge graph, and uses knowledge graph embedding and graph attention network models to perform global relationship extraction and context chain construction.
It significantly improves the consistency between the knowledge graph and the global context of the document, reduces the burden on users in document management, and improves the efficiency and value of document information utilization.
Smart Images

Figure CN120012759A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of artificial intelligence technology, and in particular relates to a global context analysis method and system for complex documents. Background Art
[0002] With the rapid development of information technology, electronic documents have become an important tool for people to obtain information and exchange ideas. However, faced with massive amounts of document data, how to efficiently extract, organize and utilize the information in it has become an urgent problem to be solved. Traditional document processing methods are often limited to simple text extraction and keyword retrieval, and are unable to deeply explore the inherent semantic structure and knowledge association of documents, thus limiting the effective use of information.
[0003] Building a knowledge graph for documents has emerged as a cutting-edge document analysis strategy that not only helps generate concise summary content, but also enables efficient content retrieval. However, it is worth noting that the process of extracting entities and relationships from documents to build a knowledge graph often faces the challenge of complex and confusing entity relationships. Summary of the invention
[0004] In view of the above-mentioned deficiencies in the prior art, the present invention provides a global context analysis method and system for complex documents to solve the above-mentioned technical problems.
[0005] In a first aspect, the present invention provides a global context analysis method for a complex document, comprising: Generate a directory for the target document by performing layout analysis on the target document; Identify text information of the image in the target document, and establish an association relationship between the image and the text content by performing semantic analysis on the text information and adjacent text content; Building a basic knowledge graph based on the titles of the catalog, and gradually performing semantic analysis on the paragraphs corresponding to the titles, adding entities and relationships to the basic knowledge graph, and obtaining a global knowledge graph; The knowledge graph embedding method is used to map entities and relations in the global knowledge graph into low-dimensional vectors, and the graph attention network model is used to extract global relations from the embedded global knowledge graph. A global context chain is constructed based on the global relationship, and the relationship of the global knowledge graph is updated based on the global context chain.
[0006] In an optional implementation, generating a directory for the target document by performing layout analysis on the target document includes: Analyzing the layout features of the target document using a layout analysis model, wherein the layout features include paragraph distribution, title format parameters, title content, and paragraph format parameters; Using a rule algorithm, construct a title hierarchy system based on the layout features; Constructing a correspondence between titles and paragraphs based on the title hierarchy system, and generating a directory based on the correspondence; Monitor the updated content of the target document, and synchronously update the directory based on the updated content.
[0007] In an optional implementation, identifying text information of an image in a target document and establishing an association relationship between the image and text content by performing semantic analysis on the text information and adjacent text content includes: Use optical character recognition technology to extract text information from images; Extracting text paragraphs adjacent to the image according to the position of the image in the target document; Using keyword matching technology to select a target paragraph matching the text information from adjacent text paragraphs, and setting an associated tag for the image and the target paragraph; Performing semantic analysis on the text information and the target paragraph, and selecting one or more sentences from the target paragraph that match the semantics of the text information; Adding an independent identifier to the text information and the sentence with semantic matching, wherein the independent identifier is used to indicate that the text information and the sentence with semantic matching are an indivisible integrated content body; When generating the document hierarchy, the specific hierarchical position of the image relative to its associated text is determined so that the image and text content are displayed together as a hierarchical unit.
[0008] In an optional implementation, a basic knowledge graph is constructed based on the titles of the catalog, and semantic analysis is gradually performed on the paragraphs corresponding to the titles, and entities and relationships are added to the basic knowledge graph to obtain a global knowledge graph, including: Extract entities and entity relationships from the titles of the catalog, and build a basic knowledge graph based on the extracted entities and entity relationships; Obtain the lowest level title in the directory, and obtain a local knowledge graph corresponding to the lowest level title, wherein the local knowledge graph contains an entity corresponding to the lowest level title; Obtaining a target paragraph corresponding to the lowest-level title based on the directory, extracting entities and relationships from the target paragraph, and updating the local knowledge graph based on the extracted entities and relationships; Traverse all the lowest-level titles to obtain the global knowledge graph.
[0009] In an optional implementation, a knowledge graph embedding method is used to map entities and relationships in a global knowledge graph into low-dimensional vectors, and a graph attention network model is used to extract global relationships from the embedded global knowledge graph, including: A translation model is used to map entities and relations in the global knowledge graph into low-dimensional vectors, and a loss function of the translation model is set to negative log-likelihood loss; The low-dimensional vectors of entities and relations are used as inputs to the graph attention network model, which extracts the relations between entities in the global knowledge graph by calculating the attention coefficients between entities and aggregating the information of neighboring nodes; The graph attention network model is used to calculate the similarity of entities in different local knowledge graphs, and the association relationship between entities in different local knowledge graphs is constructed based on the calculation results; The relationships between entities in the global knowledge graph and the association relationships between entities in different local knowledge graphs are integrated into global relationships.
[0010] In an optional embodiment, the training method of the graph attention network model includes: Extract semantically similar or related sentences and paragraphs from the document to form positive sample pairs, and select semantically irrelevant sentences or paragraphs from the document to form negative sample pairs; Use the pre-trained hierarchical attention network model to encode sentences or paragraphs in positive and negative sample pairs to generate high-dimensional semantic vectors; A contrastive loss function is used to make the model closer to the positive sample pairs and farther away from the negative sample pairs in the semantic space; The encoded positive and negative sample pairs are used to train the graph attention network model.
[0011] In an optional implementation, a global context chain is constructed based on the global relationship, and the relationship of the global knowledge graph is updated based on the global context chain, including: Using a graph wheel algorithm to construct a global context chain based on the global relationship; The entities and relationships involved in the global context chain are matched with the global knowledge graph to complete the missing entity relationships for the global knowledge graph based on the matching results.
[0012] In a second aspect, the present invention provides a global context analysis system for complex documents, comprising: A directory generation module, used to generate a directory for the target document by performing layout analysis on the target document; The image association module is used to identify the text information of the image in the target document and establish an association relationship between the image and the text content by performing semantic analysis on the text information and adjacent text content; A graph construction module is used to construct a basic knowledge graph based on the titles of the directory, and gradually perform semantic analysis on the paragraphs corresponding to the titles, add entities and relationships to the basic knowledge graph, and obtain a global knowledge graph; The global analysis module is used to map entities and relationships in the global knowledge graph into low-dimensional vectors using the knowledge graph embedding method, and to extract global relationships from the embedded global knowledge graph using the graph attention network model; The context enhancement module is used to construct a global context chain based on the global relationship and update the relationship of the global knowledge graph based on the global context chain.
[0013] The beneficial effect of the present invention is that the global context analysis method and system for complex documents provided by the present invention significantly improves the consistency between the knowledge graph and the global context of the document by integrating technical means such as automatic catalog generation, image and text association, knowledge graph construction and relationship extraction, and global context chain and relationship update. This not only reduces the burden on users in document management, but also improves the utilization efficiency and value of document information.
[0014] In addition, the invention has a reliable design principle, a simple structure and a very broad application prospect. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0016] Figure 1 is a schematic flow chart of a method according to an embodiment of the present invention.
[0017] Figure 2 is a schematic block diagram of a system according to an embodiment of the present invention. DETAILED DESCRIPTION
[0018] In order to enable those skilled in the art to better understand the technical solutions in the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.
[0019] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art of the present invention. The terms used in the specification of the present invention herein are only for the purpose of describing specific embodiments and are not intended to limit the present invention.
[0020] The global context analysis method for complex documents provided by the embodiment of the present invention is executed by a computer device, and accordingly, the global context analysis system for complex documents runs in the computer device.
[0021] Figure 1 is a schematic flow chart of a method according to an embodiment of the present invention. Figure 1 The execution subject may be a global context analysis system for complex documents. According to different requirements, the order of the steps in the flowchart may be changed, and some may be omitted.
[0022] like Figure 1 As shown, the method includes: S1. Generate a directory for the target document by performing layout analysis on the target document.
[0023] Perform a detailed layout analysis step on the target document. This process involves identifying the layout structure of the document, such as titles, paragraphs, headers and footers, to accurately segment the document content. Then, based on the recognition results, a directory containing indexes of each level of titles and their corresponding paragraphs is automatically generated for the target document. This directory not only allows users to quickly browse the document structure, but also provides a convenient index framework for subsequent information retrieval and processing.
[0024] S2. Identify the text information of the image in the target document, and establish an association relationship between the image and the text content by performing semantic analysis on the text information and adjacent text content.
[0025] Using advanced OCR (Optical Character Recognition) technology, we can identify text information in images embedded in target documents. Based on this, we can conduct in-depth semantic analysis by combining text content near or related to the image. Through strategies such as comparison and association, we can establish the relationship between the image and the surrounding text content, which helps us better understand the meaning and role of the image in the document.
[0026] S3. Construct a basic knowledge graph based on the titles of the directory, and gradually perform semantic analysis on the paragraphs corresponding to the titles, add entities and relationships to the basic knowledge graph, and obtain a global knowledge graph.
[0027] Based on the generated catalog, we first construct the initial basic knowledge graph framework with titles at all levels as nodes. Then, we perform semantic analysis one by one according to the index order of the paragraphs in the catalog. This process includes entity recognition, relationship extraction, etc., which aims to transform the key information in the paragraphs into entities and relationships in the graph, thereby gradually enriching and improving the basic knowledge graph and finally forming a global knowledge graph. The global knowledge graph not only covers the core content of the document, but also reveals the intrinsic connections between information.
[0028] S4. Use the knowledge graph embedding method to map the entities and relations in the global knowledge graph into low-dimensional vectors, and use the graph attention network model to extract global relations from the embedded global knowledge graph.
[0029] In order to more effectively utilize the information in the global knowledge graph, the knowledge graph embedding technology is used to map entities and relationships into a low-dimensional vector space. This step reduces the complexity of data processing while retaining the key structure and semantic information in the graph. Subsequently, the graph attention network model (GAN) is used to extract global relationships from the embedded global knowledge graph. The GAN model can focus on the importance of different nodes and relationships in the graph, thereby more accurately capturing the global association patterns.
[0030] S5. Construct a global context chain based on the global relationship, and update the relationship of the global knowledge graph based on the global context chain.
[0031] Based on the extracted global relations, a global context chain is constructed. This chain not only reflects the flow path of information in the document, but also reveals the deep connection between information. On this basis, the relations in the global knowledge graph are reviewed and updated to ensure that each relation in the graph conforms to the logic of the global context chain, further improving the accuracy and practicality of the knowledge graph. Through this series of steps, a global knowledge graph that fully and accurately reflects the content of the target document is finally obtained.
[0032] In an embodiment of the present invention, based on step S1, a possible embodiment is given below to illustrate its specific implementation scheme in a non-limiting manner.
[0033] S101. Layout feature analysis: First, the target document is deeply analyzed using the trained LayoutLMv3 layout analysis model to extract its layout features. These features not only cover the intuitive paragraph distribution, but also include more detailed title format parameters (such as font size, bold, center alignment, etc.), title content (i.e. specific text information), and paragraph format parameters (such as paragraph indentation, line spacing, etc.). This process is the basis for building all subsequent steps, ensuring that subsequent processing can accurately reflect the original structure and content of the document.
[0034] S102. Construction of title hierarchy system: After obtaining detailed layout features, a set of carefully designed rule algorithms are used to construct a title hierarchy system based on these features. This system aims to clearly reflect the hierarchical relationship between titles at all levels in the document, such as main title, subtitle, section title, etc. This step is crucial to understanding the structure, content hierarchy and logical relationship of the document.
[0035] The following is a simplified example of a rule algorithm to illustrate the basic principles and steps of this process: 1. Input data preparation Layout feature data: including paragraph distribution, title format parameters (such as font size, bold, color, center alignment, etc.), title content, paragraph format parameters, etc.
[0036] Document Content: The complete document text, including all headings and body content.
[0037] 2. Title Identification Title recognition based on format parameters: First, all possible titles are identified based on preset format parameter thresholds (such as font size thresholds, bold marks, etc.).
[0038] Content analysis aids identification: For titles whose format features are not obvious, content analysis (such as keyword matching and sentence structure analysis) can be used to assist in identification.
[0039] 3. Determination of hierarchical relationships Title level judgment: Determine the level of each title based on the title’s format parameters (such as the decreasing pattern of font size, boldness, etc.) and content importance (such as the keyword level contained in the title).
[0040] Hierarchical relationship construction: Through logical analysis, determine the hierarchical relationship between titles at all levels, such as main title and subtitle, subtitle and section title, etc.
[0041] 4. Error handling and verification Conflict detection: Detect whether there is a hierarchical relationship conflict (for example, a title is simultaneously identified as a title of two different levels).
[0042] Correction and adjustment: Detected conflicts are corrected through manual intervention or automatic adjustment algorithms.
[0043] Consistency check: Ensure that the heading hierarchy of the entire document is logically consistent.
[0044] 5. Output title hierarchy system Structured output: Output the constructed title hierarchy system in a structured manner, such as a tree structure, list structure, etc.
[0045] Suppose there is a simple document with the following heading hierarchy: Main Title: The main title of the document Subheading 1: The first subheading under the main heading Section title 1: The first section title under subtitle 1 Section title 2: The second section title under subtitle 1 Subtitle 2: The second subtitle under the main title Section title 3: The first section title under subtitle 2 According to the above rule algorithm, such a title hierarchy system can be automatically identified and constructed.
[0046] S103. Establishing the correspondence between titles and paragraphs and generating a directory: Based on the established title hierarchy system, further establish the correspondence between titles and paragraphs. Subsequently, based on these correspondences, a detailed directory is automatically generated. The directory not only lists the titles at all levels, but also provides the corresponding page numbers or paragraph positions, which greatly facilitates document browsing and retrieval.
[0047] Catalog generation methods include: After identifying the hierarchical relationship of the titles, use the catalog generation tool to automatically generate the catalog. The catalog can include the names of the titles at each level, page numbers or paragraph positions, etc.
[0048] S104. Catalog synchronization update mechanism: Considering that the target document may be continuously updated over time, a monitoring and synchronization update mechanism is designed to ensure the accuracy and timeliness of the catalog. This mechanism can monitor the updated content of the document in real time, including but not limited to newly added paragraphs, modified titles or paragraph content, etc. Once an update is detected, the mechanism will automatically trigger the synchronization update process of the catalog to ensure that the catalog and document content are always consistent.
[0049] In an embodiment of the present invention, based on step S2, a possible embodiment is given below to illustrate its specific implementation scheme in a non-limiting manner.
[0050] S201. Extracting text from images using optical character recognition technology Technical details: First, the image in the document is scanned pixel by pixel through high-precision optical character recognition (OCR) technology to identify and extract the text information embedded or attached in the image. This process requires ensuring the accuracy and robustness of OCR technology to cope with images and texts in different fonts, sizes and backgrounds.
[0051] S202. Extracting adjacent text paragraphs of the image according to the position of the image in the target document 1. Document format recognition PDF parsing: If the document is in PDF format, use a PDF parsing library (such as PyMuPDF, PDFMiner, etc.) to read the document's page content, font, size, position and other detailed information.
[0052] Word document parsing: For Word documents, you can use libraries such as python-docx to parse the position and properties of document elements such as paragraphs, tables, and pictures.
[0053] 2. Extraction of adjacent text paragraphs Position-based proximity search: After determining the position of the image, use the document layout information to perform a proximity search above, below, or on the left and right sides of the image to find the text paragraphs adjacent to the image.
[0054] S203. Using keyword matching technology to select a target paragraph that matches the text information from adjacent text paragraphs, and setting an associated tag for the image and the target paragraph Keyword matching: The extracted image text information is matched with the adjacent text paragraphs by keywords to find paragraphs containing relevant or similar information. The matching algorithm can be based on strategies such as word frequency, TF-IDF weight, and semantic similarity.
[0055] Association tag: Once a matching target paragraph is found, a unique association tag is set for the image and the target paragraph to identify their association relationship in subsequent processing.
[0056] S204. Performing semantic analysis on the text information and the target paragraph, and selecting one or more sentences from the target paragraph that match the semantics of the text information Semantic analysis: Using natural language processing (NLP) technology, deep semantic analysis is performed on the extracted text information and target paragraphs to identify the semantic connections between them. This may involve advanced NLP tasks such as syntactic analysis, entity recognition, and semantic role labeling.
[0057] Sentence screening: Based on the results of semantic analysis, one or more sentences that best match the semantics of the image and text information are screened from the target paragraph. These sentences should be able to accurately reflect the relationship between the image and text.
[0058] S205. Adding an independent identifier to the text information and the semantically matching sentence, the independent identifier is used to indicate that the text information and the semantically matching sentence are an indivisible integrated content body. Independent identification: Add unique identifiers or tags to the filtered semantically matching sentences and text information. These identifiers should be unique in the document to identify them as an indivisible comprehensive content body.
[0059] Content integration: After adding independent tags, text information and semantically matching sentences can be further integrated into a comprehensive content unit for easy subsequent processing and display.
[0060] S206. When generating the document hierarchy, determine the specific hierarchical position of the image relative to its associated text so that the image and text content are displayed together as a hierarchical unit Hierarchy generation: When constructing the hierarchical structure of a document, a specific hierarchical position is determined for the image based on the logical relationship between the image and its associated text (including semantically matching sentences).
[0061] Integrated presentation of images and text: Ensure that in the final document, images and their associated text are presented as a whole hierarchical unit to reflect the close relationship between them. This may require adjusting the layout and format of the document to accommodate the needs of integrated images and text.
[0062] The associated image and text information is structured with the corresponding text content in the document to facilitate subsequent data management and information extraction. Specific operations include: Image and text labeling: adding label information to each image, including image number, title, associated text paragraph, etc., for traceability. Synchronization of associated paragraphs: Synchronize the text content recognized by OCR with the corresponding paragraph in the document to ensure that the image information is integrated as part of the document structure. Hierarchical storage: The system stores the structured image and text data in the hierarchical structure of the document, and specifies the hierarchical position of the image content in the document structure to ensure the integrity of the document structure.
[0063] By matching the text information of the image twice, namely, paragraph matching based on keywords and sentence matching based on semantic analysis, the related text content can be accurately screened out for the image, and the data processing volume can be reduced compared to direct semantic analysis. In addition, by setting independent tags for the image and the related text content, the image and text content are bound together, ensuring that the image and text content are not separated in subsequent analysis, thus avoiding serious deviations in subsequent semantic analysis.
[0064] In an embodiment of the present invention, based on step S3, a possible embodiment is given below to illustrate its specific implementation scheme in a non-limiting manner.
[0065] S301. Extract entities and entity relationships from the catalog titles and build a basic knowledge graph based on the extracted entities and entity relationships Entity and relationship extraction: First, conduct a detailed analysis of the titles at all levels in the directory, and use natural language processing (NLP) techniques, such as named entity recognition (NER) and relationship extraction algorithms, to identify key entities (such as names of people, places, organization names, etc.) and potential relationships between these entities (such as hierarchical relationships, inclusion relationships, etc.) from the title text.
[0066] Basic knowledge graph construction: Based on the extracted entities and relationships, a knowledge graph containing basic entities and relationships is constructed using graph construction technology, such as Neo4j and other graph databases. This graph will serve as the basic framework for subsequent local knowledge graph construction and global knowledge graph integration.
[0067] S302. Obtain the lowest level title in the directory, and obtain the local knowledge graph corresponding to the lowest level title, wherein the local knowledge graph contains the entity corresponding to the lowest level title Identify the lowest-level headings: Based on the directory hierarchy, identify the lowest-level headings, which usually correspond to the most specific and detailed content in the document.
[0068] Local knowledge graph generation: For each lowest-level title, a local knowledge graph containing the entities corresponding to the title is generated based on its position and related entities in the basic knowledge graph. This local graph will focus on displaying entities and relationships that are closely related to the content of the lowest-level title.
[0069] S303. Obtain the target paragraph corresponding to the lowest level title based on the directory, extract entities and relationships from the target paragraph, and update the local knowledge graph based on the extracted entities and relationships Get the target paragraph: Use the directory to quickly locate the target paragraph corresponding to each bottom-level heading.
[0070] Paragraph content analysis: Perform detailed text analysis on the target paragraph and use NLP technology to extract entities and relationships in the paragraph.
[0071] Local knowledge graph update: compare and integrate the extracted entities and relationships with the existing local knowledge graph, update the information in the graph, and ensure the accuracy and completeness of the graph.
[0072] S304. Traverse all the lowest-level titles to obtain the global knowledge graph Traversal process: Traverse all the lowest-level titles in order according to the order of the directory, execute the steps in S302 and S303 for each title, and update the local knowledge graph.
[0073] Global knowledge graph integration: During the traversal process, the entities and relationships in each local knowledge graph are gradually integrated into the global knowledge graph. Through graph merging technology, the entities and relationships in the global graph are ensured to remain consistent and form a complete and coherent knowledge system.
[0074] Verification and optimization: Finally, the global knowledge graph is verified and optimized to check whether the entities and relationships in the graph are correct and ensure the quality and usability of the graph.
[0075] Through the above steps, we extract entities and relationships from the document directory, build a basic knowledge graph, generate and update local knowledge graphs based on the lowest-level titles, and finally integrate the global knowledge graph. This process not only improves the efficiency of understanding and analyzing document content, but also provides strong support for subsequent knowledge mining and application.
[0076] In an embodiment of the present invention, based on step S4, a possible embodiment is given below to illustrate its specific implementation scheme in a non-limiting manner.
[0077] S401. Use a translation model to map entities and relationships in the global knowledge graph into low-dimensional vectors, and set the loss function of the translation model to negative log-likelihood loss.
[0078] Step description: First, use translation models (such as TransE, TransR, etc.) to embed entities and relations in the global knowledge graph into low-dimensional vector space. These low-dimensional vectors can capture the semantic information between entities and relations and facilitate subsequent graph attention network model processing.
[0079] Loss function setting: In order to optimize the parameters of the translation model, the loss function is set to Negative Log Likelihood Loss, which aims to minimize the difference between the scores of positive samples and negative samples, thereby ensuring the accuracy and robustness of the embedding vector.
[0080] S402. The low-dimensional vectors of entities and relationships are used as inputs of a graph attention network model, wherein the graph attention network model extracts relationships between entities in a global knowledge graph by calculating attention coefficients between entities and aggregating information of neighboring nodes; The graph attention network model includes: Input layer: The initial embedding vectors of the entities are used as input to the GAT model. These vectors will be used as feature representations of the nodes.
[0081] Attention mechanism layer: Construct a GAT layer where each node aggregates the information of its neighbor nodes through the attention mechanism. The attention coefficient is determined by the similarity or correlation between nodes and can be calculated using dot product, bilinear transformation or neural network.
[0082] Multi-head attention: In order to enhance the stability and expressiveness of the model, a multi-head attention mechanism can be introduced. This means that each GAT layer will contain multiple independent attention heads, each of which will independently calculate the attention coefficient and aggregate neighbor information. Finally, the outputs of these heads will be concatenated or averaged to form the final node representation.
[0083] Multi-layer GAT: In order to capture global context information, multiple GAT layers can be stacked. Each layer updates the representation of the node based on the output of the previous layer. As the number of layers increases, nodes will be able to aggregate information from more distant neighbors, thereby achieving semantic information propagation and integration across entities and relations.
[0084] The training methods of the graph attention network model include: 1. Data preparation stage Document preprocessing: First, the input document is preprocessed, including removing stop words, punctuation, stemming or word form restoration, etc., to reduce noise and standardize text data.
[0085] Positive sample pair construction: From the preprocessed documents, semantic analysis or syntactic analysis techniques are used to extract semantically similar or related sentences and paragraphs, and they are combined into positive sample pairs. The sentences or paragraphs in the positive sample pairs should have high similarity or relevance in terms of content, theme or sentiment.
[0086] Negative sample pair construction: Similarly, semantically irrelevant sentences or paragraphs are randomly selected from the document or selected based on a certain strategy (such as random selection, irrelevant topics, etc.) to form negative sample pairs. The sentences or paragraphs in the negative sample pairs should have significant differences in content.
[0087] 2. Feature extraction stage Application of pre-trained hierarchical attention network model: Use the pre-trained hierarchical attention network model to encode sentences or paragraphs in positive and negative sample pairs. The hierarchical attention network model can capture the attention information within and between sentences, thereby generating a high-dimensional vector representation rich in semantic information.
[0088] Vector representation optimization: During the encoding process, the generated vector representation may need to be further optimized, such as reducing redundant information through dimensionality reduction techniques (such as PCA, t-SNE, etc.), or ensuring comparability between vectors through normalization.
[0089] 3. Model training phase Contrastive loss function design: The NT-Xent contrastive loss function is used, which is designed to measure the distance between positive and negative sample pairs in the semantic space. For positive sample pairs, the loss function should encourage the model to bring them closer; while for negative sample pairs, the loss function should encourage the model to distance them further.
[0090] Model parameter optimization: Using optimization algorithms such as gradient descent and combined with contrast loss function, the parameters of the graph attention network model are iteratively updated. In each iteration, the loss value of the model under the current parameters needs to be calculated, and the parameters are updated according to the gradient of the loss value.
[0091] Early stopping and validation: During the training process, set a validation set to monitor the performance of the model. When the performance on the validation set no longer improves, use an early stopping strategy to prevent overfitting. At the same time, you can also use the validation set to adjust the model's hyperparameters (such as learning rate, batch size, etc.).
[0092] 4. Model evaluation and optimization Performance evaluation: After training, use the test set to evaluate the performance of the model. Evaluation indicators can include accuracy, recall, F1 score, etc. These indicators can reflect the model's ability to handle semantic similarity and relevance tasks.
[0093] Model optimization: Based on the evaluation results, further optimize the model. This may include adjusting the model structure, adding attention mechanisms, introducing external knowledge, etc.
[0094] S403. Calculate the similarity of entities in different local knowledge graphs using the graph attention network model, and construct the association relationship between entities in different local knowledge graphs based on the calculation results; After obtaining the embedding vector of the global knowledge graph, the trained graph attention network model is used to calculate the similarity of entities in different local knowledge graphs. By calculating indicators such as cosine similarity or Euclidean distance between entities, the similarity and correlation between entities can be evaluated.
[0095] Relationship construction: Based on the similarity calculation results, the relationship between entities in different local knowledge graphs is constructed. These relationships can be represented as edges or links between entities for subsequent global relationship integration.
[0096] S404. Integrate the relationships between entities in the global knowledge graph and the association relationships between entities in different local knowledge graphs into global relationships.
[0097] Integrate the relationships between entities in the global knowledge graph and the associations between entities in different local knowledge graphs to form a more complete and comprehensive global knowledge graph. This global knowledge graph will contain direct relationships in the global knowledge graph and indirect relationships in the local knowledge graphs obtained through similarity calculations.
[0098] The integration method can use graph merging algorithms or graph fusion algorithms to merge and integrate relationships from different sources. At the same time, different weights are set for direct relationships and indirect relationships to distinguish them, ensuring the accuracy and reliability of the integrated global knowledge graph.
[0099] In an embodiment of the present invention, based on step S5, a possible embodiment is given below to illustrate its specific implementation in a non-limiting manner.
[0100] S501. Use a graph wheel algorithm to construct a global context chain based on the global relationship.
[0101] After obtaining the integrated global knowledge graph, the Graph Wheel Algorithm is further used to construct the global context chain. The Graph Wheel Algorithm is a graph-structure-based algorithm that constructs a series of interrelated, logically ordered context chains by analyzing the nodes (entities) and edges (relationships) in the graph. These context chains can capture the complex relationships between entities and reveal their dynamic changes in the global context.
[0102] Specific implementation: Node selection and sorting: First, a set of core nodes are selected from the global knowledge graph as the starting point. These core nodes can be key entities in the global knowledge graph or representative entities obtained by similarity calculation in the local knowledge graph. Then, the nodes are sorted according to the association relationship and weight between entities to form a logically coherent context chain.
[0103] Relationship links: Based on the node sorting, these nodes are connected using the relationship links in the global relationship. These relationship links can be direct relationships in the global knowledge graph or indirect relationships obtained through similarity calculation in the local knowledge graph. Through relationship links, a complete context chain is constructed, which can reflect the complex relationships and dynamic changes between entities.
[0104] Context chain optimization: In order to improve the accuracy and reliability of the context chain, the constructed context chain is optimized. This includes removing redundant nodes and relationships, adjusting the order of nodes, merging similar context chains, etc. Through optimization, a more concise and clear global context chain is obtained.
[0105] S502. Match the entities and relationships involved in the global context chain with the global knowledge graph to complete the missing entity relationships for the global knowledge graph based on the matching results.
[0106] Matching algorithm selection: First, select a suitable matching algorithm to match the global context chain with the global knowledge graph. These algorithms can be similarity-based matching algorithms, rule-based matching algorithms, or machine learning-based matching algorithms. By selecting a suitable matching algorithm, the accuracy and efficiency of matching can be improved.
[0107] Matching process: During the matching process, the entities and relationships in the global context chain are compared one by one with the entities and relationships in the global knowledge graph. For matching entities and relationships, just skip them. For unmatched entities and relationships, you need to analyze the reasons and use them as objects for completion. For example, if an entity relationship exists in the global context chain but does not exist in the global knowledge graph, the corresponding entity relationship is supplemented in the global knowledge graph.
[0108] Completing missing relationships: After the matching is completed, the missing entity relationships are completed for the global knowledge graph based on the matching results. This includes adding new entity relationships, updating existing entity relationships, etc. By completing missing relationships, the structure and content of the global knowledge graph can be further improved, and its effectiveness and value in practical applications can be improved.
[0109] Verification and optimization: Finally, the completed global knowledge graph needs to be verified and optimized. This includes checking the accuracy and rationality of the completed relationships, adjusting the relationship weights and importance, etc. Verification and optimization can ensure the accuracy and reliability of the global knowledge graph and provide strong support for its subsequent applications.
[0110] Through the above steps, the graph wheel algorithm is used to build global context chains, and the entities and relationships in these chains are matched with the global knowledge graph to complete the missing entity relationships and achieve the establishment of relationships across paragraph content. This helps to further improve the structure and content of the global knowledge graph and improve its effectiveness and value in practical applications.
[0111] After completing global context tracking and semantic enhancement, the system enters the content summarization and information extraction phase. This step aims to extract and summarize multi-level key information from the document and provide users with a concise and accurate content summary.
[0112] The specific implementation process is as follows: Hierarchical structure recognition: Using the previously established document hierarchy and knowledge graph, the system identifies different levels of content in the document, such as problems, requirements, goals, tasks, solutions, etc.
[0113] Key content positioning: Through semantic representation and entity relationships, the system accurately locates the key content of each level. In the knowledge graph, nodes represent entities (such as problems and requirements) and edges represent relationships (such as "lead to", "satisfy", and "achieve"), thereby clarifying the semantic position of content at each level.
[0114] Summary model application: Use a pre-trained summary generation model (such as a Transformer-based text summary model) to generate concise summaries for each level of content. Context information fusion: When generating summaries, combine global context information to ensure that the summary content is consistent with the overall semantics of the document. The model will consider context and related content when generating summaries to avoid one-sided or distorted information.
[0115] Hierarchical content organization: Organize the extracted content in a hierarchical relationship to form a structure from high to low. For example, from problems to requirements, and then to tasks, expand step by step to build a clear content architecture. Structured data storage: Store the summarized content in a structured form to facilitate subsequent retrieval and analysis. Data formats such as JSON and XML can be used to reflect the hierarchy and association of the content.
[0116] Logical chain construction: Based on the knowledge graph and global context tracking, the system constructs logical chains between different levels of content. For example, what requirements are raised by a certain problem, and what tasks and solutions do these requirements correspond to. Relationship display: When presenting content, highlight the relationship between different levels of content to help users understand the logical relationship of the content. Use visual methods such as flowcharts or tree diagrams to show the mutual relationship between content at each level.
[0117] Redundancy detection: By calculating semantic similarity, redundant information in the content is identified to avoid repeated descriptions. The system will merge semantically similar or repeated content. Information integration: Similar or related content is integrated to form a more comprehensive and unified expression to ensure the integrity and consistency of information.
[0118] Visual presentation: The summarized content is presented visually through charts, tree structures, etc., reflecting the hierarchy and relationship of the content. Users can quickly browse and locate the content of interest through the interactive interface.
[0119] Language fluency: When generating summaries and summaries, natural language generation technology is used to ensure that the output text is clear, fluent, and easy to understand.
[0120] User customization: Allows users to set the length, level of detail and other parameters of the summary to generate content that meets user needs. Users can choose the level or topic of interest to obtain personalized content summaries.
[0121] Content quality assessment: The system conducts quality assessment on the generated summaries and summaries to detect information completeness, accuracy, and language quality to ensure the high quality of the output content.
[0122] This embodiment provides a global context analysis method for a complex document, including the following process: 1. Generate document directory The layout analysis model is used to analyze the layout features of the document, and a title hierarchy system is constructed through a rule-based algorithm to generate a table of contents. The layout analysis model is based on a machine learning algorithm to learn and classify features such as paragraph distribution and title format parameters of the document.
[0123] process: Suppose the document layout feature vector is x=(x1,x2,⋯,x n ), where x i Represents the i-th layout feature, such as paragraph spacing, title font size, etc.
[0124] Use layout analysis models (such as support vector machines) to classify feature vectors and obtain preliminary classification results of titles. The decision function of the support vector machine is: Among them, α i is the Lagrange multiplier, y i is the sample label, K(x i ,x) is a kernel function (such as the radial basis kernel function K(x i ,x)=exp(−γ∥x i −x∥2)), b is the bias term.
[0125] According to the classification results, a rule-based algorithm (such as a hierarchical relationship rule based on titles) is used to construct a title hierarchy system. i The level is l i , determined by the following rules:
[0126] Build the correspondence between titles and paragraphs based on the title hierarchy system and generate a table of contents.
[0127] 2. Establish the relationship between pictures and text content Optical character recognition (OCR) technology is used to extract text information from images, and the association between images and adjacent text content is established through keyword matching and semantic analysis.
[0128] Use OCR technology to extract text information from images image .
[0129] According to the position of the image in the document, extract the adjacent text paragraph P adjacent .
[0130] Use keyword matching technology (such as TF-IDF algorithm) to calculate the matching degree between text information and adjacent text paragraphs. Let t be a keyword, W image The frequency of t in the word is tf t,image , the inverse document frequency of t in the document is idf t , then t is in W imageThe TF-IDF value in is: Similarly, we can get t in P adjacent TF-IDF value in tf−idf t,adjacent Then W image With P adjacent The matching degree S match Defined as:
[0131] Perform semantic analysis on the text information and the target paragraph (e.g., using the cosine similarity of word vectors). Let W image and P adjacent The word vectors are v image and v adjacent , then their semantic similarity S semantic for: Filter out the target paragraphs based on the matching degree and similarity, and set associated tags for the images and target paragraphs.
[0132] 3. Build a global knowledge graph Extract entities and relationships from the directory titles to build a basic knowledge graph, gradually perform semantic analysis on the paragraphs corresponding to the titles, add entities and relationships, and obtain a global knowledge graph.
[0133] Extract entities and entity relationships from the title of the directory. Suppose the extracted entity set is E={e1,e2,⋯,e m}, the relation set is R={r1,r2,⋯,r n The basic knowledge graph is represented as a triple set G0={(e i ,r j ,e k )|e i ,e k ∈E,rj∈R}.
[0134] Get the lowest level title in the directory and the local knowledge graph G corresponding to the lowest level title local .
[0135] Based on the directory, obtain the target paragraph corresponding to the lowest level title, extract entities and relationships from the target paragraph, and set the newly extracted entity set as E' and the relationship set as R'. Update the local knowledge graph: G local =G local ∪{(e i ,r j ,e k )|e i ,e k ∈E∪E′,r j ∈R∪R′} Traverse all the lowest-level titles and integrate all local knowledge graphs into the global knowledge graph Gglobal .
[0136] 4. Global Relation Extraction The knowledge graph embedding method is used to map the entities and relations in the global knowledge graph into low-dimensional vectors, and the graph attention network model is used to extract global relations.
[0137] A translation model (such as TransE) is used to map entities and relations in the global knowledge graph into low-dimensional vectors. Let the vector of entity e be denoted by e, and the vector of relation r be denoted by r. For a triple (e i ,r,e j ), TransE aims to make e i +r≈e j . The loss function is the negative log-likelihood loss:
[0138] Where σ is the sigmoid function, , is the set of negative samples.
[0139] The low-dimensional vectors of entities and relations are used as the input of the graph attention network model. Let the feature vector of node i be h i , the graph attention network calculates the attention coefficient α of node i to node j ij for: Among them, W is the learnable weight matrix, a is the attention vector, and N i is the set of neighbor nodes of node i.
[0140] Aggregate the information of neighboring nodes to obtain the updated feature h of node i i ′:
[0141] The graph attention network model is used to calculate the similarity of entities in different local knowledge graphs. Let entity e i and e j The eigenvectors of ei and h ej , then their similarity S entity for: The association relationships between entities in different local knowledge graphs are constructed based on similarity, and the relationships between entities in the global knowledge graph and the association relationships between entities in different local knowledge graphs are integrated into global relationships.
[0142] 5. Training the Graph Attention Network Model The graph attention network model is trained through positive and negative sample pairs, enabling it to distinguish relevant and irrelevant content in the semantic space.
[0143] Extract semantically similar or related sentences and paragraphs from the document to form positive sample pairs (p1 + ,p2 + ), select semantically irrelevant sentences or paragraphs from the document to form negative sample pairs (p1 − ,p2 − ).
[0144] Use the pre-trained hierarchical attention network model to encode the sentences or paragraphs in the positive sample pairs and negative sample pairs to generate a high-dimensional semantic vector v1 + ,v2 + ,v1 − ,v2 − .
[0145] Using contrast loss function (such as Triplet Loss): Here, m is the margin parameter.
[0146] The encoded positive and negative sample pairs are used to train the graph attention network model, and the model parameters are updated by minimizing the loss function.
[0147] 6. Update global knowledge graph relationships Build a global context chain based on global relationships, and complete the missing entity relationships by matching the global context chain with the global knowledge graph.
[0148] The graph wheel algorithm is used to construct the global context chain C based on the global relationship.
[0149] Match the entities and relations involved in the global context chain with the global knowledge graph. Let the set of triples in the global knowledge graph be G global , the set of triples in the global context chain is C, then the updated global knowledge graph G updated for:
[0150] The Graph Wheel Algorithm is an algorithm for constructing cyclic association paths in graph structures. Its core is to discover multi-hop associations between entities through a wheel-like structure (a closed loop consisting of a central node and multiple spoke nodes). This algorithm has important application value in scenarios such as knowledge graph completion and relational reasoning. The knowledge graph is represented as a directed graph G=(V,E), where: V={v1,v2,...,v n} is a collection of entity nodes; E={(v i ,v j ,r ij )} is the edge set, r ij Represents entity vi to v j The relationship type.
[0151] The graph wheel algorithm constructs a wheel-like structure through the following steps: Central node selection: select a core entity c as the axle; Spoke node screening: Screen entities s1, s2, ..., s directly connected to c k as spokes; Closed loop formation: Spoke nodes are connected through multi-hop paths to form a wheel-shaped cycle path centered on c.
[0152] The process of path weight calculation includes: the weight of each path is dynamically adjusted according to the relationship confidence and path length:
[0153] Among them, m is the path length; αri is the confidence of relationship ri; β∈(0,1) is the path attenuation coefficient.
[0154] In some embodiments, the global context analysis system for complex documents may include multiple functional modules composed of computer program segments. The computer programs of the various program segments in the global context analysis system for complex documents may be stored in a memory of a computer device and executed by at least one processor to perform (see Figure 1 Description) Function of global context analysis of complex documents.
[0155] In this embodiment, the global context analysis system for complex documents can be divided into multiple functional modules according to the functions it performs, such as Figure 2 As shown. The functional modules of the system 200 may include: a catalog generation module 210, a picture association module 220, a graph construction module 230, a global analysis module 240 and a context enhancement module 250. The module referred to in the present invention refers to a series of computer program segments that can be executed by at least one processor and can complete fixed functions, which are stored in a memory. In this embodiment, the functions of each module will be described in detail in subsequent embodiments.
[0156] A directory generation module, used to generate a directory for the target document by performing layout analysis on the target document; The image association module is used to identify the text information of the image in the target document and establish an association relationship between the image and the text content by performing semantic analysis on the text information and adjacent text content; A graph construction module is used to construct a basic knowledge graph based on the titles of the directory, and gradually perform semantic analysis on the paragraphs corresponding to the titles, add entities and relationships to the basic knowledge graph, and obtain a global knowledge graph; The global analysis module is used to map entities and relationships in the global knowledge graph into low-dimensional vectors using the knowledge graph embedding method, and to extract global relationships from the embedded global knowledge graph using the graph attention network model; The context enhancement module is used to construct a global context chain based on the global relationship and update the relationship of the global knowledge graph based on the global context chain.
[0157] Optionally, as an embodiment of the present invention, the directory generation module includes: A layout analysis unit, used for analyzing the layout features of the target document using a layout analysis model, wherein the layout features include paragraph distribution, title format parameters, title content, and paragraph format parameters; A system construction unit, used to construct a title hierarchy system based on the layout features by using a rule algorithm; A directory generating unit, configured to construct a correspondence between titles and paragraphs based on the title hierarchy system, and generate a directory based on the correspondence; The directory updating unit is used to monitor the updated content of the target document and synchronously update the directory based on the updated content.
[0158] Optionally, as an embodiment of the present invention, the picture association module includes: A text extraction unit, used to extract text information from the image using optical character recognition technology; A paragraph extraction unit, used for extracting adjacent text paragraphs of the image according to the position of the image in the target document; A tag association unit, used to select a target paragraph matching the text information from adjacent text paragraphs by using a keyword matching technique, and to set an association tag for the image and the target paragraph; A semantic matching unit, configured to perform semantic analysis on the text information and the target paragraph, and select one or more sentences from the target paragraph that are semantically matched with the text information; A sentence marking unit, used to add an independent identifier to the sentence whose text information and semantics match each other, wherein the independent identifier is used to indicate that the sentence whose text information and semantics match each other is an indivisible integrated content body; The position determination unit is used to determine the specific hierarchical position of the image relative to its associated text when generating the document hierarchy, so that the image and text content can be displayed together as a hierarchical unit.
[0159] Optionally, as an embodiment of the present invention, a basic knowledge graph is constructed based on the title of the directory, and semantic analysis is gradually performed on the paragraphs corresponding to the title, and entities and relationships are added to the basic knowledge graph to obtain a global knowledge graph, including: Extract entities and entity relationships from the titles of the catalog, and build a basic knowledge graph based on the extracted entities and entity relationships; Obtain the lowest level title in the directory, and obtain a local knowledge graph corresponding to the lowest level title, wherein the local knowledge graph contains an entity corresponding to the lowest level title; Obtaining a target paragraph corresponding to the lowest-level title based on the directory, extracting entities and relationships from the target paragraph, and updating the local knowledge graph based on the extracted entities and relationships; Traverse all the lowest-level titles to obtain the global knowledge graph.
[0160] Optionally, as an embodiment of the present invention, a knowledge graph embedding method is used to map entities and relationships in a global knowledge graph into low-dimensional vectors, and a graph attention network model is used to extract global relationships from the embedded global knowledge graph, including: A translation model is used to map entities and relations in the global knowledge graph into low-dimensional vectors, and a loss function of the translation model is set to negative log-likelihood loss; The low-dimensional vectors of entities and relations are used as inputs to the graph attention network model, which extracts the relations between entities in the global knowledge graph by calculating the attention coefficients between entities and aggregating the information of neighboring nodes; The graph attention network model is used to calculate the similarity of entities in different local knowledge graphs, and the association relationship between entities in different local knowledge graphs is constructed based on the calculation results; The relationships between entities in the global knowledge graph and the association relationships between entities in different local knowledge graphs are integrated into global relationships.
[0161] Optionally, as an embodiment of the present invention, the training method of the graph attention network model includes: Extract semantically similar or related sentences and paragraphs from the document to form positive sample pairs, and select semantically irrelevant sentences or paragraphs from the document to form negative sample pairs; Use the pre-trained hierarchical attention network model to encode sentences or paragraphs in positive and negative sample pairs to generate high-dimensional semantic vectors; A contrastive loss function is used to make the model closer to the positive sample pairs and farther away from the negative sample pairs in the semantic space; The encoded positive and negative sample pairs are used to train the graph attention network model.
[0162] Optionally, as an embodiment of the present invention, a global context chain is constructed based on the global relationship, and the relationship of the global knowledge graph is updated based on the global context chain, including: Using a graph wheel algorithm to construct a global context chain based on the global relationship; The entities and relationships involved in the global context chain are matched with the global knowledge graph to complete the missing entity relationships for the global knowledge graph based on the matching results.
[0163] Although the present invention has been described in detail with reference to the accompanying drawings and in combination with the preferred embodiments, the present invention is not limited thereto. Without departing from the spirit and essence of the present invention, a person of ordinary skill in the art may make various equivalent modifications or substitutions to the embodiments of the present invention, and these modifications or substitutions shall be within the scope of the present invention. Any person of ordinary skill in the art may easily think of changes or substitutions within the technical scope disclosed by the present invention, and these shall be within the scope of protection of the present invention.
Claims
1. A global context analysis method for complex documents, characterized in that: include: Generate a directory for the target document by performing layout analysis on the target document; Identify text information of the image in the target document, and establish an association relationship between the image and the text content by performing semantic analysis on the text information and adjacent text content; Building a basic knowledge graph based on the titles of the catalog, and gradually performing semantic analysis on the paragraphs corresponding to the titles, adding entities and relationships to the basic knowledge graph, and obtaining a global knowledge graph; The knowledge graph embedding method is used to map entities and relations in the global knowledge graph into low-dimensional vectors, and the graph attention network model is used to extract global relations from the embedded global knowledge graph. A global context chain is constructed based on the global relationship, and the relationship of the global knowledge graph is updated based on the global context chain.
2. The method according to claim 1, characterized in that: By performing layout analysis on the target document, a directory is generated for the target document, including: Analyzing the layout features of the target document using a layout analysis model, wherein the layout features include paragraph distribution, title format parameters, title content, and paragraph format parameters; Using a rule algorithm, construct a title hierarchy system based on the layout features; Constructing a correspondence between titles and paragraphs based on the title hierarchy system, and generating a directory based on the correspondence; Monitor the updated content of the target document, and synchronously update the directory based on the updated content.
3. The method according to claim 1, characterized in that Identify text information of the image in the target document, and establish an association relationship between the image and the text content by performing semantic analysis on the text information and adjacent text content, including: Use optical character recognition technology to extract text information from images; Extracting text paragraphs adjacent to the image according to the position of the image in the target document; Using keyword matching technology to select a target paragraph matching the text information from adjacent text paragraphs, and setting an associated tag for the image and the target paragraph; Performing semantic analysis on the text information and the target paragraph, and selecting one or more sentences from the target paragraph that match the semantics of the text information; Adding an independent identifier to the text information and the sentence with semantic matching, wherein the independent identifier is used to indicate that the text information and the sentence with semantic matching are an indivisible integrated content body; When generating the document hierarchy, the specific hierarchical position of the image relative to its associated text is determined so that the image and text content are displayed together as a hierarchical unit.
4. The method according to claim 1, characterized in that A basic knowledge graph is constructed based on the titles of the catalog, and semantic analysis is gradually performed on the paragraphs corresponding to the titles, and entities and relationships are added to the basic knowledge graph to obtain a global knowledge graph, including: Extract entities and entity relationships from the titles of the catalog, and build a basic knowledge graph based on the extracted entities and entity relationships; Obtain the lowest level title in the directory, and obtain a local knowledge graph corresponding to the lowest level title, wherein the local knowledge graph contains an entity corresponding to the lowest level title; Obtaining a target paragraph corresponding to the lowest-level title based on the directory, extracting entities and relationships from the target paragraph, and updating the local knowledge graph based on the extracted entities and relationships; Traverse all the lowest-level titles to obtain the global knowledge graph.
5. The method according to claim 4, characterized in that The knowledge graph embedding method is used to map entities and relations in the global knowledge graph into low-dimensional vectors, and the graph attention network model is used to extract global relations from the embedded global knowledge graph, including: A translation model is used to map entities and relations in the global knowledge graph into low-dimensional vectors, and a loss function of the translation model is set to negative log-likelihood loss; The low-dimensional vectors of entities and relations are used as inputs to the graph attention network model, which extracts the relations between entities in the global knowledge graph by calculating the attention coefficients between entities and aggregating the information of neighboring nodes; The graph attention network model is used to calculate the similarity of entities in different local knowledge graphs, and the association relationship between entities in different local knowledge graphs is constructed based on the calculation results; The relationships between entities in the global knowledge graph and the association relationships between entities in different local knowledge graphs are integrated into global relationships.
6. The method according to claim 5, characterized in that The training method of the graph attention network model includes: Extract semantically similar or related sentences and paragraphs from the document to form positive sample pairs, and select semantically irrelevant sentences or paragraphs from the document to form negative sample pairs; Use the pre-trained hierarchical attention network model to encode sentences or paragraphs in positive and negative sample pairs to generate high-dimensional semantic vectors; A contrastive loss function is used to make the model closer to the positive sample pairs and farther away from the negative sample pairs in the semantic space; The encoded positive and negative sample pairs are used to train the graph attention network model.
7. The method according to claim 5, characterized in that Building a global context chain based on the global relationship, and updating the relationship of the global knowledge graph based on the global context chain, including: Using a graph wheel algorithm to construct a global context chain based on the global relationship; The entities and relationships involved in the global context chain are matched with the global knowledge graph to complete the missing entity relationships for the global knowledge graph based on the matching results.
8. A global context analysis system for complex documents, characterized in that: include: A directory generation module, used to generate a directory for the target document by performing layout analysis on the target document; The image association module is used to identify the text information of the image in the target document and establish an association relationship between the image and the text content by performing semantic analysis on the text information and adjacent text content; A graph construction module is used to construct a basic knowledge graph based on the titles of the directory, and gradually perform semantic analysis on the paragraphs corresponding to the titles, add entities and relationships to the basic knowledge graph, and obtain a global knowledge graph; The global analysis module is used to map entities and relationships in the global knowledge graph into low-dimensional vectors using the knowledge graph embedding method, and to extract global relationships from the embedded global knowledge graph using the graph attention network model; The context enhancement module is used to construct a global context chain based on the global relationship and update the relationship of the global knowledge graph based on the global context chain.
9. The system according to claim 8, characterized in that The directory generation module includes: A layout analysis unit, used for analyzing the layout features of the target document using a layout analysis model, wherein the layout features include paragraph distribution, title format parameters, title content, and paragraph format parameters; A system construction unit, used to construct a title hierarchy system based on the layout features by using a rule algorithm; A directory generating unit, configured to construct a correspondence between titles and paragraphs based on the title hierarchy system, and generate a directory based on the correspondence; The directory updating unit is used to monitor the updated content of the target document and synchronously update the directory based on the updated content.
10. The system according to claim 8, characterized in that The picture association module includes: A text extraction unit, used to extract text information from the image using optical character recognition technology; A paragraph extraction unit, used for extracting adjacent text paragraphs of the image according to the position of the image in the target document; A tag association unit, used to select a target paragraph matching the text information from adjacent text paragraphs by using a keyword matching technique, and to set an association tag for the image and the target paragraph; A semantic matching unit, configured to perform semantic analysis on the text information and the target paragraph, and select one or more sentences from the target paragraph that are semantically matched with the text information; A sentence marking unit, used to add an independent identifier to the sentence whose text information and semantics match each other, wherein the independent identifier is used to indicate that the sentence whose text information and semantics match each other is an indivisible integrated content body; The position determination unit is used to determine the specific hierarchical position of the image relative to its associated text when generating the document hierarchy, so that the image and text content can be displayed together as a hierarchical unit.
Citation Information
Patent Citations
Cross-language knowledge graph link prediction method based on graph attention mechanism
CN114564596A
Named entity recognition method based on comparative learning and multi-modal semantic interaction
CN117574904A
Enhanced document generation and retrieval method based on knowledge graph
CN119646178A
Multi-source knowledge graph fusion-oriented entity alignment method and apparatus, and system
WO2023273182A1
Method for calculating chinese-english semantic similarity of paragraph-level texts
WO2024164580A1
Cited By
Retrieval enhancement generation method based on document knowledge base and knowledge graph
CN120780849A
PDF text extraction method and device based on frame coordinates
CN121033854A
A multi-agent collaborative standard document review system
CN122713208A