Method and system for cross-document information aggregation and automatic summarization of new knowledge

By constructing a semantic structure graph using a semantic embedding model and a graph convolutional network, the problem of information fragmentation and semantic association breakage in cross-file information processing is solved, enabling efficient automatic knowledge summarization and generation, which is suitable for enterprise knowledge base construction and domain knowledge discovery.

CN121052345BActive Publication Date: 2026-03-20GUIZHOU BLUE DREAM FACTORY TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511588713.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-03
Publication Date
2026-03-20
Estimated Expiration
2045-11-03

AI Technical Summary

Technical Problem

In existing technologies, cross-file information processing suffers from problems such as information fragmentation, limited semantic understanding, lack of knowledge generation, and weak decision support. It cannot effectively penetrate the barriers between different file formats, resulting in the break of semantic associations between cross-system documents and a lack of dynamic knowledge inference capabilities.

Method used

A semantic embedding model is used to vectorize and semantically cluster document file data, construct a semantic structure graph, and perform semantic clustering through a hierarchical adjustable graph convolutional network. Incomplete nodes are identified and filled in, generating structured knowledge units. A context-aware mechanism is used to complete the content, and finally, new knowledge text is generated.

Benefits of technology

It significantly improves the automation of knowledge discovery, overcomes the limitations of knowledge fragmentation between traditional documents and the low efficiency of manual summarization, and enhances knowledge expression and decision support capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121052345B_ABST
    Figure CN121052345B_ABST
Patent Text Reader

Abstract

The application discloses a cross-file information summarization and new knowledge automatic summarization method and system, and belongs to the technical field of file data processing. The method specifically comprises the following steps: collecting document file data, performing vectorization on the document file data based on a semantic embedding model, reconstructing and performing semantic clustering on the semantic vectors of the document file data, constructing a semantic structure graph, wherein the semantic structure graph represents the logical relationship between the document file data, performing content completion on the nodes with incomplete association in the semantic structure graph, extracting structured knowledge units from the completed semantic structure graph, and generating a new knowledge text based on the structured knowledge units. The application solves the problems of knowledge fragmentation between traditional document files and low efficiency of manual summarization, significantly improves the automation degree of knowledge discovery, and is suitable for scenarios such as knowledge base construction, field rule extraction, field knowledge discovery and the like.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of file data processing, and particularly relates to a cross-file information summarization and new knowledge automatic summarization method and system. BACKGROUND

[0002] In the prior art, the management of knowledge has systematic defects such as information fragmentation, semantic understanding limitation, knowledge generation deficiency and weak decision support. In the information processing level, the traditional keyword retrieval tool can only realize surface text matching and cannot penetrate the barrier formed by different file formats (such as PDF / DOC / mail / database log) to form a semantic association between cross-system documents, resulting in a broken semantic association between cross-system documents.

[0003] In the knowledge generation link, although the business intelligence system can visualize structured data, it cannot discover the association relationship in unstructured documents. Finally, in the decision support layer, the existing scheme generally lacks dynamic knowledge deduction capability, cannot generate a risk prediction model based on historical documents, and is difficult to automatically extract the technology evolution trend from the competitor technology documents, so that the enterprise decision still depends on experience and intuition. Therefore, a cross-file information summarization and new knowledge automatic summarization method is urgently needed. SUMMARY

[0004] In view of the deficiencies of the prior art, the application provides a cross-file information summarization and new knowledge automatic summarization method and system.

[0005] To achieve the above purpose, the application provides the following technical scheme:

[0006] The cross-file information summarization and new knowledge automatic summarization method comprises the following steps:

[0007] Collecting document file data;

[0008] Vectorizing the document file data based on a semantic embedding model, reconstructing and semantically clustering the semantic vectors of the document file data, and constructing a semantic structure graph, wherein the semantic structure graph represents the logical relationship between the document file data;

[0009] Completing the content of the nodes with incomplete associations in the semantic structure graph, and extracting structured knowledge units from the completed semantic structure graph;

[0010] Generating a new knowledge text based on the structured knowledge units.

[0011] Specifically, the vectorization of the document file data based on the semantic embedding model, the reconstruction and semantic clustering of the semantic vectors of the document file data, and the construction of the semantic structure graph comprise the following steps:

[0012] The collected document file data is preprocessed, including content segment identification and splitting, content format standardization and denoising;

[0013] The preprocessed document file data is input into a pre-trained feature extraction model, and the extracted features are mapped into a unified semantic space to obtain a document file data semantic vector;

[0014] The document file data semantic vector is weighted and reconstructed, and an initial semantic structure graph is constructed based on the enhanced document file data semantic vector;

[0015] The initial semantic structure graph is clustered, and a semantic structure graph is constructed according to the clustering result.

[0016] Specifically, the document file data semantic vector is weighted and reconstructed, and an initial semantic structure graph is constructed based on the enhanced document file data semantic vector, including:

[0017] The associated context metadata is extracted from the preprocessed document file data, including folder structure data, document reference relationship and time sequence data;

[0018] The associated context metadata is context encoded to obtain a unified context vector;

[0019] A context modulation network is constructed, and a self-attention mechanism is used to calculate the cross-correlation between the document file data semantic vector and the unified context vector to generate a context-sensitive weight matrix;

[0020] The document file data semantic vector is weighted and reconstructed based on the context-sensitive weight matrix to obtain an enhanced document file data semantic vector;

[0021] Based on the enhanced document file data semantic vector, an initial semantic structure graph is constructed, the nodes of the initial semantic structure graph are the semantic vectors of individual document file data in the enhanced document file data semantic vector, and the edges are the logical relationships between the semantic vectors of individual document file data.

[0022] Specifically, the initial semantic structure graph is clustered, and a semantic structure graph is constructed according to the clustering result, including:

[0023] A pre-set edge filtering strategy is used to prune the initial semantic structure graph;

[0024] In the pruned initial semantic structure graph, the input feature vector of each node is initialized;

[0025] A hierarchical adjustable graph convolution network is used to recursively propagate node features, and the recursive propagation is to weight and combine the current node and its neighbor node feature vectors at each layer;

[0026] clustering the recursive propagation results of the hierarchical adjustable graph convolutional network to divide semantic clusters;

[0027] building connection edges between the cluster nodes of the semantic clusters to obtain a semantic structure graph.

[0028] Specifically, the content completion is performed on the nodes with incomplete association in the semantic structure graph, and a structured knowledge unit is extracted from the completed semantic structure graph, including:

[0029] identifying nodes or semantic clusters with context information discontinuity, logical path disconnection, or incomplete semantic expression in the semantic structure graph;

[0030] performing context generation on the identified defective areas, the context generation including reasoning and completing missing transitional sentences or explanation fragments based on the contents of the neighbor nodes in the semantic structure graph;

[0031] performing semantic boundary segmentation on the completed semantic structure graph to identify semantic clusters as knowledge units;

[0032] representing the divided knowledge units in a structured form.

[0033] Specifically, the semantic boundary segmentation is performed on the completed semantic structure graph to identify semantic clusters as knowledge units, including:

[0034] evaluating the semantic consistency within each semantic cluster in the completed semantic structure graph;

[0035] detecting semantic turning points within the semantic cluster, and positioning preliminary boundary nodes based on syntactic changes and path mutation behaviors in the completed semantic structure graph;

[0036] analyzing each semantic node in the semantic cluster to determine whether it constitutes a complete semantic structure;

[0037] based on the complete semantic structure determination result, optimizing the preliminary boundary nodes to obtain semantic boundaries;

[0038] based on the semantic boundaries, segmenting semantic clusters as knowledge units.

[0039] Specifically, based on the complete semantic structure determination result, the preliminary boundary nodes are optimized to obtain semantic boundaries, including:

[0040] performing adjacent merging on the semantic nodes marked as incomplete semantic structures to form semantic complete units;

[0041] performing aggregation on the semantic nodes with repeated or parallel expressions to remove redundant semantic nodes;

[0042] Based on the edge weight value in the semantic cluster and the semantic node connection, the preliminary boundary node is adjusted, the semantic node with complete semantic structure is selected as the boundary node, and the semantic boundary is obtained.

[0043] Specifically, the structured knowledge unit is used to generate a new knowledge text, including:

[0044] A new knowledge generation target model is established.

[0045] The structured knowledge unit is path planned, and the path is selected as the main content of the new knowledge.

[0046] The knowledge unit corresponding to the selected path is input into the new knowledge generation target model, and the new knowledge content is generated according to the set generation parameters.

[0047] The cross-file information summary and new knowledge automatic summary system is used to realize the cross-file information summary and new knowledge automatic summary method, including a data acquisition module, a semantic graph construction module, a content completion module and a new knowledge generation module.

[0048] The data acquisition module is used to acquire document file data.

[0049] The semantic graph construction module is used to vectorize the document file data based on a semantic embedding model, reconstruct and semantically cluster the document file data semantic vector, and construct a semantic structure graph.

[0050] The content completion module is used to complete the content of the associated incomplete nodes in the semantic structure graph, and extract structured knowledge units from the completed semantic structure graph.

[0051] The new knowledge generation module is used to generate a new knowledge text based on the structured knowledge unit.

[0052] Specifically, the semantic graph construction module includes a preprocessing unit, a semantic vector unit and a graph construction unit.

[0053] The preprocessing unit is used to preprocess the acquired document file data, including content segment identification and splitting, content format standardization and denoising.

[0054] The semantic vector unit is used to input the preprocessed document file data into a pre-trained feature extraction model, and map the extracted features to a unified semantic space to obtain a document file data semantic vector.

[0055] The graph construction unit is used to weight and reconstruct the document file data semantic vector, construct an initial semantic structure graph, and cluster to construct a semantic structure graph.

[0056] Compared with the prior art, the present application has the beneficial effects of:

[0057] The present application proposes a cross-file information aggregation and new knowledge automatic summarization method and system, introduces a context-aware semantic structure graph construction mechanism, uses the folder level relationship, document reference relationship and time sequence characteristics between documents to construct a semantic structure graph, combines a hierarchical adjustable graph convolution network to realize structured semantic cluster division, accurately captures the knowledge unit boundary and semantic logic structure, and then through semantic repair and structure optimization, fills in the knowledge gaps and repairs the logical breaks, improves the integrity and expression ability of the graph structure, finally combines path planning and new knowledge generation target model to intelligently output new knowledge text, which solves the limitations of traditional document knowledge fragmentation and low efficiency of manual summarization, significantly improves the automation degree of knowledge discovery, and is suitable for enterprise knowledge base construction, domain rule extraction, domain knowledge discovery and other scenes. BRIEF DESCRIPTION OF DRAWINGS

[0058] Figure 1 A cross-file information aggregation and new knowledge automatic summarization method is provided for the present application.

[0059] Figure 2 An initial semantic structure graph and a pruning schematic diagram are provided for the present application.

[0060] Figure 3 A semantic structure graph is provided for the present application.

[0061] Figure 4 A cross-file information aggregation and new knowledge automatic summarization system architecture diagram is provided for the present application. DETAILED DESCRIPTION

[0062] The present application will be described in detail below with specific embodiments. The following embodiments will help those skilled in the art to further understand the present application, but do not limit the present application in any form. It should be noted that for those skilled in the art, without departing from the concept of the present application, a number of modifications and improvements can be made. These all belong to the protection scope of the present application.

[0063] In order to make the purpose, technical scheme and advantages of the present application more clear and obvious, the present application will be further described in detail below in combination with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and do not limit the present application.

[0064] It should be noted that the various features of the embodiments of the present application can be combined with each other, and all within the protection scope of the present application, if there is no conflict. In addition, although the functional modules are divided in the device schematic diagram, and the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order from the module division in the device or the order in the flowchart. In addition, the "first", "second", "third" and the like used in the present application do not limit the data and execution order, but only distinguish the same items or similar items with basically the same function and effect.

[0065] Unless otherwise defined, all technical and scientific terms used in the present application have the same meaning as commonly understood by one of ordinary skill in the art to which the present application belongs. The terms used in the present application are only for the purpose of describing the specific embodiments of the present application, and are not used to limit the present application. The term "and / or" used in the present application includes any and all combinations of one or more related listed items.

[0066] Embodiment 1

[0067] Please refer to Figures 1-3 The present application provides an embodiment: a cross-file information summary and new knowledge automatic summary method, comprising the following specific steps:

[0068] Step S1: collecting document file data.

[0069] In the present embodiment, the document file data is the initial source of fragmented information, which is usually in the form of unstructured text, such as reports, papers, web pages, conversation records, etc., which lack standard semantic tags and structure division, and are difficult to be directly used for semantic analysis or knowledge extraction.

[0070] During the collection process, different file format encoding methods, content structures such as headers, directories, footnotes, etc., and noise interference such as advertisements, template content, etc. are considered, and a format recognition module is used to automatically parse the document format type, and the document content is converted into a unified internal text representation, i.e. the document file data.

[0071] Through this step, pure text content with semantic value can be extracted from multiple sources, heterogeneous, and different styles of document files, so as to realize the standardized input of fragmented information.

[0072] Step S2: vectorizing the document file data based on a semantic embedding model, reconstructing and semantically clustering the document file data semantic vector, and constructing a semantic structure graph, wherein the semantic structure graph represents the logical relationship between the document file data.

[0073] The specific steps of step S2 are:

[0074] Step S201: preprocessing the collected document file data, including content segment identification and splitting, content format standardization and denoising.

[0075] In this embodiment, content segment identification and splitting refers to dividing a complete text into smaller, semantically independent or semi-independent content segments according to certain language structures, such as paragraphs, sentences, and phrases; content format standardization refers to uniformly processing non-uniform expressions in the original text due to document format differences, such as line breaks, indents, spaces, special symbols, list symbols, etc.; denoising refers to removing invalid or interfering content in the text, such as headers and footers, chart descriptions, directory information, footnotes, links, watermarks, etc. It should be noted that the preprocessing is performed using existing technologies, which will not be described in detail in this application.

[0076] Step S202: inputting the preprocessed document file data into the pre-trained feature extraction model, and mapping the extracted features to a unified semantic space to obtain a document file data semantic vector.

[0077] In this embodiment, each document file data preprocessed by step S201 is input into a pre-trained language model, such as a BERT or Transformer model, as an independent unit. The context of each input document file data is modeled to obtain its word-level or sentence-level semantic representation, and finally a fixed-dimension dense semantic vector is generated for each document file data.

[0078] Step S203: weighting and reconstructing the document file data semantic vector, and constructing an initial semantic structure graph based on the enhanced document file data semantic vector.

[0079] The specific steps of step S203 are as follows:

[0080] Step S2031: extracting associated context metadata from the preprocessed document file data, including folder structure data, document reference relationship, and time sequence data.

[0081] In this embodiment, the original carrier information of the preprocessed document file is parsed from the file system and version management system, and its storage path, access permission, editing history, directory level, reference relationship, event sequence, editing record, associated document subject label, and source information are extracted; the folder structure data is extracted according to the file path, and its hierarchical directory structure is parsed into an ordered node string as the document file path; the document reference relationship includes reference chain and cross-reference; the time sequence data is standardized into a unified format and a time sequence is constructed.

[0082] Step S2032: context encoding processing of the associated context metadata to obtain a unified context vector.

[0083] In the embodiment, the metadata type is identified and divided into: category type, such as directory path and user role; sequence type, such as time sequence; and graph structure type. Different types of metadata are processed by selecting different encoders, the concatenated or attention weighted combined sub-vectors after encoding are spliced or attention weighted combined, and a full connection layer or projection layer is used to map to the same dimension as the text semantic vector to obtain a unified context vector.

[0084] Step S2033: A context modulation network is constructed, a self-attention mechanism is used to calculate the cross-correlation between the document file data semantic vector and the unified context vector, and a context sensitive weight matrix is generated.

[0085] In the embodiment, a cross-modal attention network is constructed, the text semantic vector is used as Query, the unified context vector is used as Key and Value, a self-attention mechanism is used to calculate the attention distribution of the Query vector to the Key vector, and a context sensitive weight matrix is constructed based on the attention scoring result, which represents the response degree of each semantic segment in the current context, wherein Query represents query, Key represents key, and Value represents value.

[0086] Step S2034: The document file data semantic vector is weighted and reconstructed based on the context sensitive weight matrix to obtain an enhanced document file data semantic vector.

[0087] In the embodiment, the context sensitive weight matrix is used as an attention mask to perform weighted summation on the context semantic vectors of different dimensions, the context vector after the weighted summation is spliced or residual fused with the original semantic vector, a full connection layer is used to perform nonlinear mapping and dimension compression on the fusion result to obtain the final enhanced semantic vector.

[0088] Step S2035: Based on the enhanced document file data semantic vector, an initial semantic structure graph is constructed, the nodes of the initial semantic structure graph are the semantic vectors of individual document file data in the enhanced document file data semantic vector, and the edges are the logical relationships between the semantic vectors of individual document file data.

[0089] The advantage of the step is that by introducing multi-source context information and a self-attention mechanism, the context enhancement of the document file semantic representation is realized, the initial semantic structure graph is constructed based on the enhanced semantic vector, the problem of lack of context of fragmented information is effectively solved, the perception ability of the semantic representation to the real scene context is improved, and the logical association between different document fragments is enhanced.

[0090] Step S204: The initial semantic structure graph is clustered, and a semantic structure graph is constructed according to the clustering result.

[0091] The specific steps of step S204 are as follows:

[0092] like Figure 2 As shown, step S2041: Preset edge filtering strategy to prune the initial semantic structure graph.

[0093] In this embodiment, the edge filtering strategy includes: setting an edge weight threshold, setting a maximum edge count limit, and setting an edge information entropy standard and a semantic uncertainty standard.

[0094] exist Figure 2 In the initial semantic structure graph, there are multiple nodes ( Figure 2 (multiple circles in the middle) and edges ( Figure 2 The lines connecting the circles in the middle form a semantic unit (such as a sentence or fragment) extracted from the document and its semantic relationships. Figure 2 The lines marked with an "×" indicate that redundant or weakly correlated edges are removed using a preset edge filtering strategy, thus eliminating noisy connections.

[0095] Specifically, each edge in the initial semantic structure graph is traversed, its semantic relevance weight is calculated, and the edge is determined whether to be retained according to the preset edge filtering strategy. Edges that do not meet the conditions are removed, and only key semantic relationships are retained to generate a graph structure that is more concise and semantically more relevant.

[0096] Step S2042: In the pruned initial semantic structure graph, initialize the input feature vector of each node.

[0097] In this embodiment, all semantic vectors are standardized. If there are dimensional differences, they are unified to the input dimension required by the graph neural network through linear mapping or projection dimensionality reduction. The semantic vector of each document file is loaded into the corresponding node in the graph to form the input node feature vector of the graph neural network.

[0098] Step S2043: Use a hierarchical adjustable graph convolutional network to recursively propagate node features. The recursive propagation involves weighted merging of the feature vectors of the current node and its neighboring nodes at each layer.

[0099] In this embodiment, a network hierarchical structure is first set up, and a GCN framework with multi-layer propagation capability is designed. Each layer is responsible for aggregating the information of first-order neighbor nodes. The number of propagation layers is dynamically determined according to the actual needs of the nodes in the graph. In each layer, the current node performs weighted fusion of its own features and the features of its neighbor nodes. The weights are dynamically calculated based on the edge similarity score, self-attention mechanism or structural relevance. The merged result is non-linearly activated to obtain the updated node feature vector. As the number of layers increases, the node gradually aggregates semantic information from farther distances.

[0100] Step S2044: clustering the recursive propagation result of the hierarchical adjustable graph convolution network to divide semantic clusters.

[0101] In this embodiment, a clustering model is selected, and the embedding vectors of all nodes are input into the clustering model. The model divides the nodes into semantic clusters according to their positions in the semantic space, and each cluster represents a type of semantically consistent document file fragments.

[0102] Step S2045: constructing connection edges between the cluster nodes of the semantic clusters to obtain a semantic structure graph.

[0103] In this embodiment, as shown in the upper right side of FIG. 14, each dashed box represents a semantic cluster, and in the semantic structure graph, these semantic clusters constitute the semantic structure graph. Figure 2 Figure 3

[0104] This step realizes the construction of a high-quality semantic structure graph from an original semantic graph through pruning, graph convolution feature propagation, and semantic clustering of the initial semantic structure graph. Specifically, the edge filtering strategy eliminates weakly related or redundant connections, enhancing the clarity and semantic purity of the graph structure. Feature initialization and graph convolution recursive propagation enable each node to obtain a multi-level, multi-neighbor semantic fusion embedding representation, reflecting its context role in the overall semantic network. Semantic clustering further divides nodes with similar structures and consistent semantics into clusters, mining potential knowledge unit grouping. Finally, the connection edges are reconstructed within the clusters to obtain a complete semantic structure graph, improving the association modeling capability between information fragments and significantly enhancing the understanding and knowledge reconstruction capability of fragmented document files.

[0105] Step S3: content completion is performed on nodes with incomplete associations in the semantic structure graph, and structured knowledge units are extracted from the completed semantic structure graph.

[0106] The specific steps of step S3 are as follows:

[0107] Step S301: identifying nodes or semantic clusters with context information breaks, disconnected logical paths, or incomplete semantic expressions in the semantic structure graph.

[0108] ​​In the embodiment, the semantic vector of each node is analyzed for similarity with its neighbor nodes, and if the similarity between the node and its neighbor nodes is significantly low, or the node lacks bidirectional connection, the node is marked as a node that may have context break; the average edge density and the shortest path length between the semantic clusters in the graph and other semantic clusters are calculated, and if the edge density between a certain semantic cluster and other clusters tends to zero, or the path is longer than a set threshold, it is determined to be logically disconnected; language role labeling or event graph template is used to detect whether each node has predicate center, argument role, causal relationship and other key semantic components, and the node that lacks important components is marked as incomplete expression; through the above three detection methods, the nodes or semantic clusters with semantic quality defects in the semantic structure graph can be effectively located.

[0109] Step S302: For the identified defect area, context generation is performed, which includes reasoning and completing the missing transition sentence or explanation segment based on the content of the neighbor nodes in the semantic structure graph.

[0110] In the embodiment, for each node or semantic cluster marked as a defect area, the semantic vectors of the neighbor nodes in the graph are extracted, and the semantic type and context logic are analyzed; based on the semantic edge type between the neighbor nodes, the shortest semantic path or causal link is constructed to determine the semantic function requirement of the node or semantic cluster of the defect area, such as whether causal explanation or logical transition needs to be supplemented; the above neighbor node content is input into the pre-trained context generation model as input, and a logical completion sentence or explanatory segment is output; it should be noted that in addition to using the content of the neighbor nodes in the semantic structure graph, external context metadata (such as folder level, time sequence, reference chain information) is also introduced, and statistical reasoning or causal reasoning is combined to generate a logically complete transition sentence or explanatory segment, ensuring the accuracy of the completion content.

[0111] As shown in FIG. 3, Figure 3 the bold nodes on the right side of the graph are the nodes that are completed for the identified defect area.

[0112] Step S303: The semantic structure graph after completion is subjected to semantic boundary segmentation, and a semantic cluster serving as a knowledge unit is identified.

[0113] The specific steps of step S303 are as follows:

[0114] Step S3031: The semantic consistency in each semantic cluster in the completed semantic structure graph is evaluated.

[0115] Step S3032: The semantic turning point in the semantic cluster is detected, and the preliminary boundary node is located based on the syntactic change and the path mutation behavior in the completed semantic structure graph.

[0116] In this embodiment, the node content in each semantic cluster is syntactically analyzed to extract features such as grammatical structure, conjunction type, number of clauses, and detect significant grammatical transition features in the sentence; at the same time, based on the completed semantic structure graph, the context path of the node in the graph is tracked, the semantic embedding change rate and path length change between the two nodes are calculated, if there is a semantic jump or path surge, it is judged as a semantic jump behavior, and the above two dimensions (grammatical transition + graph path mutation) are fused to score the transition possibility of each node, and the nodes with scores higher than the preset threshold are determined as preliminary boundary nodes.

[0117] Step S3033: Analyze each semantic node in the semantic cluster to determine whether it constitutes a complete semantic structure.

[0118] In this embodiment, the semantic role labeling technology is used to identify the semantic label of each semantic node, to determine whether it has common semantic role elements such as action subject, action verb, object, and result; the logical coherence between sentences is analyzed to determine whether there are reference errors and logical breaks; a completeness scoring model is constructed to evaluate the semantic role coverage, logical structure completeness, and context consistency of the node in three dimensions, and the nodes with scores lower than the set threshold are determined as incomplete semantic structure nodes.

[0119] Step S3034: Based on the complete semantic structure judgment result, the preliminary boundary nodes are optimized to obtain the semantic boundary.

[0120] The specific steps of step S3034 are:

[0121] Step S30341: Merge adjacent nodes that are marked as incomplete semantic structure to form a complete semantic unit.

[0122] In this embodiment, based on the semantic vector cosine similarity and other measurement methods, the forward or backward neighbor nodes of the current incomplete node in the semantic structure graph are compared in terms of semantics, and a set of candidate nodes with similar semantics are selected, and each candidate node is assigned a structure proximity score based on the context path weight. The candidate nodes are prioritized, the current incomplete node is concatenated or vector fused with the optimal candidate node at the text level, and the reconstructed text is generated through simplification rules or a lightweight language model.

[0123] Step S30342: Aggregate semantic nodes with repeated or parallel expressions and remove redundant semantic nodes.

[0124] Step S30343: Based on the edge weight and semantic node connection in the semantic cluster, adjust the preliminary boundary nodes, select the semantic nodes with complete semantic structure as boundary nodes, and obtain the semantic boundary.

[0125] In this embodiment, firstly, the connection edges of all the preliminary boundary nodes in the semantic cluster are traversed, the edge weights are extracted and the edge weight vectors are constructed, and the semantic connectivity strength of the node and the surrounding nodes is evaluated; the semantic integrity of the preliminary boundary nodes and the neighbor nodes is checked, including the syntactic structure integrity and logical independence check, such as whether there are front and back dependent words or a shortage of association; if the preliminary boundary node has a weak edge connection or a semantic incompleteness feature, the candidate boundary node is offset forward or backward, that is, the adjacent node, and the node with higher edge weight connection and higher semantic integrity is selected as the boundary node; after the node replacement is completed, the contribution degree of the node to the division of the semantic cluster is recalculated, the stability of the semantic boundary is analyzed, and the position of the semantic boundary is finally determined.

[0126] This step effectively improves the accuracy and robustness of semantic boundary recognition by merging adjacent nodes with incomplete semantic structures, aggregating semantic duplicate nodes, and optimizing boundary nodes based on edge weight values and connection relationships, which can eliminate redundant information, correct boundary misjudgment, and ensure the logical and semantic integrity of knowledge units.

[0127] Step S3035: Based on the semantic boundary, the semantic cluster as a knowledge unit is divided.

[0128] This step accurately identifies and divides knowledge units with logical closure and semantic integrity by evaluating the internal consistency of the semantic clusters in the completed semantic structure graph, detecting semantic turning points, judging semantic integrity, and optimizing boundaries, which can effectively remove semantic impurities and significantly improve the accuracy of knowledge construction.

[0129] As shown in Figure 3 , for the divided semantic boundary, redundant nodes are removed. It should be noted that Figure 3 the semantic structure graph in the form of nodes is used to facilitate understanding of the technical principles.

[0130] Step S304: The divided knowledge unit is represented in a structured form.

[0131] This step identifies broken or incomplete areas in the semantic structure graph, generates context content using neighbor node semantic reasoning, realizes intelligent completion of the semantic graph, and performs semantic boundary division and structured representation on this basis, which not only enhances the accuracy and systematicness of knowledge expression, but also significantly improves the ability to automatically extract knowledge units from unstructured text.

[0132] Step S4: Based on the structured knowledge unit, a new knowledge text is generated.

[0133] The specific steps of step S4 are:

[0134] Step S401: Establish a new knowledge generation target model.

[0135] It should be noted that based on the generative pre-training model, the knowledge generation target model is designed, and the model in the prior art is used, such as the Transformer model.

[0136] Step S402: Path planning is performed on the structured knowledge unit, and a path is selected as the main content of the new knowledge.

[0137] In this embodiment, the structured knowledge unit is subjected to graph abstraction processing, the relationship type and node weight between nodes are extracted, a path scoring mechanism is set, scoring is performed through the dimensions of semantic coherence, information density, knowledge contribution degree and user intention relevance, the structured knowledge unit graph is traversed, all candidate paths are scored, and the optimal or suboptimal path is selected as the main content of the new knowledge according to the scoring result.

[0138] Step S403: The knowledge unit corresponding to the selected path is input into the new knowledge generation target model, and the new knowledge content is generated according to the set generation parameter.

[0139] In this embodiment, the knowledge unit selected in the path planning stage is encoded into an input format acceptable to the model in a serialized manner, for example, the semantic nodes and edges are converted into natural language prompts or templated structures; the decoder part of the new knowledge generation target model is used, the encoded knowledge path and the set generation parameter are used as the basis to recursively generate the output content, that is, to generate the new knowledge content.

[0140] The advantage of the step is that the step realizes the conversion from the graph structure knowledge to the new knowledge by planning the path of the structured knowledge unit and making the new knowledge generation target model generate the new knowledge text, retains the logical rigor of the knowledge structure, and significantly improves the quality of the new knowledge content.

[0141] Embodiment 2

[0142] Please refer to Figure 4 Another embodiment provided by the application is a cross-file information summary and new knowledge automatic summary system, which comprises a data acquisition module, a semantic graph construction module, a content completion module and a new knowledge generation module.

[0143] The data acquisition module is configured to acquire document file data.

[0144] The semantic graph construction module is configured to vectorize the document file data based on a semantic embedding model, reconstruct and semantically cluster the semantic vectors of the document file data, and construct a semantic structure graph.

[0145] The content completion module is configured to complete the content of the nodes with incomplete association in the semantic structure graph, and extract the structured knowledge units from the completed semantic structure graph.

[0146] The new knowledge generation module is configured to generate a new knowledge text based on the structured knowledge units.

[0147] The semantic graph construction module comprises a preprocessing unit, a semantic vector unit and a graph construction unit.

[0148] The preprocessing unit is configured to preprocess the collected document file data, including content segment identification and splitting, content format standardization and denoising.

[0149] The semantic vector unit is configured to input the preprocessed document file data into a pre-trained feature extraction model, and map the extracted features to a unified semantic space to obtain a document file data semantic vector.

[0150] The graph construction unit is configured to weight and reconstruct the document file data semantic vector, construct an initial semantic structure graph, and perform clustering to construct a semantic structure graph.

[0151] In addition, the parts of the above technical solutions provided in the embodiments of the present application that are consistent with the implementation principles of the corresponding technical solutions in the prior art are not described in detail to avoid excessive repetition.

[0152] The specific embodiments described above further detail the purposes, technical solutions and beneficial effects of the present application. It should be understood that the above description is only a specific embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should be included in the protection scope of the present application.

Claims

1. A method for summarizing cross-file information and automatically summarizing new knowledge, characterized in that: include: Collect document file data; The document file data is vectorized based on a semantic embedding model. After reconstructing and semantically clustering the semantic vectors of the document file data, a semantic structure graph is constructed, which represents the logical relationship between the document file data. Complete the content of nodes with incomplete associations in the semantic structure graph, and extract structured knowledge units from the completed semantic structure graph; Based on the structured knowledge units, new knowledge text is generated; The process of vectorizing the document file data based on a semantic embedding model, reconstructing and semantically clustering the semantic vectors of the document file data, and constructing a semantic structure graph includes: The collected document file data is preprocessed, including content fragment recognition and segmentation, content format standardization, and noise reduction; The preprocessed document file data is input into a pre-trained feature extraction model, and the extracted features are mapped to a unified semantic space to obtain the semantic vector of the document file data. The semantic vectors of the document file data are weighted and reconstructed, and an initial semantic structure graph is constructed based on the enhanced semantic vectors of the document file data. Cluster the initial semantic structure graph and construct a new semantic structure graph based on the clustering results; The weighted reconstruction of the semantic vector of the document file data, and the construction of an initial semantic structure graph based on the enhanced semantic vector of the document file data, includes: Extract the associated context metadata from the preprocessed document file data. The associated context metadata includes: the folder structure data, document association references, and time series data. Context encoding is performed on the associated context metadata to obtain a unified context vector; A context modulation network is constructed, and a self-attention mechanism is used to calculate the relationship between the semantic vector of the document file data and the unified context vector to generate a context-sensitive weight matrix; The semantic vector of the document file data is reconstructed by weighting the context-sensitive weight matrix to obtain the enhanced semantic vector of the document file data. Based on the enhanced semantic vectors of the document file data, an initial semantic structure graph is constructed. The nodes of the initial semantic structure graph are the semantic vectors of individual document file data in the enhanced semantic vectors of the document file data, and the edges are the logical relationships between the semantic vectors of individual document file data. The generation of new knowledge text based on the structured knowledge units includes: Establish a new knowledge generation target model; Path planning is performed on structured knowledge units, and the path is selected as the core content of new knowledge; The knowledge units corresponding to the selected path are input into the new knowledge generation target model, and new knowledge content is generated according to the set generation parameters.

2. The method for cross-file information aggregation and automatic summarization of new knowledge as described in claim 1, characterized in that, The step of clustering the initial semantic structure graph and constructing a semantic structure graph based on the clustering results includes: A preset edge filtering strategy is used to prune the initial semantic structure graph; In the initial semantic structure graph after pruning, the input feature vector of each node is initialized; A hierarchical adjustable graph convolutional network is used to recursively propagate node features. The recursive propagation involves weighted merging of the feature vectors of the current node and its neighboring nodes at each layer. Cluster the recursive propagation results of hierarchical adjustable graph convolutional networks to divide them into semantic clusters; By constructing connecting edges between cluster nodes of semantic clusters, a semantic structure graph is obtained.

3. The method for cross-file information aggregation and automatic summarization of new knowledge as described in claim 1, characterized in that, The process of completing the content of incompletely associated nodes in the semantic structure graph and extracting structured knowledge units from the completed semantic structure graph includes: Identify nodes or semantic clusters in the semantic structure graph that have broken contextual information, disconnected logical paths, or incomplete semantic expression; For the identified defective areas, context generation is performed. The context generation includes reasoning and completing the missing transition sentences or explanatory fragments based on the content of neighbor nodes in the semantic structure graph. The semantic boundary is segmented on the completed semantic structure graph to identify semantic clusters as knowledge units; Represent the divided knowledge units in a structured form.

4. The method for cross-file information aggregation and automatic summarization of new knowledge as described in claim 3, characterized in that, The step of segmenting the semantic boundary of the completed semantic structure graph to identify semantic clusters as knowledge units includes: The semantic consistency within each semantic cluster in the completed semantic structure graph is evaluated. Semantic turning points are detected within the semantic cluster, and preliminary boundary nodes are located based on path mutation behavior in the semantic structure graph after syntactic changes and completion. Each semantic node in the semantic cluster is analyzed to determine whether it constitutes a complete semantic structure. Based on the complete semantic structure judgment result, the preliminary boundary nodes are optimized to obtain the semantic boundary; Based on the aforementioned semantic boundaries, semantic clusters are segmented as knowledge units.

5. The method for cross-file information aggregation and automatic summarization of new knowledge as described in claim 4, characterized in that, The initial boundary nodes are optimized based on the judgment result of the complete semantic structure to obtain the semantic boundary, including: Semantic nodes marked as having incomplete semantic structures are merged into adjacent nodes to form semantically complete units; Aggregate semantic nodes that are semantically repetitive or listed in a combined manner, and remove redundant semantic nodes; Based on the edge weights and semantic node connections in the semantic cluster, the initial boundary nodes are adjusted, and the semantic nodes with complete semantic structures are selected as boundary nodes to obtain the semantic boundary.

6. A system for automatically summarizing cross-document information and new knowledge, used to implement the method for automatically summarizing cross-document information and new knowledge as described in any one of claims 1-5, characterized in that, include: The module includes a data acquisition module, a semantic graph construction module, a content completion module, and a new knowledge generation module. The data acquisition module is used to acquire document file data; The semantic graph construction module vectorizes the document file data based on the semantic embedding model, reconstructs and performs semantic clustering on the semantic vectors of the document file data, and then constructs a semantic structure graph. The content completion module is used to complete the content of incomplete nodes in the semantic structure graph and extract structured knowledge units from the completed semantic structure graph. The new knowledge generation module generates new knowledge text based on the structured knowledge units.

7. The cross-document information aggregation and new knowledge automatic summarization system as described in claim 6, characterized in that, The semantic graph construction module includes: a preprocessing unit, a semantic vector unit, and a graph construction unit; The preprocessing unit is used to preprocess the collected document file data, including content fragment recognition and segmentation, content format standardization, and noise reduction. The semantic vector unit is used to input the preprocessed document file data into the pre-trained feature extraction model and map the extracted features into a unified semantic space to obtain the semantic vector of the document file data. The graph construction unit is used to perform weighted reconstruction of the semantic vector of document file data, construct an initial semantic structure graph, and perform clustering to construct a semantic structure graph.

Citation Information

Patent Citations

  • Knowledge reasoning method based on multi-modal knowledge graph

    CN112288091A

  • Project research and development data key information processing method and device

    CN120216748A