AI-Driven Multilingual Intelligent Document Summarization and Mind Map Generation System

Through multimodal semantic feature generation and dynamic semantic flow modeling, combined with the generation adversarial network optimization map structure, the problem of inaccurate content and single logical presentation in multimodal document processing is solved, and efficient and logically clear document summary and mind map generation are achieved.

CN119494394BActive Publication Date: 2025-07-22BEIJING SOIN TECH CORP LTD

Patent Information

Application Number
CN202510076741.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-07-22
Estimated Expiration
2045-01-17

AI Technical Summary

Technical Problem

The prior art has problems in multimodal document processing that the generated content is too broad or does not conform to user intentions, the logic is single and inefficient, especially when it is precise semantic extraction is weak.

Method used

Through the multimodal semantic feature generation module, text, pictures, tables and other data are uniformly represented, and a dynamically optimized causal semantic map is constructed, combined with dynamic semantic flow modeling and generation of adversarial network optimization map structure, and through user interaction adjustment, a document summary and mind map with clear logic and accurate semantics are generated.

Benefits of technology

It realizes deep integration and intelligent presentation of multimodal data, improves the accuracy and semantic consistency of generated content, meets the needs of diverse users, and improves the flexibility and efficiency of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119494394B_ABST
    Figure CN119494394B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of artificial intelligence technology, and particularly to an AI-driven multi-language intelligent document summarization and mind map generation system. The system realizes the unified representation of data such as multi-language texts, images, and tables through a multi-modal semantic feature generation module, and constructs a dynamically optimized causal semantic graph; generates a logically clear semantic flow path based on the dynamic semantic flow modeling method for multi-language summary generation; uses a generative adversarial network to optimize the structure of the mind map, and further improves the content of the mind map through user interaction adjustment and logical consistency verification. The system of the present invention supports users to efficiently extract the core information of the document, visually displays the logical structure, significantly improves the accuracy, logic, and efficiency of document generation, and is applicable to the generation of complex multi-modal documents and decision-making support.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and particularly to an AI-driven multilingual intelligent document summarization and mind map generation system. Background Art

[0002] With the popularization of big data and multilingual texts, users need to efficiently extract the core information of documents and present logical relationships in a visual way for easy understanding and decision-making. However, traditional methods often have deficiencies in multimodal semantic integration and logical presentation and cannot meet the processing requirements of complex documents. The application of AI technology in document generation and semantic analysis provides new possibilities for improving document processing efficiency and accuracy.

[0003] In the prior art (Chinese invention patent, publication number: CN118551740B, title: Document Generation Method, System and Electronic Device), a large language model is used to process document generation and mind map construction. Although slice technology is used to slice the original document and combined with prompt information to generate the target document and mind map, there are the following main defects: the generated content is often too broad or does not conform to the user's intention, especially when precise semantic extraction is required; the prior art adjusts the target document by modifying the mind map feedback, but due to the single logic of mind map generation, it is easy to cause the adjusted content to still not fully meet the user's needs; the large language model needs to reprocess all the original documents when regenerating the target document, with a large workload and low efficiency. Summary of the Invention

[0004] In view of the many problems existing in the above prior art, the present invention provides an AI-driven multilingual intelligent document summarization and mind map generation system. The present invention uses a multimodal semantic feature generation module to uniformly represent data such as text, pictures, tables, etc., and constructs a dynamically optimized causal semantic graph data; the summary generation module uses a dynamic semantic flow modeling method to generate multilingual summaries; the mind map generation and optimization module optimizes the mind map structure based on a generative adversarial network, and further improves the system function through user interaction and logical verification. The system of the present invention can efficiently generate document summaries and mind maps with clear logic and accurate semantics, and realize the deep integration and intelligent presentation of multimodal data.

[0005] An AI-driven multilingual intelligent document summarization and mind map generation system, comprising:

[0006] A multimodal semantic feature generation module, configured to receive multimodal document data, extract and fuse the semantic features of text content, picture data, table data, and video caption data, and generate unified multimodal semantic feature data;

[0007] The semantic graph construction module constructs dynamic optimized causal semantic graph data representing the causal relationships between semantic nodes based on multi-modal semantic feature data, including causal relationship extraction and multi-modal relationship calibration, and generates calibrated causal relationship data;

[0008] The abstract generation module extracts the content of key semantic nodes based on the dynamic optimized causal semantic graph data, generates a semantic flow path using the dynamic semantic flow modeling method, and generates multi-lingual abstract data in combination with the multi-lingual generation method;

[0009] The mind map generation and optimization module generates an initial mind map based on the dynamic optimized causal semantic graph data and the multi-lingual abstract data, and generates the final optimized mind map data through hierarchical relationship optimization and user interaction adjustment.

[0010] Preferably, in the multi-modal semantic feature generation module:

[0011] The text processing unit decomposes semantic units by performing a word segmentation algorithm on the text content, establishes a syntactic tree structure using a syntactic parsing method, and extracts the text theme through a semantic analysis model;

[0012] The image processing unit extracts the region boundaries and semantic categories in the image through an image semantic segmentation method, and locates the coordinates and category information of each object in the image in combination with an object detection model;

[0013] The table processing unit identifies the association relationships between column headers, row headers, and cell contents through a table parsing algorithm, and generates table hierarchical structure data;

[0014] The video subtitle processing unit transcribes the video audio content into text through speech recognition technology, and marks the timestamps and semantic levels of the subtitle text using a time series analysis method.

[0015] Preferably, the semantic graph construction module extracts the causal relationships between semantic nodes through the following method:

[0016] Use a Bayesian network model to model the statistical correlation of multi-modal semantic feature data and establish potential causal relationships between nodes;

[0017] Use the conditional independence test method to verify the effectiveness of the causal relationship, calculate the causal strength between nodes, and eliminate the edges with strength lower than a predetermined threshold;

[0018] Record the causal relationships that have passed the conditional independence test as causal relationship data for further construction of the semantic graph.

[0019] Preferably, the multi-modal relationship calibration in the semantic graph construction module includes:

[0020] Optimize the semantic alignment degree between different modality nodes using a cross-modal relational graph network, where the feature representation of nodes is calibrated by calculating the semantic vector similarity between modalities;

[0021] Adjust the low-confidence edges in the graph based on context consistency, increase the weights of edges with high confidence, decrease the weights of edges with low confidence, and eliminate invalid edges;

[0022] Construct the calibrated nodes and edges into calibrated causal relationship data for the further generation of dynamically optimizing the causal semantic graph.

[0023] Preferably, the abstract generation module generates a semantic flow path through a dynamic semantic flow modeling method, including the following steps:

[0024] Initialize the semantic flow matrix and calculate the semantic flow intensity according to the node weights in the causal semantic graph;

[0025] Optimize the directionality of the semantic flow path based on the reinforcement learning method, where the target reward function for the semantic flow includes path semantic coverage rate, consistency of the flow direction, and maximization of node weights;

[0026] Extract the key node content along the semantic flow path to generate semantic flow path data for the generation of multilingual abstracts.

[0027] Preferably, the multilingual abstract data generated by the abstract generation module is verified through a consistency verification method, including the following steps:

[0028] Use a multilingual natural language inference model to verify the logical consistency of the abstract content in different language versions, where the verification model uses the semantic similarity and logical coherence between sentences as evaluation indicators;

[0029] Mark the inconsistent abstract paragraphs and feedback the inconsistent parts to the abstract generation module to regenerate the optimized multilingual abstract content.

[0030] Preferably, the mind map generation and optimization module generates the hierarchical structure of the mind map through the following method:

[0031] Sort the semantic weights of the mind map nodes and divide the priority levels of the nodes according to the sorting results;

[0032] Associate the low-level nodes to the high-level nodes according to the causal relationship between the nodes to form a well-defined mind map structure;

[0033] During the mind map generation process, ensure that the semantic association path between each node and its upper and lower nodes is complete and error-free.

[0034] Preferably, the user interaction adjustment function of the mind map generation and optimization module includes:

[0035] Provide a visual interface for the mind map, displaying the semantic content of nodes and their causal association paths;

[0036] Support users to adjust the position of nodes by dragging, modify node content or add annotations through input interaction, and correct the weights and edge relationships between adjusted nodes through real-time feedback;

[0037] Automatically update the mind map structure modified by the user into the mind map data optimized by the user's adjustment.

[0038] Preferably, the mind map generation and optimization module verifies the consistency between the mind map and the abstract content through a natural language inference model, including the following steps:

[0039] Use the inference model to verify whether the content of the mind map nodes matches the semantics of the abstract, where the inference model uses the node semantic accuracy rate and the abstract matching rate as evaluation criteria;

[0040] Verify whether the edge relationships in the mind map conform to the logical chain in the abstract content, and mark the inconsistent nodes or edges.

[0041] Preferably, the mind map generation and optimization module optimizes the mind map structure through a generative adversarial network, including the following steps:

[0042] The generator generates multiple candidate mind map structures, each of which contains reallocated node weights and edge relationships;

[0043] The discriminator conducts a logical consistency evaluation and semantic coherence verification on the candidate mind maps, and selects the optimal mind map structure as the final optimized mind map data.

[0044] Compared with the prior art, the advantages and beneficial effects of the present invention are as follows:

[0045] Through multi-modal data fusion technical means, the present invention realizes the unified semantic representation of multi-language texts and non-text data such as images and tables, improving the accuracy and semantic consistency of the generated content;

[0046] Through dynamic semantic flow modeling and reinforcement learning optimization techniques, the present invention realizes the clear expression of abstract logic and the dynamic optimization of semantic flow paths, significantly improving the logic and coherence of abstract generation;

[0047] By optimizing the mind map structure through a generative adversarial network, the present invention realizes the optimization of a visual mind map with strong logical consistency and clear levels, solving the problem of single logical presentation in the comparative solution;

[0048] Through user interaction adjustment and real-time optimization, the present invention significantly improves the flexibility of the system and meets diverse user needs. Description of the Drawings

[0049] Figure 1 This is the structural block diagram of the system of the present invention;

[0050] Figure 2 This is the internal process schematic diagram of the semantic graph construction module in the present invention;

[0051] Figure 3 This is the internal process schematic diagram of the abstract generation module in the present invention;

[0052] Figure 4 This is the internal process schematic diagram of the mind map generation and optimization module in the present invention. Detailed Embodiments

[0053] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present disclosure. In the following detailed description, for the purpose of explanation, many specific details are set forth in order to provide a comprehensive understanding of the embodiments of the present disclosure.

[0054] As Figure 1 shown, an AI-driven multilingual intelligent document abstract and mind map generation system includes:

[0055] A multimodal semantic feature generation module, configured to receive multimodal document data, extract and fuse the semantic features of text content, picture data, table data, and video subtitle data, and generate unified multimodal semantic feature data;

[0056] Preferably, in the multimodal semantic feature generation module:

[0057] The text processing unit decomposes semantic units by performing a word segmentation algorithm on the text content, establishes a syntactic tree structure using a syntactic parsing method, and extracts the text theme through a semantic analysis model;

[0058] The picture processing unit extracts the region boundaries and semantic categories in the picture through an image semantic segmentation method, and combines an object detection model to locate the coordinates and category information of each object in the picture;

[0059] The table processing unit identifies the association relationships between column headers, row headers, and cell contents through a table parsing algorithm, and generates table hierarchical structure data;

[0060] The video subtitle processing unit transcribes the video audio content into text through speech recognition technology, and uses a temporal analysis method to mark the timestamps and semantic levels of the subtitle text.

[0061] The core objective of the multi-modal semantic feature generation module is to extract and fuse multi-modal semantic features from document data of different modalities, forming a unified high-dimensional semantic representation, which provides a basis for subsequent semantic analysis, abstract generation, and mind map construction. The multi-modal document data includes text content, image data, table data, and video caption data. Different modalities have their own structures and characteristics. Therefore, appropriate processing methods need to be used to extract their semantic features respectively, and cross-modal fusion methods are used to map these features into a high-dimensional semantic space uniformly, generating unified multi-modal semantic feature data.

[0062] The processing of text content is mainly achieved through word segmentation, syntax parsing, and semantic analysis models. The word segmentation algorithm is used to divide the text content into basic semantic units (such as words or phrases); subsequently, a syntax parsing method is used to generate a syntactic tree structure, which includes nodes (words or phrases) and their syntactic dependency relationships; finally, the main theme of the text is extracted through a semantic analysis model.

[0063] For example, when processing a text describing the weather, the word segmentation algorithm divides it into word units such as "today", "weather", and "sunny"; the syntax parsing generates a syntactic tree of "today - subject - weather" and "weather - modifier - sunny"; the semantic analysis model identifies "the weather is sunny" as the theme feature.

[0064] The processing of image data is completed by combining image semantic segmentation and object detection. The image semantic segmentation method is used to segment different regions in the image, extracting region boundaries and semantic categories; the object detection model further identifies the specific categories of each object in the image (such as cars, buildings) and their coordinate positions.

[0065] For example, when processing a city landscape image, the image semantic segmentation method divides the image into regions such as "sky", "road", and "building"; the object detection model identifies specific objects such as "high-rise buildings" or "bridges" in the "building" region and records the coordinates and categories of these objects.

[0066] The processing of table data is completed by a table parsing algorithm. The algorithm first identifies column headers and row headers, determines the relationship between cell content and headers through association analysis; then generates hierarchical table structure data based on cell content and its position in the table.

[0067] For example, for a sales table, the algorithm identifies "product name" as the column header and "2024" as the row header. The data "1000" in the cell is parsed as "the sales volume of a certain product in 2024 is 1000" through position association, and the generated hierarchical structure associates the headers and cell content as complete semantic information.

[0068] The processing of video subtitle data is completed by combining speech recognition and temporal analysis. The speech recognition model transcribes the audio content in the video into text and records the timestamp corresponding to each piece of text; the temporal analysis method further extracts the contextual semantic relationships and temporal hierarchies of the subtitle text to generate structured subtitle semantic features.

[0069] For example, when processing a speech video, the speech recognition model transcribes the audio into multiple sentences of subtitle text and marks "0:00 - 0:10" for the first sentence and "0:10 - 0:20" for the second sentence; the temporal analysis method identifies the logical relationships (such as causal or progressive) between these sentences to generate subtitle features with time and semantic associations.

[0070] Map the semantic features of the above-mentioned multi-modal data to a unified high-dimensional semantic space. The cross-modal fusion method is implemented through a Cross-modal Transformation Network (CMTN), which receives multi-modal semantic features as input and generates a unified semantic vector representation. Specifically, each modal feature vector is transformed into a high-dimensional semantic vector through a modality-specific encoder, and then the semantic differences between modalities are aligned through an attention mechanism in the fusion layer to generate unified multi-modal semantic feature data.

[0071] Achieve semantic alignment of different modal data. The generated unified multi-modal semantic feature data has high-dimensional semantic representation capabilities, providing a basis for subsequent modules to be compatible with multi-modal data. Through technical means such as word segmentation, semantic segmentation, object detection, and temporal analysis, the core semantic information of text, pictures, tables, and subtitles can be accurately extracted to avoid information loss. Using the cross-modal transformation network and efficient pre-trained models significantly reduces the complexity of processing multi-modal data and improves the efficiency of semantic feature extraction and fusion.

[0072] In an embodiment, the user uploads a business report containing text descriptions, sales data tables, product pictures, and video explanations. The goal is to extract and unify the semantic features of these modalities for generating a business decision summary and a mind map. The processing steps include:

[0073] The text processing unit decomposes the descriptive text in the report into words through a word segmentation algorithm and extracts "quarterly sales performance" and "product market feedback" as themes through a semantic analysis model.

[0074] The picture processing unit performs semantic segmentation and object detection on the product pictures to extract semantic features such as "product appearance", "product performance display", and their corresponding regional positions.

[0075] The table processing unit analyzes the sales table data to generate a hierarchical semantic structure such as "The sales volume of Product A in 2024 is 1000 pieces".

[0076] The video subtitle processing unit transcribes the audio content of the commentary video into text subtitles, marks the timestamps, and simultaneously extracts the logical relationships (such as cause explanations) between the subtitle texts.

[0077] The cross-modal fusion method inputs the feature vectors of text, pictures, tables, and video subtitles into a cross-modal transformation network, aligns the semantic differences through an attention mechanism, and generates a unified semantic vector representation.

[0078] The generated multi-modal semantic feature data covers text topics, picture object features, table hierarchical data, and subtitle logical information. The unified high-dimensional semantic vector can support the generation of clear business decision summaries and logically rigorous mind maps.

[0079] As Figure 2 shown, the semantic graph construction module constructs a dynamically optimized causal semantic graph data representing the causal relationships between semantic nodes based on the multi-modal semantic feature data, including causal relationship extraction and multi-modal relationship calibration, and generates calibrated causal relationship data;

[0080] The semantic graph construction module aims to construct a dynamically optimized causal semantic graph based on the multi-modal semantic feature data to represent the causal relationships between semantic nodes. This module mainly includes two key parts: causal relationship extraction and multi-modal relationship calibration. By extracting the semantic nodes and their causal relationships in the multi-modal data and performing multi-modal calibration and optimization on them, a causal semantic graph with dynamic update capabilities can be generated. This graph can not only accurately describe the logical relationships between semantic nodes but also adapt to the dynamic changes of multi-modal features, ensuring efficient cooperation with subsequent modules (such as summary generation and mind map optimization).

[0081] Based on the multi-modal semantic feature data, a causal reasoning method is used to extract the causal relationships between semantic nodes. Causal reasoning models the correlation between nodes through a Bayesian network and generates a preliminary structure of the causal relationship. Specifically, the feature vector of each semantic node is mapped to a node in the network, and the edges between the nodes represent possible causal relationships. Conditional independence tests are used to further verify the reliability of the edges and calculate the strength of the causal relationship.

[0082] Multi-modal relationship calibration aims to optimize the semantic consistency of different modal nodes in the initial causal relationship graph. Through the Cross-modal Relational Graph Network (CR-GNN), the alignment and enhancement of node semantic representations are achieved. The calibration process includes the following steps:

[0083] Cross-modal alignment: Map the feature vectors of nodes in different modalities to the same semantic space, and adjust the weights of the nodes by calculating the cosine similarity between modalities, so that the cross-modal node representations have high consistency.

[0084] Relationship optimization: Adjust the edges with low confidence in the initial causal relationship graph, eliminate the wrong causal relationship edges through context consistency verification, and at the same time enhance the weights of the edges with high confidence.

[0085] Through the above two core steps, the generated calibrated causal relationship data can accurately and comprehensively describe the causal relationship between semantic nodes, while ensuring its consistency with the characteristics of multi-modal data.

[0086] Preferably, the semantic graph construction module extracts the causal relationship between semantic nodes through the following method:

[0087] Use the Bayesian network model to model the statistical correlation of multi-modal semantic feature data and establish the potential causal relationship between nodes;

[0088] Use the conditional independence test method to verify the effectiveness of the causal relationship, calculate the causal strength between nodes, and eliminate the edges with strength lower than the predetermined threshold;

[0089] Record the causal relationship after conditional independence test as causal relationship data for further construction of the semantic graph.

[0090] The causal relationship extraction function in the semantic graph construction module aims to accurately identify the causal relationship between semantic nodes based on multi-modal semantic feature data to support the dynamic optimization of semantic graph generation. This process uses the Bayesian network model to model the statistical correlation between semantic nodes to construct potential causal relationships; subsequently, the conditional independence test method is used to verify the effectiveness of the potential relationships and calculate the strength of the causal relationships; finally, the high-confidence causal relationships after testing are recorded as causal relationship data for further optimization and application of the graph.

[0091] In the present invention, the Bayesian network is used to establish the preliminary structure of the causal relationship of multi-modal semantic feature data. Each semantic node represents a random variable, and its feature vector is provided by the semantic feature generation module (such as the topic feature vector of the text node, the object feature vector of the picture node). The Bayesian network constructs possible causal relationship edges by calculating the joint probability distribution between nodes.

[0092] For example, in the scenario of user comment semantic analysis, there may be a statistical correlation between the node "stable performance" and the node "sales volume increase". Through Bayesian network modeling, it can be inferred that "stable performance" is a potential cause of "sales volume increase".

[0093] Potential causal relationships need to be verified through conditional independence tests. The test process analyzes whether there is still statistical correlation between other nodes when the conditions of a certain node are known, and then judges the authenticity of the causal relationship. The conditional independence test method used is based on calculating the causal strength between nodes:

[0094] ,

[0095] wherein, represents the causal strength of node with respect to node ; represents the probability of node when the condition is only node ; represents the probability of node and context node when the condition is node . When is less than the predetermined threshold, the corresponding edge is removed to reduce invalid relationships.

[0096] For the high-confidence causal relationships verified through conditional independence tests, they are recorded as causal relationship data in the form of node-edge-weight. Among them, the node represents the semantic feature, the edge represents the causal relationship, and the weight is the causal strength. The generated causal relationship data serves as an important basis for semantic graph optimization and summary generation.

[0097] Through double screening of Bayesian network modeling and conditional independence tests, it is ensured that only high-confidence causal relationships are recorded, reducing misleading or noisy relationships. The extraction of causal relationships provides global semantic association information for the semantic graph, facilitating efficient utilization by subsequent modules in summary generation and mind map construction. By removing low-strength edges and readjusting weights, the generated causal relationship data can adapt to the dynamic changes of multi-modal semantic features.

[0098] Example: A user uploads a multi-modal document data containing product review texts, pictures, and sales statistics data. The goal is to extract the causal relationship between product features and sales volume changes for generating data-driven marketing decisions. Processing steps:

[0099] (1) Bayesian network modeling: Taking "stable performance" and "appearance design" extracted from the reviews as semantic nodes, combining object features in the pictures (such as "high-gloss product appearance") and the "sales volume increase" node in the sales data, a preliminary causal network is constructed; possible causal edges are generated by calculating the joint probability distribution between nodes: such as "stable performance → sales volume increase", "appearance design → user satisfaction".

[0100] (2)Conditional independence test: Conduct a conditional independence test on each edge in the preliminary causal network and calculate the causal strength. . For the edge "Design → User satisfaction", if = 0.4 and is lower than the predetermined threshold (e.g., 0.5), then remove this edge; for the edge "Stable performance → Increased sales", if = 0.8 which is higher than the threshold, then retain it and assign a weight.

[0101] (3)Generate causal relationship data: Record the nodes and edge weights of the final causal network. The generated causal relationship data includes "Stable performance → Increased sales (weight 0.8)", "User experience → User satisfaction (weight 0.7)", providing support for subsequent semantic graph optimization and decision summary generation.

[0102] Preferably, the multi-modal relationship calibration in the semantic graph construction module includes:

[0103] Use a cross-modal relational graph network to optimize the semantic alignment degree between different modal nodes, where the feature representation of nodes is calibrated by calculating the semantic vector similarity between modalities;

[0104] Adjust the low-confidence edges in the graph based on context consistency, increase the weights of high-confidence edges, decrease the weights of low-confidence edges, and remove invalid edges;

[0105] Construct the calibrated nodes and edges into calibrated causal relationship data for the further generation of dynamically optimizing the causal semantic graph.

[0106] The multi-modal relationship calibration of the semantic graph construction module aims to optimize the feature representation between semantic nodes of different modalities and the credibility of causal relationships, so as to improve the overall consistency and accuracy of the semantic graph. The calibration method optimizes the alignment of semantic vectors of different modal nodes through a cross-modal relational graph network and adjusts the low-confidence causal relationship edges based on context consistency, thereby removing invalid edges and strengthening credible relationships.

[0107] The semantic features of multi-modal data come from different modalities, such as text, pictures, tables, and video captions, and their semantic expressions have significant modal differences. The cross-modal relational graph network (CR-GNN) realizes semantic alignment by constructing a relational graph of multi-modal nodes and mapping the features of different modal nodes to a unified high-dimensional semantic space.

[0108] The semantic feature vector of each modal node is represented as , whose initial value is provided by the semantic feature generation module. The edges between nodes represent the possible semantic association relationships between modalities, and their weights are obtained by calculating the cosine similarity of the node feature vectors:

[0109] ,

[0110] Among them, represents the semantic similarity between nodes and , and is the Euclidean norm of the vector.

[0111] Through the multi-layer message passing mechanism, CR-GNN shares context information among multi-modal nodes and updates the feature vector representation of nodes, gradually aligning the semantic features of different modalities.

[0112] After optimizing the node semantic alignment, further calibrate the edges with low confidence in the causal relationship graph. The context consistency calibration includes the following steps:

[0113] Re-evaluate each edge according to the context consistency, and the weights of the edges with high confidence are increased, and the weights of the edges with low confidence are decreased. The context consistency is calculated by the following formula:

[0114] ,

[0115] Among them, represents the context consistency score of edge , and is the set of context edges associated with edge . When the calibrated weight of the edge is lower than the preset threshold (such as 0.3), it is directly removed to reduce noise.

[0116] After calibration, the optimized nodes and their edge relationships are recorded as calibrated causal relationship data in the structure of node-edge-weight. The calibrated causal relationship data contains the optimized semantic node representation, edge weights and their causal strengths, which are used to dynamically optimize the causal semantic graph.

[0117] The cross-modal relationship graph network realizes the alignment of different modality semantic features in a unified space, enabling multi-modal nodes to have consistent semantic expression capabilities. The context consistency calibration strengthens the weights of high-confidence edges, removes low-confidence invalid edges, and improves the reliability of causal relationships in the semantic graph. The optimized calibrated causal relationship data provides high-quality input for dynamically optimizing the causal semantic graph and can quickly adapt to changes in multi-modal data.

[0118] In an embodiment, a user submits a set of product review data, including text evaluations, product pictures, and sales table data. The goal is to analyze the causal relationship between features such as "stable performance" and "appearance design" and "sales increase" through semantic graph construction. Processing steps:

[0119] The "appearance design" node in the review text generates a feature vector through semantic analysis . The object feature vector of the "highlight appearance" area in the product picture is generated through object detection. The feature vector of the "sales increase" node in the table data is generated by table parsing. CR-GNN updates these node feature vectors through message passing and aligns them in the semantic space to make the semantic correlation between "appearance design" and "sales increase" higher.

[0120] In the initial causal relationship graph, the edge weight of "appearance design → user satisfaction" is 0.4. After context consistency calibration, due to insufficient support for the relevant edge, its weight is reduced to 0.2, which is lower than the threshold of 0.3, so it is removed; the edge weight of "stable performance → sales increase" is increased from 0.7 to 0.85. Due to the strong consistency of its context support (such as "user experience"), it is finally retained as a high-confidence edge.

[0121] The finally generated calibrated causal relationship data records the high-confidence causal relationship of "stable performance → sales increase (weight 0.85)"; low-confidence edges such as "appearance design → user satisfaction" are removed, providing optimized input for the causal semantic graph.

[0122] The calibrated semantic graph accurately reflects the key causal relationships in the multimodal data, clearly revealing how product features affect user experience and sales changes, providing a reliable basis for subsequent marketing strategy decisions.

[0123] As Figure 3 shown, the summary generation module, based on dynamically optimizing the causal semantic graph data, extracts the content of key semantic nodes, uses the dynamic semantic flow modeling method to generate semantic flow paths, and combines the multilingual generation method to generate multilingual summary data;

[0124] The core of the summary generation module is to extract the content of key semantic nodes based on dynamically optimizing the causal semantic graph data, generate semantic flow paths through dynamic semantic flow modeling, and combine the multilingual generation method to generate accurate and coherent multilingual summary data. The implementation of this module involves three key parts: semantic node extraction, semantic flow modeling, and multilingual generation.

[0125] Dynamically optimizing causal semantic graph data is a high-quality structured representation of semantic nodes and their causal relationships. The abstract generation module first extracts key semantic nodes based on node weights and causal relationship strengths. The calculation of node weights is based on semantic importance and centrality metrics in the causal network. For example, the PageRank value of a node is used to measure its importance. During the extraction process, high-weight nodes are preferentially selected while ensuring that the extracted nodes have logical coherence, that is, maintaining the integrity of the causal relationship chain.

[0126] After the extraction of key nodes, the dynamic semantic flow modeling method is used to generate semantic flow paths, clarifying the flow direction and logical sequence of semantic nodes. Semantic flow modeling constructs a semantic flow matrix through path optimization techniques in the fluid dynamics model and further optimizes the semantic flow paths through reinforcement learning algorithms to ensure that the semantic flow direction is consistent with the abstract generation goals, such as abstract length limits and topic focus.

[0127] The multilingual generation method is implemented based on multilingual pre-trained generation models (such as mT5 or mBART). The extracted semantic flow paths are used as the input of the model, and multilingual abstract content is generated in combination with the target language and style requirements. The model uses a semantic alignment mechanism during the generation process to ensure the consistency of semantic expressions and logical sequences in the abstracts generated in different languages. To improve the generation quality, a consistency verification mechanism is used to check the logical consistency between multilingual versions and optimize the generation results.

[0128] Preferably, the abstract generation module generates semantic flow paths through the dynamic semantic flow modeling method, including the following steps:

[0129] Initialize the semantic flow matrix and calculate the semantic flow intensity according to the node weights in the causal semantic graph;

[0130] Optimize the directionality of the semantic flow path based on the reinforcement learning method, where the target reward function for the semantic flow includes path semantic coverage rate, consistency of the flow direction, and maximization of node weights;

[0131] Extract the key node content along the semantic flow path to generate semantic flow path data for the generation of multilingual abstracts.

[0132] The abstract generation module generates semantic flow paths through the dynamic semantic flow modeling method to construct abstract content with clear logic and compact semantics. The core of dynamic semantic flow modeling is to initialize the semantic flow matrix based on the node weights in the causal semantic graph, optimize the path directionality through reinforcement learning, and extract the key node content to generate semantic flow path data, providing high-quality input for multilingual abstract generation. The whole process covers three key steps: initializing the semantic flow matrix, optimizing the semantic flow path, and extracting the key node content along the path.

[0133] Initialize the semantic flow matrix based on the node weights and edge weights in the causal semantic graph. The semantic flow matrix is used to describe the semantic flow intensity between nodes. Node weights represent the importance of a node in the causal graph, usually provided by the semantic feature generation module and calculated in combination with the semantic centrality index of the node (such as the PageRank value). Node to node The semantic flow intensity

[0134] ,

[0135] where represents the causal relationship intensity from node to ; represents the weight of node ; represents the set of neighbor nodes of node . The flow intensity matrix is generated by calculating the flow intensity of all node pairs one by one, laying a foundation for subsequent path optimization.

[0136] Use the reinforcement learning method to optimize the semantic flow path to ensure that the logical order of semantic nodes is consistent with the abstract generation goal. The state, action, and reward function of the reinforcement learning model are defined as follows:

[0137] State: The node sequence of the current semantic flow path;

[0138] Action: Select the next node to add to the path;

[0139] Reward function: Comprehensively consider the path semantic coverage rate, the consistency of the flow direction, and the maximization of node weights:

[0140] ,

[0141] where represents the path semantic coverage rate, indicating the total semantic weight of the selected nodes; represents the consistency of the flow direction, calculated based on the coherence of the flow intensity between nodes; represents the sum of the weights of the nodes in the path; , , represent the balance coefficients used to adjust the relative importance of each index. The optimization process gradually adjusts the node selection strategy of the path through the policy gradient method, and finally generates a semantic flow path that meets the target requirements.

[0142] The optimized semantic flow path is used to extract the core content of each node in the path and generate semantic flow path data. The content of each node is represented by its semantic feature vector, which is directly obtained through the node representation of the causal semantic graph. The extracted semantic flow path data has logical coherence and semantic coverage, providing accurate input for multi-language summary generation.

[0143] Dynamic semantic flow modeling optimizes the path direction through reinforcement learning, ensuring the logical order and semantic coherence of the summary content, and enhancing the readability of the summary. The key node content extracted based on the semantic flow path has high semantic importance and information coverage, ensuring that the summary can convey the core information of the document. The initialization and optimization process of the flow matrix can adapt to the document content with different semantic structures and support the summary generation requirements of multi-modal data.

[0144] Example: Process a market analysis report, the content of which includes a sales data table, a competitive analysis text, and a market trend picture. The goal is to generate a multi-language summary covering the core viewpoints of the market performance. Processing steps:

[0145] Initialize the semantic flow matrix: Extract nodes from the semantic graph, including "Sales Growth", "Market Share Change", and "Competitive Analysis". The node weights are respectively 、 、 . The flow intensity matrix is calculated as:

[0146] ,

[0147] Calculate the flow intensity of all node pairs in a similar way to generate the flow matrix.

[0148] Optimize the semantic flow path: Initialize the path starting from "Sales Growth". Reinforcement learning optimizes the path direction according to the reward function and selects the path "Sales Growth → Market Share Change → Competitive Analysis". Reward function calculation:

[0149] Path semantic coverage ; Flow direction consistency ; Sum of node weights .

[0150] Extract the key node content: Extract the node content along the optimized path. "Sales Growth": The growth rate shown in the sales data table is 15%; "Market Share Change": The share change shown in the market trend chart increases from 20% to 25%; "Competitive Analysis": The impact of a certain competitor withdrawing from the market mentioned in the competitive analysis text. The generated semantic flow path data describes the market growth logical chain.

[0151] The generated multilingual abstract content expresses the logical main line of market performance: sales growth drives an increase in market share, and competitive analysis further explains the reasons for share changes. The English abstract is: "Sales growth of 15% led to a market share increase from 20% to 25%, supported by competitive market dynamics." The Chinese abstract is: "Sales growth of 15% led to a market share increase from 20% to 25%, with competitive dynamics further supporting this change."

[0152] Preferably, the multilingual abstract data generated by the abstract generation module is verified by a consistency check method, including the following steps:

[0153] Use a multilingual natural language inference model to verify the logical consistency of the abstract content in different language versions, where the verification model uses the semantic similarity and logical coherence between sentences as evaluation indicators;

[0154] Mark the inconsistent abstract paragraphs, and feedback the inconsistent parts to the abstract generation module to regenerate the optimized multilingual abstract content.

[0155] The goal of the multilingual abstract consistency check method is to verify the logical consistency and semantic similarity of the abstract content in different language versions, and ensure that the generated multilingual abstracts are unified in terms of expression logic, content consistency, and information coverage. This method is implemented through a multilingual natural language inference model (Multilingual Natural Language Inference, M-NLI), which evaluates the semantic similarity and logical coherence between sentences of the multilingual version of the abstract; feedback the content of the marked inconsistent parts to the abstract generation module to regenerate the optimized multilingual abstract content.

[0156] The multilingual natural language inference model (M-NLI) is based on multilingual semantic embedding technology and can verify the logical consistency of different language versions of the abstract. The specific verification process is as follows:

[0157] Generate semantic embedding vectors for each abstract sentence through a multilingual embedding model (such as XLM-R) , which is used to represent the semantic features of the sentence. Calculate the semantic similarity of the corresponding sentences in different language versions. The formula is as follows:

[0158] ,

[0159] where, represents the semantic similarity of the source language sentence and the target language sentence ; and are the semantic embedding vectors of the sentences and respectively; is the Euclidean norm of the vector. The logical relationship (such as causality, progression, contrast) between sentences is checked through an inference model to verify the coherence of the "causal chain", for example.

[0160] When the semantic similarity between sentences is lower than a preset threshold (such as 0.8) or the logical coherence check fails, the corresponding abstract paragraph is marked as "inconsistent". The marked inconsistent parts are transmitted to the abstract generation module through a feedback mechanism to regenerate and optimize the content to improve consistency. The feedback optimization process combines a reinforcement learning model to adjust the abstract generation parameters (such as node selection strategy or path weight assignment) to ensure that the optimized abstract meets the consistency requirements.

[0161] Through semantic embedding and semantic similarity verification, it is ensured that the multi-lingual abstract content has the same core information and semantic expression among different languages. The logical check verifies the causal chain and logical order between sentences through an inference model to avoid logical inconsistencies caused by information biases during translation or generation. The annotation and feedback mechanism efficiently repairs inconsistent content, realizing the dynamic interaction and optimization closed-loop between the abstract generation module and the verification module.

[0162] In an embodiment, the system generated a multi-lingual abstract of a market analysis report, aiming to verify the logical and semantic consistency of the Chinese and English abstracts through a consistency verification method to ensure that the multi-lingual abstract can accurately express the key information of the market performance. Processing steps:

[0163] The Chinese abstract sentence "A 15% sales increase led to a market share increase" generates a semantic embedding vector ; The English abstract sentence "Sales increased by 15%, driving market share growth" generates a semantic embedding vector .

[0164] Calculate the similarity through the formula , and the result is 0.92, higher than the threshold of 0.8, and it is determined to be semantically consistent.

[0165] Verify whether the abstract sentence "Market share increase" is caused by "Sales increase". The inference model confirms through semantic analysis that there is no deviation in the logical chain, and the verification passes. If the logical verification fails (such as "missing causal chain"), it is marked as "inconsistent" and fed back to the abstract generation module.

[0166] For paragraphs that fail the logical verification, mark and regenerate the content. Assume that the original English abstract sentence "Marketshare grew independently" does not conform to the logic, and the system generates an optimized new sentence "Market sharegrowth was driven by sales increase".

[0167] The consistency verification method successfully verifies and optimizes the multi-language abstract content, ensuring the consistency of the logical order and semantic expression between the Chinese and English abstracts. The final abstract example is as follows: Chinese abstract: A 15% sales increase led to a market share boost and a significant improvement in market performance. English abstract: Sales increased by 15%, driving market share growthand significantly improving market performance.

[0168] As Figure 4 shown, the mind map generation and optimization module generates an initial mind map based on the dynamically optimized causal semantic graph data and multi-language abstract data, and generates the final optimized mind map data through hierarchical relationship optimization and user interaction adjustment.

[0169] The goal of the mind map generation and optimization module is to generate an initial mind map with clear logic and distinct levels based on the dynamically optimized causal semantic graph data and multi-language abstract data, and generate the final optimized mind map data through hierarchical relationship optimization and user interaction adjustment. This module combines the causal relationships between nodes in the causal semantic graph, the semantic information of multi-language abstracts, and the real-time input of users to gradually improve the mind map structure and ensure its logic and user-friendliness.

[0170] The generation of the initial mind map is based on the dynamically optimized causal semantic graph data. The initial structure of the mind map is constructed through a Recursive Graph Neural Network (R-GNN): The nodes in the semantic graph serve as the basic nodes of the mind map. Each node represents a semantic concept (such as "stable performance" or "sales growth"), and its attributes include semantic weight, modality source, and abstract semantics. The causal relationships in the semantic graph are mapped as edges in the mind map. Each edge contains a weight and a directionality, which are used to represent the causal or logical association between nodes. The initial mind map is presented as a hierarchical directed graph in structure, reflecting the core logic and semantic chain of the document content.

[0171] The hierarchical relationship optimization is based on the importance ranking of nodes and the logical analysis of semantic relationships, and adjusts the structure of the initial mind map: using node weights Sort the nodes by importance, and the nodes with higher weights are preferentially promoted to higher levels. The weight calculation formula is:

[0172] ,

[0173] where, represents the semantic weight of the node, provided by the semantic graph; represents the centrality index of the node, such as the PageRank value; , represent the weight adjustment parameters, used to balance the influence of the two indicators.

[0174] According to the weight sorting of the nodes, the nodes are divided into different levels. The nodes with high weights are located at higher levels, and the nodes with low weights are assigned to sub-levels to ensure the clarity of the mind map's hierarchy. Adjust the edge relationships across levels to reduce the complexity of the cross-level paths and ensure the simplicity of the mind map's logic.

[0175] Support users to adjust the mind map structure through the interactive interface to enhance the customization and user-friendliness of the mind map: Provide dynamic visual displays of nodes and edges, and users can clearly view the hierarchical structure and logical relationships of the mind map. Allow users to drag nodes to adjust their positions, modify node content, or add annotations. Support users to adjust the weights of edges or delete invalid edges, and update the mind map logic in real time. Based on user operations, the system adjusts the causal relationships and hierarchical divisions between nodes through an interactive real-time correction algorithm to ensure the consistency and accuracy of the mind map's logic.

[0176] After the user interaction adjustment, further optimize the mind map structure through the Generative Adversarial Network (GAN): Generator: Generate multiple candidate mind maps to explore possible node layouts and edge relationships. Discriminator: Evaluate the candidate mind maps according to logical consistency and semantic coherence, and select the optimal structure from them as the data of the finally optimized mind map.

[0177] Preferably, the mind map generation and optimization module generates the hierarchical structure of the mind map through the following method:

[0178] Sort the semantic weights of the mind map nodes, and divide the priority levels of the nodes according to the sorting results;

[0179] Associate the low-level nodes with the high-level nodes according to the causal relationships between the nodes to form a clearly hierarchical mind map structure;

[0180] During the mind map generation process, ensure that the semantic association paths between each node and its upper and lower nodes are complete and error-free.

[0181] The hierarchical structure generation method of the mind map generation and optimization module aims to hierarchically divide the nodes in the dynamically optimized causal semantic map according to semantic weights and causal relationships, and construct a mind map structure with distinct levels and clear logic. In the specific implementation process, first sort the semantic weights of the nodes, and divide the priority levels of the nodes according to the sorting results; then associate the low-level nodes with the high-level nodes according to the causal relationships between the nodes, and ensure that the semantic association paths between each node and its upper and lower nodes are complete and error-free.

[0182] The semantic weight of a node reflects the importance of the node in the entire mind map, and its calculation combines the semantic feature weight of the node and the network characteristics in the map (such as centrality indicators). The specific steps are as follows:

[0183] Node weight Is obtained by weighted summation of the semantic feature weight and the centrality indicator, and the formula is:

[0184] ,

[0185] Among them, The semantic feature weight of the node, provided by the dynamically optimized causal semantic map; The centrality indicator of the node (such as PageRank value); , The weight adjustment coefficient, used to balance the influence of semantic features and centrality indicators. Sort all nodes in descending order according to the calculated weights, assign high-weight nodes to higher levels, and assign nodes with lower weights to corresponding sub-levels. The hierarchical division ensures that high-weight nodes form the backbone structure, and low-weight nodes form sub-levels around high-weight nodes.

[0186] Based on the causal relationships in the semantic map, associate low-level nodes with high-level nodes through edge relationships to construct a mind map structure with distinct levels: the weight of each edge Represents the strength of the causal relationship between node And node . If node Is at a high level, and Is higher than the preset threshold, then associate node To node . If the edge weight is low or does not conform to the causal logic, the corresponding edge is removed to reduce invalid associations.

[0187] To ensure the integrity of the semantic association paths between each node and its upper and lower nodes, it is necessary to verify whether all causal chains between the node and its upper and lower levels are connected during the mind map generation process. For interrupted chains, the module will supplement missing edges or adjust the node positions through a path repair algorithm to make the logical chains complete.

[0188] Through semantic weight sorting and causality-driven association, the generated mind map structure is clear, highlighting the core semantic nodes and logical main lines in the document. Path integrity verification ensures that the causal chain between nodes is accurately presented in the mind map, avoiding problems such as information breaks or unclear logic. Hierarchical division and path optimization improve the readability and usability of the mind map, enabling users to quickly understand the core content and logical framework of the document.

[0189] Example: Process a product report containing multimodal semantic information such as "sales growth", "user satisfaction improvement", and "market share increase". The goal is to generate a clear mind map to display product performance and its driving factors. Processing steps:

[0190] Extract semantic nodes including "Sales Growth", "User Satisfaction", and "Market Share". Calculate node weights: (because it has the highest frequency of occurrence and the most associated nodes in the report); ; . Divide "Sales Growth" into the highest level, and divide "Market Share" and "User Satisfaction" into secondary levels.

[0191] Causality-driven hierarchical association, determine edge weights: The edge weight of "User Satisfaction → Sales Growth" is , retain this edge; The edge weight of "Market Share → Sales Growth" is , retain this edge. Associate "User Satisfaction" and "Market Share" with "Sales Growth" to build a backbone structure with "Sales Growth" as the core.

[0192] Path integrity guarantee, check chain integrity: Verify whether the path of "User Satisfaction → Sales Growth → Market Share" is connected. Supplement missing edges. If it is found that the logic between "Market Share" and "User Satisfaction" is not connected, add auxiliary edges through the path repair algorithm to ensure the integrity of the chain.

[0193] The generated mind map shows the driving factors of "Sales Growth", including the logical relationship between "User Satisfaction" and "Market Share". The mind map has clear levels and logical coherence, providing intuitive support for marketing decisions.

[0194] Preferably, the user interaction adjustment function of the mind map generation and optimization module includes:

[0195] Provide a visual interface for the mind map, displaying the semantic content of nodes and their causal association paths;

[0196] Support users to adjust the position of nodes by dragging, modify the node content or add annotations through input interaction, and correct the weights and edge relationships between nodes after adjustment through real-time feedback;

[0197] Automatically update the mind map structure modified by the user into the data of the optimized mind map adjusted by the user.

[0198] The user interaction adjustment function of the mind map generation and optimization module aims to provide users with the ability to visually adjust the mind map structure through the interaction interface, ensuring that the mind map better meets the user's needs in terms of logical integrity and content customization. Users can adjust the position of nodes by dragging, modify the node content or add annotations. At the same time, the system automatically updates the weights and edge relationships between nodes according to the user's adjustments and generates an optimized mind map structure. The entire interaction process includes three key steps: visual interface presentation, real-time adjustment and correction, and dynamic update and optimization.

[0199] The visual interface is the core window for users to operate the mind map, showing the semantic content of nodes and their causal association paths. Node display: Each node shows its semantic content and weight, and the node color or size is dynamically adjusted according to the weight. The higher the weight, the more prominent the node. Edge relationship display: The edges between nodes are represented by arrows to indicate the direction of the causal relationship, and the thickness of the edge represents the weight intensity. High-weight edges are thicker, and low-weight edges are thinner or shown as dashed lines. Real-time response: When the user clicks or hovers over a node, the interface highlights all the directly associated nodes and paths of that node, facilitating the user to understand the context relationship of the node.

[0200] Users can modify the mind map through interactive operations, and the system corrects the weights and edge relationships between nodes in real time to maintain logical consistency. Users can adjust the position of nodes in the mind map by dragging, and the adjusted position triggers the redistribution of weights between nodes. The weight adjustment formula is:

[0201] ,

[0202] where, represents the weights of nodes and node after adjustment; , represents the position coordinates of nodes and after adjustment; represents the position distance between nodes and after adjustment. Users can directly input to modify the semantic content of nodes or add annotations to nodes, and the annotation content is distinguished by color or special marks. Users can manually delete invalid edges or input the weight of a specified edge, and the system updates the edge relationship in real time.

[0203] After the user's adjustment is completed, the system dynamically updates and optimizes the mind map: deleting isolated nodes or invalid edges to maintain the logical integrity of the causal relationship; recalculating the node weights according to the user's modification and adjusting the node hierarchical positions to ensure that the mind map is well-structured; if the user's adjustment causes the causal chain to break, the system completes the interrupted path through a path repair algorithm to ensure the integrity of the logical chain between nodes.

[0204] Provide an interactive function to enable users to freely adjust the mind map structure and content to meet the personalized needs of different scenarios. The system automatically corrects the node weights and edge relationships after the user's adjustment to ensure the logical consistency of the mind map and the integrity of the causal chain. The visual interface intuitively displays the node and edge relationships, enabling users to quickly understand the logical structure of the mind map and improving the practicality and readability of the mind map.

[0205] Example, process a mind map showing the product marketing logic, and the nodes include "sales growth", "improvement of user satisfaction", and "increase in market share". The user hopes to highlight the importance of "improvement of user satisfaction" through interactive adjustment and add annotations to supplement background information. Processing steps:

[0206] Presented in the visual interface, node display: the node weight of "sales growth" is 0.85, with a darker color and a larger size; the node weight of "improvement of user satisfaction" is 0.65, with a lighter color and a medium size. Edge relationship display: the edge weight of "improvement of user satisfaction → sales growth" is 0.8, shown as a thick solid line; the edge weight of "increase in market share → sales growth" is 0.6, shown as a thin solid line.

[0207] User adjustment and real-time correction, the user drags the "improvement of user satisfaction" node to a position closer to "sales growth", triggering a reallocation of weights. After adjustment, the weight of "improvement of user satisfaction → sales growth" increases from 0.8 to 0.85. The user modifies the content of the "improvement of user satisfaction" node to "Improve user satisfaction (mainly achieved through product optimization)" and adds a note to the node "Data source: 2024 user feedback survey". The user deletes the edge of "increase in market share → sales growth", and the system automatically updates the causal relationship.

[0208] Dynamic update and optimization, the system recalculates the node weights: the weight of "improvement of user satisfaction" increases from 0.65 to 0.75, and the hierarchy is promoted to the same level as "sales growth". The system repairs the chain: due to the deletion of the edge of "increase in market share → sales growth", the system supplements the path "improvement of user satisfaction → increase in market share".

[0209] The adjusted mind map highlights the core role of "improvement of user satisfaction" and provides additional background information through annotations, with a complete logical chain, clear hierarchy, and is convenient for decision-making and analysis.

[0210] Preferably, the mind map generation and optimization module verifies the consistency between the mind map and the abstract content through a natural language inference model, including the following steps:

[0211] Use the inference model to verify whether the content of the mind map nodes matches the semantics of the abstract, where the inference model uses the node semantic accuracy rate and the abstract matching rate as evaluation criteria;

[0212] Verify whether the edge relationships in the mind map conform to the logical chain in the abstract content, and mark the inconsistent nodes or edges.

[0213] The mind map generation and optimization module performs consistency verification on the generated mind map and abstract content through a natural language inference model to ensure that the node and edge relationships of the mind map are consistent with the semantic content and logical chain of the abstract. The verification process includes verifying whether the content of the mind map nodes matches the semantics of the abstract and whether the edge relationships in the mind map conform to the logical chain in the abstract. The model is evaluated through the node semantic accuracy rate and the abstract matching rate, and the inconsistent nodes or edges are identified and marked.

[0214] Use the Natural Language Inference (NLI) model to perform semantic matching on the mind map nodes and the abstract content. The verification process includes the following steps:

[0215] Generate semantic embedding vectors for the content of the mind map nodes and the relevant paragraphs of the abstract and , representing their semantic features. The embedding vectors are calculated through a pre-trained model (such as XLM-R).

[0216] Calculate the matching rate between the node content and the abstract semantics :

[0217] ,

[0218] When the matching rate is lower than the threshold (such as 0.8), mark the node as an "inconsistent node" to prompt the need for further optimization.

[0219] Statistically calculate the matching rates of all nodes, and the node semantic accuracy rate is defined as:

[0220] ,

[0221] This value is used to comprehensively evaluate the consistency of the semantics of the mind map nodes.

[0222] Use the inference model to verify whether the edge relationships in the mind map conform to the logical chain of the abstract content: Generate a relationship diagram for the logical chain clearly described in the abstract (such as causal, progressive, etc. relationships) as the target chain; Compare the edge relationships in the mind map to check whether they conform to the target chain. For example, for the causal relationship "User satisfaction improvement → Sales growth" described in the abstract, it is necessary to verify whether the edge "User satisfaction improvement → Sales growth" exists in the mind map and the edge weight meets the requirements of the logical chain strength.

[0223] Score the logical matching degree of each edge. The scoring criteria include logical consistency in the abstract and consistency between the edge weight and the relationship strength in the abstract. Mark the edges with a score lower than the threshold (such as 0.75).

[0224] Mark the inconsistent nodes or edges and feedback them to the mind map generation module or the abstract generation module for adjustment and optimization: Node marking, add annotations to the "inconsistent nodes", such as "Semantics does not match the abstract", and prompt the user to modify or regenerate the node content. Edge marking, add annotations to the "inconsistent edges", such as "Logical chain interrupted", and prompt that the edge weight needs to be regenerated or adjusted.

[0225] Through semantic matching and logical chain verification, ensure that the mind map nodes and edge relationships are highly consistent with the abstract content, so that the mind map can accurately reflect the core logic of the document. The marking and feedback mechanism supports dynamic adjustment and optimization. Through continuous verification and repair, ensure the logical integrity of the mind map structure. Consistency verification improves the credibility and user trust of the mind map content, and provides a more reliable tool for the analysis and decision-making of complex documents.

[0226] Preferably, the mind map generation and optimization module optimizes the mind map structure through a generative adversarial network, including the following steps:

[0227] The generator generates multiple candidate mind map structures, each of which contains reallocated node weights and edge relationships;

[0228] The discriminator conducts logical consistency evaluation and semantic coherence verification on the candidate mind maps, and screens out the optimal mind map structure as the final optimized mind map data.

[0229] The Mind Map Generation and Optimization Module optimizes the mind map structure through a Generative Adversarial Network (GAN) with the goal of generating an optimal mind map structure with strong logical consistency and semantic coherence. The GAN model consists of a generator and a discriminator: The generator generates multiple candidate mind map structures, each containing reallocated node weights and edge relationships; The discriminator evaluates the logical consistency and verifies the semantic coherence of these candidate mind maps, and finally selects the optimal mind map as the optimization result. The entire process combines the information of the semantic graph and the abstract content to ensure that the optimization effect of the mind map meets the logical and semantic requirements.

[0230] The task of the generator is to generate multiple candidate mind map structures based on the initial mind map and explore possible allocations of node weights and edge relationships. The specific steps are as follows:

[0231] The generator uses a random perturbation function to adjust the node weights as follows:

[0232] ,

[0233] where, is the adjusted node weight; is the original node weight; is the random perturbation value, ; is the hyperparameter of the weight adjustment amplitude. The generator generates different candidate structures by adjusting the edge weights and adding / deleting edges to explore the optimization possibilities of the mind map.

[0234] The discriminator verifies the logical consistency and semantic coherence of the candidate mind maps generated by the generator to ensure that the selected mind map meets the system requirements:

[0235] The discriminator evaluates the consistency of the candidate mind maps based on the logical structure of the mind map (such as causal chains and hierarchical relationships). The logical consistency scoring formula is:

[0236] ,

[0237] where "valid chain" refers to a chain that conforms to the logic of causal relationships.

[0238] The discriminator verifies the semantic similarity between the mind map nodes and the abstract content, and calculates the similarity between the node semantic embeddings and the abstract content embeddings: If both the semantic similarity and the logical consistency score are higher than the preset threshold, it is determined as the optimal mind map.

[0239] The discriminator scores all candidate mind maps and selects the mind map with the highest score as the finally optimized mind map according to the comprehensive score (the weighted sum of the logical consistency score and the semantic coherence score).

[0240] Through GAN optimization, the generated mind maps are significantly improved in terms of logical consistency and semantic coherence, and are more in line with user needs and the core logic of the document. The generator explores various mind map structures, and the discriminator screens the optimal solutions, achieving a balance between the diversity and optimality of mind map optimization. The GAN model can adapt to different document contents and abstract styles and continuously improve the mind map optimization effect through learning.

[0241] Example: An initial mind map was generated for a product analysis report, including three major nodes "Market demand growth", "User feedback improvement", and "Product sales increase" and their relationships. The user hopes to optimize the mind map structure to make it more logical and highlight the key nodes. Processing steps:

[0242] The generator generates candidate mind maps, and the initial node weights are: , , . The generator generates candidate weights: Candidate 1: [0.75, 0.72, 0.88]; Candidate 2: [0.82, 0.65, 0.92]. Adjust the edge relationships: Candidate 1 deletes the edge "User feedback improvement → Market demand growth"; Candidate 2 strengthens the edge "Market demand growth → Product sales increase", and the weight is increased from 0.8 to 0.9.

[0243] The discriminator evaluates the candidate mind maps. Logical consistency evaluation: The logical consistency score of Candidate 1 ; The logical consistency score of Candidate 2 .

[0244] The semantic similarity score of Candidate 1 ; The semantic similarity score of Candidate 2 .

[0245] Screen the optimal mind map. The discriminator's comprehensive score (weighted sum of logical consistency and semantic coherence): The comprehensive score of Candidate 1 is 0.83; the comprehensive score of Candidate 2 is 0.91. Finally, Candidate 2 is selected as the optimized mind map.

[0246] The optimized mind map highlights the key logic of "Market demand growth" and "Product sales increase", deletes redundant edges, and the mind map has clearer levels and more coherent semantics, providing more efficient support for report analysis.

[0247] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.

Claims

1. An AI-driven multilingual intelligent document summarization and mind map generation system, characterized in that, Including: A multi-modal semantic feature generation module, which is used to receive multi-modal document data, extract and fuse the semantic features of text content, picture data, table data and video subtitle data, and generate unified multi-modal semantic feature data; In the multi-modal semantic feature generation module: The text processing unit decomposes semantic units by performing a word segmentation algorithm on the text content, establishes a syntactic tree structure using a syntactic parsing method, and extracts the text theme through a semantic analysis model; The picture processing unit extracts the regional boundaries and semantic categories in the picture through an image semantic segmentation method, and combines an object detection model to locate the coordinates and category information of each object in the picture; The table processing unit identifies the association relationships of column headers, row headers and cell contents through a table parsing algorithm, and generates table hierarchical structure data; The video subtitle processing unit transcribes the video audio content into text through speech recognition technology, and uses a time series analysis method to mark the timestamps and semantic levels of the subtitle text; A semantic graph construction module, based on the multi-modal semantic feature data, constructs dynamic optimization causal semantic graph data representing the causal relationships between semantic nodes, including causal relationship extraction and multi-modal relationship calibration, and generates calibrated causal relationship data; The semantic graph construction module extracts the causal relationships between semantic nodes through the following methods: Using a Bayesian network model to model the statistical correlation of multi-modal semantic feature data, and establishing potential causal relationships between nodes; Using a conditional independence test method to verify the effectiveness of the causal relationship, calculate the causal strength between nodes, and remove the edges with strength lower than a predetermined threshold; Recording the causal relationships that have passed the conditional independence test as causal relationship data for further construction of the semantic graph; The multi-modal relationship calibration in the semantic graph construction module includes: Using a cross-modal relationship graph network to optimize the semantic alignment degree between different modal nodes, and calibrating the feature representation of nodes by calculating the semantic vector similarity between modalities; Adjusting the low-confidence edges in the graph based on context consistency, increasing the weights of the high-confidence edges, decreasing the weights of the low-confidence edges, and removing the invalid edges; Constructing the calibrated nodes and edges into calibrated causal relationship data for further generation of the dynamic optimization causal semantic graph; An abstract generation module, based on the dynamic optimization causal semantic graph data, extracts the content of key semantic nodes, uses a dynamic semantic flow modeling method to generate a semantic flow path, and combines a multi-language generation method to generate multi-language abstract data; The abstract generation module generates a semantic flow path through a dynamic semantic flow modeling method, including the following steps: Initializing a semantic flow matrix, and calculating the semantic flow intensity according to the node weights in the causal semantic graph; Optimizing the directionality of the semantic flow path based on a reinforcement learning method, where the target reward function for the semantic flow includes path semantic coverage, consistency of the flow direction, and maximization of node weights; Extracting the content of key nodes along the semantic flow path to generate semantic flow path data for the generation of multi-language abstracts; The mind map generation and optimization module generates an initial mind map based on dynamically optimized causal semantic graph data and multilingual abstract data, and generates the finally optimized mind map data through hierarchical relationship optimization and user interaction adjustment; The mind map generation and optimization module verifies the consistency between the mind map and the abstract content through a natural language inference model, including the following steps: Use the inference model to verify whether the content of the mind map nodes matches the semantics of the abstract, where the inference model uses the node semantic accuracy rate and the abstract matching rate as evaluation criteria; Verify whether the edge relationships in the mind map conform to the logical chain in the abstract content, and mark the inconsistent nodes or edges; The mind map generation and optimization module optimizes the mind map structure through a generative adversarial network, including the following steps: The generator generates multiple candidate mind map structures, each of which contains reallocated node weights and edge relationships; The discriminator performs logical consistency evaluation and semantic coherence verification on the candidate mind maps, and selects the optimal mind map structure as the finally optimized mind map data.

2. The AI-driven multilingual intelligent document abstract and mind map generation system according to claim 1, wherein The multilingual abstract data generated by the abstract generation module is verified through a consistency verification method, including the following steps: Use a multilingual natural language inference model to verify the logical consistency of the abstract content in different language versions, where the verification model uses the semantic similarity and logical coherence between sentences as evaluation indicators; Mark the inconsistent abstract paragraphs, feedback the inconsistent parts to the abstract generation module, and regenerate the optimized multilingual abstract content.

3. The AI-driven multilingual intelligent document summarization and mind map generation system according to claim 1, wherein The mind map generation and optimization module generates the hierarchical structure of the mind map through the following method: Sort the semantic weights of the mind map nodes, and divide the priority levels of the nodes according to the sorting results; Associate the low-level nodes with the high-level nodes according to the causal relationship between the nodes to form a well-structured mind map; During the mind map generation process, ensure that the semantic association path between each node and its upper and lower nodes is complete and error-free.

4. The AI-driven multilingual intelligent document summarization and mind map generation system according to claim 1, characterized in that, The user interaction adjustment function of the mind map generation and optimization module includes: Provide a visual interface for the mind map, displaying the semantic content of the nodes and their causal association paths; Support users to adjust the position of the nodes by dragging, modify the node content or add annotations through input interaction, and correct the weights and edge relationships between the adjusted nodes through real-time feedback; Automatically update the user-modified mind map structure to the user-adjusted and optimized mind map data.

Citation Information

Patent Citations

  • Document generation method, system and electronic device

    CN118551740B

  • Abstract generation method, device and equipment

    CN115080727A

  • Document generation method and system and electronic equipment

    CN118551740A

  • Multi-modal document structured processing and knowledge extraction method based on large language model

    CN119227794A

Cited By

  • Intelligent office plug-in system and office method based on privately deployed large language model

    CN121764552A