A method and system for generating a trusted traceability presentation based on document content

Through a trusted traceability presentation generation method based on document content, using a hierarchical semantic segmentation model and a visual mapping engine, presentations are intelligently constructed, solving the problem of low efficiency in traditional presentation production and improving work efficiency and information processing capabilities.

CN120493894BActive Publication Date: 2025-09-16珠海必优科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510990988.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-18
Publication Date
2025-09-16
Estimated Expiration
2045-07-18

AI Technical Summary

Technical Problem

Traditional presentation production relies on manual operations, resulting in a lot of time and energy being consumed in non-creative work. This is especially inefficient in high-frequency usage scenarios and makes it impossible to focus on core innovative work.

Method used

A trusted traceability presentation generation method based on document content is adopted. The hierarchical semantic segmentation model is used to perform semantic recognition and information slicing on the processed documents, a content structure tree is constructed, and a visual mapping engine is used to generate presentations to achieve intelligent construction.

Benefits of technology

It improves the efficiency of presentation generation, optimizes workflows, and assists personnel in related positions to improve work efficiency, especially in scenarios such as financial analysis and academic reporting, significantly reducing manual intervention and improving information processing and extraction efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120493894B_ABST
    Figure CN120493894B_ABST
Patent Text Reader

Abstract

The embodiment of the present application provides a method and system for generating a trustworthy traceable presentation based on document content. The method includes: obtaining a document to be processed input by a user; performing semantic recognition on the content information of the document to be processed through a hierarchical semantic segmentation model, and slicing the information of the document to be processed according to the semantic recognition result to obtain a content structure tree; the content structure tree contains document content slices that match key content information, logical relationships between each document content slice, and traceable information corresponding to each document content slice; using a structure tree-based visual mapping engine model, the document content slices and traceable information in the content structure tree are mapped to corresponding presentation content according to the logical relationship to obtain a target presentation, and the target presentation is displayed to the user. The present application can improve the efficiency of presentation generation, optimize the workflow, and assist in improving the work efficiency of personnel in related positions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of artificial intelligence. More specifically, the embodiments of the present application relate to a method and system for generating a trusted traceable presentation based on document content. Background Art

[0002] At present, the production of traditional presentations (i.e. PPT) is highly dependent on manual operations, and its core pain point is the excessive consumption of professional human resources on non-creative labor.

[0003] In the past, when creating presentations, professionals spent a great deal of time and energy on mechanical tasks like copying and pasting text, adjusting layout spacing, and drawing basic charts. This inefficiency is particularly prominent in high-frequency PPT usage scenarios such as financial analysis and academic presentations.

[0004] For example, in the financial sector, investment bank analysts are required to produce an average of three project proposals of more than 50 pages per week, with nearly half of their work time spent on typesetting. University instructors spend half of their total lesson preparation time converting course notes into courseware. This systemic loss of efficiency not only results in significant economic losses but also distracts key personnel from core innovation, significantly reducing work efficiency.

[0005] In summary, it is urgent to design a new presentation production method to improve the efficiency of presentation production and help improve work efficiency. Summary of the Invention

[0006] In this context, the embodiments of the present application hope to provide a method and system for generating a trusted traceable presentation based on document content, which can realize the intelligent construction of presentations, improve the efficiency of presentation generation, optimize the workflow, and assist in improving the work efficiency of personnel in related positions.

[0007] In a first aspect of the embodiments of the present application, a method for generating a trusted traceable presentation based on document content is provided, comprising:

[0008] Obtaining a document to be processed input by a user; wherein the content information of the document to be processed includes at least: text, image, and / or rich text information mixed with text and image;

[0009] Using a hierarchical semantic segmentation model, semantic recognition is performed on the content information of the document to be processed. Based on the semantic recognition results, the document to be processed is sliced ​​to obtain a content structure tree for the document to be processed. The content structure tree includes: document content slices that match key content information, logical relationships between each document content slice, and traceability information corresponding to each document content slice. The slicing granularity of each document content slice is dynamically configured based on the semantic recognition results matched in the document to be processed.

[0010] Using a structure tree-based visual mapping engine model, the document content slices and traceable information in the content structure tree are mapped to corresponding presentation content according to logical relationships to obtain a target presentation;

[0011] Show the target presentation to the user, and help the user understand the key content information in the document to be processed through the target presentation.

[0012] In a second aspect of the embodiments of the present application, a system for generating a trusted traceable presentation based on document content is provided, comprising:

[0013] An acquisition module is used to acquire a document to be processed input by a user; wherein the content information of the document to be processed includes at least: text, image, and / or rich text information mixed with text and image;

[0014] A construction module is used to perform semantic recognition on the content information of the document to be processed using a hierarchical semantic segmentation model, and to slice the document to be processed based on the semantic recognition results to obtain a content structure tree for the document to be processed. The content structure tree includes: document content slices matching key content information, logical relationships between each document content slice, and traceability information corresponding to each document content slice. The slicing granularity of each document content slice is dynamically configured based on the semantic recognition results matching the document to be processed.

[0015] The display module is used to adopt a structure tree-based visual mapping engine model to slice the document content and traceable information in the content structure tree and map them into corresponding presentation content according to logical relationships to obtain a target presentation; the target presentation is displayed to the user to assist the user in understanding the key content information in the document to be processed through the target presentation.

[0016] In a third aspect of the implementation of the present application, a terminal device is provided, comprising: at least one processor, a memory, and an input-output unit; wherein the memory is used to store a computer program, and the processor is used to call the computer program stored in the memory to execute the method for generating a trusted traceability presentation based on document content as described in any one of the first aspects.

[0017] In a fourth aspect of the embodiments of the present application, a computer-readable storage medium is provided, which includes instructions that, when executed on a computer, enable the computer to execute the method for generating a trusted traceability presentation based on document content as described in any one of the first aspects.

[0018] In a fifth aspect of the embodiments of the present application, a computer program product is provided, comprising a computer program, which, when executed by a processor, implements the method for generating a trusted traceability presentation based on document content as described in any one of the first aspects.

[0019] According to the document content-based trusted traceability presentation generation method and system of the embodiment of the present application, it is possible to obtain a document to be processed input by a user; wherein the content information of the document to be processed includes at least: text, image, and / or rich text information mixed with text and image. Furthermore, through a hierarchical semantic segmentation model, semantic recognition is performed on the content information of the document to be processed, and information slicing is performed on the document to be processed based on the semantic recognition results to obtain a content structure tree of the document to be processed; the content structure tree includes: document content slices matching key content information, logical relationships between each document content slice, and traceability information corresponding to each document content slice; the slice granularity of each document content slice is dynamically configured based on the semantic recognition results matched in the document to be processed. Then, a structure tree-based visual mapping engine model is used to map the document content slices and traceability information in the content structure tree to corresponding presentation content according to the logical relationship to obtain a target presentation. Finally, the target presentation is displayed to the user to assist the user in understanding the key content information in the document to be processed through the target presentation. According to the implementation mode of the present application, it is possible to realize the intelligent construction of presentations, improve the efficiency of presentation generation, optimize the workflow, and thus help improve the work efficiency of personnel in related positions. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 A flowchart of a method for generating a trusted traceability presentation based on document content provided in one embodiment of the present application;

[0021] Figure 2 A schematic diagram of the structure of a document content-based trusted traceability presentation generation system provided in one embodiment of the present application;

[0022] Figure 3 The structural diagram of a medium in an embodiment of the present application is schematically shown. DETAILED DESCRIPTION

[0023] Reference below Figure 1 , Figure 1 This is a flowchart of a method for generating a presentation with trusted provenance based on document content provided in an embodiment of the present application. It should be noted that the implementation of the present application can be applied to any applicable presentation production scenario.

[0024] Figure 1The process of the method for generating a trusted traceability presentation based on document content provided in an embodiment of the present application includes:

[0025] Step S101: Obtain the document to be processed input by the user.

[0026] In the embodiment of the present application, the content information types of the document to be processed are rich and varied, including different forms such as pure text, pure image, and rich text mixed with text and image. The content information of the document to be processed includes at least: text, image, and / or rich text mixed with text and image.

[0027] For example, plain text documents include formats like Word, PDF, TXT, email, and web page text. They are commonly found in academic papers, research reports, meeting minutes, contracts, press releases, and book chapters. They are primarily text-based, requiring the extraction of paragraph themes, logical hierarchy, and key data. Image-based documents include JPG, PNG, GIF images, scanned copies, and chart images. They are used in scenarios like poster designs, data visualization charts, hand-drawn drawings, certificate images, and product manual illustrations. They require OCR technology to extract text or computer vision to analyze visual elements. Rich text documents include PPT, Excel reports with charts, Markdown documents, HTML web pages with text and graphics, and illustrated e-books. They are often found in multimedia reports, marketing plans with illustrations and data tables, teaching materials with formulas and diagrams, and news and information web pages with mixed text and graphics. They require the simultaneous processing of text, images, tables, formulas, and other elements, as well as analysis of layout logic.

[0028] In terms of intelligently importing and acquiring documents to be processed, in step S101, documents to be processed are acquired from multiple channels. For example, local documents can be uploaded through the file selection window and the format can be automatically identified and parsed. Cloud files can also be acquired synchronously through mainstream cloud storage platforms such as network disks and shared cloud documents. Compliant content can also be crawled from the web page URL entered by the user and irrelevant information can be filtered out. Alternatively, a text box or editor can be provided for users to input and collect fragmented document content in real time.

[0029] During the format analysis and content extraction phase, natural language processing techniques are used to extract titles, keywords, data indicators, and identify logical structures for text documents. For image documents, OCR technology is used to convert text-containing images while preserving the layout. Computer vision analysis is used to convert charts and images into structured data. For rich text documents, markup language or typesetting formats are analyzed to separate elements and record the original typesetting logic.

[0030] In step S101, the obtained documents to be processed can further be intelligently pre-processed and error-checked to automatically detect duplicate content, redundant images, invalid data and prompt for processing, process old or non-standard format files through the format conversion engine, and connect to the rights management system to verify document access rights.

[0031] Furthermore, step S101 also introduces interactive import assistance. After the user uploads the file, the type is automatically determined and the parsing strategy is recommended. A visual interface is provided for the user to manually mark and correct key content. It supports batch import of documents and grouping by category and indexing.

[0032] Step S101 enables unified cross-format processing, converting documents of varying types and formats into structured content, breaking down format barriers. This significantly reduces manual intervention and the time users spend manually organizing documents, making it particularly suitable for processing large amounts of data or urgent tasks. The system boasts high compatibility and stability, supporting mainstream document formats as well as specialized file formats like encrypted PDFs and scanned documents, while ensuring accurate content capture through error checking.

[0033] Step S102 : semantically identifying the content information of the document to be processed by using a hierarchical semantic segmentation model, and slicing the information of the document to be processed according to the semantic identification result to obtain a content structure tree of the document to be processed.

[0034] It's understandable that the hierarchical semantic segmentation model is essentially a multi-granularity document understanding engine that mimics human reading logic. Its core principle is to achieve deep semantic parsing of documents through a layered processing pipeline. In one alternative example, a bidirectional long short-term memory network and conditional random fields are first used for primary semantic annotation to identify key semantic units in the document, such as core arguments and supporting data, similar to highlighting important content when reading. Next, a logical topology is constructed based on dependency parsing and graph neural networks to establish a network of logical relationships between entities, clarifying the supporting, causal, or progressive relationships between elements such as arguments and data, and problems and solutions. Finally, the slicing granularity is dynamically adjusted based on semantic density, using fine-grained single-sentence slicing in high-value information areas (such as the conclusion paragraph) and coarse-grained merging in low-information background sections, achieving intelligent compression and hierarchical understanding of document information. For other implementation methods, see the examples below, which are not elaborated here.

[0035] The hierarchical semantic segmentation model breaks through the semantic fragmentation problem caused by traditional word segmentation or fixed paragraph segmentation, integrates semantic value assessment with structural topology discovery, and makes document parsing closer to human cognitive logic. For example, in scientific paper testing, it can greatly improve the retention rate of key argument relationships and significantly enhance semantic integrity.

[0036] In the embodiment of the present application, the content structure tree includes: document content slices matching the key content information, the logical relationship between each document content slice, and the traceability information corresponding to each document content slice.

[0037] The content structure tree is a skeletal representation of document semantics, explicitly presenting the document's cognitive structure in a tree-like structure. Document content slices are dynamically segmented semantic units within the document being processed. These are cognitive modules aggregated based on semantic integrity, rather than mechanical paragraph segmentation. For example, a title and its supporting arguments form a slice, ensuring the logical coherence of the information within the slice. For example, a slice consisting of the title "Analysis of the Reasons for Q3 Market Growth" and its three supporting arguments.

[0038] Logical relationships are used to describe cognitive association patterns between slices, including parallel, progressive, and general-specific structures. For example, the point-by-point statement of product advantages is a parallel structure, the deduction from problem discovery to solution is a progressive structure, and the core argument leading the arguments is a general-specific structure. These relationships clearly demonstrate the hierarchy and context of the document content.

[0039] Traceable information is used to bind each document content slice to its digital coordinates in the source document (such as "P12-L5~L8"), forming a traceable pointer network. Users can jump to the corresponding position in the original text by clicking the slice title in the PPT, solving the black box problem of traditional generation tools, improving efficiency and reducing communication losses in scenarios such as audit compliance and collaborative editing.

[0040] The slicing granularity is dynamically configured based on semantic density. Fine-grained slicing is used to focus on key information areas, and coarse-grained merging and compressing information is used in secondary content areas. For example, the core data table in the financial annual report is sliced ​​independently, and the auxiliary explanatory text is merged as a whole, which greatly improves the information compression rate. In medical industry applications, the slicing accuracy of key indicators in pathology reports can even be improved to the data point level.

[0041] It should be noted that in the embodiment of the present application, the slice granularity is dynamic. Further optionally, the slice granularity of each document content slice is obtained by dynamic configuration based on the semantic recognition results matched in the document to be processed. Simply put, when the model detects a high-value area (such as a conclusion paragraph), it automatically switches to fine-grained mode (single-sentence slice). In low-information-density areas (such as background introductions), coarse-grained merging is used to achieve intelligent information compression. It can be seen that the slice granularity is not a simple division of paragraph length, but is obtained by dynamic division using an intelligent compression algorithm based on semantic density calculation. In financial reports, the dynamic division of slice granularity can condense a whole page of financial report data into three key indicator slices or one icon slice based on the importance of the content information or the contextual logical relationship. In the scenario where lawyers handle legal documents, the legal terms can be split into two liability limitation slices and associated with the evidence slices in the case file.

[0042] In summary, by constructing a content structure tree through a hierarchical semantic segmentation model, semantic value assessment and structural topology discovery are integrated to avoid semantic fragmentation caused by traditional word segmentation or fixed paragraph segmentation, and to achieve a leap in document processing from the character operation level to the cognitive understanding level. This model can not only accurately identify key semantic units and logical relationships in documents, avoiding the loss of logical chains caused by traditional manual combing, but also intelligently compress information through a dynamic slicing granularity mechanism to solve the problem of over-segmentation or under-segmentation caused by fixed granularity slicing. At the same time, traceable information provides a verifiable and traceable basis for document content. These features enable the system to effectively extract key information when processing massive documents. For example, when a pharmaceutical company parses 100,000 pages of clinical reports, it can automatically extract neglected key drug reaction data as independent slices, thereby providing an important reference for research and development, fully demonstrating the significant effect of this technology in improving information processing efficiency and value mining.

[0043] As an optional embodiment, in step S102, semantic recognition is performed on the content information of the document to be processed using a hierarchical semantic segmentation model, and information slicing is performed on the document to be processed based on the semantic recognition result to obtain a content structure tree of the document to be processed, including:

[0044] Through the multimodal parsing network, the document to be processed is parsed according to the data type to obtain a graph structure network; wherein the graph structure network includes: the graph nodes matching each content information, the association relationship between each graph node, and the page position of each graph node in the document to be processed;

[0045] Through the semantic recognition layer, a Bi-LSTM-CRF model is used to identify the entities contained in the graph structure network and the logical relationships between the entities, and the semantic roles corresponding to each entity are marked; the logical relationships between the entities include at least: parallel structure, progressive structure, general-specific structure, forward pyramid, and inverted pyramid;

[0046] The dynamic slicing layer obtains the semantic density of each local area based on the logical relationship between each entity and the semantic role. The slicing granularity corresponding to each local area is adjusted in combination with the semantic density according to the semantic elastic partitioning mechanism. The document content slice corresponding to each local area is constructed based on the slicing granularity to obtain the content structure tree.

[0047] Through the traceability index layer, combined with the page position of each entity in the document to be processed, a traceability pointer for each document content slice is generated, and each traceability pointer is bound to the metadata of the corresponding document content slice in the content structure tree to obtain a traceable pointer network that matches the content structure tree; the traceability pointer includes at least: the page number, paragraph offset, single sentence offset, and / or page layout offset associated with each document content slice in the document to be processed.

[0048] Each local region contains at least one or more adjacent entities. Furthermore, the semantic density of the local region is determined based on the semantic role importance of all entities within the local region. The greater the semantic density of the local region, the greater the semantic role importance of all entities within the local region, and the smaller the slice granularity of the local region.

[0049] Specifically, the hierarchical semantic segmentation model builds an intelligent processing pipeline from original documents to structured knowledge bodies.

[0050] First, a multimodal parsing network is used to convert the document into a graph-structured network. During this stage, elements such as text, images, and tables are deconstructed into graph nodes with spatial coordinates. The relationships between nodes map to the positional logic of the document's original layout (such as the proximity of charts and captions). For example, when processing a company's annual report, elements such as the financial report title, data tables, and market trend charts are encoded as nodes, while the visual flow from "title → data table → analysis conclusion" is converted into connecting edges. This layer of processing transcends the limitations of traditional OCR, which only extracts characters, and establishes a topological mapping of content elements in the two-dimensional page space.

[0051] At the semantic recognition layer, the Bi-LSTM-CRF model performs deep semantic annotation on the graph network, acting like a neural network injecting semantic units into the skeleton. It not only identifies the entity type of each node (such as "technology patent data point" or "market risk warning"), but also captures the logical connections between entities. For example, in patent documents, the Bi-LSTM-CRF model can accurately annotate the document's pyramidal structure: "Innovation A → Patent Protection Scope B → Commercialization Path C," or parse legal clauses into a general-specific network: "General Clause Limitations → Detailed Explanations → Exceptions." This understanding doesn't rely on a fixed rule base, but rather on contextual awareness developed through deep learning from hundreds of thousands of industry documents.

[0052] The dynamic slicing layer can implement flexible segmentation based on semantic density. The calculation of semantic density in local areas is essentially a quantitative assessment of the value of the content. For example, even if the "Experimental Method" and "Conclusion" sections of a technical white paper have the same physical length, the latter will receive a higher density rating because it contains core results. For example, when a key indicator area in a medical report is detected (such as "tumor regression rate = 78%"), it immediately switches to single-sentence granularity segmentation. Auxiliary content such as background literature reviews are compressed as a whole section. This dynamic granularity control ensures that high-value information is preserved intact during the knowledge distillation process and redundant information is intelligently aggregated.

[0053] The traceability index layer is used to establish tamper-proof links between discrete knowledge units and the document being processed. The traceability pointer bound to each slice contains multi-level coordinates (such as page number, paragraph offset, and inline positioning), forming a digital anchor with character-level accuracy. For example, if a user questions the statistical caliber of a conclusion in a PPT, clicking the traceability indicator immediately retrieves the corresponding paragraph in the source document. This design ensures the transparency of the automated system's output and is particularly suitable for forensic audits, medical reporting, and specialized data analysis scenarios where information integrity verification is required.

[0054] In summary, at the level of deep semantic understanding, the hierarchical semantic segmentation model has improved its accuracy in identifying logical relationships, particularly for implicit logic such as summaries and inverted pyramid relationships. Dynamic information compression capabilities enable dimensionality reduction-level optimization. For example, in a test of information extraction from an academic paper, the system converted a 47-page document into a 12-page condensed PPT, preserving the core findings while eliminating redundant descriptions. This further enhances the adaptive granularity of the hierarchical semantic segmentation model. For example, when processing clinical reports, key indicators (such as "5-year survival rate") are highlighted in separate slices, while demographic data on thousands of patients is aggregated into summary charts. This capability reduces the reading time of a 50-page industry analysis report from 130 minutes to 15 minutes, accelerating decision-making by more than eight times. Furthermore, the traceable knowledge system helps preserve information integrity, preventing information loss after extraction. For example, after a system upgrade, the average regulatory review time for a securities company's financial report PPTs decreased, and auditors were able to directly jump to the original data page for cross-verification using the slice traceability chain. Furthermore, in the enterprise knowledge graph constructed by slicing metadata, each node has its own content value weight and coordinate mapping, allowing new employees to accurately locate the experience data of similar cases in history, thereby greatly improving the efficiency of knowledge graph acquisition.

[0055] Exemplarily, in an optional embodiment of the above model, it is assumed that the multimodal parsing network includes at least a preprocessing layer, a text parsing layer, an image parsing layer, and a multimodal fusion layer. Based on this structure, the document to be processed is parsed according to the data type through the multimodal parsing network to obtain a graph structure network, including: through the preprocessing layer, the document to be processed is subjected to paging transcoding to generate a sequence of documents to be processed; through the text parsing layer, a pre-deployed Transformer-based pre-trained model is used to parse the document layout and semantic role features of text elements in the sequence of documents to be processed to obtain text feature information; the document layout includes at least: titles, paragraphs, lists; the text elements include sentences and / or word segments; the semantic role features include at least: key sentences, keywords and their corresponding statistical features; through Through the image parsing layer, the pre-deployed CLIP-ViT model is used to extract the global image features of the document sequence to be processed, and the text and / or formulas contained in the charts and images in the document sequence to be processed are detected and located and annotated to obtain the local image features of the document sequence to be processed; through the multimodal fusion layer, the adaptive attention mechanism is used to cross-modally associate text feature information, global image features, and local image features; with text elements and image elements as nodes, the connection edges are constructed based on the association results between text elements and image elements, and spatial relationship modeling is performed to obtain the graph structure network corresponding to the document sequence to be processed.

[0056] It is understandable that the multimodal parsing network achieves refined parsing and cross-modal association of multiple types of data such as text and images in documents through a layered and collaborative structural design. Its principles and functions are closely centered around the structured understanding of document content. As a basic link, the preprocessing layer performs paging transcoding operations on the documents to be processed, uniformly converting documents of various formats into standardized document sequences that can be processed by the system, laying the foundation for subsequent parsing. This process solves the compatibility problems brought about by the diversity of document formats, ensuring that elements such as text and images can enter the subsequent processing flow in an orderly sequence, providing guarantees for the stability and consistency of the overall parsing.

[0057] The text parsing layer uses a pre-trained model based on Transformer to deeply analyze the text content in document sequences. This model can identify structures such as titles, paragraphs, and lists in the document layout, clearly dividing the content hierarchy. At the same time, it processes text elements into sentences and words, and annotates key sentences, keywords, and their statistical features to extract the core semantic information in the text. For example, when parsing academic papers, it can accurately locate key parts such as the abstract and conclusion, extract arguments and data support, and make the logical context and key information of the text content explicit, providing a rich text feature foundation for subsequent multimodal fusion.

[0058] The image parsing layer uses the CLIP-ViT model to process image elements in documents, with both global and local analysis capabilities. On the one hand, it extracts global image features to grasp the theme and style of the image as a whole, such as determining whether a chart is a line chart or a bar chart. On the other hand, it detects and locates the text and formulas in charts and images to generate local image features, just like labeling each key element in the image to clarify its position and content. This dual processing mechanism enables the system to fully understand image information, knowing the overall meaning of the image while identifying detailed content. For example, when parsing a report containing data charts, it can accurately extract data points and text descriptions in the charts, providing accurate image information support for cross-modal associations.

[0059] The multimodal fusion layer uses an adaptive attention mechanism to cross-modally associate text feature information, global image features, and local image features, constructing a graph-structured network with text and image elements as nodes and the association results as edges. This process is similar to building a bridge connecting text and images, enabling information from different modalities to communicate with each other. The adaptive attention mechanism acts as an intelligent coordinator, dynamically allocating attention based on the importance of information, strengthening the association of key information and weakening the interference of secondary information. Through spatial relationship modeling, the system can clearly identify the positional associations and logical support between text and image elements. For example, it can identify whether a certain text paragraph explains a certain chart, thereby integrating scattered multimodal information into a structured graph network. This network structure not only clearly demonstrates the inherent connections between document content but also lays a solid foundation for the subsequent generation of logically clear and coherent presentations.

[0060] The multimodal parsing network thus improves the depth and accuracy of document parsing through layered processing and cross-modal fusion. The preprocessing layer achieves format unification and eliminates format barriers. The text parsing layer and the image parsing layer extract key features from the text and image dimensions respectively, avoiding the one-sidedness of single-modal parsing. The graph structure network constructed by the multimodal fusion layer fully preserves the semantic associations and spatial relationships of multiple types of elements in the document, upgrading the system's understanding of the document from scattered information fragments to an organic holistic cognition. This parsing method is particularly advantageous when processing rich text that is a mixture of text and images. For example, when parsing marketing plans, it can accurately identify the logical relationship between text descriptions, accompanying images, and data tables. The generated graph structure network can fully reflect the core points and supporting information of the plan, providing a rich and accurate content foundation for the subsequent intelligent generation of presentations, effectively improving the efficiency and quality of document processing, and promoting information processing from simple content extraction to deep semantic understanding.

[0061] Furthermore, the core function of the multimodal fusion layer is to break down the semantic barriers between different data types, such as text and images, and achieve deep association and integration of cross-modal information. Its working principle is based on an adaptive attention mechanism, which can dynamically allocate computing resources according to the semantic importance of the input content, focusing on capturing the potential associations between data of different modalities. Specifically, it first numerically aligns the text feature information output by the text parsing layer (such as the semantic role features of keywords and key sentences) with the global image features (such as the overall layout style of the chart) and local image features (such as the specific text and formula positions in the chart) extracted by the image parsing layer, and converts them into feature vectors of unified dimensions. Subsequently, the similarity weights between different modal features are calculated through the attention mechanism, for example, to determine whether a text paragraph has an explanatory relationship with a specific data area in a chart, or whether the diagram in the image is used to assist in explaining the core idea of ​​a paragraph of text.

[0062] Specifically, text feature information (such as semantic role features and keyword / sentence vectors extracted by Transformer) and image features (such as global image semantics and local chart element features extracted by CLIP-ViT) are transformed into a feature vector space with a unified dimension, ensuring comparability between data from different modalities. For example, the keyword vector for "market growth rate" in the text and the feature vector for the growth trend of a bar chart in the image are mapped into the same semantic space. Next, the similarity between the two types of features is calculated using a cross-modal interaction module. This process is similar to "feature matching," where each semantic unit in the text (such as a sentence) is paired with each visual element in the image (such as a chart area or the main image body) to calculate the semantic relevance between the two. For example, the sentence vector describing "sales increase" in the text is dot-producted with the features of the line chart area showing sales growth in the image to obtain a preliminary relevance score. A dynamic weight adjustment mechanism is then introduced to optimize the initial relevance score. The mechanism analyzes the importance of features within each modality (such as the weight of key sentences in text and the attention weight of high-information-density regions in images) and, combined with the global cross-modal correlation distribution, generates adaptive weights for each pair of feature matches. For example, if a key sentence in text also corresponds to a key graphic region in an image, the weight of the association between the two will be significantly strengthened, while vice versa, the weight of less important associations will be weakened. Finally, through normalization, the similarity weights of all cross-modal feature pairs are converted into a probability distribution, which is used to guide the multimodal feature fusion process. For example, when generating edges in a graph network, high-weighted feature pairs are prioritized for connection, forming closer semantically related edges, while low-weighted associations may be weakened or ignored, achieving the effect of "strengthening key associations and suppressing noisy associations." The essence of this process is to enable the model to dynamically focus on key associations in multimodal data, avoiding the semantic confusion caused by simply superimposing features from different modalities in traditional fusion methods, ultimately improving the accuracy and logical coherence of cross-modal information integration.

[0063] In terms of execution function, the multimodal fusion layer mainly completes two tasks: one is cross-modal association modeling, that is, identifying the semantic correspondence between text elements and image elements, such as establishing a connection between the "sales growth curve" text and the line chart area in the chart, and clarifying the "keyword-image area" mapping; the second is spatial relationship modeling, using text elements and image elements as nodes and cross-modal association results as connecting edges to construct a graph structure network containing spatial location information, such as marking the contextual dependency when a certain text is located above a chart. This modeling method not only retains the visual layout characteristics of the document, but also captures the logical dependencies between modalities, such as the order in which text interprets images or the evidential support relationship between images and text.

[0064] As a result, the multimodal fusion layer significantly improves the comprehensiveness and accuracy of document parsing. On the one hand, it solves the problem that traditional single-modal parsing cannot handle text-image linkage information. For example, it can accurately identify complex associations in the combination of "text description + table data + schematic diagram", avoiding missing key information implicit in cross-modal interactions. On the other hand, by dynamically focusing on high-value associations through the adaptive attention mechanism, the accuracy of semantic matching between images and text can be improved (compared to traditional rule matching methods), especially in document scenarios containing a large number of charts, such as scientific reports and financial analysis, which can effectively reduce information misunderstandings or logical faults caused by modal fragmentation, and provide a more accurate semantic basis for subsequent content structuring and visualization conversion.

[0065] In an optional embodiment of the above model, a Bi-LSTM-CRF model is used in the semantic recognition layer to identify the entities contained in the graph network and the logical relationships between them, and to label the semantic roles corresponding to each entity. The logical relationships between entities include at least parallel structure, progressive structure, general-specific structure, forward pyramid structure, and inverted pyramid structure.

[0066] Specifically, the semantic recognition layer first performs entity recognition, precisely locating entities with independent semantics from the nodes (text and image elements) of the graph network. This includes things like market growth rates and user profile data in a report. Semantic role labeling is then performed, assigning specific functional labels to each entity, such as core arguments, data support, and case studies, to clarify its role within the document. Finally, logical relationship extraction is performed. By analyzing the semantic association patterns between entities, logical structures such as parallel, progressive, and general-to-specific structures can be identified. For example, "Product Advantages One, Two, and Three" is a parallel structure, while "Problem Raised → Solution → Effect Verification" is a progressive structure. This process is like explicitly breaking down the "semantic context" of a document, transforming implicit logical relationships into computable and actionable structured information.

[0067] The core principle of the semantic recognition layer is to simulate the human understanding of language logic. By combining a bidirectional long short-term memory (Bi-LSTM) network with a conditional random field (CRF), it achieves structured analysis of entities and their logical relationships within a document. As a recursive neural network, Bi-LSTM simultaneously captures the semantic dependencies between contexts within a text, similar to how humans naturally connect context to understand word meaning when reading. It processes sequence data in both forward and reverse directions, extracting contextual semantic features for each entity. CRF further optimizes the annotation results based on this foundation. By analyzing the label transition probabilities between adjacent entities, it ensures logical consistency across the entire sequence, preventing errors in labeling a single entity from propagating to the global structure.

[0068] Furthermore, a topological reasoning engine is used to identify logical relationships between entities. This engine works in conjunction with a graph attention network (GAT) via a dependency syntax tree. When parsing legal contracts, dependency analysis identifies the explicit grammatical structure of "Party A's obligations → Party B's rights," while the graph attention network discovers the implicit logical correspondence between the "Breach of Liability Exemption" clause and the "Force Majeure Description" (a cross-page margin is still established even if the distance is greater than 5 pages). The system pre-defines five logical relationship templates: parallel structure (product features listed in parallel), progressive structure (market demand → capacity expansion → profit improvement), general-specific structure (patent claims and sub-items), pyramid-like structure (convergent derivation from phenomenon → essence → solution), and inverted pyramid (conclusion-first → data-supported news structure). This multi-modal coverage enables lossless logical mapping of complex documents.

[0069] For example, the semantic recognition layer breaks through the limitations of traditional NER, which only recognizes general categories, and realizes industry-customized role labeling. In financial scenarios, entities are classified into professional semantic roles such as "risk warning indicators," "cash flow core data," and "policy favorable descriptions." Taking the prospectus as an example, "inventory turnover rate of 1.8 times / year" is accurately labeled as a "key indicator of operating efficiency" rather than simply classified as a "numerical value." This depth of labeling comes from pre-training in vertical fields - the model learned from 100,000 financial documents that "debt ratio > 70%" in the "risk factors" section should be marked with a red warning label (priority > normal data).

[0070] The semantic recognition layer can also enable the construction of cross-modal logical relationships. It achieves semantic collaboration between images and text within a graph-structured network. For example, when the "User Satisfaction Survey Conclusion" text node and the adjacent line graph are detected, a "Data Visualization Support" logical edge is automatically constructed. A more groundbreaking innovation lies in solving typical document problems. For example, scattered elements are linked together to form a complete chain of evidence. In a scientific paper, the experimental design description (P5) in the "Methods" section, the data charts (P8) in the "Results" section, and the comparative analysis (P10) in the "Discussion" section are identified as a progressive relationship trio, forming a golden logical chain throughout the entire text.

[0071] The technical effect of this layer is reflected in the improvement of the depth and accuracy of the semantic understanding of documents. In traditional methods, document structure analysis often relies on fixed rules or simple statistics, which makes it difficult to capture complex logic (such as causal relationships across paragraphs). The semantic recognition layer can discover implicit associations that are easily overlooked by humans by dynamically learning contextual features, such as identifying the indirect causal relationship between "policy adjustments" and "changes in consumption trends." Experimental data show that in the processing of scientific and technological literature, this layer improves the accuracy of identifying key logical relationships, especially in long document scenarios, and can effectively avoid the problem of broken logical chains. In addition, its annotated semantic roles and logical structures provide a standardized semantic framework for subsequent content generation, knowledge graph construction and other tasks. For example, when automatically generating PPT, it can quickly organize the slide hierarchy according to the "general-specific structure", or prioritize the extraction of key charts based on the "data support" role, significantly improving the intelligence level of document processing and business application efficiency.

[0072] In an optional embodiment of the above model, the semantic density of each local area is obtained based on the logical relationship between each entity and the semantic role through the dynamic slicing layer, including:

[0073] Obtain the semantic role corresponding to each entity; dynamically configure the corresponding semantic role importance coefficient for different semantic roles based on the logical relationship between each entity in the graph structure network, the document topic, and / or the document type; adopt an adaptive sliding window algorithm, take each entity as a node, establish an initial sliding window with a preset size, traverse the graph structure network, and calculate the semantic density corresponding to the local area in the sliding window in real time based on the semantic role importance coefficient; wherein the real-time size of the sliding window in the process of traversing the graph structure network is adaptively adjusted based on the calculated semantic density and the document structure relationship between adjacent entities.

[0074] In addition to the configuration method in the above embodiment, the present application can further configure corresponding semantic role importance coefficients for the semantic roles of different entities based on the semantic similarity between the entity and the core topic.

[0075] Furthermore, according to the semantic elastic division mechanism, the slice granularity corresponding to each local area is adjusted in combination with the semantic density, and the document content slice corresponding to each local area is constructed based on the slice granularity to obtain the content structure tree, including: establishing an inverse association rule between semantic density and slice granularity and a boundary optimization strategy, wherein the higher the semantic density in the local area, the smaller the slice granularity; obtaining the semantic density of each local area, and determining the initial slice granularity of each local area based on the inverse association rule; adopting the boundary optimization strategy, combining the initial slice granularity to check whether the logical relationship between the entities on both sides of the slice boundary is an inseparable relationship; if the entities on both sides of the slice boundary are inseparable, the elastic retraction of the initial slice granularity is triggered until the entities on both sides of the slice boundary are in a separable relationship and the coherence constraint is met, thereby obtaining the target slice granularity of each local area; based on the target slice granularity, the entities contained in each local area are restructured according to their respective semantic roles to obtain the document content slices of each local area; wherein, the entities with a semantic density greater than a set threshold are set as single sentence slices, and the entities with a semantic density less than the set threshold are constructed into paragraph slices or chapter slices through dynamic aggregation.

[0076] For example, the semantic roles corresponding to each entity are one of the following: title, argument, key chapter, key supporting data, and auxiliary content. The slicing granularity corresponding to title, argument, key chapter, and key supporting data is single sentence slicing, while the slicing granularity corresponding to auxiliary content is paragraph slicing or chapter slicing.

[0077] Specifically, the core principle of the dynamic slicing layer is to dynamically and intelligently slice entities in a graph network based on the semantic structure and logical relationships of document content, achieving intelligent matching of the document's value density with presentation granularity. First, the dynamic slicing layer analyzes the semantic roles of entities (such as keywords and key sentences) and the logical relationships between entities (such as parallelism and progression). Combined with the document's topic and type, the dynamic slicing layer assigns importance coefficients to different semantic roles, thereby quantifying the weight of different content in the semantic representation. Next, an adaptive sliding window algorithm is used to traverse the graph network, constructing an initial window with entities as nodes. The semantic density within the window is calculated in real time based on the semantic role importance coefficients. During this process, the sliding window size is dynamically adjusted based on the semantic density of the current area and the structural relationships between adjacent entities. For example, the window is narrowed in semantically dense areas to focus on details, while it is expanded in semantically sparse areas to integrate information, thereby achieving a refined scan of the document content.

[0078] For example, in the adaptive sliding window algorithm, assuming that the sliding window is set to , indicating the A sliding window of time steps, including To Entity The sliding window is the basic unit of semantic density calculation and slicing, covering all the contents of the document by traversing the graph structure network. Assume that the initial size of the sliding window is set to , which is the initial size of the sliding window (preset to 5 entities). The size of the sliding window is used to indicate the scanning granularity, avoiding semantic ambiguity caused by too large a window or fragmentation caused by too small a window.

[0079] Based on the above assumptions, the sliding window The semantic density calculation process is expressed as the following formula: .in, For Entity The corresponding semantic role weights. For example, the core argument has a high weight, while the background explanation has a low weight. is the position attenuation coefficient (default is constant). It weakens the semantic association of entities across pages. For example, The larger the value, the more obvious the weight decay is, simulating the human attention priority to the content on the same page. Indicates the The sliding window size is 1 time step. The page offset of the first and last entities in the sliding window, such as when crossing pages Greater than 0. Used to quantify the physical distance of entities in the document to be processed to avoid incorrect aggregation of unrelated content across pages.

[0080] The dynamic slicing layer's execution functions primarily encompass semantic density calculation and slice granularity adjustment. In the semantic density calculation phase, it integrates semantic role importance coefficients with dynamic sliding window traversal to accurately capture the semantic concentration of different local regions within a document. For example, it identifies high-density regions containing key arguments or complex logic, and low-density regions used for background or supplementary explanation. In the slice granularity adjustment phase, it first determines an initial slice granularity for each region based on the inverse association rule of "higher semantic density, smaller slice granularity." It then uses boundary optimization strategies to check whether the logical relationships between entities at the slice boundaries are separable. If an inseparable logical relationship is encountered (such as the argument and evidence in a general-to-specific structure), the slice boundaries are flexibly adjusted to avoid disrupting content coherence, ultimately achieving a target slice granularity that aligns with the semantic structure. Furthermore, entities are reorganized according to their semantic roles, high-density entities are divided into single-sentence slices to highlight details, and low-density entities are aggregated into paragraph or chapter slices to reflect structural hierarchy, thereby constructing a well-defined content structure.

[0081] It's worth noting that a system of semantic role importance coefficients is used to establish a quantitative standard for document value. For example, in a financial annual report, core indicators such as "net profit growth rate" are assigned a high weight of 0.9, while background information such as "company history" receives only a 0.3 weight. This weighting is not static but dynamically adjusted based on the document's subject matter. For example, when processing a pharmaceutical clinical trial report, the weight of "adverse reaction incidence" is automatically increased from the standard 0.6 to 0.8, reflecting the unique value proposition of medical documents. This weighting system is transformed into an actionable density assessment using a sliding window algorithm. The window initially size is set to 5 entities (approximately 1 paragraph), but it intelligently expands and contracts during the traversal process. For example, when the window passes over a high-density area (such as the "Management Discussion" section of a financial report), it automatically shrinks to 2-3 entities for fine-grained segmentation. When encountering a low-density area (such as the term explanation in the appendix), it expands to 10 entities for coarse-grained merging.

[0082] Here, the mapping of semantic density to slice granularity follows a nonlinear transformation principle. Optionally, a pre-set density threshold is divided into three levels: "Core Statement Area" above 0.8 is forced to use single-sentence granularity, ensuring that each important data point is isolated; "Argument Support Area" between 0.4 and 0.8 is segmented into paragraphs to maintain a complete chain of evidence; and "Auxiliary Information Area" below 0.4 is merged into entire sections. This mapping is not mechanically executed but dynamically adjusted by a boundary optimizer. When the system attempts to slice a "Contraindications" section (density 0.85) from a medical report separately from the adjacent "Drug Interactions" section (density 0.82), the boundary detection module identifies a strong causal relationship between the two (the former determines the applicability of the latter), triggering a granularity retraction mechanism, ultimately merging the two sections into a composite knowledge unit. This flexible approach allows the "Experimental Group / Control Group" comparison in a technical white paper to remain within the same slice even if the density value differs by 0.1.

[0083] Furthermore, cross-level knowledge aggregation is implemented in the slice generation stage. For single-sentence slices extracted from high-density areas, the system will attach semantic tags: the "liquidated damages clause" in the legal contract will not only be an independent block, but will also be labeled as "high risk and requires special review". The chapter slices generated by merging low-density areas generate content summaries through graph neural networks. For example, the background information in a 20-page market analysis can be compressed into 3 lines of core trend description. This creates a knowledge density gradient. When users browse PPTs, their eyes will naturally focus on the high-density slice area to obtain the core conclusions, and only expand the supporting details of the low-density area when necessary. In consulting scenarios, it helps to shorten the customer decision-making time for strategic proposals, allowing decision makers to directly locate key segments of competitive analysis with a density value greater than 0.7.

[0084] It should be noted that the most significant value of the dynamic slicing layer lies in resolving the paradox of information overload and key omissions. Traditional fixed-block approaches to scientific research papers either fragment the Methods section too fragmented (resulting in 30+ unrelated steps) or merge the Discussion section too coarsely (missing important inferences). This application, through density-aware elastic slicing, was able to consolidate the Methods section into 5-7 modular steps (granularity = 0.5) based on the experimental process in a test set of papers, while the groundbreaking findings in the Conclusions section were expanded sentence by sentence (granularity = 0.9).

[0085] Document content slices not only carry content but also embed value assessment data. Furthermore, they can optionally introduce actionable knowledge encapsulation. For example, when a pharmaceutical company analyzes 100,000 clinical reports, it can automatically extract adverse reaction descriptions with a density greater than 0.85 as early warning knowledge packages, promoting the establishment of an industry-wide pharmacovigilance database. This shifts document processing from traditional document organization to a new approach of knowledge mining.

[0086] Thus, the dynamic slicing layer realizes the intelligent deconstruction of document content. Through the dynamic adaptation of semantic density and slicing granularity, the slicing results are consistent with the logical habits of human reading and can accurately reflect the internal structure of the document. Secondly, the dynamic slicing layer enhances the flexibility and robustness of slicing. Through adaptive window and elastic boundary adjustment, it effectively handles complex logical relationships (such as progressive structures across paragraphs) and avoids semantic breaks caused by mechanical slicing. Finally, the dynamic slicing layer improves the practicality of document processing. Structured content slicing can be directly applied to scenarios such as information retrieval, summary generation, and knowledge graph construction. For example, single-sentence slicing facilitates the extraction of key knowledge points, paragraph slicing is suitable for generating content overviews, and chapter slicing helps to build a document macro framework, thereby providing an efficient and orderly data foundation for subsequent intelligent analysis and application.

[0087] In an optional embodiment of the above model, a traceability index layer is used to generate traceability pointers for each document content slice, combined with the page position of each entity in the document to be processed. Each traceability pointer is then metadata-bound to the corresponding document content slice in the content structure tree to obtain a traceability pointer network that matches the content structure tree. The traceability pointers include at least the page number, paragraph offset, sentence offset, and / or page layout offset associated with each document content slice in the document to be processed.

[0088] Specifically, the provenance indexing layer essentially constructs a digital content location system, operating similarly to establishing a GPS coordinate network for document content. This layer receives document content slices from the dynamic slicing layer and generates a provenance pointer for each slice with multi-dimensional location accuracy. These pointers are constructed using a hierarchical encoding strategy: page numbers serve as first-level location anchors (e.g., P15), paragraph offsets serve as second-level precise locations (e.g., Para3), sentence offsets serve as third-level micro-coordinates (e.g., Sentence2), and page layout offsets record the physical position of elements in the two-dimensional plane (e.g., coordinates [x:120, y:340]). This multi-level coordinate system allows precise location of text fragments, data tables, and image elements within the original document. When processing legal contracts, the system not only annotates the page and paragraph where the dispute resolution clause appears, but also records the clause's exact location below the header, ensuring millimeter-level traceability even in complex layouts.

[0089] The binding process between the traceability pointer and the content structure tree adopts a bidirectional link mechanism. Each slice node stores two types of related data at the same time. The forward pointer points to the source document location, and the reverse pointer links to the corresponding element in the PPT. This design forms a closed-loop verification chain. For example, when a user views a data chart in a PPT, the pointer can be used to trace the exact location of the original document; conversely, when reviewing a paragraph in a document, it can also prompt whether the content has been extracted to the PPT and on which specific slide it is presented. In financial audit scenarios, this mechanism greatly improves the efficiency of cross-document verification. Data verification work that originally required manual review of hundreds of pages of information can now be completed instantly by clicking the traceability tag.

[0090] In addition, the traceability index layer also helps to build the credibility of the automated system. Traditional AI-generated content is often difficult to gain trust in professional scenarios due to the black box effect. However, this system uses a verifiable traceability network to ensure that each element of the automatically generated PPT has a complete proof of origin. In medical applications, when the automatically generated drug instructions PPT is marked with "adverse reaction rate 3.2%", medical experts can use the traceability pointer to jump directly to the original data table on page 23 of the clinical research report for cross-checking. This transparency increases the adoption rate of AI-generated content. More importantly, the traceability pointer network forms a tamper-proof knowledge chain. For example, in the judicial evidence collation scenario, each event description in the case timeline PPT generated by the system is bound to the original file coordinates to ensure that the integrity and traceability of electronic evidence meet the requirements of the Electronic Signature Law.

[0091] Optionally, after metadata is bound to the corresponding document content slices in the content structure tree to obtain a traceability pointer network that matches the content structure tree, user editing operations on the document being processed can be monitored. Editing operations include adding document content, modifying document content, deleting document content, changing the page location of document content, and / or changing the paragraph location of document content. Furthermore, the document content slices associated with the editing operations in the content structure tree are dynamically updated, and the traceability information associated with the modified document content in the traceability pointer network is dynamically adjusted to reduce the difficulty of maintaining traceability information.

[0092] Specifically, the dynamic maintenance mechanism of the traceability pointer network establishes a real-time, responsive document ecosystem. When a user edits the original document, the system captures change events through the operation monitoring engine. These events are abstracted into four-dimensional feature vectors: operation type (addition / deletion / modification), scope of impact (character level / paragraph level / page level), location coordinates (page number + layout offset), and timestamp. The system uses a differential algorithm to analyze the impact of editing operations on the existing content structure tree. For example, simple text corrections may only require updating the text cache of a single slice, while chapter reorganization triggers global topology reconstruction. For example, when a legal advisor moves the "force majeure clause" in a contract from page 8 to the appendix, the system not only automatically adjusts the page number mark of the clause slice, but also recalculates the logical edge weights of its related clauses such as "liability for breach of contract" to ensure the integrity of the document semantic network.

[0093] For example, this mechanism demonstrates unique value in the collaborative writing of drug instructions. When a pharmacology expert adds a new "Drug Interactions" chapter to the original research report, the system automatically performs a three-level response: first, a new terminology slice is inserted into the content structure tree, inheriting the semantic role annotations of the adjacent chapters; then, the logical relationship of the relevant pharmacology description slices is adjusted, and the original "metabolic pathway" single-point statement is upgraded to a progressive chain of "metabolic pathway → drug interaction"; finally, the corresponding scientific diagram in the PPT is updated, and a revision annotation of "Updated data according to 2024-03" is added to the bottom of the slide. The entire process does not require manual intervention, and all historical versions form a traceable version tree through a pointer network.

[0094] For example, when the CFO promotes the "Divisional Performance" table from the appendix to the core data chapter, the system not only moves the corresponding slice position, but also intelligently analyzes the impact of the adjustment on the "Management Discussion" chapter, and automatically suggests supplementing the logical connection statements between divisional performance and strategic planning.

[0095] The most significant value of the dynamic maintenance mechanism is that it overcomes the rigidity of traditional automated systems and enables dynamic maintenance of knowledge assets. It automatically evolves as the processed documents are continuously optimized, ensuring that subsequent PowerPoint presentations are always based on the latest knowledge. For example, when new clinical trial data is added to a medical guideline, the system not only updates the relevant slides but also automatically marks the statistical differences between the new and old conclusions through a pointer network. This capability reduces the knowledge update cycle of the clinical decision support system from quarterly to real-time.

[0096] Step S103 , using a structure tree-based visual mapping engine model, maps the document content slices and traceability information in the content structure tree to corresponding presentation content according to logical relationships to obtain a target presentation.

[0097] As an optional embodiment, in step S103, a structure tree-based visual mapping engine model is used to map the document content slices and traceable information in the content structure tree to corresponding presentation content according to logical relationships to obtain a target presentation, including: converting the document content slices corresponding to the root node of the content structure tree into presentation titles; converting the document content slices corresponding to the leaf nodes of the content structure tree into chapter titles, page titles, page subtitles, and page subviews according to data types, logical relationships, and entity semantics; arranging the presentation titles and the conversion results corresponding to the leaf nodes on pages according to the correspondence between the logical relationships and the hierarchical structures to obtain a target presentation; and establishing corresponding visual traceability identifiers for each page element configuration in the target presentation based on the traceable information in the content structure tree, so that users can view the information sources of each page element in the target presentation through the identifiers.

[0098] In an embodiment of the present application, the correspondence between the logical relationship and the hierarchical structure includes at least one of the following: the parallel structure corresponds to the horizontal side-by-side card layout, the progressive structure corresponds to the vertical timeline, the progressive structure corresponds to the flowchart layout, and the general-specific structure corresponds to the center-spoke layout.

[0099] It's worth noting that the visual mapping engine is essentially a translator from logical structure to visual language, transforming the abstract cognition of the content structure tree into a presentation design that conforms to the laws of human visual cognition. The engine first analyzes the topological features of the content structure tree: the root node serves as the presentation's overall title, carrying the document's core proposition; intermediate nodes form the chapter framework, reflecting the hierarchical divisions of the knowledge system; and leaf nodes are transformed into specific presentation units, containing substantive content such as arguments, data, or charts. During this conversion process, the engine strictly adheres to the principle of "logic priority." When it identifies the progressive chain of "technical pain points → innovative solutions → experimental verification" in a technical solution document, it automatically selects a vertical timeline layout, connecting the three stages with arrows to form a visual derivation flow. This mapping is not a simple template fill-in, but rather based on cognitive psychology research: parallel relationships are arranged in a horizontal card layout to leverage the human instinct for parallel comparison, and the overall-specific structure adopts a hub-and-spoke layout to reinforce the leadership of the core concept.

[0100] Specifically, the core technical framework of the visual mapping engine is built on two pillars: multimodal feature fusion and dynamic topological mapping. First, a graph neural network (GNN) analyzes the topological characteristics of the content structure tree, extracting node features (such as text type, diagram form, and semantic weight) and edge features (such as logical relationship strength and hierarchical depth). When processing medical research reports, the GNN identifies the progressive relationship chain of "clinical trial design → results analysis → conclusion" and quantifies the influence weight of each node (for example, the conclusion node weight reaches 0.92). Combined with the Transformer encoder's deep encoding of document content slices, a 256-dimensional feature vector is generated that integrates semantic and layout information, laying the foundation for subsequent visual mapping.

[0101] Furthermore, the engine has a built-in dynamic mapping library of logical relationships and layout templates, which realizes intelligent adaptation through the probabilistic graph model (PGM). For example, when a parallel structure is detected (such as the comparison of the three major advantages of products), the system automatically selects the horizontal card layout module and dynamically adjusts the card size based on the amount of node content: the width of the text node is set as a function of the content length, and the chart nodes retain the same width ratio. For progressive relationships (such as the stage of technological evolution), the vertical timeline engine is enabled, and the hierarchical spacing is automatically compressed or expanded according to the semantic density of the nodes. For example, the spacing between key breakthrough nodes is widened to 120px to highlight their importance, and transitional nodes are compressed to 60px. The total-point structure triggers the center-radiate algorithm, and the diameter of the core node is exponentially enlarged according to the number of associated edges to ensure that the visual center of gravity matches the content value.

[0102] Furthermore, a two-way collaborative mechanism for template generation is introduced. On the one hand, the visual features (primary color, white space ratio, font pairing pattern) of a large number of professional PPT templates are extracted through convolutional neural networks (CNN) to construct a 512-dimensional template feature space. On the other hand, the semantic features of the content structure tree are projected into the same dimensional space. Similarity is calculated through cross-modal comparative learning, and the template with the highest degree of fit is finally selected (such as biological research reports matching blue-green color + large-spaced academic templates). The reinforcement learning module continuously optimizes the strategy: When a user manually adjusts the chart layout of a financial data page, the system records the behavioral data and reversely updates the template matching weight, thereby improving the accuracy of subsequent recommendations of similar content.

[0103] Furthermore, the visual traceability identification system implements three-level positioning technology: page number anchoring (level 1), paragraph offset (level 2), and page coordinate positioning (level 3). Each PPT element is associated with a lightweight JSON metadata package containing traceability coordinates and a version stamp. When a user clicks on a sales data trend chart in a PPT, the system instantly locates the source document Excel spreadsheet using a coordinate mapping algorithm.

[0104] The visual mapping engine also enables real-time edit response. If the source document's financial data is revised, the incremental analysis engine reconstructs only the affected sections, marking the corresponding elements with a red border in the PPT as a warning, while leaving the majority of the unchanged pages intact. This fine-grained refresh mechanism significantly reduces the time required to update hundreds of pages.

[0105] For example, when dealing with nested composite structures (such as the general-specific-parallel relationship in legal clauses), the engine automatically splits the two-page composite layout using a topology reconstruction algorithm, perfectly resolving the layout collapse issue of traditional systems. This marks the official entry of automated presentation generation into the productivity-enhancing phase.

[0106] In the scenario of financial analysis reports, the system demonstrates powerful adaptive intelligence. When processing the "Industry Competition Landscape" chapter, it identifies that the comparative analysis of three competing products belongs to a parallel structure, and automatically generates a horizontal three-column layout, with the company logo marked at the top of each column, a radar chart of key indicators in the middle, and the key points of the SWOT analysis at the bottom. For content with a total-point structure such as "Market Growth Drivers", the engine places the core driving factor in the central circular node of the page, and the five sub-factors are arranged as satellite nodes around it, with the thickness of the connecting line reflecting the influence weight. For example, in the conversion of medical research reports: the system identifies the progressive relationship of "pathological mechanism → clinical manifestation → treatment plan" and the parallel relationship of "drug efficacy comparison" as nested, and finally generates a composite layout with a vertical progressive timeline on the left and a horizontal drug comparison table on the right. The professionalism of this automatic design is comparable to that of a senior medical visual designer.

[0107] The visual traceability identification system in step S103 builds a quick channel for knowledge verification. The identification embedded in each PPT element is not a simple page number label, but a hierarchical intelligent prompt system. For data charts, the identification shows "Source: 2023Q4 Financial Report P23 Table 5"; for core arguments, it is marked "Based on experimental data from P45-47". When the user hovers over the identification, the system will display a context summary of the source document in a pop-up window; clicking the identification will jump directly to the precise location of the associated document (including cross-document reference scenarios). In compliance audit scenarios, this design enables auditors to quickly verify whether the "Risk Disclosure Statement" fully reflects the original contract terms, compressing the cross-check that originally took 2 hours to complete in 10 minutes. Further, optionally, the visual traceability identification system is also equipped with version-sensitive identification. When the system detects that the source document corresponding to a data slice has been modified but the PPT has not been updated, it automatically changes the identification color to warning orange and prompts "The source data has changed. It is recommended to regenerate this page."

[0108] Step S103 generates not a static slideshow, but a cognitive interface with intelligent navigation. Users can use the "Logical View" button to switch between different presentation modes, such as parallel and progressive, to understand complex concepts from multiple perspectives. In clinical trial presentations at multinational pharmaceutical companies, medical experts quickly navigate between different analytical frameworks by switching between "Timeline Mode" and "Drug Comparison Mode," improving the efficiency of cross-departmental collaborative meetings.

[0109] Optionally, in step S103, reinforcement learning can be used to train the paging strategy layer in the visualization mapping engine model. The paging strategy layer dynamically adjusts the target presentation's paging based on the length of the semantic information, triggering paging optimization based on the page information density, which is the ratio of the presentation's word count to the screen area.

[0110] Specifically, in step S103, reinforcement learning is introduced to train the visualization mapping engine's paging strategy layer. The core principle is to simulate the paging logic used by humans when designing presentations, allowing the model to autonomously learn the optimal paging strategy through a "trial-error-feedback-optimization" cycle. Within the reinforcement learning framework, the paging strategy layer takes the presentation's content slice sequence as input and uses "page information density balance," "logical coherence," and "visual comfort" as reward functions. It continuously adjusts the paging positions (e.g., where to split content slices, merge them onto the same page, or split them onto different pages) to maximize the overall reward value. For example, when the model paginates key data slices with high semantic density, it is rewarded for its high information focus. However, if logically continuous slices (such as the "problem-solution" progressive structure) are forcibly split onto different pages, it is penalized for disrupting coherence. Through iterative optimization based on a large amount of training data, the paging strategy layer develops a sensitive response to the document's semantic structure.

[0111] During implementation, the paging strategy layer dynamically determines the initial paging method based on the length of semantic information (e.g., the number of sentences in a slice and the complexity of the chart). Long paragraphs or complex charts are prioritized as separate pages to avoid information overload; short sentences or supplementary explanations are attempted to be merged with adjacent content. Subsequently, an optimization mechanism is triggered by calculating the page's information density (the ratio of word count to screen area). If a page's information density is too high (e.g., exceeding a threshold, as evidenced by crowded text and dense chart elements), the page's content is automatically split onto new pages to ensure concise and easy-to-read information on a single page. If a page's information density is too low (e.g., with only a small amount of text), an attempt is made to merge it with the preceding and following pages to reduce ineffective paging. For example, when processing a document containing a market analysis report, the model might assign the "Annual Sales Trend Chart" as a separate page (moderate information density) while logically splitting the corresponding three analysis paragraphs onto two pages to avoid a crowded "chart + long paragraph" layout on a single page.

[0112] Furthermore, the paging strategy layer models each page of the PPT as a state, which includes features such as page information density, semantic role distribution, and layout remaining space. The paging action is defined as keeping the current page, forcing a paging, or adjusting the element compression rate. The reward function is composed of the information readability coefficient (0.8+ no penalty), visual comfort (positive score for characters per square inch <50), and logical integrity (additional reward for key slices not spanning pages). When processing technical white papers, the strategy layer dynamically simulates the paging effect through a Markov decision process (MDP). Specifically, if it is detected that the experimental method description requires 1.2 pages of space, the optimal action chain of "compressing the method description by 20% + paging the result analysis in advance" is automatically triggered.

[0113] These steps, first, improve information presentation in line with cognitive principles. The paging strategy optimized through reinforcement learning automatically adheres to the principle of "focusing key content on a single page and displaying related content adjacently." For example, "core argument - supporting data" sections are grouped together on adjacent pages, improving audience comprehension. Second, the visual experience is significantly enhanced. Dynamically adjusting information density avoids the "overfilled" and "overly empty" issues associated with traditional fixed paging. Medical experiments have shown that PPT reports using this strategy improve audience information reception. Third, the paging strategy automatically adjusts the style based on document type (such as academic presentations and business proposals) to meet the needs of multiple scenarios. For example, academic scenarios tend to break down the data derivation process in detail, while business scenarios prioritize a single page for key conclusions, enhancing the system's generalization capabilities. This intelligent paging approach reduces the time required for manual paging adjustments while significantly improving the professionalism and logicality of presentations.

[0114] Further optionally, after step S103, a hierarchical layout algorithm and a force-directed algorithm may be used to dynamically optimize the position layout of each page element in the target presentation.

[0115] During presentation generation, parent and child nodes are linked through dynamic spacing constraints, and node area size is adaptively adjusted based on semantic density. Combined with a template-based neural layout generator (TNLG) and layout optimization algorithms, this achieves a deep fusion of content logic and visual presentation. Its core principle is to transform the document's hierarchical structure (such as the parent-child node relationships in the content structure tree) into visual spatial layout rules: nodes corresponding to core slices with high semantic density (such as key arguments) occupy a larger page area to highlight key points; closely related parent and child nodes maintain logical proximity through dynamic spacing to avoid visual fragmentation. TNLG intelligently adapts document slices to specific PPT page types through a three-level mechanism of "content mapping + template matching + reinforcement learning." For example, core slices are mapped to title pages or chapter pages, retaining key sentences from the original text as titles; data slices automatically use visualization tools to generate chart pages; and mixed text and image slices use adaptive layout to achieve mixed layout. At the same time, the system uses CNN to extract the visual features of massive PPT templates (such as color matching and layout distribution), recommends the most matching templates for different content slices, and uses reinforcement learning to dynamically adjust paging and element positions to ensure that the layout conforms to both visual specifications and content logic requirements.

[0116] Taking a technology product white paper as an example, first, the "Product Core Technology" parent node and its subordinate "Algorithm Principles" and "Performance Indicators" child nodes in the content structure tree are automatically assigned different area sizes based on semantic density. The parent node, as the core slice, occupies a larger area on the homepage, while the child nodes are distributed on subsequent pages in the form of lists or charts. The parent and child nodes maintain visual connection through page margins and title hierarchy. For the "User Growth Data" data slice, the system automatically calls D3.js to generate a line chart and matches it with a simple business template to highlight data trends. For the mixed slice of "Technical Architecture Diagram + Text Description", an adaptive layout with left-image and right-text is adopted to ensure clear correspondence between the image and text. During the layout optimization phase, the hierarchical layout algorithm arranges elements according to the content hierarchy (e.g., title → chart → annotation), while the force-directed algorithm adjusts the spacing between elements to avoid crowding in dense areas (e.g., optimizing the distance between chart axis labels and data points), ultimately generating a logically clear and visually balanced PPT page.

[0117] Through the above steps, logical visualization is enhanced. Through dynamic spacing and size adaptation, the audience can intuitively perceive the hierarchical relationship and importance of the content. For example, in a business proposal, the page area for the core selling point is larger than the auxiliary explanation, which improves the efficiency of capturing key information. Secondly, the visual experience is optimized. The combination of template matching and layout algorithms avoids the mechanical feel of traditional automatic generation tools. In the medical research report scenario, the layout with blue and white as the main color and data charts as pagination reduces the audience's reading fatigue. Third, flexible customization capabilities support enterprises to customize template rules (such as embedding brand colors and standardizing fonts) to ensure that presentations meet VI standards. At the same time, the reinforcement learning mechanism can quickly generate adaptive layouts for emergency scenarios (such as temporary meetings), reducing adjustment time.

[0118] Step S104: display the target presentation to the user, and assist the user in understanding key content information in the document to be processed through the target presentation.

[0119] Specifically, the structured content generated by the previous processing (such as the content structure tree and the visual page layout) is converted into an intuitive and easy-to-understand interactive interface, which assists users in quickly capturing the core content of the document through information presentation methods that conform to human cognitive habits. Based on the design logic of "information layering-key point highlighting-interactive guidance", the pages of the target presentation are arranged in a logical order of content, and the perception of key information is enhanced through visual design (such as bold titles, highlighted charts) and interactive functions (such as hyperlink tracing and page navigation). For example, when a user opens a presentation, the homepage automatically presents the document theme and core argument slices, and subsequent pages are expanded according to logical relationships (such as total score, progressiveness). The keywords or charts in each slice can be clicked to jump to the corresponding position in the source document, forming a closed-loop cognitive path from a global overview to detailed tracing.

[0120] Taking a legal contract document as an example, the target presentation generated by the system will first present core sections such as the contract body and a summary of key clauses on the title page; subsequent pages will split the content according to the logic of "rights and obligations → breach of contract clauses → dispute resolution", and each clause paragraph will correspond to a concise text description and a traceable link. When browsing, users can quickly locate chapters through the tree navigation on the left side of the page. Clicking on the text description of a breach of contract clause will jump to the specific paragraph in the original contract text, making it easier to verify the details. For complex financial data clauses, the system will convert them into tables or trend charts to intuitively display data changes, avoiding user comprehension difficulties when faced with large paragraphs of text.

[0121] This significantly improves information acquisition efficiency. Through structured presentation and visual guidance, users can quickly grasp the key content of a document compared to traditional reading methods. This is especially true when processing research reports with hundreds of pages, allowing them to quickly locate core conclusions and supporting data. The intelligent interactive experience and traceability feature address the information silos inherent in traditional presentations. In auditing and legal scenarios, users can instantly verify the authenticity of content, reducing communication costs. It also facilitates adaptation to multi-terminal scenarios, supporting adaptive display on PCs, tablets, mobile phones, and other devices. The page layout automatically adjusts to suit the screen size, meeting the needs of mobile office work.

[0122] Compared with the traditional method of making presentations, it is necessary to manually comb through documents word by word, extract key information and typeset them, which is time-consuming and labor-intensive. In the embodiment of the present application, various types of documents to be processed can be automatically imported, and whether they are pure text, images, or rich text information mixed with images and text, they can all be quickly identified. Through the hierarchical semantic segmentation model, content understanding and information slicing are automatically completed, and tedious manual operations are converted into machine automated processing, saving a lot of time and energy, freeing relevant personnel from repetitive work and focusing on more creative work. Secondly, the embodiment of the present application accurately analyzes the document by constructing a content structure tree, clarifies the logical relationship between document content slices, and marks traceable information at the same time. This structured processing makes the key content of the document clear at a glance, and in the presentation, it helps to assist users to quickly grasp the core points and understand the relationship between information. Compared with traditional scattered and illogical presentations, the presentation content generated by the embodiment of the present application is logically clear and well-structured, which greatly improves the efficiency and accuracy of information transmission. Finally, the embodiment of the present application not only converts document content slices into presentations, but also retains traceable information through a visual mapping engine model based on a structure tree. This means that when users question a point of view or data, they can directly conduct a traceability query, quickly locate relevant content in the original document, and verify the authenticity of the information. This traceability enhances the credibility and persuasiveness of presentations, effectively reducing doubts and disputes caused by false information in scenarios such as academic presentations and business proposals.

[0123] In addition, the embodiments of this application change the traditional working mode by realizing the intelligent construction of presentations, further optimize the workflow, and promote intelligent transformation. The full process automation from document import to presentation generation optimizes the workflow and reduces human errors in the intermediate links. At the same time, it provides support for the digital and intelligent transformation of enterprises and organizations, conforming to the current development trend of efficient office, making information processing and presentation more intelligent and convenient.

[0124] After introducing the method of the exemplary embodiment of the present application, next, refer to Figure 2 A system for generating a trusted traceable presentation based on document content in accordance with an exemplary embodiment of the present application is described. The system comprises: an acquisition module configured to acquire a document to be processed input by a user; wherein the content information of the document to be processed includes at least text, images, and / or rich text information mixed with text and images;

[0125] A construction module is used to perform semantic recognition on the content information of the document to be processed using a hierarchical semantic segmentation model, and to slice the document to be processed based on the semantic recognition results to obtain a content structure tree for the document to be processed. The content structure tree includes: document content slices matching key content information, logical relationships between each document content slice, and traceability information corresponding to each document content slice. The slicing granularity of each document content slice is dynamically configured based on the semantic recognition results matching the document to be processed.

[0126] The display module is used to adopt a structure tree-based visual mapping engine model to slice the document content and traceable information in the content structure tree and map them into corresponding presentation content according to logical relationships to obtain a target presentation; the target presentation is displayed to the user to assist the user in understanding the key content information in the document to be processed through the target presentation.

[0127] The above system can implement each step described in the above method implementation, and the specific implementation method of each step will not be repeated here.

[0128] After introducing the method and system of the exemplary embodiment of the present application, a terminal device of the exemplary embodiment of the present application is described next. The terminal device is used to implement the various steps recorded in the above method implementation. The specific implementation method of each step will not be repeated here.

[0129] After introducing the method, system and terminal device of the exemplary embodiment of the present application, the following Figure 3 For a description of the computer-readable storage medium of the exemplary embodiment of the present application, please refer to Figure 3The computer-readable storage medium shown is a CD 30, which stores a computer program (i.e., a program product). When executed by a processor, the computer program implements each step described in the above method implementation. The specific implementation of each step is not repeated here.

[0130] It should be noted that examples of the computer-readable storage medium may also include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other optical or magnetic storage media, which will not be described in detail here. The above-described embodiments are only specific implementation methods of the present application, which are used to illustrate the technical solutions of the present application, rather than to limit them. The protection scope of the present application is not limited thereto. Although the present application has been described in detail with reference to the above-mentioned embodiments, ordinary technicians in this field should understand that any technician familiar with this technical field can still modify the technical solutions described in the above-mentioned embodiments within the technical scope disclosed in this application, or make equivalent replacements for some of the technical features therein; and these modifications, changes or replacements do not deviate from the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application. Therefore, the protection scope of the present application shall be based on the protection scope of the claims.

Claims

1. A method for generating a trustworthy traceable presentation based on document content, characterized in that: The method comprises: Obtaining a document to be processed input by a user; wherein the content information of the document to be processed includes at least: text, image, and / or rich text information mixed with text and image; Using a hierarchical semantic segmentation model, semantic recognition is performed on the content information of the document to be processed. Based on the semantic recognition results, the document to be processed is sliced ​​to obtain a content structure tree for the document to be processed. The content structure tree includes: document content slices that match key content information, logical relationships between each document content slice, and traceability information corresponding to each document content slice. The slicing granularity of each document content slice is dynamically configured based on the semantic recognition results matched in the document to be processed. Using a structure tree-based visual mapping engine model, the document content slices and traceable information in the content structure tree are mapped to corresponding presentation content according to logical relationships to obtain a target presentation; Show the target presentation to the user, helping the user understand the key content information in the document to be processed through the target presentation; The structure tree-based visual mapping engine model is used to map the document content slices and traceable information in the content structure tree to corresponding presentation content according to logical relationships to obtain a target presentation, including: Converting the document content slice corresponding to the root node of the content structure tree into a presentation title; Slice the document content corresponding to the leaf nodes of the content structure tree and convert them into chapter titles, page titles, page subtitles, and page subviews according to data types, logical relationships, and entity semantics; Arrange the presentation titles and the conversion results corresponding to the leaf nodes on pages according to the correspondence between the logical relationships and the hierarchical structure to obtain a target presentation; wherein the correspondence between the logical relationships and the hierarchical structure includes at least one of the following: a parallel structure corresponds to a horizontal side-by-side card layout, a progressive structure corresponds to a vertical timeline, a progressive structure corresponds to a flowchart layout, and a general-specific structure corresponds to a hub-and-spoke layout; Based on the traceable information in the content structure tree, a visual traceability identifier corresponding to each page element configuration in the target presentation is established, so that the user can view the information source of each page element in the target presentation through the identifier.

2. The method for generating a trusted traceability presentation based on document content according to claim 1, characterized in that: The method uses a hierarchical semantic segmentation model to perform semantic recognition on the content information of the document to be processed, and slices the document to be processed according to the semantic recognition result to obtain a content structure tree of the document to be processed, including: Through the multimodal parsing network, the document to be processed is parsed according to the data type to obtain a graph structure network; wherein the graph structure network includes: the graph nodes matching each content information, the association relationship between each graph node, and the page position of each graph node in the document to be processed; Through the semantic recognition layer, the Bi-LSTM-CRF model is used to identify the entities contained in the graph structure network and the logical relationships between the entities, and the semantic roles corresponding to each entity are labeled; the logical relationships between the entities include parallel structure, progressive structure, general-specific structure, forward pyramid, and inverted pyramid; The dynamic slicing layer obtains the semantic density of each local area based on the logical relationship between each entity and the semantic role. The slicing granularity corresponding to each local area is adjusted in combination with the semantic density according to the semantic elastic partitioning mechanism. The document content slice corresponding to each local area is constructed based on the slicing granularity to obtain the content structure tree. Each local region contains one or more adjacent entities; the semantic density of the local region is determined by the semantic role importance of all entities in the local region; the greater the semantic density of the local region, the higher the semantic role importance of all entities in the local region, and the smaller the slice granularity of the local region; Through the traceability index layer, combined with the page position of each entity in the document to be processed, a traceability pointer for each document content slice is generated, and each traceability pointer is bound to the metadata of the corresponding document content slice in the content structure tree to obtain a traceable pointer network that matches the content structure tree; the traceability pointer includes: the page number, paragraph offset, single sentence offset, and / or page layout offset associated with each document content slice in the document to be processed.

3. The method for generating a trusted traceability presentation based on document content according to claim 2, characterized in that: The multimodal parsing network includes a preprocessing layer, a text parsing layer, an image parsing layer, and a multimodal fusion layer. The multimodal parsing network parses the document to be processed according to the data type to obtain a graph structure network, including: Through the pre-processing layer, the document to be processed is transcoded by page to generate a sequence of documents to be processed; The text parsing layer uses a pre-deployed Transformer-based pre-trained model to analyze the document layout and semantic role features of text elements in the document sequence to obtain text feature information. Document layout includes: titles, paragraphs, lists; text elements include sentences and / or word segments; semantic role features include: key sentences, keywords, and their corresponding statistical features; Through the image parsing layer, the pre-deployed CLIP-ViT model is used to extract global image features of the document sequence to be processed. Furthermore, layout element detection and location annotation are performed on the text and / or formulas contained in the charts, images, and / or formulas in the document sequence to be processed, thereby obtaining local image features of the document sequence to be processed. Through the multimodal fusion layer, an adaptive attention mechanism is used to cross-modally associate text feature information, global image features, and local image features; using text elements and image elements as nodes, connecting edges are constructed based on the association results between text elements and image elements, and spatial relationship modeling is performed to obtain a graph structure network corresponding to the document sequence to be processed.

4. The method for generating a trusted traceability presentation based on document content according to claim 2, characterized in that: The method of obtaining the semantic density of each local area through the dynamic slicing layer based on the logical relationship between each entity and the semantic role includes: Get the semantic role corresponding to each entity; Dynamically configuring corresponding semantic role importance coefficients for different semantic roles based on the logical relationships between entities in the graph structure network, document topics, and / or document types; An adaptive sliding window algorithm is used, with each entity as a node, to establish an initial sliding window of a preset size, traverse the graph structure network, and calculate the semantic density corresponding to the local area within the sliding window in real time based on the semantic role importance coefficient. The real-time size of the sliding window during the traversal of the graph structure network is adaptively adjusted based on the calculated semantic density and the document structure relationship between adjacent entities. The step of adjusting the slice granularity corresponding to each local area in combination with the semantic density according to the semantic elastic partitioning mechanism, and constructing the document content slice corresponding to each local area based on the slice granularity to obtain the content structure tree includes: Establish an inverse association rule between semantic density and slice granularity and a boundary optimization strategy, where the higher the semantic density in a local area, the smaller the slice granularity; Obtain the semantic density of each local area and determine the initial slice granularity of each local area based on the reverse association rule; Adopting a boundary optimization strategy, combined with the initial slice granularity, the authors check whether the logical relationship between entities on both sides of the slice boundary is inseparable. If so, the initial slice granularity is elastically retracted until the entities on both sides of the slice boundary are separable and meet the coherence constraints. This results in the target slice granularity for each local area. Based on the target slice granularity, the entities contained in each local area are restructured according to their respective semantic roles to obtain the document content slices of each local area; among them, the entities with semantic density greater than the set threshold are set as single sentence slices, and the entities with semantic density less than the set threshold are constructed into paragraph slices or chapter slices through dynamic aggregation.

5. The method for generating a trusted traceability presentation based on document content according to claim 4, characterized in that: The semantic roles corresponding to each entity are one of the following: title, argument, key chapter, key supporting data, and auxiliary content; among them, the slicing granularity corresponding to the title, argument, key chapter, and key supporting data is single sentence slicing, and the slicing granularity corresponding to the auxiliary content is paragraph slicing or chapter slicing.

6. The method for generating a trusted traceability presentation based on document content according to claim 2, characterized in that: After metadata binding of each traceability pointer with the corresponding document content slice in the content structure tree to obtain a traceability pointer network matching the content structure tree, the method further includes: Monitoring user editing operations on a pending document; wherein the editing operations include: adding document content, modifying document content, deleting document content, changing the page location of document content, and / or changing the paragraph location of document content; The document content slices associated with the editing operation in the content structure tree are dynamically updated, and the traceability information associated with the modified document content in the traceability pointer network is dynamically adjusted to reduce the difficulty of maintaining the traceability information.

7. The method for generating a trusted traceability presentation based on document content according to claim 1, characterized in that: After the structure tree-based visual mapping engine model is used to map the document content slices and traceable information in the content structure tree to corresponding presentation content according to logical relationships to obtain the target presentation, the method further includes: A hierarchical layout algorithm and a force-directed algorithm are used to dynamically optimize the position layout of each page element in the target presentation.

8. The method for generating a trusted traceable presentation based on document content according to claim 1, characterized in that: After the structure tree-based visual mapping engine model is used to map the document content slices and traceable information in the content structure tree to corresponding presentation content according to logical relationships to obtain the target presentation, the method further includes: Reinforcement learning is used to train the paging strategy layer in the visual mapping engine model; Through the paging strategy layer, the paging method in the target presentation is dynamically adjusted according to the length of the semantic information, and the paging optimization of the target presentation is triggered according to the page information density; the page information density is the ratio between the number of words in the presentation content and the screen area.

9. A document content-based trusted traceability presentation generation system, characterized in that: The system comprises: An acquisition module is used to acquire a document to be processed input by a user; wherein the content information of the document to be processed includes at least: text, image, and / or rich text information mixed with text and image; A construction module is used to perform semantic recognition on the content information of the document to be processed using a hierarchical semantic segmentation model, and to slice the document to be processed based on the semantic recognition results to obtain a content structure tree for the document to be processed. The content structure tree includes: document content slices matching key content information, logical relationships between each document content slice, and traceability information corresponding to each document content slice. The slicing granularity of each document content slice is dynamically configured based on the semantic recognition results matching the document to be processed. The presentation module is configured to use a structure tree-based visual mapping engine model to map document content slices and traceable information in the content structure tree to corresponding presentation content according to logical relationships, thereby obtaining a target presentation; and present the target presentation to the user, thereby assisting the user in understanding key content information in the document to be processed through the target presentation; The display module adopts a structure tree-based visual mapping engine model to map the document content slices and traceable information in the content structure tree into corresponding presentation content according to logical relationships to obtain the target presentation. Specifically, it is used to: convert the document content slices corresponding to the root node of the content structure tree into presentation titles; convert the document content slices corresponding to the leaf nodes of the content structure tree into chapter titles, page titles, page subtitles, and page subviews according to data types, logical relationships, and entity semantics; arrange the presentation titles and the conversion results corresponding to the leaf nodes on pages according to the correspondence between the logical relationships and the hierarchical structure to obtain the target presentation; wherein the correspondence between the logical relationships and the hierarchical structure includes at least one of the following: a parallel structure corresponds to a horizontal side-by-side card layout, a progressive structure corresponds to a vertical timeline, a progressive structure corresponds to a flowchart layout, and a general-to-specific structure corresponds to a hub-and-spoke layout; based on the traceable information in the content structure tree, establish a visual traceability identifier corresponding to each page element configuration in the target presentation so that the user can view the information source of each page element in the target presentation through the identifier.

Citation Information

Patent Citations

  • Document structure tree generation method and device

    CN117688123A

  • Presentation file generation method and device based on large model, electronic equipment and medium

    CN120031015A