Semi-structured file processing method based on LLMs large language model

By adopting a hierarchical cross-modal attention mechanism, dynamic structure analysis network and multimodal generation inference network methods in semi-structured file processing, the diversity and complexity, computing efficiency and scalability, and accuracy and reliability of information extraction in semi-structured file processing are solved, and efficient and intelligent file processing and data management are achieved.

CN120218236APending Publication Date: 2025-06-27CHINA IND INTERNET RES INST
View PDF 0 Cites 12 Cited by

Patent Information

Application Number
CN202510272195.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The prior art faces the challenges of diversity and complexity processing, insufficient computing efficiency and scalability, and the accuracy and reliability of information extraction when processing semi-structured files.

Method used

The hierarchical cross-modal attention mechanism HCMA combines the multi-head attention mechanism with the graph neural network GNN to fusion of text and images, and designs a dynamic structure analysis network DPN to combine reinforcement learning and generative models for adaptive structure transformation, and realizes the joint generation and reasoning of text and images through the multi-modal generation inference network MGRN combined with diffusion model and neural symbol reasoning.

Benefits of technology

It significantly improves the intelligence level of semi-structured file processing, improves the automation and intelligence level of file processing, enhances the development of data management and application, improves the accuracy and robustness of document understanding, and promotes cross-modal intelligent reasoning and generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005303082930000041
    Figure BDA0005303082930000041
  • Figure BDA0005303082930000042
    Figure BDA0005303082930000042
  • Figure BDA0005303082930000051
    Figure BDA0005303082930000051
Patent Text Reader

Abstract

The invention provides a semi-structured file processing method based on an LLMs (Language Language Model), which is used for efficiently analyzing a semi-structured file and converting the semi-structured file, and belongs to the technical field of intelligent document processing and multi-modal learning. The method is characterized by comprising the following steps: multi-modal fusion analysis: adopting a hierarchical cross-modal attention mechanism HCMA, and combining a multi-head attention mechanism with a graph neural network GNN to realize multi-level feature fusion of a text and an image; adaptive structured conversion: adopting a dynamic structure analysis network DPN, and combining reinforcement learning and a generation model to realize adaptive analysis of a document structure; and model generation and multi-modal reasoning: adopting a multi-modal generation reasoning network MGRN, and combining a diffusion model and neural symbol reasoning to realize joint generation and reasoning of the text and the image. According to the method, a remarkable technical breakthrough is brought to the aspects of document analysis, information extraction and intelligent reasoning, and powerful support is provided for data-driven innovation of various industries.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of intelligent document processing and multimodal learning, and particularly relates to a semi-structured file processing method based on large language models (LLMs) for efficiently parsing semi-structured files and converting them into structured files. Background Art

[0002] Analysis of the domestic research status: The research in the field of semi-structured file parsing based on large language models in China started relatively late. However, in recent years, with the rapid development of domestic artificial intelligence technology, especially the strong rise of domestic large language models represented by Baidu Wenxin Yiyan and Alibaba Cloud Tongyi Qianwen, the related research has shown a good trend of rapid growth. In the academic community, domestic universities and research institutions have focused on the fine-tuning and optimization of LLMs for information extraction tasks. In Prompt Engineering, more effective prompts are carefully designed for specific domains and file types to guide LLMs to accurately extract target information; in Few-shot / Zero-shot Learning, a small amount of labeled data or unlabeled data is used to enable LLMs to quickly adapt to new domains and new file types; in knowledge fusion, domain knowledge is integrated into LLMs to improve the accuracy and integrity of information extraction; in the field of model compression and acceleration, in response to the efficiency requirements of semi-structured file parsing, methods for model lightweight and inference acceleration are deeply studied; in the exploration of specific domain applications, for common domestic semi-structured file types such as Chinese contracts, financial statements, and electronic medical records, research on parsing methods based on LLMs and system development are carried out; some universities and research institutions have also started to build evaluation benchmarks and datasets for Chinese semi-structured file parsing tasks to promote the research progress in this field. In the industrial community, Internet giants such as Baidu, Alibaba, and Tencent, relying on their strong technical accumulation and data advantages, have actively engaged in the research and development of products and services for semi-structured file parsing based on large language models. These products and services are widely used in fields such as intelligent document processing, knowledge graph construction, and intelligent customer service. At the same time, a number of startup companies focusing on intelligent document processing and knowledge extraction have emerged. They use open-source or self-developed LLM technologies to provide enterprises with solutions for automated processing of semi-structured data. In industries such as finance, law, healthcare, and government affairs, the technology of semi-structured file parsing based on large language models has gradually been applied. For example, in the financial field, it can automatically identify and extract contract terms, invoice information, bank statements, etc.; in the legal field, it can assist lawyers in quickly analyzing judgment documents, legal instruments, etc.; in the healthcare field, it can extract key disease information, medication information, etc. from electronic medical records. In addition, some open-source tools and platforms for semi-structured file parsing based on domestic LLMs have emerged, greatly reducing the research and application thresholds for enterprises and individuals in this field.

[0003] In the domestic field of semi-structured document parsing based on large language models, it is in a stage of rapid development. Academic research closely follows the international frontiers, and the industrial application implementation is also accelerating continuously. However, it cannot be ignored that there are still many challenges in this field. For example, the ambiguity and complexity of Chinese put higher requirements on the understanding ability of LLM; general LLM lacks professional knowledge in specific fields and needs to further strengthen the integration of domain knowledge; when dealing with sensitive semi-structured documents, data security and privacy protection issues are crucial; there is also a great need to improve the robustness and generalization ability of the model for semi-structured documents in different formats and layouts.

[0004] Analysis of foreign research status: Foreign countries started early in the field of semi-structured document parsing based on large language models, with deep technical accumulation and mature development. In the academic community, there is continuous innovation in pre-trained model architectures and training methods, such as improving the Transformer architecture and applying self-supervised learning techniques; researching the use of large language models (LLM) to solve complex semi-structured information extraction tasks, like relation extraction, event extraction, co-reference resolution, etc.; exploring multimodal fusion, integrating text with other modal information such as images and tables to improve the understanding ability of rich media semi-structured documents; paying attention to the interpretability and controllability of LLM in semi-structured document parsing, and at the same time, a large number of open-source tools and frameworks have emerged, such as Hugging Face Transformers, spaCy, Stanford CoreNLP, etc., to provide support for researchers and developers. In the industrial community, technology giants such as Google, Microsoft, and OpenAI have invested a large amount of resources and launched commercial products and APIs such as Google Document AI, Microsoft Form Recognizer, and OpenAI API; this technology is widely used in vertical fields such as finance, insurance, law, and healthcare, giving birth to innovative application scenarios; it is also integrated into enterprise process automation (RPA) and intelligent automation (IA) platforms to achieve end-to-end automated processing of semi-structured data; using LLM technology to extract knowledge from a large number of semi-structured documents to build knowledge graphs to support enterprise data governance and knowledge management; some industry organizations and standard institutions have started to formulate relevant standards and specifications to promote the standardized development of the technology. Internationally, the development of this technology shows trends such as from pipeline-based methods to end-to-end models, directly inputting semi-structured documents into LLM to output structured data; instruction tuning (Instruction Tuning) of LLM through a large amount of instruction data; chain-of-thought prompting to guide LLM to reason step by step; and combining the powerful understanding ability of LLM with traditional information extraction methods.

[0005] Based on the current research status at home and abroad, the key technologies for efficient parsing and converting semi-structured documents into structured documents based on large language models include large language models as the core engine, which are responsible for understanding text semantics, etc.; natural language processing technology, including named entity recognition (NER) and other technologies; prompt engineering for designing effective prompts; few-shot / zero-shot learning using a small amount of or no labeled data; knowledge graphs for organizing structured information; multimodal fusion technology for processing multimodal information; model optimization and deployment technology to improve model efficiency and deployability. In the future, its development trend will mainly be reflected in the development of more powerful general models, more refined domain models, more intelligent parsing processes, improved robustness and credibility, and more convenient development and deployment tool platforms, and expansion to more industries and application scenarios, such as intelligent customer service, content creation, knowledge management, etc.

[0006] Current challenges: In summary, although researchers at home and abroad have made some progress in the field of efficient semi-structured document parsing and structured document conversion based on large language models, the field is still facing the following three technical challenges: Problem 1: Diversity and complexity of semi-structured data processing. Semi-structured files come in various forms, including emails, web pages, log files, contracts, etc., and the formats, styles, and nested structures of these files are different. Large language models need to have the ability to understand and generalize semi-structured data of different types and fields. However, due to the lack of unified structure and labels, these data may contain inconsistent naming, non-standard expressions, and implicit contextual relationships. How to enable the model to accurately parse and extract useful information and handle the heterogeneity and complex nested relationships of data is an important technical challenge. Problem 2: Computational efficiency and scalability of the model. Large language models usually have a large parameter scale, resulting in high consumption of computing resources and slow inference speed. This becomes a significant bottleneck when processing a large number of semi-structured files, requiring real-time parsing, or deploying in resource-constrained environments. Improving the computational efficiency and scalability of the model, such as through model compression, knowledge distillation, optimization algorithms and hardware acceleration, so that it can efficiently process large-scale data, is a key technical issue currently faced. Question 3: Accuracy and reliability of information extraction. In the process of converting semi-structured files into structured formats, it is crucial to ensure that the extracted information is accurate and complete. Large language models sometimes generate inaccurate or unreliable content, resulting in the so-called "hallucination" phenomenon, which leads to errors or deviations in the extraction results. This may have serious consequences for downstream applications that rely on high-precision data. Therefore, how to improve the accuracy of the model in information extraction tasks, avoid errors or omissions, and ensure the reliability and credibility of the conversion process is a technical challenge that needs to be solved urgently. Summary of the invention

[0007] The objective of the present invention is to address the above problems existing in the prior art and propose a semi-structured document processing method based on large language models (LLMs).

[0008] The objective of the present invention can be achieved through the following technical solutions: A semi-structured document processing method based on large language models (LLMs) for efficiently parsing semi-structured documents and converting them into structured documents is designed as follows:

[0009] Multi-modal fusion parsing: The hierarchical cross-modal attention mechanism (HCMA) is adopted, combining the multi-head attention mechanism and the graph neural network (GNN) to achieve multi-level feature fusion of text and images. The specific design is as follows: The extraction of text features uses the improved Transformer encoder, Longformer, which captures global context information by introducing a dynamic sliding window mechanism to avoid exponential growth in computational complexity. The extraction of image features uses a hybrid architecture of Vision Transformer (ViT) and convolutional neural network (CNN). HCMA is divided into two stages: low-level fusion and high-level fusion. In the low-level fusion stage, the word segmentation, part-of-speech features of text, and edge and texture features of images are interacted through the multi-head attention mechanism. In the high-level fusion stage, the semantic entities of text and the semantic regions of images are constructed into a multi-modal graph structure, and the graph attention network (GAT) model is used to model the relationships between entities to generate unified semantics.

[0010] Adaptive structured conversion: The dynamic structure parsing network (DPN) is adopted, combining reinforcement learning and generative models to achieve adaptive parsing of document structures. The specific design is as follows: The parsing strategy is dynamically selected through a policy network. The parsing process is divided into two stages: coarse-grained and fine-grained. Coarse-grained parsing uses the pre-trained layout analysis model, LayoutLMv3, to extract the overall document structure. Fine-grained parsing combines a rule extractor and a deep learning model, BERT-based NER, to process detailed information, and a diffusion model is introduced to improve the robustness of parsing.

[0011] Generative model and multi-modal reasoning: The multi-modal generative reasoning network (MGRN) is adopted, combining the diffusion model and neuro-symbolic reasoning to achieve joint generation and reasoning of text and images. The specific design is as follows: MGRN consists of a text generator and an image generator, which jointly generate images consistent with the text description through joint training. A multi-modal knowledge graph (MKG) is constructed for neuro-symbolic reasoning, and the understanding ability of multi-modal data is improved by designing multi-modal contrast learning tasks and self-supervised tasks.

[0012] In multi-modal fusion analysis, the process of text feature extraction is as follows: The text is first converted into a high-dimensional vector representation through a word embedding layer, and positional encoding is added to preserve sequence information. Subsequently, through a multi-layer self-attention mechanism, the weights of each word are dynamically adjusted to capture the key semantic information in the text. To further enhance the expressive power of text features, a context-aware word vector optimization technique is introduced, and the semantics of words in different contexts are enhanced through a bidirectional attention mechanism Bidirectional Attention; the bidirectional attention mechanism calculates the forward and backward attention weights through two independent attention heads respectively, and the outputs of the two attention heads are fused through weighted summation to generate bidirectional context-aware word vectors.

[0013] In multi-modal fusion analysis, the design for image feature extraction is as follows: ViT captures the global semantic information of the image by dividing the image into multiple image patches and inputting them into the Transformer encoder; CNN uses EfficientNet as the backbone network and extracts local details through multi-scale convolution; a Feature Pyramid Network FPN is introduced to perform multi-scale fusion of the global feature information captured by ViT and the local feature information extracted by CNN; an adaptive pooling layer is introduced to dynamically adjust the size of the feature map according to the resolution of the input image.

[0014] In multi-modal fusion analysis, the design for the low-level fusion stage is as follows: The text features are used as queries, and the image features are used as key-value pairs. By calculating the attention weights, the fusion ratio of the text and image features is dynamically adjusted. The calculation formula is as follows:

[0015]

[0016] This formula is used to calculate the attention weights between text features (as query Q) and image features (as key-value pairs K, V). This formula can dynamically adjust the fusion ratio of text and image features, thereby capturing the local associations between them. d k is the dimension of the key vector, which plays a role in standardization to prevent the numerical value from being too large during the calculation process. The function then converts the calculation result into a probability distribution so that the sum of the weights is 1, thus determining the relative importance of different features during fusion.

[0017] In multi-modal fusion analysis, the design for the high-level fusion stage is as follows: The input of GAT is a multi-modal graph structure. The feature vector of each node calculates the weights of its neighbor nodes through an attention mechanism. For node i and its neighbor node j, the calculation formula for the attention weight is:

[0018]

[0019] This formula is used to calculate the attention weight between node i and its neighbor node j in the multi-modal graph structure at the high-level fusion stage of multi-modal fusion analysis. h i and h j are the feature vectors of nodes i and j respectively. W is a learnable weight matrix used for linear transformation of node feature vectors, enabling the model to learn more appropriate feature representations from the data. a is the attention vector through which the correlation degree between different node feature vectors is calculated. || represents the concatenation operation, where h_i and h_j are concatenated and then calculated with a. N(i) is the neighbor set of node i, and the denominator is the sum over all neighbor nodes of node i. The calculated α_ij can reflect the relative importance of node j to node i.

[0020] The output feature of node i is:

[0021]

[0022] Among them, σ is the activation function ReLU, whose role is to introduce non-linearity into the model, enabling the model to learn more complex relationships. α_ij is the attention weight between the previously calculated node i and its neighbor nodes, and W·h_j is the feature vector of the neighbor node after transformation by the weight matrix. Summing over all neighbor nodes of node i and then passing through the activation function gives the final output feature of node i, which comprehensively considers the information of neighbor nodes and their correlation degree with node i.

[0023] For the feature fusion output in multi-modal fusion analysis, the following design is adopted: The fused feature is optimized by using residual connection and layer normalization methods and then output.

[0024] In the adaptive structured transformation, the dynamic policy is designed as follows: DPN evaluates the type of the input document through the Policy Network and dynamically invokes the corresponding parsing sub-module. The Policy Network is trained based on the Proximal Policy Optimization (PPO) algorithm of reinforcement learning, and the parsing policy is optimized through the reward mechanism. The objective function of the PPO algorithm is as follows:

[0025]

[0026] E t represents the expectation operation, and the subscript t indicates the expectation over time steps. Because in practical applications, the Policy Network will interact with the environment at different time steps, generating a series of data. This expectation is to average the relevant data at all time steps, which can comprehensively consider the situations of multiple time steps and make the optimization more stable.

[0027] r t(θ) is the policy ratio, which reflects the degree of change of the new policy relative to the old policy. If it equals 1, it means that the probabilities of the new policy and the old policy taking this action in this state are the same; if it is greater than 1, the new policy is more inclined to take this action; if it is less than 1, the opposite is true.

[0028] is the advantage function, which is used to measure the advantage degree of taking action at in state st compared to the average policy. If the value of the advantage function is positive, it means that taking the action can obtain better rewards than the average policy; if the value is negative, it means that the rewards are worse than the average policy. Through the advantage function, the policy network can know which actions are more conducive to obtaining high rewards, so as to optimize the policy.

[0029] clip is the clipping function, which limits the policy ratio within the interval [1 - ε, 1 + ε]. ε is a preset clipping parameter, which is a small positive number, such as 0.2. When the policy ratio is too large or too small, the clipping function can prevent the policy update from being too large and ensure the stability of training.

[0030] In adaptive structured transformation, in the coarse-grained parsing stage, in order to further improve the parsing accuracy, the Multi-Task Learning technology is introduced to learn the layout analysis and semantic understanding tasks; the Diffusion Models convert the original layout and content of the document into a structured hierarchical representation through a step-by-step denoising process. First, noise is added to the document, and then the structured information of the document is gradually restored through the reverse diffusion process.

[0031] In the generative model and multi-modal reasoning, the design is as follows: Text generator: Based on the T5 model, input multi-modal features, and generate text descriptions corresponding to the image content. By introducing a conditional diffusion process, the text generator dynamically adjusts the semantics during the generation process to ensure the consistency between the generated text and the image content;

[0032] Image generator: Based on the diffusion model, input text descriptions and multi-modal features, and generate images. Through multi-scale feature fusion, the image generator is used to capture the detailed information of the text description and generate images consistent with the text content;

[0033] When constructing the multi-modal knowledge graph MKG, the design is as follows: Represent the entities and relationships in the text and images as graph nodes and edges, and perform reasoning through the random walk graph reasoning algorithm.

[0034] In the method for efficient parsing of semi-structured documents and conversion of structured documents based on large language models (LLMs), the following are designed for model optimization and performance improvement: Model compression, quantization, and distributed computing technologies are adopted to optimize the calculation process. For ensuring the conversion quality, the following are designed: Through data verification and error correction mechanisms, the integrity and accuracy of structured data are ensured. During the conversion process, the data is verified in multiple rounds to ensure data consistency and correctness, and data loss or errors are avoided.

[0035] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0036] 1. Improve the intelligent level of semi-structured document processing: By proposing a method for parsing semi-structured documents based on large language models (LLMs), the present invention significantly improves the automation and intelligent level of document processing. Traditional methods for parsing semi-structured documents usually rely on manually written rules and templates, lacking flexibility and universality. However, the multi-modal fusion parsing and adaptive structured conversion method proposed by the present invention can effectively handle complex document formats and information extraction tasks.

[0037] 2. Promote the development of data management and applications: The present invention can efficiently process semi-structured data (such as contracts, invoices, reports, etc.), providing more efficient and accurate support for data management. This is of great significance for various enterprises and Internet companies in information collation, data analysis, and intelligent application development. Through automated data processing, enterprises can more quickly extract valuable information from massive data, thereby promoting the development of innovative applications.

[0038] 3. Improve the accuracy and robustness of document understanding: By introducing multi-modal feature fusion and hierarchical cross-modal attention mechanism (HCMA), the present invention can achieve precise fusion between multiple information modalities (such as text, images, tables, etc.). This can not only capture the local relationships between text and images but also significantly improve the accuracy and robustness of document parsing through graph neural networks (GNNs) and global semantic modeling. Especially when dealing with complex and heterogeneous documents, the system can maintain high efficiency and accuracy.

[0039] 4. Promote cross-modal intelligent reasoning and generation: The multi-modal generation reasoning network (MGRN) of the present invention combines diffusion models with neuro-symbolic reasoning, enabling joint generation and complex reasoning of text and images. While improving document parsing, this method also provides strong data support for downstream applications such as intelligent question answering, reasoning, and decision-making support. Through accurate text generation and reasoning capabilities, the system can provide high-quality intelligent decision-making assistance.

[0040] 5. Promote the construction and application of knowledge graphs: Through structured file conversion and information extraction, the present invention can effectively connect and expand the parsed data with existing knowledge graphs. This provides a technical foundation for building higher-quality and wider-coverage knowledge graphs, which can further support more accurate intelligent applications, such as intelligent question answering, recommendation systems, knowledge reasoning, etc.

[0041] 6. Facilitate the realization of efficient computing and scalability: By optimizing the model's computing efficiency and scalability, the present invention solves the performance bottleneck in large-scale data processing. For example, technologies such as model compression, knowledge distillation, and hardware acceleration are used to ensure that large language models can efficiently process large-scale semi-structured files and meet the requirements in real-time parsing and resource-constrained environments.

[0042] 7. Ensure the accuracy and reliability of information extraction: During the information extraction process, the invention designs an adaptive structure parsing network (DPN) by combining reinforcement learning and generative models, effectively ensuring the accuracy and reliability of the extracted information. Especially in avoiding the generation of inaccurate or unreliable content, the system can provide more stable and high-precision results, ensuring the credibility of the conversion process and avoiding impacts on downstream applications.

[0043] In summary, the present invention takes multi-modal fusion parsing, adaptive structure parsing, generative models, and multi-modal reasoning as innovation points to achieve efficient parsing of semi-structured files and structured file conversion based on large language models. The present invention not only brings significant technological breakthroughs in document parsing, information extraction, and intelligent reasoning, but also reduces the dependence on manually labeled data by improving the flexibility and robustness of the model, promoting innovative development in the fields of data management and application. Especially when dealing with complex and heterogeneous semi-structured data, the method of the invention has strong universality and efficiency, providing strong support for data-driven innovation in various industries. Detailed implementation manners

[0044] The following are specific embodiments of the present invention to further describe the technical solutions of the present invention, but the present invention is not limited to these embodiments.

[0045] Objective: To develop an efficient method for parsing semi-structured files and converting structured files based on large language models (LLMs). With the rapid development of information technology, semi-structured data widely exists in various application scenarios, such as Internet data, log files, configuration files, etc. Efficient processing of these semi-structured data is crucial for data management, data analysis, and the development of intelligent applications.

[0046] The design of this method mainly includes the following aspects:

[0047] (1) Multi-modal fusion parsing

[0048] A hierarchical cross-modal attention mechanism (HCMA) is proposed, which combines the multi-head attention mechanism with the graph neural network (GNN) to achieve multi-level feature fusion of text and images. Text features are extracted through an improved Transformer encoder (such as Longformer), and image features are extracted through a hybrid architecture of the vision Transformer (ViT) and the convolutional neural network (CNN). The feature pyramid network (FPN) is used to fuse global and local features to generate a unified multi-modal representation.

[0049] (2) Adaptive structured transformation

[0050] A dynamic structure parsing network (DPN) is designed, which combines reinforcement learning with a generative model to achieve adaptive parsing of document structures. The parsing strategy is dynamically selected through a policy network. The parsing process is divided into two stages: coarse-grained parsing and fine-grained parsing. Coarse-grained parsing uses a pre-trained layout analysis model (LayoutLMv3) to extract the overall structure of the document, and fine-grained parsing combines a rule extractor with a deep learning model (BERT-based NER) to process detailed information. A diffusion model is introduced to generate a structured representation of non-standardized documents, improving the robustness of parsing.

[0051] (3) Generative model and multi-modal reasoning

[0052] A multi-modal generative reasoning network (MGRN) is proposed, which combines a diffusion model with neuro-symbolic reasoning to achieve joint generation and complex reasoning of text and images. The text generator is based on the T5 model, and the image generator is based on the diffusion model. Images consistent with the text description are generated through joint training. A multi-modal knowledge graph (MKG) is constructed for complex neuro-symbolic reasoning, and the model's understanding ability of multi-modal data is improved through self-supervised learning and contrastive learning tasks.

[0053] (4) Model optimization and performance improvement

[0054] Technologies such as model compression, quantization, and distributed computing are studied to optimize the model's computational process, reduce resource consumption, improve the inference speed, and ensure that the system can efficiently process large-scale semi-structured data to meet the requirements of real-time or near-real-time data processing.

[0055] Through the above research content, the intelligent level of semi-structured file processing will be significantly improved, promoting the development of data management and applications in related fields.

[0056] (1)Efficient and accurate parsing: Achieve efficient and accurate parsing of various complex semi-structured files. By optimizing parsing algorithms and model performance, significantly improve data processing efficiency. It is expected that when processing large-scale semi-structured data, the parsing speed can be significantly improved compared to traditional methods, and at the same time, the parsing accuracy can be significantly improved to meet the requirements of real-time or near-real-time data processing.

[0057] (2)General structured conversion: Develop a highly general structured conversion method. By designing flexible data mapping and recombination strategies, enable it to adapt to application scenarios in different fields and requirements. Whether it is transaction data in the financial field, medical record data in the medical field, or equipment monitoring data in the industrial field, etc., can all be effectively structured processed through this conversion method, providing convenience for subsequent data analysis and applications.

[0058] (3)Optimize computing performance: Optimize the computing performance of the large language model during parsing and conversion. By model optimization techniques, reduce resource consumption. On the premise of maintaining model accuracy, reduce the computing resource consumption of the model, achieve real-time or near-real-time data processing capabilities, reduce data processing latency, and improve the system's response speed.

[0059] (4)Guarantee conversion quality: Guarantee the quality of the conversion results. Through data verification and error correction mechanisms, ensure the integrity and accuracy of structured data. During the conversion process, conduct multiple rounds of verification on the data to ensure data consistency and correctness, avoid data loss or errors, and meet the strict requirements for data quality in subsequent data analysis and applications.

[0060] Key scientific problems to be solved: For the structured parsing of files by large language models, the following three scientific problems will be mainly solved:

[0061] Design a general parsing framework: How to design a general parsing framework that adapts to diverse semi-structured file formats to ensure the robustness and adaptability of the model under different data structures. This requires in-depth study of the characteristics and commonalities of different semi-structured file formats, and combined with the characteristics of large language models, design a parsing framework that can automatically learn and adapt to different data structures.

[0062] Improve model efficiency: How to improve the efficiency of large language models in large-scale data processing and solve the problems of high computing resource consumption and slow processing speed. By researching technologies such as model compression, quantization, and distributed computing, optimize the computing process of the model, reduce the demand for computing resources, and at the same time improve the inference speed of the model.

[0063] Ensuring Information Integrity and Accuracy: How to ensure the integrity and accuracy of information during the parsing and conversion processes and avoid data loss or errors. This requires designing effective data verification and error correction mechanisms in all aspects of data preprocessing, parsing, and conversion.

[0064] By solving the above key scientific problems, this method will significantly improve the intelligent level of semi-structured document processing, promote the development of data management and applications in related fields, and provide strong support for data-driven innovative applications.

[0065] The specific design plan to be adopted is as follows: Centering around multi-modal fusion parsing, an adaptive structure parsing module, and a generative model and multi-modal reasoning, it realizes the efficient parsing of semi-structured documents based on large language models and the conversion of structured documents.

[0066] Multi-modal Fusion Parsing: Multi-modal feature fusion is the core part of the model. Its goal is to achieve the organic unity of text and image features through deep interaction. This method proposes a hierarchical cross-modal attention mechanism (HCMA), which combines multi-head attention and graph neural network (GNN) to achieve multi-level feature fusion. This mechanism can not only capture the local correlations between text and images but also model the global semantic relationships through the graph structure, significantly improving the accuracy and robustness of feature fusion.

[0067] (1) Text Feature Extraction: The extraction of text features uses an improved Transformer encoder, Longformer. By introducing a dynamic sliding window mechanism, it can capture global context information when processing long texts while avoiding exponential growth in computational complexity. Specifically, the text is first converted into a high-dimensional vector representation through a word embedding layer, and positional encoding is added to preserve sequence information. Subsequently, through multiple layers of self-attention mechanisms, the model can dynamically adjust the weights of each vocabulary to capture the key semantic information in the text. To further improve the expressive power of text features, this solution also introduces context-aware word vector optimization technology, enhancing the semantic representation of vocabulary in different contexts through a bidirectional attention mechanism. The bidirectional attention mechanism calculates the forward and backward attention weights through two independent attention heads respectively, and the outputs of the two attention heads are fused through weighted summation (weights are learnable) to generate a bidirectional context-aware word vector representation.

[0068] (2) Image feature extraction: The extraction of image features adopts a hybrid architecture of Vision Transformer (ViT) and Convolutional Neural Network (CNN). ViT captures the global semantic information of the image by dividing the image into multiple patches (32x32) and inputting them into the Transformer encoder; while CNN extracts local details (such as edges and textures) through multi-scale convolutions. To fuse these two types of features, this solution introduces a Feature Pyramid Network (FPN) to perform multi-scale fusion of the global features of ViT and the local features of CNN, generating rich visual representations. In addition, by introducing an adaptive pooling layer, the model can dynamically adjust the size of the feature map according to the resolution of the input image, ensuring the flexibility of feature extraction.

[0069] Specifically, the input of ViT is image patches. Each patch is converted into a vector representation after linear projection, and positional encoding is added to retain spatial information. The Transformer encoder of ViT consists of multiple layers of self-attention mechanisms and feed-forward neural networks. The number of attention heads in each layer is usually set to 12. The CNN part uses EfficientNet as the backbone network to extract local features through multi-scale convolutions. FPN fuses the global features of ViT and the local features of CNN to generate multi-scale visual representations (usually 4 scales: 1 / 4, 1 / 8, 1 / 16, 1 / 32 resolution)

[0070] (3) Hierarchical cross-modal attention mechanism: HCMA is divided into two levels: low-level fusion and high-level fusion. Low-level fusion: In the low-level fusion stage, the word segmentation, part-of-speech features of the text and the edge and texture features of the image interact through the multi-head attention mechanism. Specifically, the text features are used as queries (Query), and the image features are used as key-value pairs (Key-Value). By calculating the attention weights, the fusion ratio of the text and image features is dynamically adjusted. This mechanism can effectively capture the local correlations between the text and the image. The calculation formula of multi-head attention is as follows:

[0071]

[0072] This formula is used to calculate the attention weights between the text features (as query Q) and the image features (as key-value pairs K, V). This formula can dynamically adjust the fusion ratio of the text and image features, and then capture the local correlations between them. d k is the dimension of the key vector, which plays a role in normalization to prevent the numerical value from being too large during the calculation process. The function then converts the calculation result into a probability distribution so that the sum of the weights is 1, thus determining the relative importance of different features during fusion.

[0073] High-level fusion: In the high-level fusion stage, the semantic entities of the text (such as the results of named entity recognition) and the semantic regions of the image (such as the results of object detection) are constructed into a multi-modal graph structure. Using the Graph Attention Network (GAT), the model can model the relationships between entities and generate a unified semantic representation. The input of GAT is the multi-modal graph structure, and the feature vector of each node calculates the weights of its neighbor nodes through the attention mechanism. Specifically, for node i and its neighbor node j, the calculation formula for the attention weight is:

[0074]

[0075] This formula is used to calculate the attention weight between node i and its neighbor node j in the multi-modal graph structure during the high-level fusion stage of multi-modal fusion analysis. h i and h j are the feature vectors of nodes i and j respectively. W is a learnable weight matrix used to perform a linear transformation on the node feature vectors, enabling the model to learn more appropriate feature representations according to the data. a is the attention vector, through which the correlation degree between different node feature vectors is calculated. || represents the concatenation operation, concatenating h

[0076]

[0077] and then calculating with a. N(j) is the set of neighbors of node i, and the denominator is the sum over all neighbor nodes of node i. The calculated α

[0078] (4) Feature fusion output

[0079] The fused features are optimized through residual connections and layer normalization (Layer Normalization), and a unified multi-modal representation is output. This representation contains both the semantic information of the text and the visual information of the image, providing high-quality input for subsequent tasks.

[0080] Adaptive Structured Transformation: The adaptive structure parsing module aims to dynamically parse the complex layout and content of semi-structured documents. A Dynamic Parsing Network (DPN) is proposed, which combines reinforcement learning and generative models to achieve adaptive parsing of document structures. This module can not only dynamically adjust the parsing strategy according to the document type, but also process non-standard document structures through the generative model, significantly improving the flexibility and accuracy of parsing.

[0081] (1) Dynamic Policy Selection: DPN evaluates the type of the input document (such as tables, reports, contracts, etc.) through a policy network and dynamically invokes the corresponding parsing sub-module. The policy network is trained based on reinforcement learning (such as the PPO algorithm), and the parsing strategy is optimized through a reward mechanism. For example, for table documents, the policy network will select a table parser; for report documents, it will select a paragraph parser. The objective function of the PPO algorithm is:

[0082]

[0083] E t denotes the expectation operation, and the subscript indicates the expectation over time steps. In practical applications, the policy network interacts with the environment at different time steps, generating a series of data. Here, the expectation is to average the relevant data over all time steps, which can comprehensively consider the situations of multiple time steps and make the optimization more stable.

[0084] r t (θ) is the policy ratio, which reflects the change degree of the new policy relative to the old policy. If it equals 1, it means that the new policy and the old policy have the same probability of taking this action in this state; if it is greater than 1, the new policy is more inclined to take this action; if it is less than 1, the opposite is true.

[0085] is the advantage function, which is used to measure the advantage degree of taking action at in state st compared to the average policy. If the value of the advantage function is positive, it means that taking this action can obtain better benefits than the average policy; if the value is negative, it means that the benefits are worse than the average policy. Through the advantage function, the policy network can know which actions are more conducive to obtaining high rewards, so as to optimize the policy.

[0086] clip is the clipping function, which limits the policy ratio within the interval [1 - ε, 1 + ε]. ε is a pre-set clipping parameter, which is a small positive number, such as 0.2. When the policy ratio is too large or too small, the clipping function can prevent the policy update from being too large and ensure the stability of training.

[0087] (2) Multi-granularity parsing: The parsing process is divided into two stages: coarse-grained parsing and fine-grained parsing. Coarse-grained parsing: In the coarse-grained parsing stage, the model extracts the layout information of the document through the pre-trained layout analysis model LayoutLMv3, and generates the overall structure of the document (such as chapters, paragraphs, headings) in combination with the text semantic features. To further improve the parsing accuracy, this solution also introduces the multi-task learning (Multi-Task Learning) technology, enabling the model to learn layout analysis and semantic understanding tasks simultaneously. Fine-grained parsing: In the fine-grained parsing stage, the model parses detailed information such as sentences in paragraphs and cells in tables. The combination of a rule-based extractor (regular expression) and a deep learning-based method (BERT-based NER) is used to ensure the parsing accuracy. For example, for tabular data, the model extracts financial data by matching the form of "number + unit" and infers its semantic meaning in combination with the context information.

[0088] (3) Structure parsing based on generative models: To handle complex and non-standard document structures, this solution introduces diffusion models to generate the structured representation of the document. The diffusion model transforms the original layout and content of the document into a structured hierarchical representation through a step-by-step denoising process. Specifically, the model first adds noise to the document, and then gradually recovers the structured information of the document through the reverse diffusion process. This method can not only handle diverse document layouts, but also capture the implicit structure in the document through the generation process, significantly improving the robustness of the parsing.

[0089] Generative Model and Multimodal Reasoning: The generative model and multimodal reasoning module aims to achieve the joint generation and complex reasoning of text and images. This solution proposes a Multimodal Generative Reasoning Network (MGRN), which combines the diffusion model and neuro-symbolic reasoning to achieve efficient multimodal generation and reasoning. This module can not only generate high-quality text and image content, but also perform complex reasoning through knowledge graphs, significantly enhancing the application value of the model. (1) Multimodal Generative Model: MGRN consists of a text generator and an image generator, which are synchronously generated through joint training. Text Generator: Based on the T5 model, it inputs multimodal features and generates text descriptions corresponding to the image content. By introducing a conditional diffusion process, the text generator can dynamically adjust semantic representations during generation to ensure the consistency between the generated text and the image content. Image Generator: Based on the diffusion model, it inputs text descriptions and multimodal features to generate high-quality images. Through multi-scale feature fusion, the image generator can capture the detailed information of the text description and generate images highly consistent with the text content. (2) Neuro-Symbolic Reasoning Engine: A Multimodal Knowledge Graph (MKG) is constructed to represent entities and relationships in text and images as graph nodes and edges. Complex reasoning is performed through graph reasoning algorithms (random walk). For example, given the text description "The price of product A is 100 yuan" and product A in the image, the reasoning engine can automatically infer the corresponding relationship between "The price of product A" and the object in the image.

[0090] (3) Self-Supervised Learning and Contrastive Learning: By designing multimodal contrastive learning tasks (such as text-image alignment) and self-supervised tasks (such as image completion, text mask prediction), the model's understanding ability of multimodal data is further improved. For example, in the contrastive learning task, the model needs to match the text description with the corresponding image to learn the semantic association between text and image.

[0091] Significance: In the fields of information science and artificial intelligence, how to efficiently extract valuable information from massive unstructured and semi-structured data has always been a challenging and significant research topic. Especially in scenarios such as the Internet and enterprise internal systems, a large amount of key information exists in semi-structured forms such as contracts, reports, invoices, and technical documents. Traditional parsing methods based on rules, template matching, or shallow machine learning show significant defects such as low efficiency, weak generalization ability, and high maintenance costs when dealing with complex, heterogeneous semi-structured documents.

[0092] From the perspective of the development trend of scientific research, large language models (LLMs) centered around the Transformer architecture are profoundly revolutionizing the field of natural language processing (NLP) with their powerful context modeling capabilities, emerging zero-shot / few-shot learning capabilities, and profound understanding of natural language. They also demonstrate great potential in complex document understanding and information extraction tasks. The research of this method closely aligns with this cutting-edge trend, aiming to deeply explore how to utilize the internal mechanisms and knowledge representation capabilities of LLMs to build a new generation of efficient and robust semi-structured document parsing systems, breaking through the technical bottlenecks of traditional methods, which will contribute to the realization of the following scientific significance:

[0093] (1) Deepen the theoretical understanding of LLMs in the multi-modal understanding of complex documents: This method will deeply study how LLMs effectively model various information modalities such as intertwined text content, layout, and table structures in semi-structured documents, explore the mechanism of their internal attention mechanism in cross-modal information fusion, and reveal the advantages and limitations of LLMs when processing documents with complex spatial relationships and semantic dependencies, providing theoretical support for more effectively using LLMs for intelligent document processing.

[0094] (2) Explore a new end-to-end semi-structured information extraction paradigm based on LLMs: Different from the traditional approach that relies on manual feature engineering and multi-stage pipelines, this method will study how to utilize the powerful text generation and conditional generation capabilities of LLMs to transform the semi-structured document parsing task into an end-to-end sequence generation or structured prediction problem. For example, explore using LLMs to directly generate target structured data (such as JSON, relational database records), or using LLMs for prompt-based structured information extraction, thereby simplifying the system architecture, reducing development and maintenance costs, and enhancing the flexibility and adaptability of the system.

[0095] (3) Build a more generalizable and robust semi-structured data parsing method: Traditional methods rely on hard-coding for specific document structures, making it difficult for them to adapt to new document formats and domains. This method will study how to utilize the powerful knowledge transfer and context learning capabilities of LLMs to improve the zero-shot or few-shot adaptation ability of the parsing system to semi-structured documents in different formats and domains. For example, explore using pre-trained LLMs combined with a small number of target domain samples for fine-tuning, or using prompt engineering to guide LLMs to understand and parse new document structures, thereby improving the accuracy and efficiency of information extraction and reducing the dependence on large-scale labeled data.

[0096] (4) Promote the progress of basic technologies for knowledge graph construction and intelligent applications: Effectively integrating the parsed structured data into the knowledge graph is the key to achieving deeper knowledge mining and intelligent applications. This method will study how to utilize the semantic understanding ability of LLMs for entity linking, relation extraction, and knowledge fusion, effectively connecting and expanding the information extracted from semi-structured documents with the existing knowledge graph, thereby improving the quality and coverage of the knowledge graph and providing a more reliable data foundation for downstream applications such as intelligent question answering, reasoning, and decision support. At the same time, it will also explore using LLMs for natural language description and generation of structured data to achieve two-way information conversion and promote the intelligent development of human-computer interaction.

[0097] The results of this research will directly serve China's digital transformation strategy, provide key technical support for the informatization construction of key fields such as finance, government affairs, and healthcare, and help enhance China's core competitiveness in the field of artificial intelligence.

[0098] The specific embodiments described in this article are merely illustrative of the spirit of the present invention. Those skilled in the art to which the present invention pertains can make various modifications or supplements to the described specific embodiments or use similar methods for substitution, but will not deviate from the spirit of the present invention or exceed the scope defined by the appended claims.

Claims

1. A method for processing semi-structured files based on the LLMs large language model, characterized in that the method for processing semi-structured files based on the LLMs large language model is used to parse semi-structured files and convert them into structured files, and is designed as follows: Multimodal fusion analysis: Adopting the hierarchical cross-modal attention mechanism HCMA, combined with the multi-head attention mechanism and the graph neural network GNN, the multi-level feature fusion of text and image is realized; the specific design is as follows: the text feature extraction adopts the improved Transformer encoder Longformer, and by introducing the dynamic sliding window mechanism, the global context information is captured to avoid the exponential growth of computational complexity; the image feature extraction adopts the hybrid architecture of ViT and convolutional neural network CNN; HCMA is divided into two stages: low-level fusion and high-level fusion. In the low-level fusion stage, the word segmentation and part-of-speech features of the text interact with the edge and texture features of the image through the multi-head attention mechanism; in the high-level fusion stage, the semantic entities of the text and the semantic regions of the image are constructed as a multimodal graph structure, and the relationship between entities is modeled using the graph attention network GAT model to generate unified semantics; Adaptive structured conversion: Adopting the dynamic structure parsing network DPN, combined with reinforcement learning and generative models, to achieve adaptive parsing of document structure; the specific design is as follows: the parsing strategy is dynamically selected through the policy network, and the parsing process is divided into two stages: coarse-grained and fine-grained. The coarse-grained parsing uses the pre-trained layout analysis model LayoutLMv3 to extract the overall structure of the document, and the fine-grained parsing combines the rule extractor and the deep learning model BERT-based NER to process detailed information, and introduces the diffusion model Diffusion Models to improve the robustness of the parsing; Generative model and multimodal reasoning: The multimodal generative reasoning network MGRN is used, combined with the diffusion model and neural symbolic reasoning, to realize the joint generation and reasoning of text and images; the specific design is as follows: MGRN consists of a text generator and an image generator, which generates images consistent with the text description through joint training; a multimodal knowledge graph MKG is constructed to perform neural symbolic reasoning, and the ability to understand multimodal data is improved by designing multimodal comparative learning tasks and self-supervision tasks.

2. The semi-structured document processing method based on the LLMs large language model as claimed in claim 1, characterized in that: In multimodal fusion parsing, the process of extracting text features is as follows: the text is first converted into a high-dimensional vector representation through a word embedding layer, and position encoding is added to retain sequence information. Subsequently, the weight of each word is dynamically adjusted through a multi-layer self-attention mechanism to capture the key semantic information in the text. In order to further improve the expressiveness of text features, context-aware word vector optimization technology is introduced, and the semantics of words in different contexts are enhanced through the bidirectional attention mechanism. The bidirectional attention mechanism calculates the forward and backward attention weights respectively through two independent attention heads, and fuses the outputs of the two attention heads through weighted summation to generate a bidirectional context-aware word vector.

3. The semi-structured document processing method based on the LLMs large language model as claimed in claim 1, characterized in that: In multimodal fusion analysis, the design for image feature extraction is as follows: ViT captures the global semantic information of the image by dividing the image into multiple image blocks and inputting them into the Transformer encoder; CNN uses EfficientNet as the backbone network to extract local details through multi-scale convolution; The feature pyramid network FPN is introduced to perform multi-scale fusion of the global feature information captured by ViT and the local feature information extracted by CNN; the adaptive pooling layer is introduced to dynamically adjust the size of the feature map according to the resolution of the input image.

4. The semi-structured document processing method based on the LLMs large language model as claimed in claim 1, characterized in that: In multimodal fusion analysis, the design for the low-level fusion stage is as follows: text features are used as queries, image features are used as key-value pairs, and the fusion ratio of text and image features is dynamically adjusted by calculating the attention weights. The calculation formula is as follows: Among them, d k is the dimension of the key vector.

5. The semi-structured document processing method based on the LLMs large language model as claimed in claim 1, characterized in that: In the multimodal fusion analysis, the high-level fusion stage is designed as follows: the input of GAT is a multimodal graph structure, and the feature vector of each node calculates the weight of its neighboring nodes through the attention mechanism. For node i and its neighbor node j, the calculation formula of the attention weight is: Among them, h i and h j is the node feature vector, W is the learnable weight matrix, a is the attention vector, and N(i) is the neighbor set of node i; The output features of node i are: Where σ is the activation function ReLU.

6. The semi-structured document processing method based on the LLMs large language model as claimed in claim 1, characterized in that: The feature fusion output design in multimodal fusion analysis is as follows: the fused features are optimized and output by using residual connection and layer normalization methods.

7. The semi-structured document processing method based on the LLMs large language model as claimed in claim 1, characterized in that: In adaptive structured conversion, the dynamic strategy is designed as follows: DPN evaluates the type of input documents through the policy network Policy Network and dynamically calls the corresponding parsing submodule. The policy network Policy Network is trained based on the reinforcement learning PPO algorithm and optimizes the parsing strategy through the reward mechanism. The objective function of the PPO algorithm is as follows: Among them, r t (θ) is the strategy ratio, is the advantage function, ∈ is the trimming parameter; by maximizing the objective function, the policy network learns the optimal parsing strategy.

8. The semi-structured document processing method based on the LLMs large language model as claimed in claim 1, characterized in that: In the adaptive structured conversion, in the coarse-grained parsing stage, in order to further improve the parsing accuracy, the Multi-Task Learning technology is introduced to learn the layout analysis and semantic understanding tasks; the Diffusion Models transform the original layout and content of the document into a structured hierarchical representation through a gradual denoising process. First, noise is added to the document, and then the structural information of the document is gradually restored through the reverse diffusion process.

9. The method for processing semi-structured documents based on the LLMs large language model as claimed in claim 1, characterized in that: In the generation model and multimodal reasoning, the design is as follows: Text generator: Based on the T5 model, multimodal features are input to generate text descriptions corresponding to the image content. By introducing the conditional diffusion process, the text generator dynamically adjusts the semantics during the generation process to ensure the consistency of the generated text and the image content; Image generator: Based on the diffusion model, the image generator takes text description and multimodal features as input and generates images. Through multi-scale feature fusion, the image generator is used to capture the details of the text description and generate images consistent with the text content. The design for constructing the multimodal knowledge graph MKG is as follows: entities and relationships in text and images are represented as graph nodes and edges, and reasoning is performed using the random walk graph reasoning algorithm.

10. The semi-structured document processing method based on the LLMs large language model as claimed in claim 1, characterized in that: In the efficient semi-structured file parsing and structured file conversion method based on the LLMs large language model, the design for model optimization and performance improvement is as follows: model compression, quantization and distributed computing technology are used to optimize the calculation process; the design for ensuring the conversion quality is as follows: through data verification and error correction mechanisms, the integrity and accuracy of structured data are ensured. During the conversion process, multiple rounds of data verification are performed to ensure data consistency and correctness, and to avoid data loss or errors.

Citation Information

Cited By

  • Method, device, medium, product and system for reasoning acceleration of pre-training model

    CN120408126A

  • Zero sample template inference and document structured recognition method and device

    CN120877300A

  • Event allocation method based on collaborative reasoning of multi-modal data and knowledge graph

    CN120950699A

  • Multi-modal parallel reasoning method, electronic equipment and program product

    CN121071824A

  • Signal text data generation and bimodal fusion continuous learning method and system under long-tail distribution

    CN121145936A