Multi-modal document structured processing method and device, equipment and medium

Through the multimodal document structured processing method, for documents containing images and tables, image-to-text technology and multimodal data processing large model are used for structured processing, which solves the problems of complicated processes, poor versatility and low understanding accuracy when processing complex documents in the prior art, and achieves more efficient and accurate document understanding.

CN120068810APending Publication Date: 2025-05-30BEIJING UNISOUND INFORMATION TECH CO LTD +7
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510126948.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-27
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Existing document understanding techniques face problems such as cumbersome processes, poor versatility and low understanding accuracy when processing complex documents containing images and tables.

Method used

A multimodal document structured processing method is proposed, which extracts original text and other data types (such as images and tables) by obtaining the to-processed documents, and processes them according to the structured processing steps corresponding to each data type. For images, distinguish between text images and non-text images, and use image to text technology or multimodal data processing large models for processing; for tables, obtain table headers and row names and generate descriptive text.

Benefits of technology

It realizes deep structured processing of image and tabular data, improves the accuracy and usability of information extraction, enhances the versatility and accuracy of document understanding, and provides richer input data to support downstream tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120068810A_ABST
    Figure CN120068810A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal document structured processing method and device, equipment and a medium. For other different data types extracted from the to-be-processed document, the data of the other data types can be subjected to structured processing according to the preset structured processing steps corresponding to the other data types, so that the information carried by the data of the other data types is deeply mined, and the processing efficiency of the to-be-processed document is improved. And the information carried by the data of other data types in the to-be-processed document is presented in a unified and structured form through effective integration and utilization of the multi-modal data. For the image data, the character image type and the non-character image type are distinguished, and different processing modes are adopted, so that the information carried in the image data is extracted more accurately. For the table data, the table data is converted into an information set with rich semantics from a simple numerical matrix by obtaining a header name and a line name and generating a descriptive text for each data item.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the fields of artificial intelligence and data processing, and in particular, to a multi-modal document structured processing method, apparatus, device, and medium. Background Art

[0002] In the application scenario of document understanding, the complexity and diversity of document content pose severe challenges to document understanding. In addition to text content, modern documents often also embed rich image and table information, which is crucial for comprehensively understanding document content in specific application scenarios. For example, in legal documents, scientific research reports, or financial reports, images may contain key character information (such as photocopied signatures, charts, or data graphics), while tables are often used to display structured data, such as statistical results, financial data comparisons, etc.

[0003] Currently, there are several limitations in the commonly used processing methods for document understanding that involve images and tables. For image content, a common approach is to save it in the form of a URL (Uniform Resource Locator) link during the document structuring process. When specific images need to be referenced in subsequent understanding steps, a separate QA (Question Answering) - style multi-modal processing flow needs to be initiated to associate and parse relevant images according to the user's specific questions. Although this process realizes the utilization of image information to a certain extent, due to the need to pre-determine the relevance between images and questions and the relatively independent processing steps, its generality and flexibility in the actual production environment are limited.

[0004] As for table content, common methods tend to convert it into a text sequence in formats such as Markdown (a lightweight markup language) or HTML (Hyper Text Markup Language), and then enter a unified RAG (Retrieval Augmented Generation) - style question - answering process together with other text content in the document. However, this processing method ignores the essential characteristics of tables as two - dimensional information carriers. Reducing two - dimensional table information to a one - dimensional text sequence not only loses the original spatial structure and data relevance of the table but also causes difficulties for large language models in understanding. Since large language models are essentially trained based on one - dimensional text sequences, their ability to understand two - dimensional structure information is relatively weak, which directly leads to a decrease in the accuracy of understanding table content.

[0005] In summary, existing document understanding technologies face problems such as cumbersome processes, poor generality, and low understanding accuracy when dealing with complex documents containing images and tables. Summary of the Invention

[0006] The present application provides a multi-modal document structuring processing method, apparatus, device and medium, which are used to solve the problems of cumbersome processes, poor versatility and low understanding accuracy rate faced by existing document understanding technologies when processing complex documents containing images and tables.

[0007] In a first aspect, the present application provides a multi-modal document structuring processing method, and the method includes:

[0008] Obtain a document to be processed;

[0009] Extract the original text and data of other data types from the document to be processed; wherein, the other data types include one or more of the following: images and tables;

[0010] Perform corresponding structuring processing on the data of each of the other data types according to the structuring processing steps respectively corresponding to each of the other data types;

[0011] Execute a downstream task based on the structured data and the original text;

[0012] Among them, the structuring processing step for any image is: determine whether the image belongs to the type of text image according to the content contained in the image; wherein, the text image type is an image that expresses the image content through the displayed text information; if it is determined that the image belongs to the text image type, then through the image-to-text technology, identify the text content from the image; otherwise, through a pre-trained multi-modal data processing large model, based on the image, obtain the content description of the image; wherein, the multi-modal data large model is a large model with the ability to recognize image content and natural language understanding ability;

[0013] The structuring processing step for any table is: obtain the header name and row name in the table; for any data item in the table, form a descriptive text of the data item through the value of the data item in the table and the corresponding header name and row name; arrange the respective descriptive text corresponding to the table according to the row order and column order.

[0014] In a second aspect, the present application further provides a multi-modal document structuring processing apparatus, and the apparatus includes:

[0015] An obtaining unit, configured to obtain a document to be processed;

[0016] An extracting unit, configured to extract the original text and data of other data types from the document to be processed; wherein, the other data types include one or more of the following: images and tables;

[0017] A processing unit for performing corresponding structured processing on the data of each of the other data types according to the structured processing steps respectively corresponding to each of the other data types; wherein, the structured processing step of any image is: determining whether the image belongs to the text image type according to the content included in the image; wherein, the text image type is an image that expresses the image content through the displayed text information; if it is determined that the image belongs to the text image type, then through the image-to-text technology, the text content is recognized from the image; otherwise, through a pre-trained multi-modal data processing large model, based on the image, the content description of the image is obtained; wherein, the multi-modal data large model is a large model with image content recognition ability and natural language understanding ability; the structured processing step of any table is: obtaining the header name and row name in the table; for any data item in the table, forming a descriptive text of the data item through the value of the data item in the table and the corresponding header name and row name; arranging the respective descriptive text corresponding to the table according to the row order and column order.

[0018] A task execution unit for performing downstream tasks based on the structured processed data and the original text.

[0019] In a third aspect, the present application provides a computer device, the computer device includes a processor, and the processor is used to implement the steps of the multi-modal document structured processing method as described above when executing a computer program stored in a memory.

[0020] In a fourth aspect, the present application provides a computer-readable storage medium, which stores a computer program, and the computer program realizes the steps of the multi-modal document structured processing method as described above when executed by a processor.

[0021] The beneficial effects of the present application are as follows:

[0022] 1. For different other data types extracted from the document to be processed, the data of the other data types can be structurally processed according to the preset structured processing steps corresponding to the other data types, so as to realize in-depth mining of the information carried by the data of the other data types, as well as the effective integration and utilization of multi-modal data, and present the information carried by the data of each other data type in the document to be processed in a unified and structured form. This provides richer and more accurate input for subsequent downstream tasks such as information retrieval, text understanding, and data analysis.

[0023] 2. For image data, by distinguishing between text image types and non-text image types and adopting different processing methods, the extraction of information carried in the image data becomes more accurate. For text images, the text information hidden in the image is converted into text content using image-to-text technology, greatly improving the usability of the text information. For non-text images, with the help of a pre-trained large multi-modal data processing model, the image content can be deeply understood and a detailed description can be generated, endowing the image with semantic information and helping to more comprehensively convey the document intention.

[0024] 3. For tabular data, by obtaining the header names, row names, and generating descriptive text for each data item, the tabular data is transformed from a simple numerical matrix into an information set with rich semantics. This not only improves the readability of the tabular data but also facilitates the analysis and comparison of the tabular data, enabling a clear understanding of the specific meaning represented by each data item and the relationships between different data items, thus providing stronger support for downstream tasks.

[0025] 4. Since traditional downstream tasks often rely only on the original text information, the depth and breadth of task processing are limited. However, this application combines the structured processed data with the original text, providing a more comprehensive and rich information source for downstream tasks and significantly improving the processing effect and efficiency of downstream tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] To more clearly illustrate the technical solutions in the embodiments of this application, the following briefly introduces the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of this application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0027] Figure 1 It is a schematic diagram of the process of a multi-modal document structuring process provided by an embodiment of this application;

[0028] Figure 2 It is a schematic diagram of the specific process of multi-modal document structuring provided by an embodiment of this application;

[0029] Figure 3 It is a schematic diagram of the specific process of structuring any image provided by an embodiment of this application;

[0030] Figure 4 It is a schematic diagram of the specific process of structuring any table provided by an embodiment of this application;

[0031] Figure 5 It is a schematic diagram of the structure of a multi-modal document structuring device provided by an embodiment of this application;

[0032] Figure 6 It is a schematic structural diagram of a computer device provided by an alternative embodiment of the present application. Detailed implementation manners

[0033] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings. Apparently, the described embodiments are only a part rather than all of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.

[0034] In order to effectively achieve the comprehensive understanding of image, table and text information in a document to improve the accuracy of subsequent tasks, the present application provides a multi-modal document structuring processing method, device, equipment and medium.

[0035] Embodiment 1:

[0036] The present application provides a multi-modal document structuring processing method. Figure 1 It is a schematic diagram of the process of multi-modal document structuring processing provided by an embodiment of the present application. The process includes:

[0037] S101: Obtain a document to be processed.

[0038] The multi-modal document structuring processing method provided by the present application is applied to a computer device. The computer device can be an intelligent device, such as a mobile terminal, a computer, etc., or a server, such as an application server, a business server, etc.

[0039] First, it is necessary to obtain the document to be processed. The document to be processed can be a document from various channels, such as a file uploaded by a user, a file downloaded from the network, a file stored in a local database, etc. The document to be processed can be in various formats, including but not limited to PDF, Word document, web page document, slide document, etc.

[0040] In a possible implementation manner, in order to facilitate subsequent data extraction, the document to be processed can be parsed into a format recognizable and processable by the computer device. For documents in different formats, corresponding parsing tools or libraries can be used. For example, for a PDF document, a PDF parsing library can be used to convert it into a collection of elements such as text and images; for a web page document, a web page parsing tool can be used to extract text, images and table elements.

[0041] S102: Extract the original text and data of other data types in the document to be processed; wherein, the other data types include one or more of the following: images and tables.

[0042] After obtaining the document to be processed, it is necessary to extract the original text and data of other data types therein. Among them, the original text is the part presented in pure text form in the document to be processed, and through text extraction algorithms or tools, it can be separated from the document. For data of other data types, corresponding element detection and extraction algorithms are required for extraction. Among them, other data types include but are not limited to: images and tables.

[0043] For example, for image extraction, through image processing technology, the image data is separated from the document to be processed.

[0044] In a possible implementation manner, the extracted images may be stored in different formats, such as JPEG, PNG, BMP, etc., and appropriate format conversion and preprocessing can be performed on them for subsequent processing.

[0045] Again, for example, for table extraction, through table recognition algorithms, the boundaries and structures of the tables are found from the document to be processed, and the table elements are extracted. The extracted tables are stored in a suitable data structure, such as a two-dimensional array or a custom table data structure, for subsequent operations.

[0046] S103: According to the structured processing steps corresponding to each of the other data types, perform corresponding structured processing on the data of each of the other data types; among them, the structured processing step for any image is: according to the content contained in the image, determine whether the image belongs to the type of text image; where the text image type is an image that expresses the image content through the displayed text information; if it is determined that the image belongs to the text image type, then through image-to-text technology, the text content is recognized from the image; otherwise, through a pre-trained large multi-modal data processing model, based on the image, obtain the content description of the image; where the large multi-modal data model is a large model with the ability to recognize image content and natural language understanding ability; the structured processing step for any table is: obtain the header name and row name in the table; for any data item in the table, form a descriptive text of the data item through the value of the data item and the corresponding header name and row name; arrange the respective descriptive text corresponding to the table according to the row order and column order.

[0047] For the different other data types extracted in the above embodiments, corresponding structured processing can be performed on the data of the other data type according to the preset structured processing steps corresponding to the other data type. The following is a detailed description of the structured processing steps for different data types:

[0048] 1. The other data type is an image.

[0049] When other data types are images, there will be various different types of images. These can be mainly subdivided into the following categories:

[0050] Text image type: This type covers scanned documents, tables inserted in the form of pictures, flowcharts, etc. Its characteristic is that the content of the image is mainly expressed through the displayed text information, that is, the text is the key information carrier of such images.

[0051] Pure image type: Typical examples are CAD drawings, topographic maps, etc. The information of such images is mainly presented through visual elements such as graphics, patterns, lines, colors, etc., and does not rely on text to express the core information.

[0052] Mixed type containing text and images: This type of image contains both important text information and significant image elements, and they combine with each other to jointly constitute the complete content of the image.

[0053] Based on this, in order to achieve effective structured processing of these different types of images, it is necessary to formulate differentiated processing strategies according to the specific content of the images.

[0054] Therefore, for any extracted image, the primary task is to determine the type to which it belongs. Exemplarily, for any extracted image, it can be determined whether the image belongs to the text image type. If the image belongs to the text image type, the text information in the image can be extracted using image-to-text technology (such as optical character recognition technology OCR). In this process, the image-to-text technology will comprehensively consider the layout characteristics of the text to ensure that the extracted text information conforms to the normal reading order and layout structure, and finally extract the text information in the image completely and accurately. If the image does not belong to the text image type, that is, the image can be a mixed-type image or a pure-image type image, then the content description of the image can be obtained by combining the content of the image through a pre-trained multi-modal data processing large model. In this process, to ensure the accuracy and comprehensiveness of the description, the multi-modal data processing model needs to be pre-trained on a large-scale image-text pair dataset so that the multi-modal data processing large model has the ability to recognize image content and natural language understanding ability. These datasets contain various types of images, covering different objects, scenes, colors, textures, and spatial layouts, and each image is accompanied by rich and accurate natural language descriptions. During the pre-training process, the multi-modal data processing model will learn the complex correspondence relationship between the image and the natural language description. Through a large amount of data training, it will adjust its own parameters to minimize the difference between the predicted natural language description and the true description. For example, during the training process, for a grassland image containing red flowers, the multi-modal data processing model will continuously learn how to generate description information such as "On a green grassland, several bright red flowers are in full bloom, and the flowers are gently swaying in the breeze" based on features such as the color, shape, position of the flowers, and the color and texture of the grassland, so that it can generate accurate description information based on various features of the image.

[0055] In a possible implementation manner, obtaining the content description of the image based on the image through the pre-trained multi-modal data processing large model includes:

[0056] Obtain the context information associated with the image from the document to be processed;

[0057] Through the multi-modal data processing large model, based on the image and the context information, obtain the content description of the image.

[0058] In order to improve the understanding of the image content, in the present application, context information associated with the image can be obtained from the document to be processed. Exemplarily, the text information around the image can be found as context information based on the position of the image in the document to be processed. For example, in the case where the image is located in the middle of a paragraph, a certain length of text before and after the image is extracted as context information. For example, if the image is located in the middle of a paragraph, N characters before and after the image in the paragraph where the image is located (N can be set according to the specific situation, such as 200 characters) can be extracted as context information. These adjacent texts may contain explanations, descriptions or related descriptions of the image, which can help the multimodal data processing model better understand the role and meaning of the image in the document to be processed.

[0059] After obtaining context information associated with the image, the image and the context information are input into a pre-trained multimodal data processing model to obtain a content description of the image.

[0060] In a possible implementation, it can be determined whether the image belongs to the text image type in the following manner: the image is converted into a grayscale image, and noise reduction processing is performed to remove noise interference to improve the accuracy of subsequent analysis. The image is denoised using methods such as Gaussian filtering. Then, an edge detection algorithm (such as Canny edge detection) is used to detect lines and contours in the image. For text image types, there are usually more regular lines and more horizontal or vertical edges, which may correspond to the strokes and typesetting structure of the text. Based on the edge information, by analyzing the distribution and density of the edges, it can be preliminarily determined whether the image may be a text image type.

[0061] In another possible implementation, it can also be determined whether the image belongs to the text image type in the following manner: image classification is achieved using deep learning image classification technology. Specifically, a special binary classifier can be trained, which is obtained based on a large number of training samples and their corresponding type labels, including a rich variety of text image and non-text image samples. Among them, the selection of these samples needs to cover a variety of different scenes, styles, resolutions and image formats to ensure that the classifier has a strong generalization ability and can cope with various practical situations. In the process of training the classifier, the training sample can be input into the classifier. After the classifier processes the training sample, it can output the probability that the training sample belongs to a text image or a non-text image. Based on the probability and the type label of the training sample, the parameters of the classifier are adjusted to obtain a binary classifier that can accurately classify. Subsequently, for any extracted image, the image can be input into the pre-trained binary classifier, and the binary classifier can be used to determine whether the image belongs to the text image type.

[0062] 2. Other data types are tables.

[0063] For any extracted table, first, the header names, row names, and their corresponding data items in the table can be identified through the structural information of the table.

[0064] In a possible implementation, for the identification of header names, if it is a simple table, the header names can be determined according to the first row or the first column of the table; if it is a complex table, more in-depth analysis can be carried out according to the format and content of the table. For example, if the table uses special fonts or formats to distinguish headers, the row or column where the header is located can be identified through a font analysis tool. For the identification of row names in the table, considering that the row names may be located in the first column of the table or distinguished by specific identifiers, by analyzing the layout and content of the table, the row names of each row can be found.

[0065] In another possible implementation, the header names and row names in the table can be obtained through a large language model based on the table and a text prompt prompt used to instruct the large language model to extract the header names and row names. Among them, the large language model is an artificial intelligence model built based on deep learning technology. It is trained on a large amount of text data and can learn rich language knowledge and semantic information.

[0066] For each data item in the table, a descriptive text of the data item can be formed through the value of the data item in the table and the corresponding header names and row names. For example, through a large language model, based on the table and the header names and row names in the table, the descriptive text of each data item in the table is determined.

[0067] In an example, for any data item in the table, forming a descriptive text of the data item through the value of the data item and the corresponding header names and row names includes:

[0068] Obtain the specific cell positions of each data item in the table;

[0069] For any data item, align the value of the data item to the corresponding specific cell position; combine the value of the data item, the row name of the row where the specific cell position of the data is located, and the header name of the column where it is located to form the descriptive text of the data.

[0070] In this application, the specific cell positions of each data item in the table can be obtained. For example, the row and column coordinates of a two-dimensional array are used to represent the specific cell positions. For any data item, align the value of the data item to the corresponding specific cell position. Then, combine the value of the data item, the row name of the row where the specific cell position of the data item is located, and the header name of the column where it is located to form the descriptive text of the data item. For example, for a data item located in the 2nd row and 3rd column, the row name of its row is "Product B", the header name of its column is "Sales Volume", and the value of the data item is "100", then the formed descriptive text is "The sales volume of Product B is 100".

[0071] In a possible implementation, some description templates can be designed in advance for different types of data items or different types of tables. For example, for a student grade table, the template "The grade of student {row name} in {header name} is {data item value}" can be used. When processing a data item, fill it into the corresponding template position according to its row name, header name, and value. For example, substituting Zhang San's math grade of 85 into the template, we get "The grade of student Zhang San in math is 85". Another example, for a product sales table, the template can be "The {header name} of product {row name} is {data item value}". When processing a data item, such as the "Sales Volume" of product "Mobile Phone" is 100 units, substituting it into the template gives "The sales volume of product Mobile Phone is 100 units".

[0072] Arrange all the descriptive texts of the table in the row order and column order of the table. Starting from the first row, sequentially combine the descriptive texts of each row together in column order to form an ordered sequence of descriptive texts. This can ensure that each data in the table has a clear description and is organized in the original structure order of the table, facilitating subsequent processing and understanding.

[0073] In some possible implementations, there may be text images containing tables. Therefore, if, based on the image structuring processing steps in the above embodiments, after structuring the text image containing the table, the obtained text content is the table in the text image. In this case, based on the table structuring processing steps in the above embodiments, continue to structure the text content to convert the table in the text image into descriptive text for convenient subsequent processing.

[0074] S104: Perform downstream tasks based on the structured data and the original text.

[0075] After obtaining all the structured data based on the above embodiments, the structured data and the original text can be used as the input of the downstream task and output to the downstream task, so that the downstream task can deeply process based on the content in the document to be processed. Among them, the downstream task includes but is not limited to the following types: text classification, sentiment analysis, named entity recognition, relation extraction, event extraction, semantic understanding.

[0076] In a possible implementation manner, performing the downstream task based on the structured data and the original text includes:

[0077] Obtaining the positional relationship between the original text and the data of each other data type in the text to be processed;

[0078] Recombining the structured data and the original text into a plain text document according to the positional relationship;

[0079] Performing the downstream task on the plain text document.

[0080] For the convenience of downstream task processing, in this application, the positional relationship between the original text and the data of each other data type in the text to be processed is obtained. Through the layout information and markings of the document to be processed, the positions of images and tables in the original text are determined. For example, after which paragraph the image is located in the original text, and which part of the original text the table is located in. A document parser can be used to record this positional information and store it in a positional relationship data structure, such as stored in the form of a dictionary, with the key being the unique identifier of the other data type and data, and the value being the positional information. According to the above positional relationship, the structured data and the original text are recombined into a plain text document. Exemplarily, the content description of the image and the descriptive text of the table are inserted into the corresponding positions in the original text according to the positional relationship to form a complete plain text document. Among them, during the recombination process, appropriate delimiters and markings can be added to distinguish different other data types and contents. For example, for the content description of the image, it can be marked with "[Image description start]...[Image description end]"; for the descriptive text of the table, it can be marked with "[Table description start]...[Table description end]". Finally, the plain text document is used as the input of the downstream task and output to the downstream task.

[0081] The beneficial effects of this application are as follows:

[0082] 1. For different other data types extracted from the document to be processed, the data of this other data type can be structurally processed according to the preset structural processing steps corresponding to this other data type, so as to achieve in-depth mining of the information carried by the data of this other data type, as well as the effective integration and utilization of multimodal data, and present the information carried by the data of each other data type in the document to be processed in a unified and structured form. This provides richer and more accurate input for subsequent downstream tasks, such as information retrieval, text understanding, data analysis, etc.

[0083] 2. For image data, by distinguishing between text image types and non-text image types and adopting different processing methods, the extraction of information carried in the image data is made more accurate. For text images, the text image conversion technology is used to convert the text information hidden in the image into text content, greatly improving the usability of the text information. For non-text images, with the help of a pre-trained large multimodal data processing model, the image content can be deeply understood and a detailed description can be generated, endowing the image with semantic information, which helps to convey the document intention more comprehensively.

[0084] 3. For tabular data, by obtaining the header names, row names, and generating descriptive text for each data item, the tabular data is transformed from a simple numerical matrix into an information set with rich semantics. This not only improves the readability of the tabular data but also facilitates the analysis and comparison of the tabular data, enabling a clear understanding of the specific meaning represented by each data item and the relationship between different data items, thus providing more powerful support for downstream tasks.

[0085] 4. Since traditional downstream tasks often rely only on the original text information, the depth and breadth of task processing are limited. However, this application combines the structurally processed data with the original text, providing a more comprehensive and rich information source for downstream tasks, significantly improving the processing effect and efficiency of downstream tasks.

[0086] Example 2:

[0087] The following uses specific examples to illustrate a multimodal document structuring method provided by this application. Figure 2 It is a schematic flowchart of the specific multimodal document structuring provided by the embodiments of this application, and the process includes:

[0088] S201: Obtain the document to be processed.

[0089] S202: Extract the original text, images, and tables in the document to be processed.

[0090] S203: Perform the structural processing steps of the image for each extracted image.

[0091] Figure 3 This is a schematic flowchart of the specific process for structuring any image provided by the embodiment of the present application. The process includes:

[0092] For any extracted image, the following S203a - S203d are all executed:

[0093] S203a: Obtain the content included in the image, and based on this content, determine whether the image belongs to the text image type. If so, execute S302; otherwise, execute S303.

[0094] Among them, the text image type is an image that expresses the image content through the displayed text information.

[0095] S203b: Through the image - to - text technology, identify the text content from the image, and execute S203d.

[0096] S203c: Obtain the context information associated with the image from the document to be processed. Through a pre - trained large - model for multi - modal data processing, based on the image and the context information, obtain the content description of the image, and execute S205.

[0097] Among them, the large - model for multi - modal data is a large - model with the ability to recognize image content and natural language understanding ability.

[0098] S203d: Determine whether the text content is a table. If so, execute S204; if not, execute S205.

[0099] S204: Perform the structural processing steps for each obtained table.

[0100] Figure 4 This is a schematic flowchart of the specific process for structuring any table provided by the embodiment of the present application. The process includes:

[0101] For any extracted table, the following S204a - S204d are all executed:

[0102] S204a: Through a pre - trained large - language model, based on the table and the text prompt prompt used to indicate the large - language model to extract the table header name and row name, obtain the table header name and row name in the table.

[0103] S204b: Obtain the specific cell positions of each data item in the table.

[0104] S204c: For any data item, align the value of the data item to the corresponding specific cell position, and then combine the value of the data item, the row name of the row where the specific cell position of the data item is located, and the header name of the column where it is located to form the descriptive text of the data item.

[0105] S204d: Arrange the respective descriptive text of the table according to the row order and column order.

[0106] S205: Obtain the positional relationship between the original text in the text to be processed and the data of other data types.

[0107] S206: Recombine the structured processed data and the original text into a plain text document according to the positional relationship.

[0108] S207: Perform downstream tasks on the plain text document.

[0109] Embodiment 3:

[0110] Based on the same inventive concept, the present application also provides a multi-modal document structuring processing device. Figure 5 FIG. is a schematic structural diagram of a multi-modal document structuring processing device provided by an embodiment of the present application. The device includes:

[0111] An acquisition unit 51, configured to acquire a document to be processed;

[0112] An extraction unit 52, configured to extract the original text and data of other data types in the document to be processed; wherein, the other data types include one or more of the following: images and tables;

[0113] A processing unit 53, configured to perform corresponding structured processing on the data of each of the other data types according to the structured processing steps respectively corresponding to each of the other data types; wherein, the structured processing step for any image is: determine whether the image belongs to the text image type according to the content included in the image; wherein, the text image type is an image that expresses the image content through displayed text information; if it is determined that the image belongs to the text image type, then through the image-to-text technology, identify the text content from the image; otherwise, through a pre-trained multi-modal data processing large model, based on the image, obtain the content description of the image; wherein, the multi-modal data large model is a large model with image content recognition ability and natural language understanding ability; the structured processing step for any table is: obtain the header name and row name in the table; for any data item in the table, form the descriptive text of the data item through the value of the data item in the table and the corresponding header name and row name; arrange the respective descriptive texts corresponding to the table according to the row order and column order.

[0114] A task execution unit 54, configured to execute a downstream task based on the structured processed data and the original text.

[0115] In a possible implementation manner, the processing unit 53 is further configured to, if the text content is a table, continue to perform structured processing on the text content based on the structured processing steps of the table data.

[0116] In a possible implementation manner, the task execution unit 54 is configured to obtain the positional relationship between the original text and the data of each of the other data types in the text to be processed; according to the positional relationship, reorganize the structured processed data and the original text into a plain text document; and execute the downstream task on the plain text document.

[0117] In a possible implementation manner, the processing unit 53 is specifically configured to:

[0118] Obtain the specific cell positions of each data item in the table;

[0119] For any data item, align the value of the data item to the corresponding specific cell position; combine the value of the data item, the row name of the row where the specific cell position of the data is located, and the header name of the column where the specific cell position of the data is located to form a descriptive text of the data.

[0120] In a possible implementation manner, the processing unit 53 is specifically configured to:

[0121] Obtain the context information associated with the image from the document to be processed;

[0122] Through the multi-modal data processing large model, obtain the content description of the image based on the image and the context information.

[0123] The multi-modal document structuring processing device in this embodiment is presented in the form of functional modules. Here, the module refers to an application specific integrated circuit (ASIC), a processor and a memory that execute one or more software or fixed programs, and / or other devices that can provide the above functions.

[0124] The further function descriptions of the above-mentioned various modules and units are the same as those in the corresponding above-mentioned embodiments, and will not be elaborated here.

[0125] Embodiment 4:

[0126] Please refer to Figure 6 , Figure 6 which is a schematic structural diagram of a computer device provided by an alternative embodiment of this application. AsFigure 6 As shown, the computer device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. Each component communicates with each other using different buses and can be installed on a common motherboard or installed in other ways as needed. The processor can process instructions executed within the computer device, including instructions stored in the memory or on the memory to display graphical information of the GUI on an external input / output device (such as a display device coupled to the interface). In some alternative embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (for example, as a server array, a set of blade servers, or a multi-processor system). Figure 6 In this case, one processor 10 is taken as an example.

[0127] The processor 10 can be a central processing unit, a network processor, or a combination thereof. Among them, the processor 10 can further include a hardware chip. The above-mentioned hardware chip can be an application-specific integrated circuit, a programmable logic device, or a combination thereof. The above-mentioned programmable logic device can be a complex programmable logic device, a field programmable gate array, a generic array logic, or any combination thereof.

[0128] Among them, the memory 20 stores instructions executable by at least one processor 10, so that at least one processor 10 executes the method shown in the above embodiments.

[0129] The memory 20 can include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the computer device presented by a kind of landing page of a small program, etc. In addition, the memory 20 can include a high-speed random access memory, and can also include a non-transitory memory, such as at least one disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some alternative embodiments, the memory 20 can optionally include a memory remotely set relative to the processor 10, and these remote memories can be connected to the computer device through a network. Examples of the above-mentioned network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and a combination thereof.

[0130] The memory 20 can include a volatile memory, such as a random access memory; the memory can also include a non-volatile memory, such as a flash memory, a hard disk, or a solid-state drive; the memory 20 can also include a combination of the above types of memories.

[0131] The computer device further includes an input device 30 and an output device 40. The processor 10, the memory 20, the input device 30, and the output device 40 may be connected through a bus or other means. Figure 6 Taking connection through the bus as an example.

[0132] The input device 30 can receive input digital or character information and generate key signal inputs related to the user settings and function controls of the computer device, such as a touch screen, a keypad, a mouse, a trackpad, a touchpad, a pointing stick, one or more mouse buttons, a trackball, a joystick, etc. The output device 40 may include a display device, an auxiliary lighting device (e.g., an LED), and a haptic feedback device (e.g., a vibration motor), etc. The above display device includes but is not limited to a liquid crystal display, a light-emitting diode, a display, and a plasma display. In some alternative embodiments, the display device may be a touch screen.

[0133] Embodiment 5:

[0134] Based on the above embodiments, the embodiment of the present application further provides a computer-readable storage medium, in which a computer program executable by a processor is stored. When the program runs on the processor, the processor is caused to perform the following steps when executing:

[0135] Obtain a document to be processed;

[0136] Extract the original text and data of other data types in the document to be processed; wherein, the other data types include one or more of the following: images and tables;

[0137] According to the structured processing steps respectively corresponding to each of the other data types, perform corresponding structured processing on the data of each of the other data types;

[0138] Based on the structured processed data and the original text, perform downstream tasks;

[0139] Among them, the structured processing step for any image is: according to the content included in the image, determine whether the image belongs to the type of text image; wherein, the text image type is an image that expresses the image content through the displayed text information; if it is determined that the image belongs to the type of text image, then through the image-to-text technology, identify the text content from the image; otherwise, through a pre-trained multi-modal data processing large model, based on the image, obtain the content description of the image; wherein, the multi-modal data large model is a large model with the ability to recognize image content and natural language understanding ability;

[0140] The structured processing steps for any table are as follows: obtain the header name and row name in the table; for any data item in the table, form a descriptive text of the data item based on the value of the data item in the table and the corresponding header name and row name; arrange the respective descriptive text corresponding to the table according to the row order and column order.

[0141] Since the principle of solving problems by the above computer-readable storage medium is similar to that of the multi-modal document structuring method, the implementation of the above computer-readable storage medium can refer to the embodiments of the method, and the repeated parts will not be described again.

[0142] Obviously, those skilled in the art can make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalent technologies, this application is also intended to include these changes and modifications.

Claims

1. A method for structuring multimodal documents, characterized in that: The method comprises: Get the documents to be processed; Extracting original text and data of other data types from the document to be processed; wherein the other data types include one or more of the following: images and tables; According to the structural processing steps corresponding to each of the other data types, the data of each of the other data types are subjected to corresponding structural processing; Perform downstream tasks based on the structured processed data and the original text; The structural processing step of any image is as follows: determining whether the image belongs to a text image type according to the content contained in the image; wherein the text image type is an image that expresses the image content through displayed text information; if it is determined that the image belongs to a text image type, identifying the text content from the image through image-to-text technology; otherwise, obtaining a content description of the image based on the image through a pre-trained multimodal data processing large model; wherein the multimodal data large model is a large model with image content recognition capabilities and natural language understanding capabilities; The steps for structuring any table are as follows: obtaining the header name and row name in the table; for any data item in the table, forming a descriptive text for the data item through the value of the data item in the table and the corresponding header name and row name; arranging the descriptive texts corresponding to the table according to the row order and column order.

2. The method according to claim 1, characterized in that If the text content is a table, the text content is further structured based on the structural processing steps of the table data.

3. The method according to claim 1, characterized in that The downstream tasks are performed based on the structured processed data and the original text, including: Acquire the positional relationship between the original text and the data of each other data type in the text to be processed; According to the positional relationship, reorganizing the structured data and the original text into a plain text document; The downstream task is performed on the plain text document.

4. The method according to claim 1, characterized in that For any data item in the table, the descriptive text of the data item is formed through the value of the data item and the corresponding header name and row name, including: Get the specific cell position of each data item in the table; For any data item, align the value of the data item to the corresponding specific cell position; combine the value of the data item, the row name of the row where the specific cell position of the data is located, and the header name of the column where the data is located to form a descriptive text for the data.

5. The method according to claim 1, characterized in that The method of obtaining a content description of the image based on the image by using a pre-trained multimodal data processing large model includes: Acquire context information associated with the image from the document to be processed; The content description of the image is obtained based on the image and the context information through the multimodal data processing model.

6. A multimodal document structuring processing device, characterized in that: The device comprises: An acquisition unit, used for acquiring documents to be processed; An extraction unit, used to extract original text and data of other data types from the document to be processed; wherein the other data types include one or more of the following: images and tables; A processing unit, used to perform corresponding structural processing on the data of each other data type according to the structural processing steps corresponding to each of the other data types; wherein, the structural processing step of any image is: according to the content contained in the image, determine whether the image belongs to the text image type; wherein the text image type is an image that expresses the image content through the displayed text information; if it is determined that the image belongs to the text image type, then the text content is identified from the image through the image-to-text technology; otherwise, the content description of the image is obtained based on the image through a pre-trained multimodal data processing large model; wherein the multimodal data large model is a large model with image content recognition capabilities and natural language understanding capabilities; the structural processing step of any table is: obtain the header name and row name in the table; for any data item in the table, form a descriptive text of the data item through the value of the data item in the table and the corresponding header name and row name; arrange the descriptive texts corresponding to the table according to the row order and column order; The task execution unit is used to execute downstream tasks based on the structured processed data and the original text.

7. The device according to claim 6, characterized in that The processing unit is further configured to, if the text content is a table, continue to perform structural processing on the text content based on the structural processing steps of the table data.

8. The device according to claim 6, characterized in that The task execution unit is used to obtain the positional relationship between the original text and the data of each other data type in the text to be processed; according to the positional relationship, reorganize the structured processed data and the original text into a plain text document; and execute the downstream task on the plain text document.

9. A computer device, characterized in that: The computer device includes a processor, and the processor is used to implement the steps of the multimodal document structuring processing method as described in any one of claims 1 to 5 above when executing a computer program stored in a memory.

10. A computer-readable storage medium, characterized in that: It stores a computer program, which, when executed by a processor, implements the steps of multimodal document structuring processing as described in any one of claims 1-5 above.

Citation Information

Cited By

  • Multi-modal document automatic proofreading method and system based on artificial intelligence

    CN120782394A

  • A Multimodal Document Automatic Proofreading Method and System Based on Artificial Intelligence

    CN120782394B

  • Document analysis method and device, electronic equipment and storage medium

    CN121561098A