Document processing method and device, equipment and medium

Through multimodal large model and object detection technology, combined with cutout tools, multi-format documents are analyzed and reorganized, which solves the stability and accuracy of document processing and realizes efficient and reliable conversion and formatting of document content.

CN120145997APending Publication Date: 2025-06-13CHENGDU BOSS INNOVATION TECH CO LTD

Patent Information

Application Number
CN202510285205.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The prior art is difficult to ensure the stability and accuracy of the analysis when processing multi-format electronic documents, especially in complex documents containing images and tables, and there is inconvenience in processing sensitive information.

Method used

A multimodal large model is used to combine object detection and cutout tools to automatically process the document through layout analysis, object detection, object recognition and reorganization, and convert it into document results in the target format.

Benefits of technology

Improves the stability and accuracy of document processing, ensures the completeness and reliability of document content, adapts to a variety of document formats, supports accurate extraction and naming of images and tables.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120145997A_ABST
    Figure CN120145997A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a document processing method and device, equipment and a medium, and relates to the technical field of data processing.The method comprises the steps that a to-be-processed document is converted into a single-page atlas on the basis of the reading sequence of the to-be-processed document, layout analysis is conducted on the single-page atlas, the single-page atlas is split into a single-bar graph sequence, and the single-bar graph sequence is stored in the single-page atlas sequence; and sequentially extracting single-column pages in the single-column graph sequence according to the reading sequence, and carrying out target detection on the corresponding single-column pages according to a preset scanning sequence so as to determine page elements in the single-column pages based on a detection result. And performing target identification on the page element to determine position information of the page element in the corresponding single-column page, and converting the page element into a page object in a target format based on an identification result. The page objects are re-typeset based on the reading sequence, the scanning sequence and the position information and assembled into the document result in the target format, so that the reliability of document processing is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to a document processing method, device, equipment and medium. Background Art

[0002] Document processing is one of the important research topics in the field of large-model retrieval-augmented generation (RAG), which involves how to effectively extract, understand and utilize information in documents in document applications such as document knowledge question and answering.

[0003] However, with the acceleration of digital transformation, a large number of electronic documents have accumulated. The formats of relevant electronic documents are diverse, making it a problem that needs to be studied how to improve the reliability of document processing based on large models. Summary of the invention

[0004] One of the purposes of the present invention includes, for example, providing a document processing method, apparatus, device and medium to at least partially improve the reliability of document processing.

[0005] The embodiments of the present invention can be implemented as follows:

[0006] In a first aspect, an embodiment of the present invention provides a document processing method, including:

[0007] Determine a corresponding reading order based on the page number of the document to be processed, and convert the document to be processed into a single-page atlas according to the reading order;

[0008] Performing layout analysis on the single-page atlas, and splitting the single-page atlas into a single-column atlas sequence according to the layout analysis result, wherein the single-column atlas sequence is sequentially arranged according to the reading order;

[0009] Extracting the single-column pages in the single-column image sequence in sequence according to the reading order, and performing target detection on the corresponding single-column pages according to a preset scanning order, so as to determine the page elements therein based on the detection results;

[0010] Performing target recognition on the page element to determine the position information of the page element in the corresponding single-column page, and converting the page element into a page object in a target format based on the recognition result;

[0011] The page objects are rearranged based on the reading order, the scanning order and the position information, and assembled into a document result in a target format.

[0012] In an optional implementation, determining a corresponding reading order based on the page number of the document to be processed, and converting the document to be processed into a single-page atlas according to the reading order, includes:

[0013] Convert the content of each page in the document to be processed into a single-page image respectively;

[0014] Convert each of the single-page images into a single-page image set based on the reading order.

[0015] In an optional embodiment, perform layout analysis on the single-page image set, and split the single-page image set into a single-column image sequence according to the layout analysis result. The single-column image sequence continues in the reading order and includes:

[0016] Perform layout analysis on each of the single-page images in the single-page image set to determine the segmentation strategy for each of the single-page images;

[0017] Split the single-page image into a single-column image sequence based on the segmentation strategy. The single-column image sequence continues in the reading order.

[0018] In an optional embodiment, the page elements include images and tables;

[0019] Perform object detection on the corresponding single-column page according to the preset scanning order to determine the page elements therein based on the detection results, including:

[0020] Perform object detection on the corresponding single-column page according to the preset scanning order, and label the images and tables in the single-column page.

[0021] In an optional embodiment, perform object recognition on the page elements to determine the position information of the page elements in the corresponding single-column page, and convert the page elements into page objects in the target format based on the recognition results, including:

[0022] In the case where the page element is an image, use a matte tool to cut out the target image;

[0023] Name the target image according to the prompt word instruction of the multimodal large model and upload it, and generate a remote access hyperlink for the target image;

[0024] Based on the position information, fill the cut-out part of the single-column page with a solid color and add an image placeholder;

[0025] In the case where the page element is a table, use a matte tool to cut out the table as the target image;

[0026] Output the table content in the target image as a document format based on the prompt word instruction of the multimodal large model;

[0027] Based on the position information, fill the cut-out part of the single-column page with a solid color and add a table placeholder.

[0028] In an alternative embodiment, the document format includes the Markdown format, and the method further includes:

[0029] After filling the cropped part of the single-column page with a solid color and adding an image placeholder, based on the prompt instructions of the multimodal large model, all the content is recognized and output as a Markdown-formatted document, the original layout is retained in the document, and the image placeholder is replaced with a Markdown image remote access hyperlink;

[0030] After filling the cropped part of the single-column page with a solid color and adding a table placeholder, based on the prompt instructions of the multimodal large model, all the content is recognized and output as a Markdown-formatted document, the original layout is retained in the document, and the table placeholder is replaced with a Markdown table.

[0031] In an alternative embodiment, the re-typesetting the page object based on the reading order, the scanning order, and the position information and assembling it into a document result in the target format includes:

[0032] Based on the reading order, the scanning order, and the position information, the Markdown-formatted documents are spliced and assembled to obtain a document result in the Markdown format;

[0033] Wherein, the prompt instructions of the multimodal large model include: the role definition, rules, and workflow of the multimodal large model; when outputting the image content as a document, recognizing bold font, font color, italic font, formula, table, link, picture, multi-level heading, unordered list, and ordered list; when outputting a Markdown-formatted document, retaining the font style, color, mathematical formula, link, and table.

[0034] In a second aspect, an embodiment of the present invention provides a document processing device, including:

[0035] An information acquisition module, configured to determine a corresponding reading order based on the page numbers of the document to be processed, and convert the document to be processed into a single-page picture set according to the reading order;

[0036] An information processing module is used to perform layout analysis on the single-page atlas, and split the single-page atlas into a single-column image sequence according to the layout analysis result, and the single-column image sequence is extended according to the reading order; extract the single-column pages in the single-column image sequence in sequence according to the reading order, and perform target detection on the corresponding single-column pages according to a preset scanning order to determine the page elements therein based on the detection results; perform target recognition on the page elements to determine the position information of the page elements in the corresponding single-column pages, and convert the page elements into page objects in a target format based on the recognition results; rearrange the page objects based on the reading order, the scanning order and the position information, and assemble them into a document result in a target format.

[0037] In a third aspect, an embodiment of the present invention provides an electronic device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the document processing method described in any one of the aforementioned implementations is implemented.

[0038] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium includes a computer program, and when the computer program is executed, the electronic device where the computer-readable storage medium is located is controlled to execute the document processing method described in any one of the aforementioned implementations.

[0039] The beneficial effects of the embodiments of the present invention include, for example: after converting the document to be processed into a single-page atlas in reading order, performing layout analysis, extraction, and target detection in sequence to determine the page elements, performing target recognition and processing on each page element to obtain the page object in the target format, and typeset and reorganize the page objects to obtain the document results in the target format, thereby achieving unified and automated processing of various types of documents and ensuring the reliability of document processing. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments are briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without creative work.

[0041] Figure 1 A schematic flow chart of a document processing method provided by an embodiment of the present invention is shown.

[0042] Figure 2 A schematic diagram of a single-page atlas provided by an embodiment of the present invention is shown.

[0043] Figure 3Shows a schematic diagram of a layout analysis for a single-page atlas provided by an embodiment of the present invention for Figure 2 the single-page atlas shown.

[0044] Figure 4 Shows a schematic diagram of one of the single-page images obtained by splitting according to the layout analysis result provided by an embodiment of the present invention. Figure 3 for the single-page atlas shown.

[0045] Figure 5 Shows a schematic diagram of a single-page image of an input model provided by an embodiment of the present invention.

[0046] Figure 6 Shows a schematic diagram of a picture for naming provided by an embodiment of the present invention.

[0047] Figure 7 Shows a schematic diagram of a conversion of an image shown into Markdown format text provided by an embodiment of the present invention for Figure 5 the image shown.

[0048] Figure 8 Shows a rendering schematic diagram of Markdown provided by an embodiment of the present invention.

[0049] Figure 9 Shows a schematic diagram of the overall process of a document processing method provided by an embodiment of the present invention.

[0050] Figure 10 Shows an exemplary structural block diagram of a document processing apparatus provided by an embodiment of the present invention.

[0051] Figure 11 Shows a structural block diagram of an electronic device provided by an embodiment of the present invention.

[0052] Icons: 11 - Electronic device; 111 - Controller; 112 - ROM; 113 - RAM; 114 - Bus; 115 - I / O interface; 116 - Input unit; 117 - Output unit; 118 - Storage unit; 119 - Communication unit; 140 - Document processing apparatus; 141 - Information acquisition module; 142 - Information processing module. Detailed implementation manners

[0053] Partial term definitions:

[0054] Multimodal large model: Refers to a deep learning model that can process and understand various types of data, such as text, images, etc. It improves the understanding and processing ability of information by learning the associations between different modal data.

[0055] Layout analysis: Refers to a technology used to identify and understand the layout of a document image, including identifying the positions and arrangement order of elements such as text, images, tables, etc.

[0056] Object detection: Refers to a technology in image processing for identifying and locating page elements in an image, such as people, vehicles, specific objects, etc.

[0057] Matting tool: Refers to an image processing tool used to extract specific image elements from the background, usually achieved by generating a mask.

[0058] Instance segmentation: Refers to a technology that goes a step further than traditional object detection. It not only identifies objects in an image but also segments and labels each instance of the object.

[0059] YOLO (You Only Look Once): Refers to a popular object detection algorithm, known for its speed and high accuracy, suitable for real-time object detection tasks.

[0060] Markdown format: Refers to a lightweight markup language that allows writing documents in a human-readable and writable plain text format and then converting them into valid extensible hypertext markup languages such as XHTML and HTML documents. Its goal is to achieve "easy to read and write" and have a certain readability even without format conversion.

[0061] Prompt: Refers to a type of prompt or prompt information in a multi-modal large model used to guide the model to generate specific types of outputs.

[0062] With the acceleration of digital transformation, enterprises and society have accumulated a large number of electronic documents, such as product manuals, contracts, reports, medical records, etc. These documents contain rich information and are of great value for decision-making support, data analysis, and knowledge management.

[0063] However, parsing various types of documents based on large models faces multiple challenges.

[0064] For example, due to the diverse formats of documents, including PDF, Word, Excel, etc., parsing technologies need to be able to adapt to different file formats.

[0065] For another example, due to the complexity of document content, including tables, images, mixed text, layout, handwritten and photocopied versions, etc., it increases the difficulty of parsing.

[0066] For yet another example, in the case where a document may contain sensitive information, it is not convenient to use third-party parsing services, which puts higher requirements on the document parsing process, such as being not only efficient but also ensuring data integrity, accuracy, security, and privacy.

[0067] To address these challenges, various document parsing techniques can be considered, such as rule-based methods, machine learning methods, deep learning methods, etc.

[0068] Among them, rule-based methods rely on predefined patterns and logic. Although they work well in some cases, they often lack flexibility and scalability. Machine learning methods, such as Support Vector Machine (SVM) and Random Forest, learn features from training data and can handle more complex patterns, but require a large amount of labeled data, and the generalization ability of the model is limited by the quality and diversity of the training data.

[0069] In contrast, document parsing based on deep learning technology has stronger applicability. Deep learning models, such as Convolutional Neural Networks (CNN) and Recurrent Neural Networks (RNN), have demonstrated excellent performance in image recognition, text processing, and semantic understanding. These models can automatically learn complex patterns and features in the data, thereby improving the accuracy and efficiency of document parsing. In this context, multi-modal large models such as GPT-4V and Qwen-VL have made significant progress in the joint understanding of images and text. By deeply integrating visual and language information, they can better understand and parse complex documents containing images and text. They can not only recognize the text in images, name the images, but also convert them into structured format text, such as Markdown, thus enhancing the richness of document parsing.

[0070] It has been found through research that when dealing with unstructured data such as Portable Document Format (PDF), deep learning-based parsing methods, such as Layout-parser, show advantages in extracting tables, images, and preserving the document layout structure. When dealing with images and tables in documents, multi-modal large models, such as Qwen-VL, show advantages in image recognition and naming, table recognition, and Markdown standard format output. These methods can better understand and process complex structures in documents, such as multi-column layouts, cross-page content, and embedded images, through deep learning models.

[0071] However, in some scenarios, such as the knowledge Q&A scenario of kitchen appliances, the stability and accuracy of knowledge structure output need to be improved based on document parsing and processing in related technologies.

[0072] Based on the above research, an embodiment of the present invention provides a document processing solution. An automated document parsing and formatting processing solution is proposed based on multi-modal processing technology, covering all stages from document format conversion, layout analysis, image extraction, image naming to feature extraction and text assembly, thereby significantly improving the stability and accuracy of document processing.

[0073] Regarding the defects existing in the above solutions, they are all the results obtained by the inventor through practice and careful research. Therefore, the discovery process of the above problems and the solutions proposed by the embodiments of the present invention below for the above problems should all be the contributions made by the inventor during the invention process.

[0074] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Usually, the components of the embodiments of the present invention described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations.

[0075] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed present invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.

[0076] It should be noted that the term "including", "comprising", or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or device including a series of elements not only includes those elements but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article, or device including the said element.

[0077] It should be noted that similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.

[0078] It should be noted that, without conflict, the features in the embodiments of the present invention can be combined with each other.

[0079] Please refer to Figure 1 , which is a schematic flowchart of a document processing method provided by an embodiment of the present invention. The document processing method includes S110, S120, S130, S140, and S150.

[0080] S110, determining a corresponding reading order based on the page number of the document to be processed, and converting the document to be processed into a single-page atlas according to the reading order.

[0081] S120, performing layout analysis on the single-page atlas, and splitting the single-page atlas into a single-column image sequence according to the layout analysis result, and the single-column image sequence is extended according to the reading order.

[0082] S130, extracting the single-column pages in the single-column image sequence in sequence according to the reading order, and performing target detection on the corresponding single-column pages according to a preset scanning order, so as to determine the page elements therein based on the detection results.

[0083] S140, performing target recognition on the page element to determine the position information of the page element in the corresponding single-column page, and converting the page element into a page object in a target format based on the recognition result.

[0084] S150: rearrange the page objects based on the reading order, the scanning order and the position information, and assemble them into a document result in a target format.

[0085] For all types of documents to be processed, they are uniformly converted into single-page atlases in image format in reading order, and layout analysis, extraction, target detection, target recognition and processing are carried out in sequence. The processing results are typeset and reorganized to obtain the document results in the target format, realizing automated and unified processing of various types of documents, ensuring the convenience and reliability of document processing.

[0086] In S110 , the document to be processed may include text, images, tables, etc.

[0087] Given the variety of document formats, such as scanned versions, Word versions, PPT versions, etc., if documents of different formats all adopt the processing logic corresponding to their formats, it will increase the complexity of the document processing process and it is not easy to ensure the consistency of the processing result output. Based on this, in this embodiment, the documents to be processed are uniformly converted into image formats, and after obtaining a single-page atlas, layout analysis is performed to obtain the layout analysis results, so as to unify the format.

[0088] The conversion of the document to be processed into an image format to obtain a single-page atlas can be flexibly selected. For example, the image size can be preset, and each part of the document to be processed is converted into a single-page image in sequence (such as from top to bottom, from left to right, etc.) according to the preset image size to obtain a single-page atlas, and the reading order of each single-page image is set accordingly according to the reading order of each part of the document to be processed.

[0089] For another example, when the document to be processed includes multiple pages of content, the content of each page in the document to be processed can be converted into a single-page image, and each single-page image can be converted into a single-page atlas based on the reading order, so that the reading order of the document to be processed and the single-page atlas is consistent.

[0090] For another example, a single-page atlas may be a whole image including all the contents of multiple single-page images and set in a reading order. Each whole image has a uniform size and can be converted into one or more single-page atlases based on the amount of content in the document to be processed.

[0091] For example, please refer to Figure 2 Taking the product manual as an example, based on S110, the documents in various formats are converted into image formats according to the reading order, and the Figure 2 The single-page atlas shown includes multiple single-page images and maintains the original order of page numbers, providing a unified image basis for subsequent layout analysis and content extraction.

[0092] After the document to be processed is converted into a single-page atlas based on S110, in S120, layout analysis is performed on the single-page atlas, and the single-page atlas is split into a single-column atlas sequence according to the layout analysis result, which can be flexibly implemented.

[0093] For example, target detection algorithms such as YOLO (You Only Look Once) can be used for layout analysis to accurately segment and mark the reading order of each single-page image in a single-page atlas to ensure the accuracy of the logical order of the document content.

[0094] For example, please refer to Figure 3 , showing the Figure 2 The single-page atlas shown in the figure detects the single-page image, performs layout analysis on a page-by-page basis, and determines a segmentation strategy and a sorting scheme.

[0095] For example, layout splitting can be performed by using YOLO to perform layout analysis, and based on the layout analysis results, a single-page atlas can be split into single-page images to preserve the reading order. Figure 4 As shown, for Figure 3 The single-page atlas shown is split to obtain one of the pages of images. Figure 3 The layout analysis results shown in the figure are used to split all the single-page atlases, so that the single-page images with the arrangement order consistent with the original document can be obtained.

[0096] After obtaining the single-page image, the segmentation strategy of each single-page image can be determined according to the layout analysis results, and the single-page image can be split into a single-column image sequence based on the segmentation strategy. The single-column image sequence is postponed according to the reading order of the document to be processed to ensure that the arrangement order of the single-column image sequence is consistent with the original document.

[0097] In the case of obtaining a single-column figure sequence based on S120, in S130, the single-column pages in the single-column figure sequence are sequentially extracted in the reading order, and target detection is performed on the corresponding single-column pages according to a preset scanning order, so as to determine the page elements therein based on the detection results, which can be flexibly implemented. Taking the detection targets including tables and images as an example, target detection is performed on the corresponding single-column pages according to a preset scanning order to mark the images and tables in the single-column pages. For example, the detection targets such as tables and images are marked by means of box selection, so as to obtain page elements marked with the reading order.

[0098] For example, the multi-modal large model Prompt instruction technology can be used to identify and process each page element of the single-column page. For example, the detection model can be used to identify and mark the "picture" elements in the single-column page to automatically detect the page elements marked with the reading order.

[0099] Taking the detection target as an image as an example, the detection model can be trained through the following process:

[0100] Use an open-source image annotation tool such as labelimg to annotate. In the single-column page, select the pictures by box and save the annotation information. Organize the annotated pictures and the corresponding annotation information into the dataset structure required by YOLO. For example, create two folders: "images" for storing pictures and "labels" for storing.txt annotation files.

[0101] Further organize the two created folders into "train" and "val" sub-folders, which are used for training and validating the detection model respectively. Configure the YOLO training parameters and train the detection model until the set convergence condition is reached, such as the picture detection accuracy reaches the threshold, so as to obtain a trained detection model for prediction.

[0102] Based on the trained detection model, input the image to be detected, such as a single-page image. The detection model can output the labels and start and end coordinates of the page elements, so as to split and reorganize the single-page image into a single-column figure sequence. According to the layout analysis result, the splitting and reorganization of the single-page image are realized, and at the same time, the arrangement order consistent with the original document is maintained.

[0103] Exemplarily, after inputting the single-page image as shown in Figure 5 into the trained detection model, the detection model predicts and outputs the labels and start and end coordinates as follows:

[0104]

[0105] After determining the page element based on S130, in S140, target recognition is performed on the page element to determine the position information of the page element in the corresponding single-column page, and the page element can be flexibly converted into a page object in the target format based on the recognition result.

[0106] For example, in the case where the page element is an image, a matte tool can be used to cut out the target image. After naming the target image based on the prompt instruction of the multimodal large model and uploading it, such as naming it according to the image content using the multimodal large model Prompt instruction and uploading the image to cloud storage. Generate a remote access hyperlink for the target image. Based on the position information, fill the cut-out part of the single-column page with a solid color and add an image placeholder.

[0107] For another example, in the case where the page element is a table, a matte tool can be used to cut out the table as a target image. Based on the prompt instruction of the multimodal large model, output the table content in the target image as a document format, such as using the multimodal large model Prompt instruction to output the table content as a standard Markdown table. Based on the position information, fill the cut-out part of the single-column page with a solid color and add a table placeholder.

[0108] For another example, in the single-column page after the above image processing and table processing, the remaining content is text content. For the remaining content, the multimodal large model Prompt instruction can be used to recognize all of its content and output it in the standard Markdown format, retaining the original layout, such as multi-level headings, bold fonts, font colors, italic fonts, mathematical formulas, etc. After obtaining the Markdown text of the single-column page, replace the image placeholders in it with standard Markdown image hyperlinks and the table placeholders with standard Markdown tables, and the processed result in the document format can be obtained.

[0109] For example, after filling the cut-out part of the single-column page with a solid color and adding an image placeholder, based on the prompt instruction of the multimodal large model, recognize all of the content and output it as a Markdown-formatted document, retaining the original layout in the document and replacing the image placeholder with a Markdown image remote access hyperlink. After filling the cut-out part of the single-column page with a solid color and adding a table placeholder, based on the prompt instruction of the multimodal large model, recognize all of the content and output it as a Markdown-formatted document, retaining the original layout in the document and replacing the table placeholder with a Markdown table.

[0110] By combining target detection with a matte tool, using target detection to mark page elements and precisely extracting them through the matte tool, the extraction accuracy is improved.

[0111] Exemplarily, when recognizing an image in a single-column page and the image is marked as "image", image extraction can be achieved in the following way:

[0112] Precisely extract the marked "image" picture part through a matting tool. For example, use YOLO for instance segmentation to generate a mask for each page element (image) in the single-column page. Use YOLO to perform inference on the picture to obtain the mask of each target. Here, the mask is an array with the same size as the picture, where the area of the target is True and the background area is False. According to the mask, the target can be cut out from the original picture, which can be achieved by applying the mask to the original picture and only retaining the pixels where the mask is True. After cutting out the target, blank areas will be left in the original picture. By creating a new image with the same size as the original picture and filling it with a specified solid color, only the areas where the mask is True are replaced with the corresponding areas of the original picture to achieve background filling.

[0113] Solid color filling can be achieved based on the position coordinates of the source image. By using solid color filling, it can be avoided that because the cropped area contains text, the text in the cropped area is misrecognized into the result during the parsing process, which may damage the integrity of the knowledge structure of the source document and ensure the reliability of subsequent processing.

[0114] In this embodiment, image naming can be achieved in the following way:

[0115] Based on a multimodal large model, recognize and generate a file name describing the "picture", such as "Schematic Diagram of Product Dimensions and Installation Requirements.png". The large model Prompt can be designed to let the large model name the picture according to what it sees. Access the multimodal large model, input the picture and the Prompt, and thus obtain the image naming.

[0116] Please refer to Figure 6 , this embodiment provides a reference for the Prompt of image naming for the picture shown based on a multimodal large model Figure 6 :

[0117]

[0118]

[0119]

[0120] Compared with general image recognition models that can only give categories and cannot give more refined picture titles or names. In this embodiment, through the multimodal large model, the Prompt can be flexibly controlled to generate more refined and accurate image naming nouns for it, such as: Schematic Diagram of Product Dimensions and Installation Requirements.png.

[0121] Based on the obtained image naming, image storage and link generation can be carried out in the following way: Upload the extracted picture to the cloud storage system and generate a remote access link pointing to the picture to ensure that the picture can be remotely accessed. For example, the remote access link of the picture can be:

[0122] https: / / ai-test-platform.myroki.com / abp / manager / api / file / download / ?filepath

[0123] = / media / document / product / KZQS-40-CQ903 / Schematic diagram of product dimensions and installation requirements.png.

[0124] By generating descriptive image naming through a multimodal large model and integrating cloud storage, efficient storage and remote access of images are achieved. Image naming is equivalent to adding a retrieval tag to the image, facilitating recall. For example, there is a picture belonging to the main product picture, but the words "main picture" are not mentioned in the text outside this picture, then retrieval is relatively difficult. Cut out the part of the picture and then fill in the image name as a pure text tag for easy recall.

[0125] In the case of generating a remote access hyperlink for the image, filling the cut-out part of the single-column page with a solid color and adding an image placeholder can be achieved in the following way: Use a cutout tool to fill the blank area after cutting out the "picture" and add a descriptive label (image naming), such as "Schematic diagram of product dimensions and installation requirements.png", as the image placeholder. Given that solid color filling means covering the corresponding area, after marking the solid color filling area with a placeholder, it indicates that this area is the picture of "Schematic diagram of product dimensions and installation requirements.png", thus protecting the content integrity and visual integrity of the single-column page.

[0126] Extract key information of image features such as bold font, font color, italic font, formulas, tables, links, pictures, multi-level headings, unordered lists, ordered lists, etc. through the "image-to-text" multimodal large model and generate standard Markdown format text for subsequent processing.

[0127] Among them, taking the Figure 5 shown image as an example, in order to achieve the recognition of key information of features, the Prompt prompt words of the multimodal large model can be as follows:

[0128]

[0129]

[0130]

[0131] Please refer to Figure 7 , this embodiment provides a method for constructing a corresponding Prompt according to the key feature information for the Figure 5 shown image, and using the "image-to-text" multi-modal large model to convert the image into Markdown format text, while preserving as much source format information of the document to be processed as possible, including but not limited to font styles, colors, mathematical formulas, links, tables, etc.

[0132] Reference for Prompt prompts:

[0133]

[0134]

[0135]

[0136]

[0137] When using the multi-modal large model for image-to-text, the above process is adopted to extract the text information in the image, convert the image into text, and control the stability of image-to-text by controlling the Prompt and parameters such as temperature.

[0138] Applying the multi-modal large model in document parsing realizes the in-depth understanding and processing of various types of data in the document, such as text, images, and tables, thereby improving the accuracy and integrity of document content extraction and ensuring the reliability of document processing.

[0139] If the text after the multi-modal large model performs image recognition has pictures marked with placeholder tags, replace them with the above actual Markdown image remote access hyperlinks to ensure the correspondence between the text and the pictures.

[0140] When replacing the image placeholder with the Markdown image remote access hyperlink, the Prompt prompt words of the multi-modal large model can be as follows:

[0141] # Installation Instructions

[0142] At the set position of the cabinet, set the square hole according to the following installation diagram, and smoothly embed the appliance into the square hole, taking care not to install it obliquely. The specific opening dimensions (mm) are shown in the following table:

[0143] | Serial Number | Name | W | H | D |

[0144] |---|---|---|---|---|

[0145] | 1 | Full-embedded opening dimensions (width × height × depth) | 600 | 460 | 565 |

[0146] |2|Semi-embedded hole size (width × height × depth)|560|450|550|

[0147] ![(https: / / ai-test-platform.myroki.com / abp / manager / api / file / download / ?file path= / media / document / product / KZQS-40-CQ903 / Product size and installation requirements schematic diagram.png)](https: / / ai-test-platform.myroki.com / abp / manager / api / file / download / ?file path= / media / document / product / KZQS-40-CQ903 / Product size and installation requirements schematic diagram.png)

[0148] Installation requirements:

[0149] - The cabinet plane or countertop where the appliance is placed must be horizontal.

[0150] - Try to ensure air circulation around the appliance inside the cabinet. It is recommended that the splint and fixing plate use moisture-proof, waterproof, corrosion-resistant, and non-combustible heat-insulating materials.

[0151] - Use the two provided installation screws to fix the body to the cabinet through the fixing holes on the left and right door frames.

[0152] Power supply requirements:

[0153] - For permanent installation, the circuit must be equipped with corresponding cut-off and protection devices. The power plug and socket connections should be of the same model and comply with local relevant regulations.

[0154] - The power cord connection must be convenient to ensure that the power can be disconnected at any time after the appliance is installed. Use a single 10A or above socket alone, do not use the same power socket with several electrical appliances at the same time, and ensure that the socket is safely and effectively grounded.

[0155] - If there are other electrical appliances nearby, ensure that the installation distance is greater than 100mm.

[0156] Please refer to Figure 8 for a rendering example diagram of Markdown. The construction and application of feature constraints for Prompts.

[0157] Utilize the "image-to-text" multi-modal large model to complete the automated generation of Markdown-formatted text, achieving the automated conversion from images to structured text, improving the convenience and efficiency of document processing. By constructing Prompts corresponding to the key feature information, guiding the multi-modal large model to generate more accurate and rich Markdown-formatted text, the integrity of document processing is improved.

[0158] After obtaining the processing result in the Markdown document format based on S110 to S140, in S150, the page objects are rearranged based on the reading order, the scanning order and the position information, and assembled into a document result in the target format, which can be flexibly implemented.

[0159] For example, all generated Markdown formatted texts can be assembled together according to the original reading order, scanning order and position information of the document to be processed to complete the parsing and formatting process of the entire document. For example, based on the reading order, scanning order and position information, the Markdown formatted documents are assembled and spliced ​​to obtain the document result in Markdown format. Figure 9 , provides an overall flow chart of the document processing method.

[0160] In order to execute the corresponding steps in the above embodiments and various possible methods, a method for implementing a document processing device is provided below. Figure 10 , Figure 10 A functional module diagram of a document processing device 140 provided in an embodiment of the present invention, the document processing device 140 can be applied to Figure 11 The electronic device 11 is shown. It should be noted that the basic principle and technical effect of the document processing device 140 provided in this embodiment are the same as those of the above method embodiment. For the sake of brief description, for the parts not mentioned in this embodiment, reference can be made to the corresponding contents in the above method embodiment. The document processing device 140 includes an information acquisition module 141 and an information processing module 142.

[0161] The information acquisition module 141 is used to determine a corresponding reading order based on the page number of the document to be processed, and convert the document to be processed into a single-page atlas according to the reading order.

[0162] The information processing module 142 is used to perform layout analysis on the single-page atlas, and split the single-page atlas into a single-column image sequence according to the layout analysis result, and the single-column image sequence is arranged in the reading order; the single-column pages in the single-column image sequence are extracted in sequence according to the reading order, and the corresponding single-column pages are subjected to target detection according to a preset scanning order to determine the page elements therein based on the detection results; the page elements are subjected to target recognition to determine the position information of the page elements in the corresponding single-column pages, and the page elements are converted into page objects in the target format based on the recognition results; the page objects are rearranged based on the reading order, the scanning order and the position information, and assembled into a document result in the target format.

[0163] On the basis described above, an embodiment of the present invention further provides a computer-readable storage medium. The computer-readable storage medium includes a computer program, and when the computer program runs, it controls the electronic device where the computer-readable storage medium is located to execute the above-mentioned document processing method.

[0164] Adopting the above solution in the embodiment of the present invention, based on a multi-modal large model, it aims to capture local and global features in images to achieve efficient conversion of various format documents. The converted documents are presented in a Markdown structured form, facilitating machine access and understanding, while ensuring the preservation of the original document's layout and content integrity. Through an automated process, covering all stages from document format conversion, layout analysis, image extraction, image naming to feature extraction and text assembly, it significantly improves the comprehensiveness and accuracy of document processing, effectively improving the problem of inaccurate storage of key elements such as product main images and product structure diagrams in kitchen appliance instruction manuals in various fields.

[0165] Please refer to Figure 11 , which is a structural block diagram of the electronic device 11 for executing the document processing method provided by the embodiment of the present invention. The components, their connections and relationships, and their functions shown in this embodiment are only examples and are not intended to limit the implementation of the embodiment of the present invention described and / or required in this embodiment.

[0166] As Figure 11 shown, the electronic device 11 includes a controller 111 and a memory. There is at least one controller 111, and the memory is such as a read-only memory (ROM112), a random access memory (RAM113), etc. Among them, the memory stores computer-executable instructions that can be executed by at least one controller 111. The controller 111 can execute various appropriate actions and processes according to the computer-executable instructions stored in the read-only memory (ROM112) or the computer-executable instructions loaded from the storage unit into the random access memory (RAM113). In the RAM113, various programs and data required for the operation of the electronic device 11 can also be stored. The controller 111, the ROM112, and the RAM113 are connected to each other through a bus 114. The input / output interface (I / O interface 115) is also connected to the bus 114.

[0167] Multiple components in the electronic device 11 are connected to the I / O interface 115, including: an input unit 116, such as a keyboard, a mouse, a touch screen, etc.; an output unit 117, such as various types of displays, speaker components, etc.; a storage unit 118, such as a magnetic disk, an optical disc, etc.; and a communication unit 119, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 119 allows the electronic device 11 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks. The controller 111 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the controller 111 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The controller 111 executes the various methods and processes described above, such as a recipe creation method based on cooking audio.

[0168] In some embodiments, the recipe creation method based on cooking audio can be implemented as computer-executable instructions that are tangibly embodied in a computer-readable storage medium, such as the storage unit. In some embodiments, part or all of the computer-executable instructions can be loaded and / or installed onto the electronic device 11 via the ROM 112 and / or the communication unit 119. When the computer-executable instructions are loaded into the RAM 113 and executed by the controller 111, one or more steps of the recipe creation method based on cooking audio described above can be executed. Alternatively, in other embodiments, the controller 111 can be configured to execute the recipe creation method based on cooking audio by any other suitable means (e.g., by means of firmware).

[0169] The various embodiments of the systems and techniques described above can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGA), application-specific integrated circuits (ASIC), application-specific standard products (ASSP), system-on-chip systems (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof.

[0170] These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor. The programmable processor can be a special-purpose or general-purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0171] The computer-executable instructions for implementing the method of the embodiments of the present invention can be written in any combination of one or more programming languages. These computer-executable instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that the computer-executable instructions, when executed by the processor, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The computer-executable instructions can be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0172] In the context of the embodiments of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device.

[0173] The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the computer-readable storage medium can include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0174] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in the embodiments of the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions of the embodiments of the present invention can be achieved, and no limitation is imposed herein.

[0175] The above specific implementation manners do not constitute a limitation on the protection scope of the embodiments of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the embodiments of the present invention shall be included within the protection scope of the embodiments of the present invention.

Claims

1. A document processing method, characterized in that: include: Determine a corresponding reading order based on the page number of the document to be processed, and convert the document to be processed into a single-page atlas according to the reading order; Performing layout analysis on the single-page atlas, and splitting the single-page atlas into a single-column atlas sequence according to the layout analysis result, wherein the single-column atlas sequence is sequentially arranged according to the reading order; Extracting the single-column pages in the single-column image sequence in sequence according to the reading order, and performing target detection on the corresponding single-column pages according to a preset scanning order, so as to determine the page elements therein based on the detection results; Performing target recognition on the page element to determine the position information of the page element in the corresponding single-column page, and converting the page element into a page object in a target format based on the recognition result; The page objects are rearranged based on the reading order, the scanning order and the position information, and assembled into a document result in a target format.

2. The document processing method according to claim 1, characterized in that: The step of determining a corresponding reading order based on the page number of the document to be processed, and converting the document to be processed into a single-page atlas according to the reading order, comprises: Converting the content of each page in the document to be processed into a single page image; The single-page images are converted into a single-page atlas based on the reading order.

3. The document processing method according to claim 2, characterized in that: The performing layout analysis on the single-page atlas and splitting the single-page atlas into a single-column atlas sequence according to the layout analysis result, wherein the single-column atlas sequence is sequentially arranged according to the reading order, including: Performing layout analysis on each of the single-page images in the single-page atlas to determine a segmentation strategy for each of the single-page images; The single-page image is split into a single-column image sequence based on the segmentation strategy, and the single-column image sequence is extended according to the reading order.

4. The document processing method according to claim 1, characterized in that: The page elements include images and tables; The target detection is performed on the corresponding single-column page according to the preset scanning order to determine the page elements therein based on the detection result, including: Target detection is performed on the corresponding single-column page according to a preset scanning order, and images and tables in the single-column page are marked.

5. The document processing method according to claim 4, characterized in that: The target recognition of the page element to determine the position information of the page element in the corresponding single-column page, and converting the page element into a page object in a target format based on the recognition result, includes: When the page element is an image, using a cutout tool to cut out the target image; The target image is named and uploaded based on the prompt word instruction of the multimodal large model, and a remote access hyperlink of the target image is generated; Based on the position information, fill the cut-out portion of the single-column page with a solid color and add an image placeholder; In the case where the page element is a table, using a cutout tool to cut out the table as a target image; Outputting the table content in the target image into a document format based on the prompt word instruction of the multimodal large model; Based on the position information, the extracted portion of the single-column page is filled with a solid color and a table placeholder is added.

6. The document processing method according to claim 5, characterized in that: The document format includes a Markdown format, and the method further includes: After filling the cut-out part of the single-column page with a solid color and adding an image placeholder, all the content is recognized and output as a document in Markdown format based on the prompt word instruction of the multimodal large model, retaining the original layout in the document and replacing the image placeholder with a Markdown image remote access hyperlink; After filling the extracted part of the single-column page with a solid color and adding a table placeholder, all the content is recognized and output as a document in Markdown format based on the prompt word instructions of the multimodal large model. The original layout is retained in the document, and the table placeholder is replaced with a Markdown table.

7. The document processing method according to claim 6, characterized in that: The step of rearranging the page objects based on the reading order, the scanning order and the position information and assembling the page objects into a document result in a target format includes: Based on the reading order, the scanning order and the position information, the documents in the Markdown format are spliced ​​and assembled to obtain a document result in the Markdown format; Among them, the prompt word instructions of the multimodal large model include: the role definition, rules, and workflow of the multimodal large model; when outputting image content as a document, identifying font bold, font color, font italic, formulas, tables, links, pictures, multi-level headings, unordered lists, and ordered lists; when outputting documents in Markdown format, retaining font style, color, mathematical formulas, links, and tables.

8. A document processing device, characterized in that: include: An information acquisition module, used to determine a corresponding reading order based on the page number of the document to be processed, and convert the document to be processed into a single-page atlas according to the reading order; An information processing module, configured to perform layout analysis on the single-page atlas, and split the single-page atlas into a single-column image sequence according to the layout analysis result, wherein the single-column image sequence is sequentially extended according to the reading order; extract the single-column pages in the single-column image sequence in sequence according to the reading order, and perform target detection on the corresponding single-column pages according to a preset scanning order, so as to determine the page elements therein based on the detection result; Performing target recognition on the page element to determine the position information of the page element in the corresponding single-column page, and converting the page element into a page object in a target format based on the recognition result; The page objects are rearranged based on the reading order, the scanning order and the position information, and assembled into a document result in a target format.

9. An electronic device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the document processing method according to any one of claims 1 to 7 when executing the program.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a computer program, and when the computer program is executed, the electronic device where the computer-readable storage medium is located is controlled to execute the document processing method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Document layout analysis model training method, application method, computer device and computer readable storage medium

    CN117649670A

  • Document rearrangement method and device

    CN118095204A

  • PDF document analysis method, electronic equipment, storage medium and program product

    CN118114648A

  • Rendering process-based PDF element extraction method suitable for RAG

    CN119337817A

Cited By

  • Document processing method and device, electronic equipment and storage medium

    CN121412479A

  • Document processing method and apparatus, electronic device, and storage medium

    CN121412479B

  • Cross-end copying method and related device

    CN121597446A