Table detection structured output method and device based on multi-model collaboration and storage medium

Through the multi-model collaboration table detection method, combined with deep learning and large language models, the problems of low accuracy of table detection and difficulty in combining multi-modal information in the existing technology are solved, and efficient and accurate table detection and structured output are achieved.

CN120148045APending Publication Date: 2025-06-13HANG ZHOU GAO XIN BIN JIANG SHUI WU YOU XIAN GONG SI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510217002.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

When existing table detection methods deal with complex and variable table layouts, they have low accuracy and poor generalization capabilities, and it is difficult for a single model to effectively combine information from different modes, resulting in unsatisfactory table detection results.

Method used

The table detection method based on multi-model collaboration is adopted, and the images are targetedly detected through pre-trained deep learning table detection model and text detection model, and the table position coordinates and text content are processed in combination with the large language model to output structured data.

Benefits of technology

It realizes fast, efficient and accurate table detection and structured output, improves the accuracy and automation level of table recognition, adapts to complex and changeable table styles, and improves the efficiency of data extraction and processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120148045A_ABST
    Figure CN120148045A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of table recognition, in particular to a multi-model collaboration-based table detection structured output method and device, equipment and a storage medium, and the method comprises the steps of obtaining a table picture containing a target table in response to a detection instruction for the target table; performing target detection on the image by using a pre-trained deep learning table detection model, and obtaining position coordinates of a table according to a detection result; performing target detection on the image by using a pre-trained deep learning character detection model, and obtaining position coordinates and character types of characters according to a detection result; and calling different character recognition models for character content recognition according to character categories, processing the table position coordinates, the character position coordinates and the character content, outputting table character content and table coordinates, and sending the output into a large language model to obtain a final text key value pair matching result. According to the invention, manual intervention and misrecognition are reduced, and the efficiency and accuracy of automatic processing are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of table detection, and specifically to a method, device, equipment and storage medium for structured output of table detection based on multi-model collaboration. Background Art

[0002] Today, with the increasing development of informatization and digitization, table data in various electronic documents and images are widely used in fields such as enterprise management, financial analysis, and medical records. As an important form of structured data representation, tables usually contain a large amount of key information, such as numerical data, classification information, etc. However, since tables are usually embedded in documents or images, and their formats, layouts, and contents are diverse, traditional manual extraction or single technical methods are difficult to handle the detection and structured output of complex tables. Most existing table detection methods rely on rule-based template matching or single deep learning models for recognition, but these methods often face problems of low accuracy and poor generalization ability when dealing with complex and variable table layouts.

[0003] In addition, a single model is difficult to effectively combine different modalities of information, such as visual features in images and text content in tables, resulting in unsatisfactory table detection results, especially when there are overlaps, rotations, misalignments, etc. between tables. To overcome the above problems, a table detection method based on multi-model collaboration has emerged. By combining multiple technologies such as computer vision, natural language processing, and deep learning, it is possible to extract features from different modalities of information and perform joint learning, thereby achieving more accurate and robust table detection and structured output. This method has strong adaptability and flexibility, can handle complex and variable table styles, and improve the automation level of table recognition and data extraction. Summary of the Invention

[0004] (1) Technical Problems to be Solved

[0005] Aiming at the deficiencies of the prior art, the present invention provides a method, device, equipment and storage medium for structured output of table detection based on multi-model collaboration, which can quickly, efficiently and accurately detect and identify tables and output structured data, and solve the problem of automation of the table recognition process.

[0006] (2) Technical Solutions

[0007] To achieve the above object, the present application provides the following technical solutions: A quick sterilizer for corn silk tea, including a method, device, equipment and storage medium for structured output of table detection based on multi-model collaboration, characterized in that the method includes:

[0008] Respond to the detection instruction for the target table, and obtain the table picture containing the target table;

[0009] Perform object detection on the image using a pre-trained deep learning table detection model, and obtain the position coordinates of the table according to the detection results; then perform object detection on the image using a pre-trained deep learning text detection model, and obtain the position coordinates and text categories of the text according to the detection results; call different text recognition models according to the text categories to perform text content recognition, process the above table position coordinates, text position coordinates, and text content, output the table text content and table coordinates, and send the output to a large language model to obtain the final text-key value pair matching result.

[0010] Preferably, the detection results of using the pre-trained deep learning table detection model include:

[0011] The pixel position coordinates of the upper left and lower right corners of each rectangular block in the table.

[0012] Preferably, the results of using the pre-trained deep learning text detection model include:

[0013] The pixel position coordinates of the upper left and lower right corners of the circumscribed rectangle of each text in the input image;

[0014] The category information of each text in the input image (specifically including ordinary printed text, handwritten numbers and dates).

[0015] Preferably, the pre-trained deep learning text content recognition OCR model includes:

[0016] Before recognition, crop it according to the obtained position coordinates of the text circumscribed rectangle;

[0017] Call different deep learning text content recognition OCR models according to the category of the cropped text image. If the category of the text image is ordinary printed text, use a deep learning text content recognition OCR model with a small number of parameters. If the category of the text image is handwritten numbers and dates, use a deep learning text content recognition OCR model with a large number of parameters;

[0018] Summarize the results output by the two models to obtain the coordinates of the text circumscribed rectangle and the text information.

[0019] Preferably, the processing of the table position coordinates, text position coordinates, and text content includes:

[0020] Integrate the obtained position coordinates of the table, position coordinates of the text, and content information of the text;

[0021] According to the position coordinates of the text and the position coordinate information of the table, bind the table and the positions within the table according to the inclusion relationship of the position coordinates; finally, use the text content within the table as the content of the table, and output the position coordinates of the table and the text content of the table.

[0022] A method for structured output of table detection based on multi-model collaboration. When sending the position coordinates of the table and the text content of the table into the large language model for processing, it includes:

[0023] Reorganize the content of the position coordinates of the above-mentioned table and the text content of the table. Each small rectangle content within each table becomes an element, including the position coordinates of the rectangle and the text content.

[0024] Convert all rectangle elements (including the position coordinates of the rectangle and the text content) into text and input it into the large language model.

[0025] The large language model used this time is fine-tuned and trained on specific data. The input is text (including the position coordinates of the rectangle and the text content), and the output is the key-value pair matching result.

[0026] Through the processing of the large language model, the final key-value pair result can be obtained.

[0027] An apparatus and device for structured output of table detection based on multi-model collaboration. The apparatus and device include:

[0028] An acquisition module, configured to acquire a table picture including the target table in response to an instruction to detect the target table.

[0029] A detection module, configured to detect the table picture using a deep learning object detection model to obtain the detection results of the table and the text.

[0030] An identification module, configured to perform cropping according to the detection results of the table and the text, and then call different deep learning optical character recognition (OCR) models according to the text category to obtain the recognition results.

[0031] A fusion module, configured to reorganize the content of the position coordinates of the above-mentioned table and the text content of the table to obtain the information results of each unit in the table.

[0032] A large language model module, configured to convert all rectangle elements (including the position coordinates of the rectangle and the text content) into text, input it into the large language model, and obtain the key-value pair results of the target table.

[0033] Preferably, the electronic device includes a processor and a memory. When the processor executes the computer program stored in the memory, it implements the method for structured output of table detection based on multi-model collaboration according to any one of claims 1 to 8.

[0034] (III) Beneficial Effects

[0035] Compared with the prior art, the present invention provides a method, apparatus, device and storage medium for structured output of table detection based on multi-model collaboration, having the following beneficial effects:

[0036] 1. Improve the accuracy of table detection:

[0037] By combining a convolutional neural network (CNN) and object detection models (such as YOLO, Faster R-CNN), it is possible to accurately locate table regions in images or documents, especially having significant advantages in table detection in complex scenarios (such as overlapping, rotating, and misaligned tables). This high-precision table detection helps reduce manual intervention and misidentification, and improves the efficiency and accuracy of automated processing.

[0038] 2. Multi-modal fusion improves the effect of structured analysis:

[0039] By combining image features with text information extracted by OCR technology and through a multi-modal learning mechanism, accurate parsing of rows, columns and fields inside the table is achieved. The collaborative extraction of visual information and text information enables the model to better understand the semantic structure of the table, significantly improving the quality of structured output for complex tables, especially having obvious advantages when dealing with tables in multiple languages, multiple styles or complex layouts.

[0040] 3. Automated processing and efficient data extraction:

[0041] This technology can automatically identify the row and column relationships of tables and quickly and accurately convert table content into structured data (such as CSV, Excel, etc. formats), greatly improving the efficiency of data extraction and processing. Whether it is financial statements, statistical data, invoice processing, or other forms of documents, automated data scraping and collation can be achieved, reducing the time cost and error rate of manual intervention.

[0042] 4. Adapt to complex documents and table styles:

[0043] Since this technology combines deep learning and multi-modal learning, it can flexibly adapt to various complex document formats and table styles, such as irregular table content layouts, rotated or distorted table regions, etc. This enables the technology to still maintain a high recognition and structured output effect when facing complex or diverse tables.

[0044] 5. Wide range of extended applications:

[0045] This technology has high versatility and can be widely applied to various scenarios that require tabular data processing, such as intelligent office work, electronic document processing, financial auditing, medical records, invoice management, etc. Its highly automated feature can provide strong support for the digital transformation of all industries.

[0046] 6. Improve data quality and decision support:

[0047] Through efficient and accurate tabular data extraction, enterprises and institutions can obtain high-quality structured data. This not only helps improve the quality of data analysis and decision-making but also provides more accurate data support for subsequent data mining, machine learning, and artificial intelligence analysis.

[0048] 7. Reduce labor costs and errors:

[0049] The automated table recognition and data output system significantly reduces the need for manual input, thereby reducing labor costs, improving work efficiency, and effectively avoiding errors and omissions that may occur during the manual processing. For large-scale data processing, especially in scenarios that require processing a large number of tabular files, it can greatly improve productivity.

[0050] 8. Intelligent applications applicable to multiple fields:

[0051] Whether in finance, healthcare, education, logistics, or the administrative management of governments and enterprises, this technology can provide strong support. Whether it is automated report generation, contract data extraction, or data parsing of scanned documents, it can achieve fast table recognition and data structuring, promoting the intelligent transformation and implementation of automated applications in various industries. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those skilled in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0053] Figure 1 It is a schematic flow chart of a method for table detection and structured output based on multi-model collaboration in an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0054] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.

[0055] The flowcharts shown in the accompanying drawings are only illustrative, not necessarily including all contents and operations / steps, nor necessarily executed in the described order. For example, some operations / steps can be decomposed, combined, or partially merged, so the actual execution order may change according to the actual situation. In addition, although the functional modules are divided in the device schematic diagram, in some cases, it can be different from the module division in the device schematic diagram.

[0056] Embodiments of the present application provide a method, device, equipment, and storage medium for structured output of table detection based on multi-model collaboration. It is used for structured output of table detection based on multi-model collaboration, for table detection, text detection, and structured output. On the one hand, it provides the accuracy and speed of table detection, and on the other hand, it realizes the conversion of tables into structured data, accelerating the automated processing of table information. Exemplarily, in the process of processing wired table images, due to the influence of shooting angles and light, ordinary algorithms often cannot accurately detect tables and the text in the images, especially for the recognition of handwritten dates, which is quite difficult. The table detection can be carried out according to the method for structured output of table detection based on multi-model collaboration in the embodiments of the application. By the collaborative work of multiple models, the accuracy and speed of text recognition can be greatly improved. At the same time, for different table position arrangements, it is impossible to judge the corresponding relationship between table contents. By introducing a large language model, the corresponding relationship between tables can be converted into structured key-value pairs, which is beneficial to the subsequent process flow processing of table information.

[0057] Next, some embodiments of the present application will be described in detail in conjunction with the accompanying drawings. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0058] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of a method for structured output of table detection based on multi-model collaboration provided by an embodiment of the present application.

[0059] As Figure 1 shown, the method for structured output of table detection based on multi-model collaboration may include the following steps S110-step S50.

[0060] Step S110, obtain a document image, which can adapt to different sizes of content.

[0061] Exemplarily, a document image of the document to be detected is obtained by means of photographing, scanning or document format conversion. For example, the document to be detected is converted into a color image by photographing.

[0062] Step S120: Input the image into a table detection model and a text detection model to obtain the positions of the tables, the positions of the text, and the categories of the text.

[0063] By inputting the image into a deep learning table detection model, this model can detect the position coordinates of each table rectangular block in the image and return the position coordinates of the upper left corner and the lower right corner of each table rectangular block

[0064] By inputting the image into a deep learning text detection model, this model can detect the position coordinates of the circumscribed rectangular block of the text block in the image and return the position coordinates of the upper left corner and the lower right corner of the text rectangular block

[0065] Step S130: Crop the text image from the original image according to the position of the text, and call different text recognition models according to the category of the text (ordinary printed text and handwritten date) to output the recognized text content

[0066] According to the text position coordinates output by the text detection model, the text is cropped into small images. According to the category of each text box output by the text detection model, different text recognition models are used for processing according to different categories. Ordinary printed text uses an OCR text recognition model with a small number of parameters, and handwritten date text uses an OCR text recognition model with a large number of parameters.

[0067] The OCR text recognition model with a small number of parameters uses a lightweight deep learning model based on resnet. It has a fast response speed and low GPU resources required for calculation, and can achieve the effect of quickly recognizing text.

[0068] The OCR text recognition model with a large number of parameters uses a deep learning model based on transformer. Its response speed is slightly slower and the GPU resources required for calculation are larger, but the recognition accuracy is high. It can recognize complex handwritten text and can achieve the effect of accurately recognizing text.

[0069] Step S140: Screen out the text and content contained in each small grid in the table according to the positional inclusion relationship between the table and the text.

[0070] We can judge whether the text belongs to the table according to the positional relationship between the text position and the table, that is, whether the center point of the text falls within the rectangle. Therefore, the text can be bound to each table block.

[0071] If there are multiple texts in a table, we will sort the texts according to their position coordinates from top to bottom and from left to right, and combine them into new text content.

[0072] After binding the content according to the text and the table, the information we need is only the position information (coordinates of the upper left corner and the lower right corner) of each table block and the text information (pure text) of the table block.

[0073] Step S150: After integrating the position information and content of the table, input it into the large language model to output formatted key-value pair information.

[0074] Convert the result obtained in step S140 into a string and input it into the fine-tuned large language model, which can output the corresponding relationship of the table blocks. Here, for the sake of understanding, we give an example

[0075]

[0076] Input: [[[0,0],[1,1], “Name”],

[0077] [[1,0],[2,1], “Xiaoming”],

[0078] [[0,1],[2,2], “Evaluation for this semester is excellent”]

[0079] [[0,2],[1,3], “Total score”],

[0080] [[1,2],[2,3], “100 points”]

[0081] Output: {“Name”: “Xiaoming”,

[0082] “Evaluation for this semester is excellent”: “”,

[0083] “Total score”: “100 points”}

[0084] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principle and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A table detection structured output method, device and storage medium based on multi-model collaboration, characterized in that: The method comprises: In response to a detection instruction for a target table, acquiring a table image including the target table; Use a pre-trained deep learning table detection model to perform target detection on the image, and obtain the position coordinates of the table based on the detection results; then use a pre-trained deep learning text detection model to perform target detection on the image, and obtain the position coordinates and text category of the text based on the detection results; call different text recognition models according to the text category to perform text content recognition, process the above-mentioned table position coordinates, text position coordinates, and text content, output the table text content and table coordinates, and send the output to the large language model to obtain the final text key-value pair matching result.

2. According to the method for structured output of table detection based on multi-model collaboration according to claim 1, it is characterized in that: The detection results using the pre-trained deep learning table detection model include: The pixel position coordinates of the upper left corner and lower right corner of each rectangular block in the table.

3. The method for structured output of table detection based on multi-model collaboration according to claim 1, characterized in that: The results of using the pre-trained deep learning text detection model include: The pixel position coordinates of the upper left corner and lower right corner of the bounding rectangle of each character in the input image; Enter the category information to which each character in the image belongs.

4. The method for structured output of table detection based on multi-model collaboration according to claim 1, characterized in that: The pre-trained deep learning text content recognition OCR model used includes: Before recognition, the text circumscribed rectangle is clipped according to the position coordinates of the obtained text; The cropped text image is called different deep learning text content recognition OCR models according to its category. If the category of the text image is ordinary printed text, a deep learning text content recognition OCR model with a small parameter amount is used. If the category of the text image is a handwritten digital date, a deep learning text content recognition OCR model with a large parameter amount is used. The output results of the two models are summarized to obtain the coordinates of the text circumscribed rectangle and text information.

5. The method for structured output of table detection based on multi-model collaboration according to claim 1, characterized in that: The processing of the table position coordinates, text position coordinates, and text content includes: Integrate the obtained table position coordinates, text position coordinates, and text content information; According to the position coordinates of the text and the position coordinates of the table, and according to the inclusion relationship of the position coordinates, the table and the position in the table are bound; finally, the text content in the table is used as the content of the table, and the position coordinates of the table and the text content of the table are output.

6. The method for table detection structured output based on multi-model collaboration according to claim 1, characterized in that: The step of sending the position coordinates of the table and the text content of the table to the large language model for processing includes: The position coordinates of the above table and the text content of the table are reorganized, and each small rectangular content in the table becomes an element, including the position coordinates and text content of the rectangle; Convert all rectangular elements into text and input it into the large language model; The large language model used this time is fine-tuned and trained on specific data, with text as input and key-value pair matching results as output; Through the processing of the large language model, the final key-value pair results can be obtained.

7. According to the multi-model collaborative table detection structured output device device of claim 1, it is characterized in that , the device comprises: An acquisition module, configured to acquire a table image including the target table in response to an instruction to detect the target table; A detection module, used to detect the table image using a deep learning target detection model to obtain detection results of the table and text; A recognition module, used to cut the table and text according to the detection results, and then call different deep learning text recognition OCR models according to the text category to obtain the recognition results; A fusion module is used to reorganize the position coordinates of the table and the text content of the table to obtain the information results of each unit in the table; The large language model module is used to convert all rectangular elements into text, input it into the large language model, and obtain the key-value pair results of the target table.

8. A structured output storage medium for table detection based on multi-model collaboration according to claim 1, characterized in that The electronic device includes a processor and a memory, and the processor is used to implement the table detection structured output method based on multi-model collaboration as described in any one of claims 1 to 8 when executing the computer program stored in the memory.

9. The structured output storage medium for table detection based on multi-model collaboration according to claim 1, characterized in that The storage medium stores a computer program, and when the computer program is executed by the processor, it implements the table detection structured output method based on multi-model collaboration as described in any one of claims 1 to 6.