Picture document cross-page table merging method and device and related equipment

By disassembling and merging independent pages, combining the table header recognition model and the table body column sequence width judgment, efficient merging of complex span table tables is achieved, solving the problems of low recognition accuracy and high memory usage in the existing technology, and improving the effects of OCR recognition and data extraction.

CN120579524APending Publication Date: 2025-09-02BEIJING TAIXIN TIANCHENG TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510691606.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-09-02

AI Technical Summary

Technical Problem

When processing complex spanning tables, the recognition accuracy is low and the misalignment rate after splicing is high, which affects the OCR recognition and data extraction effect. Moreover, the memory usage is too high when processing large documents, which cannot meet real-time needs.

Method used

By disassembling the image document into independent pages of the same size, using the pre-trained header recognition model to judge the consistency of the table header, and merging it according to the number of columns and column width of the table body, automatic merging of the spreadsheet table is achieved.

Benefits of technology

It improves the accuracy of cross-page table recognition, reduces splicing misalignment rate, reduces memory usage, and meets the real-time processing needs of large documents.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120579524A_ABST
    Figure CN120579524A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a picture document cross-page table merging method and device and related equipment. The method comprises the following steps: disassembling a picture document into independent pages with the same size, and arranging the independent pages according to page numbers to obtain a page set; performing header identification on the next page of the two adjacent pages in the page set through a pre-trained header identification model, and determining whether the next page of the two adjacent pages has a header or not; if the header exists, judging whether the header of the next page of the two adjacent pages is consistent with the header of the previous page or not, if the headers are consistent, removing the header of the next page, and merging the bodies of the two pages; if the headers are inconsistent, not carrying out merging processing; if the header does not exist, judging the column numbers of the table bodies of the two adjacent pages, and carrying out merging judgment; and if the column numbers of the table bodies of the two adjacent pages are different, not merging. According to the method, whether complex cross-page tables are combined or not is judged, and automatic processing of data is promoted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the field of text processing technology, and in particular to a method, device and related equipment for merging tables across pages of an image document. Background Art

[0002] Tables that span multiple pages are a common form of content in professional image documents such as legal documents, financial statements, and scientific research reports. During the offline indexing phase of the current RAG system, the presence of tables spanning multiple pages causes the table content to be stored in different chunks after identification. This results in the incomplete table content being retrieved during the retrieval phase, which in turn affects the quality of the generation phase.

[0003] Traditional image collection processing technology has two core flaws:

[0004] Physical paging dependency: The fixed page margin-based splitting method causes the "table title on page 1, content on page 2" cross-page content to be abruptly split, destroying the integrity of the table.

[0005] Inefficient global analysis: The full-image document scanning method requires traversing all pages for layout analysis. The computational complexity increases exponentially with the number of pages in the image document. When processing image documents with more than 1,000 pages, the memory usage exceeds 8GB, which cannot meet real-time processing requirements.

[0006] The existing technology has a recognition accuracy rate of less than 70% in complex scenarios such as blurred table lines and cell merging, and the misalignment rate of the table after splicing is as high as 40%, which seriously affects subsequent OCR recognition, data extraction and other operations of the image document. Summary of the Invention

[0007] The embodiments of the present invention provide a method, apparatus and related equipment for merging tables across pages of image documents, aiming to solve the technical problems of identifying complex tables across pages and achieving overlapping splicing effects in the prior art.

[0008] In a first aspect, an embodiment of the present invention provides a method for merging a table across pages of an image document, comprising:

[0009] Split the image document into independent pages of the same size, and arrange each independent page by page number to obtain a page set;

[0010] Performing header recognition on the next page of two adjacent pages in the page set using a pre-trained header recognition model to determine whether a header exists on the next page of the two adjacent pages;

[0011] If a table header exists, determine whether the table header of the next page is consistent with the table header of the previous page. If the table headers are consistent, remove the table header of the next page and merge the table bodies of the two pages; if the table headers are inconsistent, do not merge;

[0012] If there is no table header, determine whether the number of table body columns on two adjacent pages is the same;

[0013] If the number of columns in the table bodies of two adjacent pages is the same, then the table bodies of the two adjacent pages are checked to see if their column widths are the same. If they are the same, the table bodies of the two pages are merged. If they are different, the table bodies of the two pages are not merged.

[0014] If the number of table body columns on two adjacent pages is different, they will not be merged.

[0015] In a second aspect, an embodiment of the present invention provides a device for merging tables across pages of an image document, comprising:

[0016] A disassembly module is used to disassemble the image document into independent pages of the same size, and arrange each independent page by page number to obtain a page set;

[0017] A header recognition module is used to perform header recognition on the next page of two adjacent pages in the page set using a pre-trained header recognition model to determine whether a header exists on the next page of the two adjacent pages;

[0018] The first merging module is used to determine whether the header of the next page of two adjacent pages is consistent with the header of the previous page if the headers exist. If the headers are consistent, the header of the next page is removed and the table bodies of the two pages are merged; if the headers are inconsistent, no merging is performed;

[0019] The second merging module is used to determine whether the number of table body columns on two adjacent pages is the same if there is no table header; if the number of table body columns on two adjacent pages is the same, then continue to determine whether the table body column widths on the two adjacent pages are the same; if the column widths are the same, then merge the table bodies of the two pages; if the column widths are different, then do not merge; if the number of table body columns on two adjacent pages is different, then do not merge.

[0020] In a third aspect, an embodiment of the present invention further provides a computer device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the computer program, the method for merging cross-page tables of image documents as described in the first aspect above is implemented.

[0021] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the processor executes the method for merging cross-page tables of image documents described in the first aspect above.

[0022] Embodiments of the present invention provide a method, device, and related equipment for merging tables across pages in an image document. The method breaks down an image document into independent pages of equal size and arranges each independent page by page number to obtain a page set. A pre-trained header recognition model is used to identify the header of the next page of two adjacent pages in the page set to determine whether a header exists on the next page of the two adjacent pages. If a header exists, a determination is made as to whether the header of the next page of the two adjacent pages is consistent with the header of the previous page. If the headers are consistent, the header of the next page is removed and the table bodies of the two pages are merged. If the headers are inconsistent, no merging is performed. If no header exists, a determination is made as to whether the table body columns of the two adjacent pages are the same. If the table body columns of the two adjacent pages are the same, a determination is made as to whether the table body column widths of the two adjacent pages are the same. If the column widths are the same, the table bodies of the two pages are merged. If the column widths are different, no merging is performed. If the table body columns of the two adjacent pages are different, no merging is performed. This method determines whether complex tables across pages should be merged and facilitates automated data processing. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0024] Figure 1 A flowchart of an embodiment of a method for merging tables across pages of image documents provided by an embodiment of the present invention;

[0025] Figure 2 for Figure 1 A schematic diagram of a sub-process of an embodiment of step S120;

[0026] Figure 3 A schematic diagram of a method for merging tables across pages of image documents according to an embodiment of the present invention;

[0027] Figure 4 A schematic diagram of a method for merging tables across pages of image documents according to an embodiment of the present invention;

[0028] Figure 5A schematic diagram of a sub-flow diagram of step S2211 in an embodiment of a method for merging tables across pages of image documents provided by an embodiment of the present invention;

[0029] Figure 6 A schematic diagram of a method for merging tables across pages of image documents according to an embodiment of the present invention;

[0030] Figure 7 A schematic diagram of a method for merging tables across pages of image documents according to an embodiment of the present invention;

[0031] Figure 8 A schematic diagram of a method for merging tables across pages of image documents according to an embodiment of the present invention;

[0032] Figure 9 A schematic diagram of a method for merging tables across pages of image documents according to an embodiment of the present invention;

[0033] Figure 10 A schematic block diagram of a device for merging tables across pages of image documents provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0034] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0035] It will be understood that when used in this specification and the appended claims, the terms “comprises” and “comprising” indicate the presence of described features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.

[0036] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the present invention. As used in the specification and appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise.

[0037] It should be further understood that the term "and / or" used in the present description and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.

[0038] See also Figure 1This is a flow chart of a method for merging tables across pages of an image document provided by an embodiment of the present invention. The method includes steps S110 to S140.

[0039] Step S110: Decompose the image document into independent pages of the same size, and arrange each independent page by page number to obtain a page set;

[0040] In this embodiment, a PDF parsing engine is used to break the image document into independent pages. Each independent page is converted into an image of the same size, ensuring that each independent page contains visual layout data (coordinates, fonts, and lines) and text content. The images of each independent page are arranged in order by page number to obtain a page set.

[0041] Step S120: Using a pre-trained header recognition model, perform header recognition on the next page of two adjacent pages in the page set to determine whether a header exists on the next page of the two adjacent pages.

[0042] Step S131: If a table header exists, determine whether the table header of the next page of the two adjacent pages is consistent with the table header of the previous page. If the table headers are consistent, remove the table header of the next page and merge the table bodies of the two pages; if the table headers are inconsistent, do not merge;

[0043] In this embodiment, the header recognition model can be constructed based on existing object recognition algorithms such as YOLO and SSD. The constructed header recognition model detects the text content of the first row of the table to determine whether a header exists on the next page between two adjacent pages. If a header exists on the next page between two adjacent pages, it is necessary to determine whether the header of the next page is consistent with the header of the top page. If the headers are consistent, it is determined that the table bodies of the two pages belong to a cross-page table, and the table bodies of the two pages need to be merged. If the headers are inconsistent, it is determined that the table bodies of the two pages do not belong to a cross-page table, and no merging is performed.

[0044] Step S132: If the table header does not exist, determine whether the number of table body columns on two adjacent pages is the same;

[0045] Step S141: If the number of table body columns on two adjacent pages is the same, then determine whether the table body column widths on the two adjacent pages are the same; if the column widths are the same, merge the table bodies on the two pages; if the column widths are different, do not merge;

[0046] Step S142: If the number of table body columns of two adjacent pages is different, no merging is performed.

[0047] In this embodiment, if the table header does not exist on the next page between two adjacent pages, the table on the next page may be a new table; it may also be a table spanning the previous page's table body and belonging to the same table. Therefore, when determining that the table body on the next page between two adjacent pages does not have a table header, it is necessary to further determine whether the table body columns on the two adjacent pages are the same. If the number of columns and column widths are the same, it indicates that the two table bodies belong to the same table, and the table bodies on the two pages are merged. If the column widths are different, it indicates that the two table bodies do not belong to the same table, and the merging is not performed.

[0048] like Figure 2 As shown, in one embodiment, step S120 includes:

[0049] Step S210: Check whether the table layout and column information of two adjacent rows of cells on two pages with the same table headers are consistent;

[0050] Step S221: If yes, merge the two adjacent rows of cells;

[0051] Step S222: If not, then perform splicing processing on the two adjacent rows of cells.

[0052] In this embodiment, the table bodies of two pages with consistent table headers are merged, specifically including: detecting the two rows of cells connected by two adjacent pages with consistent table headers, and if the arrangement and column information of the two rows of cells connected by two adjacent pages with consistent table headers are consistent, then the two connected rows of cells are merged. This type of situation is generally that in the two connected rows of cells, the table to which the first row of cells belongs has not ended on the previous page, and the previous page cannot fully accommodate the entire table to which it belongs, so the first row of cells is used as the starting cell of the next page to indicate that the two table bodies are connected, so the two rows of cells in this case need to be merged. If the two rows of cells connected by two adjacent pages with consistent table headers do not meet the conditions of the same cell arrangement, consistent column information, and no cross-row cells, it means that the two rows of cells belong to the same table, but are cells in different rows, so the two connected rows of cells need to be spliced.

[0053] like Figure 3 As shown in the figure, the column information in the 29th row of cells is inconsistent with the column information in the 30th row of cells before merging. Then the 29th row of cells and the 30th row of cells are spliced ​​together to obtain the following: Figure 4 The table shown.

[0054] In one embodiment, step S221 includes:

[0055] Step S2211: Take each two opposing cells in the two adjacent rows of cells as a cell group, and detect whether the cell group in the two adjacent rows of cells has segmented content based on the fine-tuned large language model. If so, merge the corresponding cell groups; if not, do not merge the corresponding cell groups.

[0056] In this embodiment, since there may be a row of cells distributed on two pages to form two divided rows of cells, and there may be a part of content in each of the two divided rows of cells, it is necessary to perform semantic recognition on each of the two opposite cells in the merged two rows of cells to determine whether each of the two opposite cells in the two rows of cells is a divided two rows of cells. For the convenience of presentation, each of the two opposite cells in the two connected rows of cells is regarded as a cell group. Determining whether each of the two opposite cells in the two rows of cells is a divided two rows of cells specifically includes: detecting whether there is divided content in each cell group in the two connected rows of cells based on a pre-trained large language model; if there is divided content in each cell group in the two connected rows of cells, merging the corresponding cell groups; if there is no divided content in each cell group in the two connected rows of cells, not merging the corresponding cell groups, and retaining the corresponding cell groups.

[0057] In one embodiment, the semantic recognition model can also use a word segmentation model to detect the presence of segmented content in the cell group of two connected rows of cells. Specifically, the text of each cell is input into the word segmentation model and disassembled into individual words, punctuation marks and stop words. Compare the word segmentation results of the text before merging and the word segmentation results of the text after merging. If the word segmentation result of the text after merging contains the word segmentation result of the text before merging, no merging process is performed; if the word segmentation result of the text after merging does not contain the word segmentation result of the text before merging, merging process is performed. For example, the word segmentation result of the first cell text is three words a, b, and c, and the word segmentation result of the text after merging is five words a, b, c, d, and e. It can be judged that the word segmentation of the first cell is independent and unrelated, and no merging process is performed. If the word segmentation result of the first cell is three words a, b, and c, and the word segmentation result of the text after merging is four words a, b, d, and e, it can be judged that the word segmentation of the first cell is non-independent and unrelated, and merging process must be performed. If the word segmentation results for the first cell are a, b, c, and c1, where c1 is a word or single character not included in the local vocabulary, and the word segmentation results for the merged text are a, b, c, d, and e, then the word segmentation results for the first cell are not independent and must be merged. If the word segmentation results for the first cell are a, b, c, and c1, where c1 is a word or single character not included in the local vocabulary, and the word segmentation results for the merged text are a, b, c, c1, d, and e, then it is impossible to determine the granularity of the vocabulary and a large language model must be used for judgment or the local vocabulary must be expanded to meet the word segmentation requirements.

[0058] like Figure 5 As shown, in one embodiment, step S2211 includes:

[0059] Step S510: Determine whether a cell group with split contents exists in two adjacent rows of cells and use it as the split cell group;

[0060] Step S520: Determine whether all cell groups except the split cell group in the two connected rows of cells have blank contents. If so, merge all cell groups in the two connected rows of cells. Otherwise, merge only the split cell group.

[0061] In this embodiment, since there is a cell group with split contents in the two connected rows of cells, there may also be a cell with blank content. This is because the content in the original cells of a row of cells is small after it spans two pages and is not automatically split into rows, resulting in a cell with blank content. In this case, it is determined that there is a cell group with split contents in the two connected rows of cells and it is used as a split cell group. It is determined whether there are cells with blank content in all cell groups other than the split cell group in the two connected rows of cells. If so, all cell groups in the two connected rows of cells are merged; if not, only the split cell group is merged.

[0062] For example Figure 6 As shown in FIG, the semantic recognition model determines that "XX Nur·" and "YY Rexiti" are cell groups with name-type segmented content. If the other cell groups except the segmented cell group in the two connected rows of cells do not all have blank content, only the segmented cells are merged, and the other cell groups are not merged, as shown in FIG. Figure 7 For example, as shown in Figure 8, the semantic recognition model determines that "XX Nur·" and "YY Rexiti" are cell groups with name-type segmented content. If there are blank contents in other cell groups, all cell groups are merged to obtain the following: Figure 9 The table shown.

[0063] In addition, the processing results of the two adjacent pages are written into the result image document in sequence, and a processing log is generated to record the splicing position and splicing results of each two adjacent pages for users to query later.

[0064] This method breaks down an image document into independent pages of equal size and arranges each independent page by page number to obtain a page set. A pre-trained header recognition model is used to identify the header of the next page between two adjacent pages in the page set to determine whether a header exists. If a header exists, the method determines whether the header of the next page is consistent with the header of the previous page. If so, the method removes the header of the next page and merges the table bodies of the two pages. If the headers do not match, the method does not merge the two pages. If the headers do not exist, the method determines whether the table body columns of the two adjacent pages are the same. If the table body columns of the two adjacent pages are the same, the method determines whether the table body column widths of the two adjacent pages are the same. If the column widths are the same, the table bodies of the two pages are merged. If the column widths are different, the method does not merge the two pages. This method determines whether to merge complex tables that span multiple pages and facilitates automated data processing.

[0065] The embodiment of the present invention further provides a device for merging a table across pages of an image document, which is used to execute any embodiment of the aforementioned method for merging a table across pages of an image document. Figure 10 , Figure 10 FIG. 1 is a schematic block diagram of an apparatus for merging a table across pages of an image document provided by an embodiment of the present invention. The apparatus 100 for merging a table across pages of an image document can be configured in a server.

[0066] like Figure 10 As shown, the apparatus 100 for merging tables across multiple pages of an image document includes a decomposing module 110 , a header identifying module 120 , a first merging module 130 , and a second merging module 140 .

[0067] The disassembly module 110 is used to disassemble the image document into independent pages of the same size, and arrange each independent page by page number to obtain a page set;

[0068] The header recognition module 120 is configured to perform header recognition on the next page of two adjacent pages in the page set using a pre-trained header recognition model to determine whether a header exists on the next page of the two adjacent pages;

[0069] The first merging module 130 is configured to determine whether the header of the next page of two adjacent pages is consistent with the header of the previous page if the headers exist. If the headers are consistent, the header of the next page is removed and the table bodies of the two pages are merged; if the headers are inconsistent, no merging is performed;

[0070] The second merging module 140 is configured to determine whether the number of table body columns on two adjacent pages is the same if no table header exists; if the number of table body columns on two adjacent pages is the same, then determine whether the table body column widths on the two adjacent pages are the same; if the column widths are the same, then merge the table bodies of the two pages; if the column widths are different, then do not merge; if the number of table body columns on two adjacent pages is different, then do not merge.

[0071] In one embodiment, the evaluation module 140 includes:

[0072] an IoU calculation unit, configured to calculate an IoU ratio between the first coordinate set and the second coordinate set to obtain an IoU ratio of corresponding tables on two adjacent pages;

[0073] A title detection unit is used to detect whether a table title and a column title exist in the table body of the next page between two adjacent pages, and obtain a title detection result;

[0074] A scoring unit is used to calculate the splicing necessity based on the intersection-over-union ratio and the title detection result to obtain the scoring result.

[0075] In one embodiment, the splicing module 150 includes:

[0076] a merging unit configured to merge the table body of the next page with the table body of the previous page of two adjacent pages when the scoring result is lower than a preset threshold, and retain the original paging corresponding to the non-table body content of the two adjacent pages, wherein the non-table body content includes text and images;

[0077] The splicing unit is configured to splice two adjacent pages into full pages when the scoring result is higher than the threshold.

[0078] In one embodiment, the merging unit includes:

[0079] An identification subunit is used to identify whether the layout width and resolution of the table body of the next page are consistent with those of the previous page between two adjacent pages;

[0080] The table merging subunit is used to merge the table body of the next page with the table body of the previous page in two adjacent pages.

[0081] The adjusting subunit is configured to adjust the layout width and resolution of the tables in the overlapping area to be consistent and then perform merging processing if no.

[0082] In one embodiment, the apparatus 100 for merging tables across multiple pages of an image document includes:

[0083] The model fine-tuning module is configured to obtain a set of sample page pairs, wherein each pair of adjacent pages in the set is labeled with a true label indicating whether a cross-page table exists; obtain a large base visual model, and input the two adjacent pages in the set of sample page pairs into the large base visual model to predict the cross-page table, thereby obtaining a predicted label; calculate the label loss between the predicted label and the corresponding true label according to a preset loss function, and adjust the model parameters of the large base visual model based on the label loss to obtain the large visual model.

[0084] In one embodiment, the apparatus 100 for merging tables across multiple pages of an image document includes:

[0085] The recording module is used to generate a processing log to record the splicing position and splicing result of each two adjacent pages.

[0086] An embodiment of the present invention also provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned method for merging tables across pages of image documents when executing the computer program.

[0087] In another embodiment of the present invention, a computer-readable storage medium is provided. The computer-readable storage medium may be a non-volatile computer-readable storage medium. The computer-readable storage medium stores a computer program that, when executed by a processor, causes the processor to perform the above-described method for merging tables across multiple pages in an image document.

[0088] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described equipment, devices and units can refer to the corresponding processes in the aforementioned method embodiments, and will not be repeated here. Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented with electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in terms of function in the above description. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.

[0089] In the several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, or units with the same function may be combined into one unit. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection through some interfaces, devices or units, or may be an electrical, mechanical or other form of connection.

[0090] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected according to actual needs to achieve the objectives of the embodiments of the present invention.

[0091] In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0092] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the existing technology, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a magnetic disk, or an optical disk.

[0093] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and such modifications or substitutions are intended to be within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be subject to the scope of protection of the claims.

Claims

1. A method for merging tables across pages of image documents, characterized in that: include: Split the image document into independent pages of the same size, and arrange each independent page by page number to obtain a page set; Performing header recognition on the next page of two adjacent pages in the page set using a pre-trained header recognition model to determine whether a header exists on the next page of the two adjacent pages; If a table header exists, determine whether the table header of the next page is consistent with the table header of the previous page. If the table headers are consistent, remove the table header of the next page and merge the table bodies of the two pages. If the headers are inconsistent, no merging will be performed; If there is no table header, determine whether the number of table body columns on two adjacent pages is the same; If the number of columns in the table bodies of two adjacent pages is the same, then the table bodies of the two adjacent pages are checked to see if their column widths are the same. If they are the same, the table bodies of the two pages are merged. If they are different, the table bodies of the two pages are not merged. If the number of table body columns on two adjacent pages is different, they will not be merged.

2. The method for merging tables across multiple pages of image documents according to claim 1, characterized in that: If the table headers are consistent, the table header of the next page is removed, and the table bodies of the two pages are merged, including: Check whether the table layout and column information of two rows of cells connected to two adjacent pages with the same header are consistent; If so, merge the two adjacent rows of cells; If not, the two adjacent rows of cells are spliced.

3. The method for merging tables across multiple pages of image documents according to claim 2, characterized in that: The merging of two adjacent rows of cells includes: Each two opposing cells in two adjacent rows of cells are regarded as a cell group. Based on the fine-tuned large language model, it is detected whether the cell group in the two adjacent rows of cells has segmented content. If so, the corresponding cell group is merged; if not, the corresponding cell group is not merged.

4. The method for merging tables across pages of picture documents according to claim 3, characterized in that: The pre-trained semantic recognition model is used to detect whether a cell group in two adjacent rows of cells has segmented content, and if so, the corresponding cell group is merged; If not, the corresponding cell group will not be merged, including: Determine whether there is a cell group with split contents in two adjacent rows of cells and use it as the split cell group; Determine whether there is blank content in all cell groups except the split cell group in the two connected rows of cells. If so, merge all cell groups in the two connected rows of cells. Otherwise, only merge the split cell group.

5. The method for merging tables across multiple pages of picture documents according to claim 1, characterized in that: If the number of table body columns of two adjacent pages is different, then they are not merged, including: Generate a processing log to record the splicing position and splicing results of each two adjacent pages.

6. A device for merging tables across pages of image documents, characterized in that: include: A disassembly module is used to disassemble the image document into independent pages of the same size, and arrange each independent page by page number to obtain a page set; A header recognition module is used to perform header recognition on the next page of two adjacent pages in the page set using a pre-trained header recognition model to determine whether a header exists on the next page of the two adjacent pages; A first merging module is configured to determine whether the header of the next page of two adjacent pages is consistent with the header of the previous page if a header exists. If the headers are consistent, the header of the next page is removed and the table bodies of the two pages are merged; If the headers are inconsistent, no merging will be performed; The second merging module is used to determine whether the number of table body columns on two adjacent pages is the same if there is no table header; if the number of table body columns on two adjacent pages is the same, then continue to determine whether the table body column widths on the two adjacent pages are the same; if the column widths are the same, then merge the table bodies of the two pages; if the column widths are different, then do not merge; if the number of table body columns on two adjacent pages is different, then do not merge.

7. The apparatus for merging tables across multiple pages of image documents according to claim 6, characterized in that: The first merging module includes: A detection unit, configured to detect whether the table arrangement and column information of two adjacent rows of cells on two pages with the same table headers are consistent; Merge cells, used to merge the two adjacent rows of cells if yes; The splicing unit is used to splice the two adjacent rows of cells if no.

8. The apparatus for merging tables across multiple pages of image documents according to claim 7, characterized in that: The merging unit comprises: The merging subunit is used to treat each two opposing cells in two connected rows of cells as a cell group, and detect whether the cell groups in the two connected rows of cells have segmented content based on the pre-trained semantic recognition model. If so, the corresponding cell groups are merged; if not, the corresponding cell groups are not merged.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method for merging cross-page tables of image documents as described in any one of claims 1 to 5 when executing the computer program.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, enables the processor to execute the method for merging cross-page tables in an image document according to any one of claims 1 to 5.