Table merging method, table merging prediction model training method and device

CN117807968BActive Publication Date: 2026-09-25SHANGHAI HUNDSUN JUYUAN DATA SERVICE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311863551.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-12-29
Publication Date
2026-09-25
Estimated Expiration
2043-12-29

AI Technical Summary

Technical Problem

[0005]本申请的目的在于,针对上述现有技术中的不足,提供一种表格合并方法、表格合并预测模型训练方法及装置,以解决现有技术中跨页面表格合并的结果不准确的问题

Benefits of technology

[0075]本申请实施例所提供的表格合并方法、表格合并预测模型训练方法及装置,将待处理图像输入预先训练得到的表格合并预测模型中,表格合并预测模型能够预测得到物理合并预测结果以及语义合并预测序列,物理合并预测结果指示了第一表格段与第二表格段在物理结构上是否可合并,语义合并预测序列指示了第一表格段与第二表格段的各对应单元格是否可语义合并,即通过语义合并预测序列能够在单元格级别上预测单元格从上下文语义角度是否可以合并。由于结合了物理结构以及单元格级别的上下文语义信息来判断表格是否可合并,因此能够使得表格合并的结果的准确性得以显著提高。另外,语义合并预测序列中的各项分别对应第一表格段以及第二表格段的一个单元格,因此,语义合并预测序列为可变长度的序列,而不需要限定在固定的长度下,因此能够灵活适用于不同列数的表格合并场景中。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117807968B_ABST
    Figure CN117807968B_ABST
Patent Text Reader

Abstract

The application provides a table merging method and a table merging prediction model training method and device. The table merging method comprises: acquiring a to-be-processed image, the to-be-processed image comprising a first table segment and a second table segment distributed across pages; inputting the to-be-processed image into a table merging prediction model for prediction processing to obtain a physical merging prediction result and a semantic merging prediction sequence, each item in the semantic merging prediction sequence corresponding to a cell of the first table segment and the second table segment, and the value of each item in the semantic merging prediction sequence being used to indicate whether the corresponding cell on the first table segment and the second table segment can be semantically merged; and performing merging processing according to the physical merging prediction result and the semantic merging prediction sequence. The table merging result is combined with the physical structure and the context semantic information at the cell level to determine whether the table can be merged, so that the accuracy of the table merging result is significantly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and more specifically, to a table merging method, a table merging prediction model training method, and an apparatus. Background Technology

[0002] In scenarios involving document recognition and parsing, such as finance, it is often necessary to merge tables across multiple pages. Specifically, due to the limited page range of a document, the same table may be displayed on different pages. When recognizing and parsing the document, it is necessary to merge the contents of these tables.

[0003] Existing technologies provide several methods for merging tables across pages. These methods primarily determine whether tables can be physically merged by analyzing the physical structure of the tables.

[0004] However, existing methods, which analyze only the physical structure of tables, may result in inaccurate results when merging tables across pages. Summary of the Invention

[0005] The purpose of this application is to address the shortcomings of the prior art by providing a table merging method, a table merging prediction model training method, and an apparatus to solve the problem of inaccurate results in cross-page table merging in the prior art.

[0006] To achieve the above objectives, the technical solutions adopted in the embodiments of this application are as follows:

[0007] In a first aspect, one embodiment of this application provides a table merging method, including:

[0008] Obtain an image to be processed, which includes a first table segment and a second table segment distributed across a page;

[0009] The image to be processed is input into a pre-trained table merging prediction model for prediction processing to obtain a physical merging prediction result and a semantic merging prediction sequence. The physical merging prediction result is used to indicate whether the first table segment and the second table segment can be merged in terms of physical structure. Each item in the semantic merging prediction sequence corresponds to a cell of the first table segment and the second table segment, and the value of each item in the semantic merging prediction sequence is used to indicate whether the cells on the corresponding first table segment and the second table segment can be semantically merged.

[0010] Based on the physical merge prediction result and the semantic merge prediction sequence, the first table segment and the second table segment are merged.

[0011] As one possible implementation, the table merging prediction model includes: a visual feature processing network, a semantic feature processing network, and a linear processing layer;

[0012] The step of inputting the image to be processed into a pre-trained table merging prediction model for prediction processing to obtain physical merging prediction results and semantic merging prediction sequences includes:

[0013] The image to be processed is input into the visual feature processing network, which encodes the visual features to obtain visual features, and then performs positional encoding on the visual features to obtain position-encoded features.

[0014] The location-encoded features are input into the semantic feature processing network, which performs semantic feature encoding to obtain semantically encoded features.

[0015] The semantically encoded features are input into the linear processing layer for linear transformation to obtain the physical merging prediction result and the semantic merging prediction sequence.

[0016] As one possible implementation, the visual feature processing network includes multiple coding layers connected in sequence;

[0017] The step of inputting the image to be processed into the visual feature processing network, and having the visual feature processing network encode visual features to obtain visual features, includes:

[0018] The image to be processed is input into the first coding layer for encoding, and the processed features are input into the next coding layer for encoding, and so on, until the last coding layer completes the encoding process.

[0019] Max pooling is performed on the processed features of each coding layer except the last coding layer to obtain the pooled features of each coding layer.

[0020] The pooled features from each coding layer and the features encoded by the last coding layer are concatenated to obtain the visual features.

[0021] As one possible implementation, the step of positionally encoding the visual features to obtain positionally encoded features includes:

[0022] Based on the two-dimensional sine and cosine coding algorithm, the visual features are encoded in both the horizontal and vertical directions to obtain the encoded features.

[0023] As one possible implementation, the merging process of the first table segment and the second table segment based on the physical merge prediction result and the semantic merge prediction result includes:

[0024] If the physical merge prediction result indicates that the first table segment and the second table segment can be merged in physical structure, then each item in the semantic merge prediction sequence is traversed. For the current item that is traversed, if the value of the current item is a preset value, then the content of the cell corresponding to the current item in the first table segment and the cell corresponding to the current item in the second table segment are merged.

[0025] As one possible implementation, before acquiring the image to be processed, the process further includes:

[0026] Based on the page segmentation marks in the electronic document, the first initial table segment and the second initial table segment are extracted from adjacent pages of the electronic document to obtain the initial image;

[0027] The initial image is cropped and the table edges are aligned to obtain the image to be processed.

[0028] Secondly, embodiments of this application provide a method for training a table merging prediction model, including:

[0029] A training dataset is constructed based on the original electronic document containing cross-page tables;

[0030] The initial merging model is trained based on the training dataset to obtain a table merging prediction model. The prediction results of the table merging prediction model include physical merging prediction results and semantic merging prediction sequences. The physical merging prediction results are used to indicate whether the first table segment and the second table segment in the input image to be processed can be merged in terms of physical structure. Each item in the semantic merging prediction sequence corresponds to a cell in the first table segment and the second table segment, and the values ​​of each item in the semantic merging prediction sequence are used to indicate whether the cells in the corresponding first table segment and the second table segment can be semantically merged.

[0031] As one possible implementation, the training dataset is constructed based on the original electronic document containing cross-page tables, including:

[0032] The cross-page tables in the original electronic document are cropped, trimmed, and their edges aligned to obtain the test dataset and the positive training dataset.

[0033] The tables in the original electronic document are split and concatenated to obtain a negative training dataset.

[0034] As one possible implementation, training the initial merging model based on the training dataset to obtain a table merging prediction model includes:

[0035] The sample data in the training dataset is input into the initial merging model to obtain the processing result of the initial merging model. The initial merging model includes: a visual feature processing network, a semantic feature processing network, and a linear processing layer.

[0036] The loss of the initial merging model is obtained by calculating the loss of the initial merging model based on the processing result of the target loss function. The target loss function includes at least an error propagation loss function, which is used to calculate the deviation between the processing result of the visual feature processing network and the processing result of the semantic feature processing network.

[0037] Based on the loss of the initial merging model, the initial merging model is iteratively corrected to obtain the table merging prediction model.

[0038] As one possible implementation, the target loss function may also include: a physical merging loss function and a semantic merging loss function;

[0039] The physical merging loss function is used to calculate the loss of the physical merging prediction result, and the semantic merging loss function is used to calculate the loss of the semantic merging prediction result.

[0040] Thirdly, embodiments of this application provide a table merging apparatus, including:

[0041] The acquisition module is used to acquire the image to be processed, which includes a first table segment and a second table segment distributed across the page.

[0042] The processing module is used to input the image to be processed into a pre-trained table merging prediction model for prediction processing, and obtain a physical merging prediction result and a semantic merging prediction sequence. The physical merging prediction result is used to indicate whether the first table segment and the second table segment can be merged in physical structure. Each item in the semantic merging prediction sequence corresponds to a cell of the first table segment and the second table segment, and the value of each item in the semantic merging prediction sequence is used to indicate whether the cells on the corresponding first table segment and the second table segment can be semantically merged.

[0043] The merging module is used to merge the first table segment and the second table segment according to the physical merging prediction result and the semantic merging prediction sequence.

[0044] As one possible implementation, the table merging prediction model includes: a visual feature processing network, a semantic feature processing network, and a linear processing layer;

[0045] The processing module is specifically used for:

[0046] The image to be processed is input into the visual feature processing network, which encodes the visual features to obtain visual features, and then performs positional encoding on the visual features to obtain position-encoded features.

[0047] The location-encoded features are input into the semantic feature processing network, which performs semantic feature encoding to obtain semantically encoded features.

[0048] The semantically encoded features are input into the linear processing layer for linear transformation to obtain the physical merging prediction result and the semantic merging prediction sequence.

[0049] As one possible implementation, the visual feature processing network includes multiple coding layers connected in sequence;

[0050] The processing module is specifically used for:

[0051] The image to be processed is input into the first coding layer for encoding, and the processed features are input into the next coding layer for encoding, and so on, until the last coding layer completes the encoding process.

[0052] Max pooling is performed on the processed features of each coding layer except the last coding layer to obtain the pooled features of each coding layer.

[0053] The pooled features from each coding layer and the features encoded by the last coding layer are concatenated to obtain the visual features.

[0054] As one possible implementation, the processing module is specifically used for:

[0055] Based on the two-dimensional sine and cosine coding algorithm, the visual features are encoded in both the horizontal and vertical directions to obtain the encoded features.

[0056] As one possible implementation, the merging module is specifically used for:

[0057] If the physical merge prediction result indicates that the first table segment and the second table segment can be merged in physical structure, then each item in the semantic merge prediction sequence is traversed. For the current item that is traversed, if the value of the current item is a preset value, then the content of the cell corresponding to the current item in the first table segment and the cell corresponding to the current item in the second table segment are merged.

[0058] As one possible implementation, the processing module is further configured to:

[0059] Based on the page segmentation marks in the electronic document, the first initial table segment and the second initial table segment are extracted from adjacent pages of the electronic document to obtain the initial image;

[0060] The initial image is cropped and the table edges are aligned to obtain the image to be processed.

[0061] Fourthly, embodiments of this application provide a training apparatus for a table merging prediction model, comprising:

[0062] The building block is used to construct the training dataset based on the original electronic document containing cross-page tables;

[0063] The training module is used to train the initial merging model based on the training dataset to obtain a table merging prediction model. The prediction results of the table merging prediction model include physical merging prediction results and semantic merging prediction sequences. The physical merging prediction results are used to indicate whether the first table segment and the second table segment in the input image to be processed can be merged in terms of physical structure. Each item in the semantic merging prediction sequence corresponds to a cell of the first table segment and the second table segment, and the values ​​of each item in the semantic merging prediction sequence are used to indicate whether the cells on the corresponding first table segment and the second table segment can be semantically merged.

[0064] As one possible implementation, the building module is specifically used for:

[0065] The cross-page tables in the original electronic document are cropped, trimmed, and their edges aligned to obtain the test dataset and the positive training dataset.

[0066] The tables in the original electronic document are split and concatenated to obtain a negative training dataset.

[0067] As one possible implementation, the training module is specifically used for:

[0068] The sample data in the training dataset is input into the initial merging model to obtain the processing result of the initial merging model. The initial merging model includes: a visual feature processing network, a semantic feature processing network, and a linear processing layer.

[0069] The loss of the initial merging model is obtained by calculating the loss of the initial merging model based on the processing result of the target loss function. The target loss function includes at least an error propagation loss function, which is used to calculate the deviation between the processing result of the visual feature processing network and the processing result of the semantic feature processing network.

[0070] Based on the loss of the initial merging model, the initial merging model is iteratively corrected to obtain the table merging prediction model.

[0071] As one possible implementation, the target loss function may also include: a physical merging loss function and a semantic merging loss function;

[0072] The physical merging loss function is used to calculate the loss of the physical merging prediction result, and the semantic merging loss function is used to calculate the loss of the semantic merging prediction result.

[0073] Fifthly, embodiments of this application provide an electronic device, including: a processor and a memory, wherein the memory stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor executes the machine-readable instructions to perform the steps of the table merging method as described in the first aspect above or the steps of the table merging prediction model training method as described in the second aspect above.

[0074] Sixthly, this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, performs the steps of the table merging method described in the first aspect above or the steps of the table merging prediction model training method described in the second aspect above.

[0075] The table merging method, table merging prediction model training method, and apparatus provided in this application input the image to be processed into a pre-trained table merging prediction model. The table merging prediction model can predict both physical merging prediction results and semantic merging prediction sequences. The physical merging prediction results indicate whether the first table segment and the second table segment can be merged physically, while the semantic merging prediction sequence indicates whether each corresponding cell in the first and second table segments can be semantically merged. That is, the semantic merging prediction sequence can predict at the cell level whether a cell can be merged from a contextual semantic perspective. Because it combines physical structure and cell-level contextual semantic information to determine whether tables can be merged, the accuracy of the table merging results can be significantly improved. In addition, each item in the semantic merging prediction sequence corresponds to a cell in the first and second table segments, so the semantic merging prediction sequence is a variable-length sequence and does not need to be limited to a fixed length, thus it can be flexibly applied to table merging scenarios with different numbers of columns. Attached Figure Description

[0076] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0077] Figure 1 This is an example diagram illustrating a scenario of merging tables across multiple pages.

[0078] Figure 2 A flowchart illustrating the table merging method provided in this application embodiment;

[0079] Figure 3 Here is an example of merging a table;

[0080] Figure 4 Example of merging another table;

[0081] Figure 5 A schematic diagram of an architecture for a table merging prediction model;

[0082] Figure 6 Another flowchart illustrating the table merging method provided in this application embodiment;

[0083] Figure 7 Example diagram of the structure of a visual feature processing network;

[0084] Figure 8 Another flowchart illustrating the table merging method provided in this application embodiment;

[0085] Figure 9 A flowchart illustrating the table merging prediction model training method provided in this application embodiment;

[0086] Figure 10 Another flowchart illustrating the table merging prediction model training method provided in this application embodiment;

[0087] Figure 11 A schematic diagram illustrating the training of the initial merging model using the error propagation loss function;

[0088] Figure 12 A module structure diagram of a table merging device provided in an embodiment of this application;

[0089] Figure 13 A module structure diagram of a table merging prediction model training device provided in an embodiment of this application;

[0090] Figure 14 This is a schematic diagram of the structure of the electronic device 140 provided in an embodiment of this application. Detailed Implementation

[0091] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. It should be understood that the accompanying drawings in this application are for illustrative and descriptive purposes only and are not intended to limit the scope of protection of this application. Furthermore, it should be understood that the schematic drawings are not drawn to scale. The flowcharts used in this application illustrate operations implemented according to some embodiments of this application. It should be understood that the operations in the flowcharts may not be implemented in sequence, and steps without logical contextual relationships may be reversed or implemented simultaneously. In addition, those skilled in the art, guided by the content of this application, may add one or more other operations to the flowcharts, or remove one or more operations from the flowcharts.

[0092] Furthermore, the described embodiments are merely some, not all, of the embodiments of this application. The components of the embodiments of this application described and illustrated herein can typically be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.

[0093] It should be noted that the term "comprising" will be used in the embodiments of this application to indicate the presence of the features declared thereafter, but does not exclude the addition of other features.

[0094] Existing methods for merging cross-page tables primarily determine whether tables can be merged by analyzing their physical structure. Examples of several existing technologies are listed below.

[0095] In the first example, a simulation-generated method is used. Two tables, table1 and table2, from different pages are selected. Table1 and table2 are divided into upper and lower parts based on their middle rows. The cell coordinates within each of the four parts are then combined pairwise to form the input format: [SEP] + table1_cell1 + table1_cell2 + ... + table1_cellm + [SEP] + table2_cell1 + table2_cell2 + ... + table2_celln + [SEP]. A binary classification model is trained using a depth-based bidirectional transformer. If the output confidence score is greater than 0.5, it indicates that table1 and table2 can be merged; otherwise, they do not need to be merged. This method only considers whether the tables can be merged physically, without considering the logical relationship between the upper and lower cells. Furthermore, the bidirectional depth transformer in this method suffers from error accumulation, which can easily lead to prediction bias.

[0096] The second example employs a two-stage method for determining cross-page table merging. First, after removing headers and footers, the table areas for the preceding and following pages are obtained. A header detection model is then used to check if headers exist on the preceding and following pages. Simultaneously, based on the rules governing the presence of headers on the preceding and following pages, it is determined whether the table bodies can be merged. For cases where merging is possible, similarity is calculated for the cell groups at the merging point. The text of the merged cell groups is then combined into a new text, and the similarity between this new text and the semantic vectors of other cells in the same column is calculated. If this similarity is greater than the similarity calculated in the same way before merging, the tables are merged. This process, combined with other prior knowledge, completes the final table merging and splicing. The merging result of this method relies on a pre-trained header recognition model; errors in header detection will lead to merging errors. Furthermore, this method does not consider contextual semantic information, making it difficult to cover a wide range of table scenarios.

[0097] In summary, existing technologies focus on whether tables can be merged in terms of physical structure, without considering the contextual semantic information of cells above and below. Therefore, the results of merging tables using existing technologies have low accuracy.

[0098] Based on the above-mentioned problems, this application proposes a table merging method and a table merging prediction model training method. A table merging prediction model is trained using a training dataset. This table merging prediction model can not only predict whether a table can be merged from the perspective of physical structure, but also predict whether a cell can be merged from the perspective of contextual semantics at the cell level. Since the physical structure and cell-level contextual semantic information are combined to determine whether a table can be merged, the accuracy of the table merging result can be significantly improved.

[0099] The embodiments of this application can be applied to scenarios involving merging tables across multiple pages. Figure 1 This is an example diagram illustrating a scenario of merging tables across multiple pages, such as... Figure 1 As shown, in a certain electronic document, due to page length limitations, the same table is displayed on adjacent pages, resulting in cross-page display. For the purpose of content analysis of this electronic document, it is necessary to merge the cross-page tables to obtain a complete table. For example, in... Figure 1 In the table below, the content of the last cell in the first row is essentially the same as the content of the last cell in the last row of the table above, and should be merged into one cell. However, the method described in this application can accurately achieve this merging of tables across multiple pages.

[0100] Before introducing the technical solutions of the embodiments of this application, it should be noted that the table merging method and the table merging prediction model training method provided in the embodiments of this application must obtain explicit user consent before collecting and generating personal information and important data during the execution process, in order to be applied to the solutions of the embodiments of this application. Furthermore, the personal information and important data collected and generated during the execution process of the table merging method and the table merging prediction model training method provided in the embodiments of this application must adhere to the principles of legality, legitimacy, and necessity, and the rules for collection and use must be publicly disclosed, clearly stating the purpose, method, and scope of information collection and use. The data in the above examples does not include personal information unrelated to the services provided in the above examples.

[0101] This application provides a table merging method and a table merging prediction model training method. The table merging prediction model training method describes the training process of the table merging prediction model used in the table merging method. The table merging method will be explained below.

[0102] Figure 2 This is a flowchart illustrating the table merging method provided in an embodiment of this application. The executing entity of this method can be any electronic device with computing capabilities. Figure 2 As shown, the method may include:

[0103] S201. Obtain the image to be processed, which includes a first table segment and a second table segment distributed across the page.

[0104] For example, when it is necessary to perform content analysis on an electronic document, the adjacent pages in the electronic document containing cross-page tables are identified based on specific markers such as page separators and table lines. All or part of the table segments in the adjacent pages are combined into an image to obtain the image to be processed.

[0105] It should be understood that for the same electronic document, there may be one or more sets of adjacent pages with cross-page tables, resulting in one or more images to be processed. When there are multiple images to be processed, the method steps of the embodiments of this application can be used to merge the tables for each image.

[0106] In addition, the image to be processed includes a first table segment and a second table segment distributed across the page. The first table segment may be part or all of the table segment in the upper page of an adjacent page, and the second table segment may be part or all of the table segment in the lower page of an adjacent page.

[0107] S202. Input the image to be processed into the pre-trained table merging prediction model for prediction processing to obtain a physical merging prediction result and a semantic merging prediction sequence. The physical merging prediction result indicates whether the first table segment and the second table segment can be merged physically. Each item in the semantic merging prediction sequence corresponds to a cell in the first table segment and the second table segment, and the value of each item in the semantic merging prediction sequence indicates whether the corresponding cells in the first and second table segments can be semantically merged.

[0108] Optionally, the above physical merge prediction result can be a 1-bit value. For example, using the tab_merge label to represent the physical merge prediction result, if the value of tab_merge is 1, it means that the first table segment and the second table segment can be merged in terms of physical structure; if the value of tab_merge is 0, it means that the first table segment and the second table segment cannot be merged in terms of physical structure.

[0109] Optionally, the semantic merge prediction sequence can be a numerical sequence, the length of which, i.e., the number of items, is the same as the minimum granularity column number of the first table segment and the minimum granularity column number of the second table segment. Each item corresponds to a cell in the first table segment and a cell in the second table segment, respectively. It should be understood that the cell corresponding to each item in the first table segment is the cell in the last row of the first table segment, and the cell corresponding to each item in the second table segment is the cell in the first row of the second table segment. Accordingly, the value of each item in the numerical sequence indicates whether the corresponding cell in the first table segment and the corresponding cell in the second table segment can be semantically merged. For example, the label `col_merge` is used to represent the semantic merge prediction sequence, the length of which is the same as the number of columns in the first and second table segments. For example, assuming the minimum granularity of columns in the first and second table segments is 8, then the length of col_merge is 8. When the value of a certain item in col_merge is 1, it means that the cell in the first table segment corresponding to that item can be semantically merged with the cell in the corresponding second table segment. When the value of a certain item in col_merge is 0, it means that the cell in the first table segment corresponding to that item cannot be semantically merged with the cell in the corresponding second table segment.

[0110] Table 1 below explains the numerical meanings of tab_merge and col_merge. Table 2 below explains the merging rules for tab_merge and col_merge.

[0111] Table 1

[0112]

[0113] Table 2

[0114]

[0115] The following two examples illustrate this.

[0116] Figure 3 Here is an example of merging a table, such as Figure 3 As shown, the minimum granularity of columns in the first and second table segments is 8. Based on the table merging prediction model described above, it is predicted that the first and second table segments are physically mergingable, and all 8 cells can be semantically merged. Therefore, the value of `tab_merge` is 1, and the value of `col_merge` is 11111111. It should be understood that the value 1 in each item of 11111111 indicates that the corresponding cell is semantically mergingable. For example, the first value 1 in 11111111 indicates that the first cell in the last row of the first table segment and the first cell in the first row of the second table segment are semantically mergingable.

[0117] Figure 4 For another example of table merging, such as Figure 4 As shown, the minimum granularity of columns in the first and second table segments is 4. Based on the table merging prediction model described above, it is predicted that the first and second table segments are physically mergingable, and all four cells are semantically unmergingable. Therefore, the value of `tab_merge` is 1, and the value of `col_merge` is 0000. It should be understood that the value 0 in each item of 0000 indicates that the corresponding cell is semantically unmergingable. For example, the first value 0 in 0000 indicates that the first cell in the last row of the first table segment and the first cell in the first row of the second table segment are semantically unmergingable.

[0118] S203. Based on the above physical merging prediction results and the above semantic merging prediction sequence, merge the above first table segment and the above second table segment.

[0119] Since the above physical merge prediction results indicate whether the first table segment and the second table segment can be merged in terms of physical structure, and the above semantic merge prediction sequence indicates whether the corresponding cells of the first table segment and the second table segment can be semantically merged, therefore, by combining the above physical merge prediction results and the above semantic merge prediction sequence, the first table segment and the second table segment can be accurately merged.

[0120] In this embodiment, the image to be processed is input into a pre-trained table merging prediction model. The table merging prediction model can predict both physical merging prediction results and semantic merging prediction sequences. The physical merging prediction results indicate whether the first table segment and the second table segment can be merged in terms of physical structure, while the semantic merging prediction sequences indicate whether the corresponding cells of the first and second table segments can be semantically merged. That is, the semantic merging prediction sequences can predict at the cell level whether cells can be merged from a contextual semantic perspective. Because both physical structure and cell-level contextual semantic information are combined to determine whether tables can be merged, the accuracy of the table merging results can be significantly improved. In addition, each item in the semantic merging prediction sequence corresponds to a cell in the first and second table segments, so the semantic merging prediction sequence is a variable-length sequence and does not need to be limited to a fixed length, thus it can be flexibly applied to table merging scenarios with different numbers of columns.

[0121] The following describes the forecasting process of the table merging forecasting model.

[0122] Figure 5 This is a schematic diagram of an architecture for a table merging prediction model, such as... Figure 5As shown, the table merging prediction model includes a visual feature processing network, a semantic feature processing network, and a linear processing layer. These three networks are connected in series. As an example, the visual feature processing network can be a residual visual feature encoder, such as ResNet50. The semantic feature processing network can be a semantic feature encoder, such as a Transformer semantic feature encoder. The linear processing layer can be a linear layer, which may include one or more fully connected layers.

[0123] Based on the above Figure 5 Given the model architecture shown, step S202 can be performed as follows.

[0124] Figure 6 Another flowchart illustrating the table merging method provided in this application embodiment is shown below. Figure 6 As shown, step S202 above may include:

[0125] S601. Input the above-mentioned image to be processed into the above-mentioned visual feature processing network, and the above-mentioned visual feature processing network performs visual feature encoding to obtain visual features, and performs position encoding on the visual features to obtain position-encoded features.

[0126] Optionally, the aforementioned visual feature processing network encodes visual features into the image to be processed to obtain the aforementioned visual features. During pre-training, the visual feature processing network can learn the ability to align table column lines based on the training dataset. Therefore, the visual features obtained after processing by the visual feature processing network can characterize whether the column lines of the first table segment and the second table segment are aligned. The higher the degree of column line alignment, the higher the probability that the first table segment and the second table segment can be merged in physical structure. As described above... Figure 3 In the example shown, all column lines of the first and second table segments are aligned, thus confirming that the first and second table segments can be physically merged.

[0127] Optionally, after obtaining the above visual features, the visual features are positionally encoded to better represent the contextual positional relationship between the first table segment and the second table segment, thereby achieving accurate prediction.

[0128] As one possible implementation, the above visual features can be encoded horizontally and vertically based on a two-dimensional sine and cosine coding algorithm to obtain the encoded features.

[0129] Specifically, by using the two-dimensional sine and cosine coding algorithm described above, the visual features can be encoded in both the horizontal and vertical directions. This allows the encoded features to further characterize the contextual positional relationship between the first and second table segments, and thus the positional relationship can be used as auxiliary information for accurate semantic analysis.

[0130] S602. Input the above-mentioned position-encoded features into the above-mentioned semantic feature processing network, and the above-mentioned semantic feature processing network performs semantic feature encoding to obtain semantically encoded features.

[0131] Optionally, during pre-training, the aforementioned semantic feature processing network can learn, based on the training dataset, whether cells in the same column are semantically merging. Therefore, the semantically encoded features obtained after processing by the semantic feature processing network can characterize whether corresponding cells in the first and second table segments are semantically merging. (As mentioned above...) Figure 3 In the example shown, the 8 cells in the last row of the first table segment and the corresponding cells in the first row of the second table segment are semantically mergeable. Therefore, it can be confirmed that the corresponding cells in the first table segment and the second table segment are semantically mergeable.

[0132] It should be understood that the semantically encoded features described above can characterize whether the corresponding cells in the first table segment and the second table segment are semantically mergeable, while also characterizing whether the column lines of the first table segment and the second table segment are aligned, that is, whether the first table segment and the second table segment can be merged in terms of physical structure.

[0133] S603. Input the semantically encoded features into the linear processing layer for linear transformation to obtain the physical merging prediction result and the semantic merging prediction sequence.

[0134] Optionally, since the semantically encoded features obtained after processing by the aforementioned semantic feature processing network cannot directly represent the prediction result, in this step, the semantically encoded features can be linearly transformed using the aforementioned linear processing layer to obtain the prediction result. This prediction result specifically includes the aforementioned physical merging prediction result and the semantic merging prediction sequence.

[0135] In this embodiment, the table merging prediction model sequentially includes a visual feature processing network, a semantic feature processing network, and a linear processing layer. The visual feature processing network can learn whether the tables can be merged in terms of physical structure. The semantic feature processing network can learn whether the corresponding cells in the first table segment and the second table segment are semantically mergeable. At the same time, the positional encoding can learn the contextual positional relationship between the first table segment and the second table segment. Through linear transformation, the learned features can be transformed into intuitive prediction results. Combining these processing processes, the table merging prediction model can ensure that it outputs accurate and intuitive prediction results.

[0136] As an optional implementation, the visual feature processing network described above may include multiple coding layers connected in sequence. Figure 7 Example diagram of the structure of a visual feature processing network, such as Figure 7 As shown, the visual feature processing network can specifically be a residual visual feature encoder, and the residual visual feature encoder includes four encoding layers connected in sequence. It should be understood that... Figure 7 The four encoding layers shown are merely an example; the number of encoding layers is not limited to these and can be flexibly set as needed. Accordingly, the visual feature processing network can encode visual features according to the following process.

[0137] Optionally, one possible method for the visual feature encoding by the visual feature processing network in step S601 above includes:

[0138] The image to be processed is input into the first coding layer for encoding, and the processed features are input into the next coding layer for encoding. This process is repeated until the last coding layer completes its encoding. Max pooling is then performed on the processed features of each coding layer except the last one to obtain the pooled features of each coding layer. The pooled features of each coding layer and the encoded features of the last coding layer are then concatenated to obtain the aforementioned visual features.

[0139] Reference Figure 7 The image to be processed is input into the first encoding layer for encoding. The first encoding layer then inputs the encoded features into the second encoding layer for further encoding, and so on, until the last encoding layer completes the encoding process. In addition, the features encoded by each of the previous encoding layers (except the last one) can be subjected to max pooling. The size of the features after max pooling is consistent with the size of the features encoded by the last encoding layer. Finally, the max-pooled features from all the previous encoding layers are concatenated with the features encoded by the last encoding layer to obtain the aforementioned visual features.

[0140] In this embodiment, by performing encoding processing through multiple encoding layers in sequence, the visual features output by the visual feature processing network can more accurately represent whether physical structures can be merged. In addition, by performing max pooling on each encoding layer, the feature sizes can be aligned, ensuring the accuracy of feature processing.

[0141] As an optional implementation, step S203 above may include:

[0142] If the above physical merge prediction result indicates that the first table segment and the second table segment can be merged in terms of physical structure, then each item in the above semantic merge prediction sequence is traversed. For the current item that is traversed, if the value of the current item is a preset value, then the content of the cell corresponding to the current item in the first table segment and the cell corresponding to the current item in the second table segment are merged.

[0143] For example, the preset value mentioned above can be the value 1 shown in step S202 above. When the value of the current item in the semantic merge prediction sequence is 1, it indicates that the semantics of the cell corresponding to the current item can be merged.

[0144] Optionally, if the first table segment and the aforementioned second table segment are physically mergeable, then the process described above can be further followed to determine whether each cell is semantically mergeable. When they are physically mergeable and semantically mergeable, the corresponding cells can be merged. For example, in the aforementioned... Figure 3 In the example above, the cells corresponding to the 8 column lines are all semantically mergeable, so the contents of the cells above and below each cell can be merged separately. When cells are physically mergeable but semantically incompatible, the corresponding cells are not merged. For example, in the aforementioned... Figure 4 In the example, the cells corresponding to the four column lines are semantically non-mergeable, so the content of the cells above and below each cell is not merged.

[0145] In another case, if the first table segment and the aforementioned second table segment cannot be merged in terms of physical structure, then the first table segment and the aforementioned second table segment will not be merged, and the semantic merging prediction sequence will no longer be judged.

[0146] In this embodiment, given that the physical structures of the first table segment and the aforementioned second table segment can be merged, cell-level merging can be accurately performed by determining whether each cell is semantically mergeable.

[0147] As an optional implementation, the above-mentioned image to be processed can be generated through the following process.

[0148] Figure 8 Another flowchart illustrating the table merging method provided in this application embodiment, as shown below. Figure 8 As shown, the procedure before step S201 includes:

[0149] S801. Based on the page segmentation marks in the electronic document, extract the first initial table segment and the second initial table segment from the adjacent pages of the electronic document to obtain the initial image.

[0150] Optionally, the pages of the electronic document can be converted into images and image recognition can be performed to identify whether preset page segmentation marks exist. When page segmentation marks exist, it indicates that they belong to adjacent pages. Therefore, it can be further determined whether there are table column lines or table row lines above and below the page segmentation marks. If so, it can be determined that there is a table spanning multiple pages. Therefore, part or all of the table in the adjacent pages can be extracted to obtain the aforementioned initial image. For example, the bottom half of the upper page and the top half of the lower page in the adjacent pages can be extracted to obtain the aforementioned initial image. It should be understood that the initial image includes the page segmentation marks between adjacent pages.

[0151] S802. The initial image is cropped and the table edges are aligned to obtain the image to be processed.

[0152] Optionally, the initial image can be cropped to retain only the table portion, without preserving the blank areas outside the table, based on the table edge lines in the initial image. Then, the lower page can be shifted to the left, using the leftmost column line of the upper page as a reference, so that the leftmost column line of the upper page aligns with the leftmost column line of the lower page.

[0153] In this embodiment, by performing table cropping, trimming, and alignment on the electronic document, the tables on the upper and lower pages of the resulting image to be processed can be aligned on one side, thereby eliminating the negative impact of the table merging prediction model and ensuring the accuracy of the prediction results.

[0154] The following describes the training method for the table merging prediction model according to the embodiments of this application.

[0155] Figure 9 This is a flowchart illustrating the table merging prediction model training method provided in an embodiment of this application. The execution entity of this method can be any electronic device with computing power. Figure 9 As shown, the method includes:

[0156] S901. Construct a training dataset based on the original electronic document containing cross-page tables.

[0157] Optionally, the number of original electronic documents can be one or more. Each original electronic document may include one or more sets of cross-page tables. By processing these cross-page tables, a training dataset of a certain size can be obtained, which may include a test set and a training set. The specific process of constructing the training dataset will be described in detail in the following embodiments.

[0158] S902. Based on the above training dataset, the initial merging model is trained to obtain the table merging prediction model.

[0159] The prediction results of the table merging prediction model include physical merging prediction results and semantic merging prediction sequences. The physical merging prediction results are used to indicate whether the first table segment and the second table segment in the input image to be processed can be merged in terms of physical structure. Each item in the semantic merging prediction sequence corresponds to a cell in the first table segment and the second table segment, and the values ​​of each item in the semantic merging prediction sequence are used to indicate whether the cells in the corresponding first table segment and the second table segment can be semantically merged.

[0160] Optionally, the initial merging model described above can be built in advance. Since the table merging prediction model is trained from the initial merging model, the model structure of the initial merging model is the same as that of the table merging prediction model. See the preceding text for details. Figure 5 as well as Figure 7 The structure shown will not be elaborated here. The initial merging model can undergo multiple rounds of iterative training and parameter adjustments. When the result of a certain iteration training meets the preset convergence condition or the number of iterations, training can be stopped, and the initial merging model trained in the last iteration can be used as the table merging prediction model.

[0161] In this embodiment, by constructing a training dataset, a table merging prediction model can be trained using the training dataset. Through training the initial merging model, the resulting table merging prediction model can predict both physical merging results and semantic merging prediction sequences. The physical merging prediction result indicates whether the first and second table segments can be merged physically, while the semantic merging prediction sequence indicates whether corresponding cells in the first and second table segments can be semantically merged. In other words, the semantic merging prediction sequence can predict at the cell level whether cells can be merged from a contextual semantic perspective. Because both physical structure and cell-level contextual semantic information are combined to determine whether tables can be merged, the accuracy of the table merging results can be significantly improved.

[0162] As an optional implementation, step S901 above may include:

[0163] The cross-page tables in the original electronic document are cropped, trimmed, and have their edges aligned to obtain the test dataset and the positive training dataset; the tables in the original electronic document are split and spliced ​​to obtain the negative training dataset.

[0164] For example, the aforementioned page segmentation markers are used to identify the table regions on the upper and lower pages, with table coordinates of (t1_x1, t1_y1, t2_x2, t2_y2) and (t2_x1, t2_y1, t2_x2, t2_y2). The images t1 and t2 of the upper and lower table regions are cropped to sizes (H0, W0, 3) and (H1, W1, 3), respectively, where H0 = t1_y2 - t1_y1, W0 = t1_x2 - t1_x1, H1 = t2_y2 - t2_y1, and W1 = t2_x2 - t2_x1. The upper and lower table region images are then left-aligned and stitched together to form an image I of size H × W × 3, where H = H0 + H1 and W = Max(W0, W1). Image I obtained in this way can be used as an image in the test dataset. Additionally, for cross-page tables in the original electronic document, firstly, using the page segmentation marker as the center, the bottom half of the upper page and the top half of the lower page are stitched together to form a new image Ipg. Then, the table regions within Ipg are detected, and corresponding cropping and stitching are performed. Images obtained in this way can be used as images in the positive training dataset. Furthermore, for tables in the original electronic document, line detection is performed, and lines at random middle positions are used to divide the same table into two different parts, tupi and tdowni, resulting in multiple image groups G. By stitching the parts from Gi and Gj, a stitched image set is obtained. Simultaneously, the images in the image set are cropped to ensure the stitching position is located at the center horizontal line of the image. Images obtained in this way can be used as images in the negative training dataset.

[0165] Based on the generated training dataset, the training dataset can also be labeled. For example, the physical merge label tab_merge value for stitched images from different groups and upper and lower pages is 0, meaning they cannot be merged. At the same time, the value of the label sequence col_merge can be marked according to the smallest unit column obtained by dividing the image based on the total number of column lines on the upper and lower pages.

[0166] In this embodiment, by processing the tables in the original electronic document, the training dataset becomes richer, thereby making the prediction results of the trained table merging prediction model more accurate.

[0167] Figure 10 Another flowchart illustrating the table merging prediction model training method provided in this application embodiment is shown below. Figure 10 As shown, step S902 above may include:

[0168] S1001. Input the sample data from the training dataset into the initial merging model to obtain the processing result of the initial merging model. The initial merging model includes: a visual feature processing network, a semantic feature processing network, and a linear processing layer.

[0169] S1002. Calculate the loss of the initial merging model based on the target loss function. The target loss function includes at least one error propagation loss function, which is used to calculate the deviation between the processing results of the visual feature processing network and the processing results of the semantic feature processing network.

[0170] Figure 11 This is a schematic diagram illustrating the training of the initial merging model using the error propagation loss function, as shown below. Figure 11 As shown, the visual feature processing network is specifically a residual visual feature encoder, and the semantic feature processing network is specifically a Transformer semantic feature encoder. The visual feature output by the residual visual feature encoder is F, and the position-encoded feature obtained after position encoding F is V. in The Transformer semantic feature encoder for V in The semantically encoded features output after semantic feature encoding are V. out Because the Transformer semantic feature encoder has propagation errors, to ensure that the semantic features learned by the initial merging model are as close as possible to the visual features, the aforementioned error propagation loss function can be set. This error propagation loss function is used to calculate the deviation between the processing results of the visual feature processing network and the processing results of the semantic feature processing network, that is, to calculate F and V. out The deviation. For example Figure 11 As shown, the loss calculated by the error propagation loss function is L0. Using L0, the propagation error of the Transformer semantic feature encoder can be mitigated or offset.

[0171] As an optional implementation, the target loss function further includes a physical merging loss function and a semantic merging loss function. The physical merging loss function is used to calculate the loss of the physical merging prediction result, and the semantic merging loss function is used to calculate the loss of the semantic merging prediction result.

[0172] Continue to refer to Figure 11 The loss calculated using the physical merging loss function is L. tab_merge The loss calculated by the semantic merging loss function is L. col_mergeThe physical merging loss function can be calculated using cross-entropy, while the semantic merging loss function can be calculated using sequence prediction methods such as ctc_loss, resulting in a sequence consisting of 0s and 1s. The loss L of the above target loss function can be expressed as: L = L tab_merge +L col_merge +L0.

[0173] S1003. Based on the loss of the initial merging model, the initial merging model is iteratively corrected to obtain the table merging prediction model.

[0174] The initial merging model can undergo multiple rounds of iterative training and parameter correction. When the result of a certain iteration training meets the preset convergence condition or the number of iterations, training can be stopped, and the initial merging model trained in the last iteration can be used as the table merging prediction model.

[0175] In this embodiment, an error propagation loss function is introduced into the loss function of the initial merging model. This error propagation loss function can calculate the deviation between the processing results of the visual feature processing network and the processing results of the semantic feature processing network. Therefore, by using the loss of this error propagation loss function to correct the initial merging model, the propagation error can be alleviated or offset, thereby making the prediction results of the trained table merging prediction model more accurate.

[0176] Based on the same inventive concept, this application also provides a table merging device corresponding to the table merging method. Since the principle of the device in this application is similar to the table merging method described above, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be described again.

[0177] Figure 12 A module structure diagram of a table merging device provided in an embodiment of this application is shown below. Figure 12 As shown, the device includes:

[0178] The acquisition module 1201 is used to acquire an image to be processed, which includes a first table segment and a second table segment distributed across the page.

[0179] The processing module 1202 is used to input the image to be processed into a pre-trained table merging prediction model for prediction processing to obtain a physical merging prediction result and a semantic merging prediction sequence. The physical merging prediction result is used to indicate whether the first table segment and the second table segment can be merged in physical structure. Each item in the semantic merging prediction sequence corresponds to a cell of the first table segment and the second table segment, and the value of each item in the semantic merging prediction sequence is used to indicate whether the cells on the corresponding first table segment and the second table segment can be semantically merged.

[0180] The merging module 1203 is used to merge the first table segment and the second table segment according to the physical merging prediction result and the semantic merging prediction sequence.

[0181] As an optional implementation, the table merging prediction model includes: a visual feature processing network, a semantic feature processing network, and a linear processing layer.

[0182] The processing module 1202 is specifically used for:

[0183] The image to be processed is input into the visual feature processing network, which encodes the visual features to obtain visual features, and then performs positional encoding on the visual features to obtain position-encoded features.

[0184] The location-encoded features are input into the semantic feature processing network, which performs semantic feature encoding to obtain semantically encoded features.

[0185] The semantically encoded features are input into the linear processing layer for linear transformation to obtain the physical merging prediction result and the semantic merging prediction sequence.

[0186] As an optional implementation, the visual feature processing network includes multiple coding layers connected in sequence.

[0187] The processing module 1202 is specifically used for:

[0188] The image to be processed is input into the first coding layer for encoding, and the processed features are input into the next coding layer for encoding, and so on, until the last coding layer completes the encoding process.

[0189] Max pooling is performed on the processed features of each coding layer except the last coding layer to obtain the pooled features of each coding layer.

[0190] The pooled features from each coding layer and the features encoded by the last coding layer are concatenated to obtain the visual features.

[0191] As an optional implementation, the processing module 1202 is specifically used for:

[0192] Based on the two-dimensional sine and cosine coding algorithm, the visual features are encoded in both the horizontal and vertical directions to obtain the encoded features.

[0193] As an optional implementation, the merging module 1203 is specifically used for:

[0194] If the physical merge prediction result indicates that the first table segment and the second table segment can be merged in physical structure, then each item in the semantic merge prediction sequence is traversed. For the current item that is traversed, if the value of the current item is a preset value, then the content of the cell corresponding to the current item in the first table segment and the cell corresponding to the current item in the second table segment are merged.

[0195] As an optional implementation, the processing module 1202 is further configured to:

[0196] Based on the page segmentation marks in the electronic document, a first initial table segment and a second initial table segment are extracted from adjacent pages of the electronic document to obtain an initial image.

[0197] The initial image is cropped and the table edges are aligned to obtain the image to be processed.

[0198] Based on the same inventive concept, this application also provides a table merging prediction model training device corresponding to the table merging prediction model training method. Since the principle of the device in this application is similar to the table merging prediction model training method described above in this application, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be described again.

[0199] Figure 13 A module structure diagram of a table merging prediction model training device provided in this application embodiment is shown below. Figure 13 As shown, the device includes:

[0200] Module 1301 is used to build a training dataset based on the original electronic document containing cross-page tables.

[0201] Training module 1302 is used to train the initial merging model based on the training dataset to obtain a table merging prediction model. The prediction result of the table merging prediction model includes a physical merging prediction result and a semantic merging prediction sequence. The physical merging prediction result is used to indicate whether the first table segment and the second table segment in the input image to be processed can be merged in terms of physical structure. Each item in the semantic merging prediction sequence corresponds to a cell of the first table segment and the second table segment, and the value of each item in the semantic merging prediction sequence is used to indicate whether the cells on the corresponding first table segment and the second table segment can be semantically merged.

[0202] As one possible implementation, the construction module 1301 is specifically used for:

[0203] The cross-page tables in the original electronic document are cropped, trimmed, and their edges aligned to obtain the test dataset and the positive training dataset.

[0204] The tables in the original electronic document are split and concatenated to obtain a negative training dataset.

[0205] As one possible implementation, the training module is specifically used for:

[0206] The sample data in the training dataset is input into the initial merging model to obtain the processing result of the initial merging model. The initial merging model includes a visual feature processing network, a semantic feature processing network, and a linear processing layer.

[0207] The loss of the initial merging model is obtained by calculating the loss of the processing result of the initial merging model based on the target loss function. The target loss function includes at least an error propagation loss function, which is used to calculate the deviation between the processing result of the visual feature processing network and the processing result of the semantic feature processing network.

[0208] Based on the loss of the initial merging model, the initial merging model is iteratively corrected to obtain the table merging prediction model.

[0209] As one possible implementation, the target loss function may further include: a physical merging loss function and a semantic merging loss function.

[0210] The physical merging loss function is used to calculate the loss of the physical merging prediction result, and the semantic merging loss function is used to calculate the loss of the semantic merging prediction result.

[0211] The processing flow of each module in the device and the interaction flow between each module can be referred to the relevant descriptions in the above method embodiments, and will not be detailed here.

[0212] This application also provides an electronic device 140, such as... Figure 14 The diagram shown is a schematic representation of the structure of an electronic device 140 provided in an embodiment of this application. It includes a processor 141, a memory 142, and optionally, a bus 143. The memory 142 stores machine-readable instructions executable by the processor 141 (e.g., ...). Figure 12 The device in question acquires the execution instructions corresponding to module 1201, processing module 1202, and merging module 1203, or... Figure 13 The device includes the execution instructions corresponding to the construction module 1301 and the training module 1302. When the electronic device 140 is running, the processor 141 and the memory 142 communicate through the bus 143. When the machine-readable instructions are executed by the processor 141, the method steps in the aforementioned method embodiment are executed.

[0213] This application also provides a computer-readable storage medium storing a computer program, which, when run by a processor, executes the steps of the above-described table merging method or table merging prediction model training method.

[0214] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems and devices described above can be referred to the corresponding processes in the method embodiments, and will not be repeated here. In the several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. Furthermore, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection can be through some communication interfaces; the indirect coupling or communication connection of devices or modules can be electrical, mechanical, or other forms.

[0215] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. If the functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0216] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.

Claims

1. A method for merging tables, characterized in that, include: Obtain an image to be processed, which includes a first table segment and a second table segment distributed across a page; The image to be processed is input into a pre-trained table merging prediction model for prediction processing to obtain a physical merging prediction result and a semantic merging prediction sequence. The physical merging prediction result is used to indicate whether the first table segment and the second table segment can be merged in terms of physical structure. Each item in the semantic merging prediction sequence corresponds to a cell of the first table segment and the second table segment, and the value of each item in the semantic merging prediction sequence is used to indicate whether the cells on the corresponding first table segment and the second table segment can be semantically merged. Based on the physical merge prediction result and the semantic merge prediction sequence, the first table segment and the second table segment are merged. The table merging prediction model includes: a visual feature processing network, a semantic feature processing network, and a linear processing layer; The step of inputting the image to be processed into a pre-trained table merging prediction model for prediction processing to obtain physical merging prediction results and semantic merging prediction sequences includes: The image to be processed is input into the visual feature processing network, which encodes the visual features to obtain visual features, and then performs positional encoding on the visual features to obtain position-encoded features. The location-encoded features are input into the semantic feature processing network, which performs semantic feature encoding to obtain semantically encoded features. The semantically encoded features are input into the linear processing layer for linear transformation to obtain the physical merging prediction result and the semantic merging prediction sequence.

2. The method according to claim 1, characterized in that, The visual feature processing network includes multiple coding layers connected in sequence; The step of inputting the image to be processed into the visual feature processing network, and having the visual feature processing network encode visual features to obtain visual features, includes: The image to be processed is input into the first coding layer for encoding, and the processed features are input into the next coding layer for encoding, and so on, until the last coding layer completes the encoding process. Max pooling is performed on the processed features of each coding layer except the last coding layer to obtain the pooled features of each coding layer. The pooled features from each coding layer and the features encoded by the last coding layer are concatenated to obtain the visual features.

3. The method according to claim 1, characterized in that, The step of performing positional encoding on the visual features to obtain position-encoded features includes: Based on the two-dimensional sine and cosine coding algorithm, the visual features are encoded in both the horizontal and vertical directions to obtain the encoded features.

4. The method according to any one of claims 1-3, characterized in that, The step of merging the first table segment and the second table segment based on the physical merging prediction result and the semantic merging prediction result includes: If the physical merge prediction result indicates that the first table segment and the second table segment can be merged in physical structure, then each item in the semantic merge prediction sequence is traversed. For the current item that is traversed, if the value of the current item is a preset value, then the content of the cell corresponding to the current item in the first table segment and the cell corresponding to the current item in the second table segment are merged.

5. The method according to any one of claims 1-3, characterized in that, Before acquiring the image to be processed, the process also includes: Based on the page segmentation marks in the electronic document, the first initial table segment and the second initial table segment are extracted from adjacent pages of the electronic document to obtain the initial image; The initial image is cropped and the table edges are aligned to obtain the image to be processed.

6. A method for training a table merging prediction model, characterized in that, include: A training dataset is constructed based on the original electronic document containing cross-page tables; The initial merging model is trained based on the training dataset to obtain a table merging prediction model. The prediction results of the table merging prediction model include physical merging prediction results and semantic merging prediction sequences. The physical merging prediction results are used to indicate whether the first table segment and the second table segment in the input image to be processed can be merged in terms of physical structure. Each item in the semantic merging prediction sequence corresponds to a cell of the first table segment and the second table segment, and the values ​​of each item in the semantic merging prediction sequence are used to indicate whether the cells on the corresponding first table segment and the second table segment can be semantically merged. The step of training the initial merging model based on the training dataset to obtain the table merging prediction model includes: The sample data in the training dataset is input into the initial merging model to obtain the processing result of the initial merging model. The initial merging model includes: a visual feature processing network, a semantic feature processing network, and a linear processing layer. The loss of the initial merging model is obtained by calculating the loss of the initial merging model based on the processing result of the target loss function. The target loss function includes at least an error propagation loss function, which is used to calculate the deviation between the processing result of the visual feature processing network and the processing result of the semantic feature processing network. Based on the loss of the initial merging model, the initial merging model is iteratively corrected to obtain the table merging prediction model.

7. The method according to claim 6, characterized in that, The training dataset, constructed based on the original electronic document containing cross-page tables, includes: The cross-page tables in the original electronic document are cropped, trimmed, and their edges aligned to obtain the test dataset and the positive training dataset. The tables in the original electronic document are split and concatenated to obtain a negative training dataset.

8. The method according to claim 6, characterized in that, The target loss function also includes: a physical merging loss function and a semantic merging loss function; The physical merging loss function is used to calculate the loss of the physical merging prediction result, and the semantic merging loss function is used to calculate the loss of the semantic merging prediction result.

9. A table merging device, characterized in that, include: The acquisition module is used to acquire the image to be processed, which includes a first table segment and a second table segment distributed across the page. The processing module is used to input the image to be processed into a pre-trained table merging prediction model for prediction processing, and obtain a physical merging prediction result and a semantic merging prediction sequence. The physical merging prediction result is used to indicate whether the first table segment and the second table segment can be merged in physical structure. Each item in the semantic merging prediction sequence corresponds to a cell of the first table segment and the second table segment, and the value of each item in the semantic merging prediction sequence is used to indicate whether the cells on the corresponding first table segment and the second table segment can be semantically merged. The merging module is used to merge the first table segment and the second table segment according to the physical merging prediction result and the semantic merging prediction sequence; The table merging prediction model includes: a visual feature processing network, a semantic feature processing network, and a linear processing layer; The processing module is specifically used for: The image to be processed is input into the visual feature processing network, which encodes the visual features to obtain visual features, and then performs positional encoding on the visual features to obtain position-encoded features. The location-encoded features are input into the semantic feature processing network, which performs semantic feature encoding to obtain semantically encoded features. The semantically encoded features are input into the linear processing layer for linear transformation to obtain the physical merging prediction result and the semantic merging prediction sequence.

10. A training device for a table merging prediction model, characterized in that, include: The building block is used to construct the training dataset based on the original electronic document containing cross-page tables; The training module is used to train the initial merging model based on the training dataset to obtain a table merging prediction model. The prediction results of the table merging prediction model include physical merging prediction results and semantic merging prediction sequences. The physical merging prediction results are used to indicate whether the first table segment and the second table segment in the input image to be processed can be merged in terms of physical structure. Each item in the semantic merging prediction sequence corresponds to a cell of the first table segment and the second table segment, and the values ​​of each item in the semantic merging prediction sequence are used to indicate whether the cells on the corresponding first table segment and the second table segment can be semantically merged. The training module is specifically used for: The sample data in the training dataset is input into the initial merging model to obtain the processing result of the initial merging model. The initial merging model includes: a visual feature processing network, a semantic feature processing network, and a linear processing layer. The loss of the initial merging model is obtained by calculating the loss of the initial merging model based on the processing result of the target loss function. The target loss function includes at least an error propagation loss function, which is used to calculate the deviation between the processing result of the visual feature processing network and the processing result of the semantic feature processing network. Based on the loss of the initial merging model, the initial merging model is iteratively corrected to obtain the table merging prediction model.

11. An electronic device, characterized in that, include: The processor and memory, the memory storing machine-readable instructions executable by the processor, wherein when the electronic device is running, the processor executes the machine-readable instructions to perform the steps of the table merging method as described in any one of claims 1 to 5 or the steps of the table merging prediction model training method as described in any one of claims 6 to 8.

12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, performs the steps of the table merging method as described in any one of claims 1 to 5 or the steps of the table merging prediction model training method as described in any one of claims 6 to 8.

Citation Information

Patent Citations

  • PDF document cross-page table merging method and device, electronic equipment and storage medium

    CN112380825A

  • Table fragment link recovery method and system based on semantic processing

    CN112632927A