Table Structure Recognition and Model Training Method, Device, Equipment and Medium

By obtaining the target features of the table image, determining the reference position and performing feature fusion processing, and using the encoder and decoder to obtain cell structure information, the problem of poor table structure recognition effect in the prior art is solved, and more efficient table structure recognition is achieved.

CN116884025BActive Publication Date: 2025-07-22BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310708376.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-14
Publication Date
2025-07-22
Estimated Expiration
2043-06-14

AI Technical Summary

Technical Problem

The existing table structure recognition methods have shortcomings in recognition effects, making it difficult to accurately identify the cell structure in the table.

Method used

By obtaining the target image features of the table image, determining the reference position of the table segmentation line, and fusing the reference characteristics corresponding to the reference position with the query characteristics of the candidate points, using the encoder and decoder to obtain the structural information of the cells, constructing a total loss function to adjust the model parameters to improve the recognition accuracy.

Benefits of technology

It improves the accuracy and efficiency of table structure identification, can better identify complex forms such as financial statements and company bills, simplify manual processing, and improve office efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116884025B_ABST
    Figure CN116884025B_ABST
Patent Text Reader

Abstract

The present disclosure provides a method, apparatus, device, and storage medium for table structure recognition and model training, which relates to the field of artificial intelligence technology, specifically in the technical fields of computer vision, image processing, deep learning, etc., and can be applied to scenarios such as intelligent finance. The table structure recognition method includes: obtaining target image features of a table image; determining a reference position of table dividing lines in the table based on the target image features; performing fusion processing on the reference features corresponding to the reference position and the query features of candidate points on the table dividing lines to obtain fusion features; performing decoding processing on the fusion features to obtain the structural information of the candidate points; and determining the structural information of cells in the table based on the structural information of the candidate points. The present disclosure can improve the effect of table structure recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technologies, specifically to technical fields such as computer vision, image processing, and deep learning, and can be applied to scenarios such as intelligent finance. In particular, it relates to a method, device, equipment, and medium for table structure recognition and model training. Background Art

[0002] With the popularization of automated office work, table data is increasingly widely used in various industries. With the wide application of table data, the task of table structure recognition has become increasingly important. Table structure recognition mainly obtains the structural information of cells in a table. Summary of the Invention

[0003] The present disclosure provides a method, device, equipment, and medium for table structure recognition and model training.

[0004] According to one aspect of the present disclosure, a method for table structure recognition is provided, including: obtaining target image features of a table image; wherein, the table image contains a table; based on the target image features, determining a reference position of table dividing lines in the table; performing fusion processing on the reference features corresponding to the reference position and the query features of candidate points on the table dividing lines to obtain fusion features; performing decoding processing on the fusion features to obtain the structural information of the candidate points; based on the structural information of the candidate points, determining the structural information of cells in the table.

[0005] According to another aspect of the present disclosure, a method for training a table structure recognition model is provided. The table structure recognition model includes: an encoder and a decoder. The method includes: using the encoder to obtain target image features of a sample image; wherein, the sample image contains a table; based on the target image features, determining a predicted reference position of table dividing lines in the table; performing fusion processing on the reference features corresponding to the predicted reference position and the query features of candidate points on the table dividing lines to obtain fusion features; using the decoder to perform decoding processing on the fusion features to obtain the predicted structural information of the candidate points; based on the predicted reference position and the predicted structural information, constructing a total loss function; using the total loss function to adjust the model parameters of the encoder and the model parameters of the decoder.

[0006] According to another aspect of the present disclosure, a table structure recognition device is provided, including: an acquisition module configured to acquire target image features of a table image; wherein, the table image contains a table; a determination module configured to determine a reference position of a table dividing line in the table based on the target image features; a fusion module configured to perform a fusion process on the reference features corresponding to the reference position and the query features of candidate points on the table dividing line to obtain fusion features; a decoding module configured to perform a decoding process on the fusion features to obtain the structural information of the candidate points; and an identification module configured to determine the structural information of cells in the table based on the structural information of the candidate points.

[0007] According to another aspect of the present disclosure, a training device for a table structure recognition model is provided. The table structure recognition model includes an encoder and a decoder. The device includes: an encoding module configured to use the encoder to acquire target image features of a sample image; wherein, the sample image contains a table; a determination module configured to determine a predicted reference position of a table dividing line in the table based on the target image features; a fusion module configured to perform a fusion process on the reference features corresponding to the predicted reference position and the query features of candidate points on the table dividing line to obtain fusion features; a decoding module configured to use the decoder to perform a decoding process on the fusion features to obtain the predicted structural information of the candidate points; a construction module configured to construct a total loss function based on the predicted reference position and the predicted structural information; and an adjustment module configured to use the total loss function to adjust the model parameters of the encoder and the model parameters of the decoder.

[0008] According to another aspect of the present disclosure, an electronic device is provided, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the method according to any one of the above aspects.

[0009] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause the computer to execute the method according to any one of the above aspects.

[0010] According to another aspect of the present disclosure, a computer program product is provided, including a computer program which, when executed by a processor, implements the method according to any one of the above aspects.

[0011] According to the technical solution of the present disclosure, the table structure recognition effect can be improved.

[0012] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. Description of the Drawings

[0013] The drawings are used to better understand the present solution and do not constitute a limitation to the present disclosure. Among them:

[0014] Figure 1 is a schematic diagram according to the first embodiment of the present disclosure;

[0015] Figure 2 is a schematic diagram of reference points and candidate points provided according to the embodiments of the present disclosure;

[0016] Figure 3 is a schematic diagram of the decoding result provided according to the embodiments of the present disclosure;

[0017] Figure 4 is a schematic diagram of the application scenario provided according to the embodiments of the present disclosure;

[0018] Figure 5 is a schematic diagram of the overall architecture of the table structure recognition model provided according to the embodiments of the present disclosure;

[0019] Figure 6 is a schematic diagram of obtaining row reference features and column reference features based on target image features provided according to the embodiments of the present disclosure;

[0020] Figure 7 is a schematic diagram of the decoder provided according to the embodiments of the present disclosure;

[0021] Figure 8 is a schematic diagram of the merging process provided according to the embodiments of the present disclosure;

[0022] Figure 9 is a schematic diagram according to the second embodiment of the present disclosure;

[0023] Figure 10 is a schematic diagram according to the third embodiment of the present disclosure;

[0024] Figure 11 is a schematic diagram according to the fourth embodiment of the present disclosure;

[0025] Figure 12 is a schematic diagram according to the fifth embodiment of the present disclosure;

[0026] Figure 13 is a schematic diagram of an electronic device for implementing the table structure recognition method or the training method of the table structure recognition model according to the embodiments of the present disclosure. Detailed Embodiments

[0027] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, descriptions of well-known functions and structures are omitted below for clarity and conciseness.

[0028] In the related art, there are some table structure recognition schemes, such as the method based on segmentation first and then combination, the method based on detection first and then classification, etc., but the recognition effect needs to be improved.

[0029] To improve the table structure recognition effect, the following embodiments are provided in the present disclosure.

[0030] Figure 1 It is a schematic diagram according to the first embodiment of the present disclosure. This embodiment provides a table structure recognition method, which includes:

[0031] 101. Obtain the target image features of the table image; wherein, the table image contains a table.

[0032] 102. Based on the target image features, determine the reference positions of the table dividing lines in the table.

[0033] 103. Perform fusion processing on the reference features corresponding to the reference positions and the query features of the candidate points on the table dividing lines to obtain fusion features.

[0034] 104. Perform decoding processing on the fusion features to obtain the structure information of the candidate points.

[0035] 105. Based on the structure information of the candidate points, determine the structure information of the cells in the table.

[0036] Among them, the table image refers to an image containing a table.

[0037] The image features of the table image can be divided into initial image features and target image features.

[0038] The initial image features refer to the image features obtained by using an image feature extraction network. The image feature extraction network can also be called a backbone network. Input the table image into the backbone network, and the output is the initial image features.

[0039] After obtaining the initial image features, the initial image features can be directly used as the target image features.

[0040] To improve the table structure recognition effect, the initial image features can also be enhanced to obtain the target image features.

[0041] The basic element of a table is a cell. Horizontally, one or more cells form a row of the table, and vertically, one or more cells form a column of the table.

[0042] Table dividing lines include row dividing lines and column dividing lines. A row dividing line is the boundary line of a row. Each row corresponds to two row dividing lines, one above and one below. A column dividing line is the boundary line of a column. Each column corresponds to two column dividing lines, one on the left and one on the right.

[0043] The reference position of a table dividing line can be identified by a reference point. The reference point of a row dividing line can be called a row reference point, and the reference point of a column dividing line can be called a column reference point.

[0044] For example, as Figure 2 shown, for the Figure 2 table image shown, the table in this image includes 7 row dividing lines and 6 column dividing lines. The position of each row dividing line can be identified by a row reference point (black dot), and there are 7 row reference points in total. The position of each column dividing line can be identified by a column reference point (white dot), and there are 6 column reference points in total.

[0045] Generally speaking, in a coordinate system, the left - right direction is the x - axis direction, and the up - down direction is the y - axis direction. Since rows are distributed vertically and columns are distributed horizontally, the reference position of a row dividing line refers to the y - coordinate of the row reference point, and the reference position of a column dividing line refers to the x - coordinate of the column reference point.

[0046] A reference feature is a feature obtained after encoding the reference position. For example, for a row reference point, after encoding the y - coordinate of each row reference point, a vector is obtained as the reference feature of this row reference point. For a column reference point, after encoding the x - coordinate of each column reference point, a vector is obtained as the reference feature of this column reference point.

[0047] The basic element of a table is a cell (cell). For example, for the Figure 2 table image shown, it has 7 row dividing lines and 6 column dividing lines in total. These table dividing lines form 6 * 5 = 30 cells. Each cell includes 4 vertices, and these vertices are distributed on the row dividing lines and column dividing lines.

[0048] To identify the cells in a table, the vertices can be identified first, and then the cells are formed by the vertices. The vertices can be determined from candidate points. For example, if vertices A, B, C, and D are identified among multiple candidate points, then a cell can be formed by these four points.

[0049] Among them, for each table dividing line, a fixed number of candidate points can be preset. For example, for each row dividing line, M (a positive integer greater than 1) first candidate points can be set, and for each column dividing line, N (a positive integer greater than 1) second candidate points can be set. As Figure 2 shown, the first candidate points (black triangles) on the uppermost row dividing line and the second candidate points (white triangles) on the leftmost column dividing line are shown.

[0050] The features of the candidate points can be called query features. Specifically, for the first candidate points, their query features can be called row query features, and for the second candidate points, their query features can be called column query features. Both the row query features and the column query features can be pre-trained.

[0051] After obtaining the corresponding reference features and the query features of the candidate points for each table dividing line, the reference features and the query features can be fused to obtain fused features. The reference features are generally in the form of vectors, which can be called reference feature vectors, and the query features of each candidate point are also in the form of vectors, which can be called query feature vectors. Therefore, the fusion process can specifically be the addition operation of the reference feature vector and the query feature vectors of each candidate point to obtain a fused feature vector.

[0052] After obtaining the fused features (fused feature vectors), they can be decoded to obtain the structural information of the corresponding candidate points.

[0053] Among them, for the fused features corresponding to the row dividing line, a row decoder can be used for decoding to obtain the structural information of the first candidate points; for the fused features corresponding to the column dividing line, a column decoder can be used for decoding to obtain the structural information of the second candidate points.

[0054] The structural information of the candidate points can include: the position coordinates and connection relationships of the candidate points.

[0055] Specifically, as Figure 3As shown, a row decoder is used to obtain the structural information of the first candidate point. The structural information of the first candidate point includes: the position coordinates of the first candidate point and the connection relationship of the first candidate point. The position coordinates of the first candidate point refer to the y coordinate of the first candidate point. The connection relationship of the first candidate point is a downward connection relationship, that is, whether there is a connection downward. Among them, having a connection downward indicates the existence of a column dividing line, otherwise, there is no column dividing line. A column decoder is used to obtain the structural information of the second candidate point. The structural information of the second candidate point includes: the position coordinates of the second candidate point and the connection relationship of the second candidate point. The position coordinates of the second candidate point refer to the x coordinate of the second candidate point. The connection relationship of the second candidate point is a rightward connection relationship, that is, whether there is a connection rightward. Among them, having a connection rightward indicates the existence of a row dividing line, otherwise, there is no row dividing line.

[0056] After obtaining the structural information of the candidate point, the table structure recognition result can be obtained based on the structural information of the candidate point. The table structure recognition result is the structural information of the cells in the table. The structural information of the cells can include the physical coordinates and logical coordinates of the cells. The physical coordinates of the cells include the physical coordinates of the four vertices of the cell, which can be represented by {(x1,y1),(x2,y2),(x3,y3),(x4,y4)}, where (x1,y1), (x2,y2), (x3,y3), and (x4,y4) are the physical coordinates of the four vertices of the cell respectively. The logical coordinates of the cells are used to characterize the logical position of the cells, which can be represented by the i-th row and the j-th column. Both i and j are positive integers. Specifically, for example, it is located in the first row and the second column.

[0057] In this embodiment, the reference position of the target table element is determined based on the target image feature. The fusion process is performed on the reference position feature corresponding to the reference position and the query feature of the candidate point to obtain the fusion feature, and the structure recognition result is obtained based on the fusion feature, which can improve the accuracy of the table structure recognition. In addition, since the target image feature can reflect the overall feature of the table image, compared with the method of first segmenting and then recognizing or first detecting and then recognizing, the accuracy and efficiency of the table structure recognition can be improved. Therefore, the effect of the table structure recognition can be improved.

[0058] To better understand the embodiments of the present disclosure, the application scenarios applicable to the embodiments of the present disclosure are described.

[0059] Figure 4 is a schematic diagram of the application scenario provided by the embodiments of the present disclosure. As Figure 4As shown in the figure, the user can send the table image to the server 402 through the user terminal 401. The server 402 processes the table image to obtain the table structure recognition result, which can specifically be the structure information of the target cell. Enter the image to be recognized at the client. The image contains a table. The client sends the image to the server, and the server recognizes the image to obtain the structure information of the table in the image. The user terminal can be a personal computer (Personal Computer), a laptop, a mobile device (such as a mobile phone), etc. The server can be a local server or a cloud server, etc. The server can be a single server or a server cluster. It can be understood that if the user terminal has the corresponding ability, the table structure recognition can also be performed locally on the user terminal.

[0060] Among them, based on the table structure recognition model, the table image can be processed to obtain the table structure recognition result.

[0061] As Figure 5 shown in the figure, the table structure recognition model can include: an encoder and a decoder. Among them, the encoder can include: an image feature extraction network 501 and an image feature enhancement network 502; the decoder can include: a row decoder 503 and a column decoder 504.

[0062] The image feature extraction network 501 is used to perform image feature extraction processing on the table image to obtain the initial image features.

[0063] The initial image features can include image features of multiple sizes. The image feature extraction network is, for example, a ResNet50 network. After inputting the table image into this ResNet50 network, feature maps of multiple resolutions (sizes) can be obtained, and these feature maps are used as the initial image features. The specific sizes and number of channels of these feature maps can be as follows: (H / 4, W / 4, 256), (H / 8, W / 8, 512), (H / 16, W / 16, 1024), (H / 32, W / 32, 2048), where H represents the height and W represents the width.

[0064] The image feature enhancement network 502 is used to perform enhancement processing on the initial image features to obtain the target image features.

[0065] Among them, the image feature enhancement network can be a network based on a Transformer encoder. The input is the initial image features of multiple sizes, and the output is the target image features after attention processing of these sizes of the initial image features.

[0066] The attention processing can specifically be multi-scale deformable attention processing.

[0067] The calculation formula is as follows:

[0068]

[0069] Among them, is the target image feature; are the initial image features of multiple sizes; h, l, k respectively represent the attention head index, the size level index of the input initial image feature, and the sampling point index; A hlk (q) represents the attention weight of vector q, and vector q is set or learned; φ l (p(q)) represents mapping the normalized coordinates to the l-th size level; Δp hlk (q) corresponds to the sampling offset of vector q; W h , W h ' represents the learnable weight in the original multi-head attention mechanism; compared with using the original attention mechanism, the deformable attention mechanism only needs to sample L×K sampling points, instead of sampling H×W sampling points on the original image, which greatly reduces the computational complexity.

[0070] After obtaining the target image feature, the row reference feature and the column reference feature can be obtained based on the target image feature.

[0071] Among them, the target image feature can be regressed to obtain the reference position, and then the reference position can be encoded to obtain the reference feature.

[0072] For example, as Figure 6 shown, the table structure recognition model can further include: a row regression network 601, a row position encoding module 602, a column regression network 603, and a column position encoding module 604.

[0073] The row regression network 601 is used to perform row regression processing on the target image feature to determine the reference position of the row dividing line, that is, the y coordinate of the row reference point.

[0074] The row position encoding module 602 is used to perform encoding processing on the reference position of the row dividing line to obtain the row reference feature, that is, convert the y coordinate of the row reference point into a vector.

[0075] The column regression network 603 is used to perform column regression processing on the target image feature to determine the reference position of the column dividing line, that is, the x coordinate of the column reference point.

[0076] The column position encoding module 604 is used to perform encoding processing on the reference position of the column dividing line to obtain the column reference feature, that is, convert the x coordinate of the column reference point into a vector.

[0077] The row regression network and the column regression network can specifically be fully connected layers.

[0078] Taking the row as an example (the column can be referred to for implementation), through the row regression network, the predicted y coordinate values of M first candidate points in each row can be obtained. Among these M predicted y coordinate values, K (which can be set) predicted y coordinate values with relatively high confidence are selected, and the mean of these K predicted y coordinate values is used as the y coordinate of the row reference point for each row.

[0079] The row position encoding module and the column position encoding module can specifically be Sine-based encoding layers, that is, they can specifically include: Sine coordinate encoding embedding layer -> linear layer -> regularization layer.

[0080] The row query feature and the column query feature can be pre-trained.

[0081] Taking the row as an example (the column can be referred to for implementation), the row reference feature corresponds to the row dividing line. That is, assuming there are N row dividing lines, then N row reference features can be obtained, and each row reference feature corresponds to a row dividing line.

[0082] The row query feature is pre-trained. Assuming M first candidate points are set on each row dividing line, then the row query feature is M-dimensional, and the elements of each dimension correspond to a first candidate point.

[0083] Both the row reference feature and the row query feature are in vector form. When performing an addition operation between the row reference feature and the row query feature, it is a vector addition, that is, the row query feature (row query feature vector) is copied N times, and each row query feature vector is added to each row reference feature vector to obtain the fusion feature (fusion feature vector) corresponding to each row dividing line, and a total of N groups of fusion feature vectors are obtained. Each group of fusion feature vectors includes M elements.

[0084] For example, referring to Figure 2 , for the uppermost row dividing line, it corresponds to a row reference point and 6 first candidate points. Then, the reference feature vector of the row reference point can be added to the query feature vector of each first candidate point respectively, and a total of 6 groups of fusion feature vectors are obtained.

[0085] After obtaining the fusion features (row fusion feature and column fusion feature), the corresponding decoder can be used to perform decoding processing on them to obtain the structural information of the corresponding candidate points.

[0086] Still taking the row as an example, a row decoder can be used to perform decoding processing on the row fusion feature to obtain the structural information of each first candidate point on each row dividing line.

[0087] The structural information of the first candidate point includes: the position coordinates of the first candidate point and the connection relationship of the first candidate point. For example Figure 3As shown, the position coordinate of the first candidate point refers to the y coordinate of the first candidate point, and the connection relationship of the first candidate point refers to whether there is a downward connection.

[0088] Similarly, for columns, a column decoder is used to decode the column fusion features to obtain the structural information of each second candidate point on each column segmentation line.

[0089] The structural information of the second candidate point includes: the position coordinate of the second candidate point and the connection relationship of the second candidate point. As Figure 3 shown, the position coordinate of the second candidate point refers to the x coordinate of the second candidate point, and the connection relationship of the second candidate point refers to whether there is a rightward connection.

[0090] The row decoder and the column decoder can have the same structure but different model parameters, and the model parameters are determined through training.

[0091] As Figure 7 shown, the structure of the decoder (row decoder and column decoder) can include: multiple decoder layers ( Figure 7 represented by the first decoder layer to the sixth decoder layer in

[0092] Each decoder layer includes: a factorized self-attention layer, a multi-scale deformable cross-attention layer, and a feed-forward fully connected layer. Different from the way of only using self-attention, in this embodiment, the factorized self-attention layer is used to perform the interaction between features to improve the expression ability of features. The factorized self-attention layer is calculated from the within-group (intra-group fusion features) attention and the between-group (different-group fusion features). Taking rows as an example, the fusion features corresponding to the same row segmentation line belong to the same group and can be called intra-group features; the fusion features corresponding to different row segmentation lines belong to different groups and can be called between-group features. Among them, the within-group attention mechanism means calculating Q, K, and V based on the intra-group features, and finally obtaining the weighted row fusion features through softmax after obtaining the attention weights; the between-group attention mechanism is to calculate the attention weights for the between-group features to obtain the weighted row fusion features. This decoder uses 6 identical sub-modules in total to improve the expression ability of features.

[0093] After the decoding process of the decoder, the structural information of the candidate points can be obtained. Then, based on the structural information of the candidate points, the final table structure recognition result, that is, the structural information of the target cell, can be obtained.

[0094] Specifically, for the row decoder, this branch outputs the y coordinates of M points (first candidate points) in each row and the downward connection relationship (whether there is a downward connection) of each point; for the column decoder, this branch outputs the x coordinates of N points (second candidate points) in each column and the rightward connection relationship (whether there is a rightward connection) of each point.

[0095] Based on the above output, the final table structure recognition result can be obtained through the following process, including the physical coordinates and logical coordinates of the target cells:

[0096] (1) First, pair up the y - coordinates of the M points in each row with the x - coordinates of the N points to obtain the coordinates (x, y) of the candidate points in each row, as well as the downward connection relationship and the right - ward connection relationship;

[0097] (2) Based on the coordinates (x, y), determine the upper - left vertex. Generally speaking, the positive x - direction is to the right, and the positive y - direction is downward. Then select the vertex with the smallest x and the smallest y as the upper - left vertex. Start traversing to the right from the upper - left vertex. When there is no downward connection for a certain candidate point, it indicates that the corresponding candidate cell needs to be merged, so merge it until there is a downward connection for a certain candidate point and then stop merging. At this time, the ending column index of the target cell is obtained;

[0098] (3) When the right - ward traversal stops, start traversing downward. When there is no right - ward connection for a certain candidate point, it indicates that the corresponding candidate cell needs to be merged, so merge it until there is a right - ward connection for a certain candidate point and then stop merging. At this time, the ending row index of the target cell is obtained;

[0099] (4) By analogy, finally, the physical coordinates and logical coordinates of all target cells in the table can be obtained. Among them, the physical coordinates refer to the coordinates (x, y) of the four vertices of the target cell finally merged, and the logical coordinates refer to the row and column where the target cell is located. The ending row index and the ending column index can be used as the corresponding logical coordinates.

[0100] For example, Figure 8As shown, after the processing by the decoder, the structural information (x, y; whether there is a connection downward; whether there is a connection to the right) of each candidate point (such as the black origin) can be obtained. Starting from the top-left vertex O and traversing, when traversing to a candidate point A and learning that there is no connection downward, a merging process is performed, that is, the candidate point A is not the vertex of the final target cell. When traversing to another candidate point B and learning that there is a connection downward, the merging stops. B is used as the vertex of a target cell, and the termination column index of this target cell is obtained. The termination column index and the termination row index can be incremented from the initial value. For example, if the initial value is 1, since this is the first time to determine the termination column index, the termination column index j = 1 at this time. Then, when traversing downward to the candidate point C and since C has a connection to the right, the merging stops. C is used as the vertex of this target cell, and the termination row index of this target cell is obtained, such as i = 1. And so on, another vertex D of this target cell can be obtained, that is, a target cell composed of OBCD can be determined, and the physical coordinates (x, y of the four points O, B, C, D) of this target cell, as well as the logical coordinates (determined based on the termination row index i and the termination column index j, such as the 1st row and the 1st column) can be obtained.

[0101] For another example, after traversing C and D and then traversing downward to the candidate point E, since E has no connection to the right, E is merged. Continuing to traverse downward to the candidate point F, since the candidate point F has a connection to the right, the merging stops, and F is used as the vertex of the target cell. Similarly, another target cell CDFG can be obtained, as well as the physical coordinates and logical coordinates of this target cell.

[0102] Through the above process, the final table structure recognition result, that is, the physical coordinates and logical coordinates of the target cell, can be obtained.

[0103] In this embodiment, determining the reference position can be regarded as the first stage, and the subsequent processing based on the reference position can be regarded as the second stage. Therefore, it is a two-stage process. Additionally, during the decoding process, processing is performed based on intra-group features and inter-group features. Therefore, it can be called a two-stage table structure recognition method based on intra-group and inter-group attention. This method can extract table data from financial statements, company bills, exam papers, etc., greatly simplifying the manpower and accelerating the work efficiency of the office. This embodiment can be used as a separate table structure recognition product or as a complete text recognition product, embedding the table structure recognition ability into the document service. It requires a pre-table area detection module and a post-optical character recognition (OCR) content detection and recognition module. Finally, the model can directly output the cell coordinate information and content information of the table end-to-end. Compared with some previous solutions, the solution proposed in this embodiment can cover more table data scenarios. By using only the same table structure recognition model, it can solve multiple types of table data simultaneously, reducing the model deployment and model size, and having obvious advantages over previous solutions.

[0104] Combined with the above application scenarios, the present disclosure also provides the following embodiments.

[0105] Figure 9 It is a schematic diagram according to the second embodiment of the present disclosure. This embodiment provides a table structure recognition method, which includes:

[0106] 901. Perform image feature extraction processing on the table image to obtain initial image features of multiple sizes.

[0107] For example, as Figure 5 shown, an image feature extraction network in the table structure recognition model can be used to process the input table image, and the output is initial image features of multiple sizes.

[0108] 902. Perform enhancement processing on the initial image features of the multiple sizes to obtain target image features.

[0109] For example, as Figure 5 shown, an image feature enhancement network in the table structure recognition model can be used to process the initial image features of the multiple sizes, and the output is target image features.

[0110] In this embodiment, by obtaining initial image features of multiple sizes and obtaining target image features based on the initial image features of multiple sizes, the expression ability of the target image features can be improved, thereby improving the table structure recognition effect.

[0111] 903. Determine the reference positions of the table dividing lines based on the target image features.

[0112] Among them, the table dividing lines include: row dividing lines and column dividing lines;

[0113] Correspondingly, row regression processing can be performed on the target image features to determine the reference positions of the row dividing lines; column regression processing can be performed on the target image features to determine the reference positions of the column dividing lines.

[0114] For example, as Figure 6 shown, the row regression network in the table structure recognition model can be used to perform row regression processing on the input target image features to determine the reference positions of the row dividing lines; the column regression network in the table structure recognition model can be used to perform column regression processing on the target image features to determine the reference positions of the column dividing lines.

[0115] In this embodiment, by performing regression processing on the target image features, accurate reference positions can be obtained, thereby improving the accuracy of table structure recognition.

[0116] 904. Perform fusion processing on the reference features corresponding to the reference positions and the query features of the candidate points on the table dividing lines to obtain fusion features.

[0117] For example, referring to Figure 6 , a row coordinate encoding module can be used to perform encoding processing on the input row reference positions, and the output is the row reference features; a column coordinate encoding module can be used to perform encoding processing on the input column reference positions, and the output is the column reference features.

[0118] As Figure 5 shown, the row query features and column query features can be pre-trained. The row query features and row reference features can be added to obtain row fusion features; the column query features and column reference features can be added to obtain column fusion features.

[0119] In this embodiment, by adding the reference features and query features, fusion features can be obtained. Based on the fusion features, the table structure recognition results can be obtained, which can improve the accuracy of table structure recognition.

[0120] 905. Perform decoding processing on the fusion features to obtain the structural information of the candidate points.

[0121] For example, as Figure 5As shown in the figure, the row decoder in the table structure recognition model is used to perform row decoding processing on the row fusion feature to obtain the structure information of the first candidate point. The structure information of the first candidate point includes: the ordinate (y coordinate) of the first candidate point, and the downward connection relationship. The column decoder in the table structure recognition model is used to perform column decoding processing on the column fusion feature to obtain the structure information of the second candidate point. The structure information of the second candidate point includes: the abscissa (x coordinate) of the second candidate point, and the rightward connection relationship.

[0122] In this embodiment, by performing decoding processing on the fusion feature, the structure information of the corresponding candidate point can be accurately obtained.

[0123] 906. Based on the structure information of the candidate point, determine the structure information of the cell in the table.

[0124] Among them, the structure information of the cell includes: the physical coordinates and the logical coordinates of the cell. Correspondingly, the determining the structure information of the cell in the table based on the structure information of the candidate point includes: determining a reference point based on the ordinate of the first candidate point and the abscissa of the second candidate point; determining a first target point and a termination column index based on the downward connection relationship of the first candidate point; determining a second target point and a termination row index based on the rightward connection relationship of the second candidate point; determining a third target point based on the first target point and the second target point; forming the cell with the reference point, the first target point, the second target point, and the third target point, and determining the physical coordinates of the cell; determining the logical coordinates of the cell based on the termination row index and the termination column index.

[0125] For example, such as Figure 8As shown, after the processing of the decoder, the structural information (x, y; whether there is a connection downward; whether there is a connection to the right) of each candidate point (such as the black origin) can be obtained. Starting from the top-left vertex O and traversing, when a candidate point A is traversed and it is known that there is no connection downward, a merging process is performed, that is, the candidate point A is not the vertex of the final target cell. When another candidate point B is traversed and it is known that there is a connection downward, the merging stops. B is used as the vertex of a target cell, and the termination column index of this target cell is obtained. The termination column index and the termination row index can be incremented from the initial value. For example, if the initial value is 1, since this is the first time the termination column index is determined, the termination column index j = 1 at this time. After that, when traversing downward to the candidate point C, since C has a connection to the right, the merging stops. C is used as the vertex of this target cell, and the termination row index of this target cell is obtained, such as i = 1. By analogy, another vertex D of this target cell can be obtained, that is, a target cell composed of OBCD can be determined, and the physical coordinates (x, y of the four points O, B, C, D) of this target cell, as well as the logical coordinates (determined based on the termination row index i and the termination column index j, such as the first row and the first column) can be obtained.

[0126] For another example, after traversing C and D and then traversing downward to the candidate point E, since E has no connection to the right, E is merged. Continuing to traverse downward to the candidate point F, since the candidate point F has a connection to the right, the merging stops, and F is used as the vertex of the target cell. Similarly, another target cell CDFG can be obtained, as well as the physical coordinates and logical coordinates of this target cell.

[0127] Through the above process, the final table structure recognition result, that is, the physical coordinates and logical coordinates of the target cell, can be obtained.

[0128] In this embodiment, obtaining the cell structure information based on the structural information of the candidate points can improve the accuracy and efficiency of table structure recognition.

[0129] Figure 10 It is a schematic diagram according to the third embodiment of the present disclosure. This embodiment provides a training method for a table structure recognition model. The table structure recognition model includes: an encoder and a decoder. The method includes:

[0130] 1001. Using the encoder to obtain the target image features of the sample image; wherein, the sample image contains a table.

[0131] 1002. Based on the target image features, determining the predicted reference positions of the table dividing lines in the table.

[0132] 1003. Perform a fusion process on the reference feature corresponding to the predicted reference position and the query feature of the candidate points on the table dividing line to obtain a fusion feature.

[0133] 1004. Use the decoder to perform a decoding process on the fusion feature to obtain the predicted structure information of the candidate points.

[0134] 1005. Based on the predicted reference position and the predicted structure information, construct a total loss function; 1006. Use the total loss function to adjust the model parameters of the encoder and the model parameters of the decoder.

[0135] In this embodiment, determining the reference position of the target table element based on the target image feature, performing a fusion process on the reference position feature corresponding to this reference position and the query feature of the candidate points to obtain a fusion feature, and training the model based on the fusion feature can improve the accuracy of the table structure recognition model; in addition, since the target image feature can reflect the overall feature of the table image, compared with the methods of first segmenting and then recognizing or first detecting and then recognizing, it can improve the accuracy and efficiency of the table structure recognition model. Therefore, the effect of the table structure recognition model can be improved.

[0136] Among them, for the total loss function, a first loss function can be constructed based on the predicted reference position and the true reference position of the table reference line; a second loss function can be constructed based on the predicted structure information and the true structure information of the candidate points; based on the first loss function and the second loss function, the total loss function is constructed.

[0137] The true reference position and the true structure information can be marked by means such as manual or automatic.

[0138] Total loss function = First loss function + Second loss function;

[0139] First loss function = Regression loss of the y coordinate of the row reference point + Regression loss of the x coordinate of the column reference point;

[0140] Second loss function = Regression loss of the y coordinate of the first candidate point + Classification loss of the downward connection relationship of the first candidate point + Regression loss of the x coordinate of the second candidate point + Classification loss of the rightward connection relationship of the second candidate point.

[0141] The above regression loss can specifically adopt L1 loss, and the classification loss can specifically adopt Focal loss.

[0142] The model is updated based on the total loss function. For example, the Back Propagation (BP) algorithm is used to adjust the model parameters and update the model parameters based on the total loss function.

[0143] The adjusted parameter can be: adjusted parameter = parameter before adjustment + learning rate * gradient value, where the gradient value is obtained by taking the derivative of the total loss function with respect to the model parameters.

[0144] In this embodiment, constructing a total loss function based on the first loss function and the second loss function can refer to information in multiple dimensions and improve the model effect.

[0145] For other steps, the table structure recognition process can be referred to.

[0146] In some embodiments, the encoder includes: an image feature extraction network and an image feature enhancement network. Using the encoder to obtain the target image features of the sample image includes:

[0147] Using the image feature extraction network to perform image feature extraction processing on the table image to obtain initial image features of multiple sizes;

[0148] Using the image feature enhancement network to perform enhancement processing on the initial image features of multiple sizes to obtain the target image features.

[0149] In this embodiment, by obtaining initial image features of multiple sizes and obtaining target image features based on the initial image features of multiple sizes, the expression ability of the target image features can be improved, and further the effect of the table structure recognition model can be improved.

[0150] In some embodiments, the table structure recognition model further includes: a row regression network and a column regression network; the table dividing lines include: row dividing lines and column dividing lines; based on the target image features, determining the predicted reference positions of the table dividing lines in the table includes: using the row regression network to perform row regression processing on the target image features to determine the predicted reference positions of the row dividing lines; using the column regression network to perform column regression processing on the target image features to determine the predicted reference positions of the column dividing lines.

[0151] In this embodiment, by performing regression processing on the target image features, accurate predicted reference positions can be obtained, and further the accuracy of the table structure recognition model can be improved.

[0152] In some embodiments, the table dividing lines include: row dividing lines and column dividing lines; the reference features include: row reference features and column reference features; the query features include: row query features and column query features; the fusion features include: row fusion features and column fusion features; the process of fusing the reference features corresponding to the reference positions and the query features of the candidate points on the table dividing lines to obtain fusion features includes: adding the row reference features and the row query features to obtain the row fusion features; adding the column reference features and the column query features to obtain the column fusion features.

[0153] In this embodiment, by adding the reference features and the query features, fusion features can be obtained, and based on the fusion features, training a table structure recognition model can improve the accuracy of the table structure recognition model.

[0154] In some embodiments, the decoder includes: a row decoder and a column decoder; the candidate points include: the first candidate points on the row dividing lines and the second candidate points on the column dividing lines; the process of using the decoder to decode the fusion features to obtain the predicted structure information of the candidate points includes: using the row decoder to perform row decoding on the row fusion features to obtain the predicted structure information of the first candidate points, and the predicted structure information of the first candidate points includes: the ordinate of the first candidate points and the downward connection relationship; using the column decoder to perform column decoding on the column fusion features to obtain the predicted structure information of the second candidate points, and the predicted structure information of the second candidate points includes: the abscissa of the second candidate points and the rightward connection relationship.

[0155] In this embodiment, by decoding the fusion features, the predicted structure information of the corresponding candidate points can be accurately obtained.

[0156] Figure 11 It is a schematic diagram according to the fourth embodiment of the present disclosure. This embodiment provides a table structure recognition device, as Figure 11 shown, the device 1100 includes: an acquisition module 1101, a determination module 1102, a fusion module 1103, a decoding module 1104, and an identification module 1105.

[0157] The acquisition module 1101 is used to acquire the target image features of the table image; wherein, the table image contains a table; the determination module 1102 is used to determine the reference positions of the table dividing lines in the table based on the target image features; the fusion module 1103 is used to perform a fusion process on the reference features corresponding to the reference positions and the query features of the candidate points on the table dividing lines to obtain fusion features; the decoding module 1104 is used to perform a decoding process on the fusion features to obtain the structural information of the candidate points; the recognition module 1105 is used to determine the structural information of the cells in the table based on the structural information of the candidate points.

[0158] In this embodiment, table structure recognition is performed based on the target image features of the table image. Due to the target image features, the reference positions of the target table elements are determined based on the target image features. The reference position features corresponding to the reference positions and the query features of the candidate points are fused to obtain fusion features, and the structure recognition result is obtained based on the fusion features, which can improve the accuracy of table structure recognition. In addition, since the target image features can reflect the overall features of the table image, compared with the methods of segmenting first and then recognizing or detecting first and then recognizing, the accuracy and efficiency of table structure recognition can be improved. Therefore, the effect of table structure recognition can be improved.

[0159] In some embodiments, the acquisition module 1101 is further configured to: perform image feature extraction processing on the table image to obtain initial image features of multiple sizes; perform enhancement processing on the initial image features of multiple sizes to obtain the target image features.

[0160] In this embodiment, by acquiring the initial image features of multiple sizes and obtaining the target image features based on the initial image features of multiple sizes, the expression ability of the target image features can be improved, and further the table structure recognition effect can be improved.

[0161] In some embodiments, the table dividing lines include: row dividing lines and column dividing lines;

[0162] The determination module 1102 is further configured to: perform row regression processing on the target image features to determine the reference positions of the row dividing lines; perform column regression processing on the target image features to determine the reference positions of the column dividing lines.

[0163] In this embodiment, by performing regression processing on the target image features, accurate reference positions can be obtained, and further the accuracy of table structure recognition can be improved.

[0164] In some embodiments, the table dividing lines include: row dividing lines and column dividing lines; the reference features include: row reference features and column reference features; the query features include: row query features and column query features; the fusion features include: row fusion features and column fusion features; the fusion module 1103 is further configured to: perform an addition process on the row reference features and the row query features to obtain the row fusion features; perform an addition process on the column reference features and the column query features to obtain the column fusion features.

[0165] In this embodiment, by performing an addition process on the reference features and the query features, fusion features can be obtained, and based on the fusion features, a table structure recognition result can be obtained, which can improve the accuracy of table structure recognition.

[0166] In some embodiments, the candidate points include: first candidate points on the row dividing lines and second candidate points on the column dividing lines; the decoding module 1104 is further configured to: perform row decoding on the row fusion features to obtain the structural information of the first candidate points, where the structural information of the first candidate points includes: the ordinate of the first candidate points, and the downward connection relationship; perform column decoding on the column fusion features to obtain the structural information of the second candidate points, where the structural information of the second candidate points includes: the abscissa of the second candidate points, and the rightward connection relationship.

[0167] In this embodiment, by performing a decoding process on the fusion features, the structural information of the corresponding candidate points can be accurately obtained.

[0168] In some embodiments, the structural information of the cell includes: the physical coordinates of the cell and the logical coordinates of the cell; the recognition module 1105 is further configured to: determine a reference point based on the ordinate of the first candidate point and the abscissa of the second candidate point; determine a first target point and a termination column index based on the downward connection relationship of the first candidate point; determine a second target point and a termination row index based on the rightward connection relationship of the second candidate point; determine a third target point based on the first target point and the second target point; form the cell with the reference point, the first target point, the second target point, and the third target point, and determine the physical coordinates of the cell; determine the logical coordinates of the cell based on the termination row index and the termination column index.

[0169] In this embodiment, obtaining the structural information of the cell based on the structural information of the candidate points can improve the accuracy and efficiency of table structure recognition.

[0170] Figure 12 It is a schematic diagram according to the fifth embodiment of the present disclosure. This embodiment provides a training device for a table structure recognition model, where the model includes an encoder and a decoder, asFigure 12 As shown in the figure, the device 1200 includes: an encoding module 1201, a determination module 1202, a fusion module 1203, a decoding module 1204, a construction module 1205, and an adjustment module 1206.

[0171] The encoding module 1201 is used to obtain the target image features of the sample image by using the encoder; wherein, the sample image contains a table; the determination module 1202 is used to determine the predicted reference position of the table dividing line in the table based on the target image features; the fusion module 1203 is used to perform a fusion process on the reference features corresponding to the predicted reference position and the query features of the candidate points on the table dividing line to obtain fusion features; wherein, the candidate points are the vertices of the candidate cells corresponding to the table dividing line; the decoding module 1204 is used to perform a decoding process on the fusion features by using the decoder to obtain the predicted structure information of the candidate points; the construction module 1205 is used to construct a total loss function based on the predicted reference position and the predicted structure information; the adjustment module 1206 is used to adjust the model parameters of the encoder and the model parameters of the decoder by using the total loss function.

[0172] In this embodiment, by determining the reference position of the target table element based on the target image features, performing a fusion process on the reference position features corresponding to the reference position and the query features of the candidate points to obtain fusion features, and training the model based on the fusion features, the accuracy of the table structure recognition model can be improved; in addition, since the target image features can reflect the overall features of the table image, compared with the methods of segmenting first and then recognizing or detecting first and then recognizing, the accuracy and efficiency of the table structure recognition model can be improved. Therefore, the effect of the table structure recognition model can be improved.

[0173] In some embodiments, the encoder includes: an image feature extraction network and an image feature enhancement network; the encoding module 1201 is further used to: perform image feature extraction processing on the table image by using the image feature extraction network to obtain initial image features of multiple sizes; perform enhancement processing on the initial image features of multiple sizes by using the image feature enhancement network to obtain the target image features.

[0174] In this embodiment, by obtaining the initial image features of multiple sizes and obtaining the target image features based on the initial image features of multiple sizes, the expression ability of the target image features can be improved, and further the effect of the table structure recognition model can be improved.

[0175] In some embodiments, the table structure recognition model further includes: a row regression network and a column regression network; the table dividing lines include: row dividing lines and column dividing lines; the determining module 1202 is further configured to: use the row regression network to perform row regression processing on the target image features to determine the predicted reference positions of the row dividing lines; use the column regression network to perform column regression processing on the target image features to determine the predicted reference positions of the column dividing lines.

[0176] In this embodiment, by performing regression processing on the target image features, accurate reference positions can be obtained, thereby improving the accuracy of the table structure recognition model.

[0177] In some embodiments, the table dividing lines include: row dividing lines and column dividing lines; the reference features include: row reference features and column reference features; the query features include: row query features and column query features; the fusion features include: row fusion features and column fusion features; the fusion module 1203 is further configured to: perform an addition process on the row reference features and the row query features to obtain the row fusion features; perform an addition process on the column reference features and the column query features to obtain the column fusion features.

[0178] In this embodiment, by performing an addition process on the reference features and the query features, fusion features can be obtained, and training the table structure recognition model based on the fusion features can improve the accuracy of the table structure recognition model.

[0179] In some embodiments, the decoder includes: a row decoder and a column decoder; the candidate points include: the first candidate points on the row dividing lines and the second candidate points on the column dividing lines; the decoding module 1204 is further configured to: use the row decoder to perform row decoding processing on the row fusion features to obtain the predicted structure information of the first candidate points, and the predicted structure information of the first candidate points includes: the ordinate of the first candidate points, and the downward connection relationship; use the column decoder to perform column decoding processing on the column fusion features to obtain the predicted structure information of the second candidate points, and the predicted structure information of the second candidate points includes: the abscissa of the second candidate points, and the rightward connection relationship.

[0180] In this embodiment, by performing decoding processing on the fusion features, the predicted structure information of the corresponding candidate points can be accurately obtained.

[0181] In some embodiments, the construction module 1205 is further configured to: construct a first loss function based on the predicted reference positions and the true reference positions of the table reference lines; construct a second loss function based on the predicted structure information and the true structure information of the candidate points; construct the total loss function based on the first loss function and the second loss function.

[0182] In this embodiment, a total loss function is constructed based on the first loss function and the second loss function, which can refer to information in multiple dimensions and improve the model effect.

[0183] It can be understood that in the embodiments of the present disclosure, the same or similar content in different embodiments can be referred to each other.

[0184] It can be understood that in the embodiments of the present disclosure, "first", "second", etc. are only used for distinction and do not indicate the level of importance, the sequence of time, etc.

[0185] In the technical solution of the present disclosure, the processing of the collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0186] According to the embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0187] Figure 13 FIG. shows a schematic block diagram of an exemplary electronic device 1300 that can be used to implement the embodiments of the present disclosure. The electronic device 1300 is intended to represent various forms of digital computers, such as, for example, a laptop computer, a desktop computer, a workbench, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device 1300 can also represent various forms of mobile devices, such as, for example, a personal digital assistant, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0188] As Figure 13 shown, the electronic device 1300 includes a computing unit 1301, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1302 or a computer program loaded from a storage unit 1308 into a random access memory (RAM) 1303. In the RAM 1303, various programs and data required for the operation of the electronic device 1300 can also be stored. The computing unit 1301, the ROM 1302, and the RAM 1303 are connected to each other via a bus 1304. An input / output (I / O) interface 1305 is also connected to the bus 1304.

[0189] Multiple components in the electronic device 1300 are connected to the I / O interface 1305, including: an input unit 1306, such as a keyboard, a mouse, etc.; an output unit 1307, such as various types of displays, speakers, etc.; a storage unit 1308, such as a magnetic disk, an optical disc, etc.; and a communication unit 1309, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1309 allows the electronic device 1300 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0190] The computing unit 1301 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1301 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1301 executes the various methods and processes described above, such as the table structure recognition method or the training method of the table structure recognition model. For example, in some embodiments, the table structure recognition method or the training method of the table structure recognition model can be implemented as a computer software program, which is tangibly included in a machine-readable medium, such as the storage unit 1308. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 1300 via the ROM 1302 and / or the communication unit 1309. When the computer program is loaded into the RAM 1303 and executed by the computing unit 1301, one or more steps of the table structure recognition method or the training method of the table structure recognition model described above can be executed. Alternatively, in other embodiments, the computing unit 1301 can be configured to execute the table structure recognition method or the training method of the table structure recognition model by any other suitable means (e.g., by means of firmware).

[0191] The various embodiments of the systems and technologies described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs, which can be executed and / or interpreted on a programmable system including at least one programmable processor, and the programmable processor can be a special or general programmable processor, which can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0192] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program codes can be executed entirely on the machine, partially on the machine, executed partially on the machine as an independent software package and partially on a remote machine, or executed entirely on a remote machine or server.

[0193] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0194] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0195] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.

[0196] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services ("Virtual Private Server", or simply "VPS"). The server can also be a server of a distributed system, or a server combined with blockchain.

[0197] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitations are imposed herein.

[0198] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.

Claims

1. A method for recognizing a table structure, comprising: Obtaining target image features of a table image; wherein, the table image contains a table; Based on the target image features, determining reference positions of table dividing lines in the table; Performing fusion processing on the reference features corresponding to the reference positions and query features of candidate points on the table dividing lines to obtain fusion features; Performing decoding processing on the fusion features to obtain structural information of the candidate points; Based on the structural information of the candidate points, determining structural information of cells in the table; The table dividing lines include: row dividing lines and column dividing lines; The reference features include: row reference features and column reference features; The query features include: row query features and column query features; The fusion features include: row fusion features and column fusion features; The performing fusion processing on the reference features corresponding to the reference positions and query features of candidate points on the table dividing lines to obtain fusion features includes: Performing addition processing on the row reference features and the row query features to obtain the row fusion features; Performing addition processing on the column reference features and the column query features to obtain the column fusion features; The candidate points include: first candidate points on the row dividing lines and second candidate points on the column dividing lines; The performing decoding processing on the fusion features to obtain structural information of the candidate points includes: Performing row decoding processing on the row fusion features to obtain structural information of the first candidate points, and the structural information of the first candidate points includes: the ordinate of the first candidate points and downward connection relationships; Performing column decoding processing on the column fusion features to obtain structural information of the second candidate points, and the structural information of the second candidate points includes: the abscissa of the second candidate points and rightward connection relationships.

2. The method according to claim 1, wherein The obtaining target image features of a table image includes: Performing image feature extraction processing on the table image to obtain initial image features of multiple sizes; Performing enhancement processing on the initial image features of the multiple sizes to obtain the target image features.

3. According to the method of claim 1, wherein, The table dividing lines include: row dividing lines and column dividing lines; The based on the target image features, determining reference positions of table dividing lines in the table includes: Performing row regression processing on the target image features to determine reference positions of the row dividing lines; Performing column regression processing on the target image features to determine reference positions of the column dividing lines.

4. According to the method of claim 1, wherein, The structural information of the cells includes: physical coordinates of the cells and logical coordinates of the cells; The based on the structural information of the candidate points, determining structural information of cells in the table includes: Based on the ordinate of the first candidate points and the abscissa of the second candidate points, determining a reference point; Based on the downward connection relationships of the first candidate points, determining first target points and a termination column index; Based on the rightward connection relationships of the second candidate points, determining second target points and a termination row index; Determine a third target point based on the first target point and the second target point; Form the cell with the reference point, the first target point, the second target point, and the third target point, and determine the physical coordinates of the cell; Determine the logical coordinates of the cell based on the end row index and the end column index.

5. A training method for a table structure recognition model, the table structure recognition model comprising: An encoder and a decoder, the method comprising: Using the encoder to obtain target image features of a sample image; wherein, the sample image contains a table; Determine a predicted reference position of a table dividing line in the table based on the target image features; Perform a fusion process on the reference features corresponding to the predicted reference position and the query features of candidate points on the table dividing line to obtain fusion features; Using the decoder to perform a decoding process on the fusion features to obtain predicted structural information of the candidate points; Construct a total loss function based on the predicted reference position and the predicted structural information; Using the total loss function to adjust the model parameters of the encoder and the model parameters of the decoder; The table dividing line includes: a row dividing line and a column dividing line; The reference features include: row reference features and column reference features; The query features include: row query features and column query features; The fusion features include: row fusion features and column fusion features; The performing a fusion process on the reference features corresponding to the reference position and the query features of candidate points on the table dividing line to obtain fusion features includes: Perform an addition process on the row reference features and the row query features to obtain the row fusion features; Perform an addition process on the column reference features and the column query features to obtain the column fusion features; The decoder includes: a row decoder and a column decoder; The candidate points include: first candidate points on the row dividing line and second candidate points on the column dividing line; The using the decoder to perform a decoding process on the fusion features to obtain predicted structural information of the candidate points includes: Using the row decoder to perform a row decoding process on the row fusion features to obtain predicted structural information of the first candidate points, and the predicted structural information of the first candidate points includes: the ordinate of the first candidate points, and the downward connection relationship; Using the column decoder to perform a column decoding process on the column fusion features to obtain predicted structural information of the second candidate points, and the predicted structural information of the second candidate points includes: the abscissa of the second candidate points, and the rightward connection relationship.

6. The method according to claim 5, wherein, The encoder includes: an image feature extraction network and an image feature enhancement network; The using the encoder to obtain target image features of a sample image includes: Using the image feature extraction network to perform an image feature extraction process on the table image to obtain initial image features of multiple sizes; Using the image feature enhancement network to perform an enhancement process on the initial image features of multiple sizes to obtain the target image features.

7. The method according to claim 5, wherein, The table structure recognition model further includes: a row regression network and a column regression network; The table dividing lines include: row dividing lines and column dividing lines; Determining the predicted reference positions of the table dividing lines in the table based on the target image features includes: Using the row regression network to perform row regression processing on the target image features to determine the predicted reference positions of the row dividing lines; Using the column regression network to perform column regression processing on the target image features to determine the predicted reference positions of the column dividing lines.

8. The method according to any one of claims 5-7, wherein Constructing the total loss function based on the predicted reference positions and the predicted structure information includes: Constructing a first loss function based on the predicted reference positions and the true reference positions of the table reference lines; Constructing a second loss function based on the predicted structure information and the true structure information of the candidate points; Constructing the total loss function based on the first loss function and the second loss function.

9. A table structure recognition device, including: An acquisition module, configured to acquire the target image features of a table image; wherein, the table image contains a table; A determination module, configured to determine the reference positions of the table dividing lines in the table based on the target image features; A fusion module, configured to perform fusion processing on the reference features corresponding to the reference positions and the query features of the candidate points on the table dividing lines to obtain fusion features; A decoding module, configured to perform decoding processing on the fusion features to obtain the structure information of the candidate points; An identification module, configured to determine the structure information of the cells in the table based on the structure information of the candidate points; The table dividing lines include: row dividing lines and column dividing lines; The reference features include: row reference features and column reference features; The query features include: row query features and column query features; The fusion features include: row fusion features and column fusion features; The fusion module is further configured to: Perform an addition process on the row reference features and the row query features to obtain the row fusion features; Perform an addition process on the column reference features and the column query features to obtain the column fusion features; The candidate points include: first candidate points on the row dividing lines and second candidate points on the column dividing lines; The decoding module is further configured to: Perform row decoding processing on the row fusion features to obtain the structure information of the first candidate points, and the structure information of the first candidate points includes: the vertical coordinates of the first candidate points, and the downward connection relationship; Perform column decoding processing on the column fusion features to obtain the structure information of the second candidate points, and the structure information of the second candidate points includes: the horizontal coordinates of the second candidate points, and the rightward connection relationship.

10. The apparatus according to claim 9, wherein, The acquisition module is further configured to: Perform image feature extraction processing on the table image to obtain initial image features of multiple sizes; Perform enhancement processing on the initial image features of the multiple sizes to obtain the target image features.

11. According to the device of claim 9, wherein, The table dividing lines include: row dividing lines and column dividing lines; The determination module is further configured to: Perform row regression processing on the target image features to determine the reference position of the row dividing line; Perform column regression processing on the target image features to determine the reference position of the column dividing line.

12. The apparatus according to claim 9, wherein, The structural information of the cell includes: the physical coordinates of the cell and the logical coordinates of the cell; The recognition module is further configured to: Determine a reference point based on the ordinate of the first candidate point and the abscissa of the second candidate point; Determine a first target point and a termination column index based on the downward connection relationship of the first candidate point; Determine a second target point and a termination row index based on the rightward connection relationship of the second candidate point; Determine a third target point based on the first target point and the second target point; Form the cell by combining the reference point, the first target point, the second target point, and the third target point, and determine the physical coordinates of the cell; Determine the logical coordinates of the cell based on the termination row index and the termination column index.

13. A training device for a table structure recognition model, the table structure recognition model comprising: An encoder and a decoder, the apparatus includes: An encoding module, configured to use the encoder to obtain the target image features of the sample image; wherein, the sample image contains a table; A determination module, configured to determine a predicted reference position of the table dividing line in the table based on the target image features; A fusion module, configured to perform fusion processing on the reference features corresponding to the predicted reference position and the query features of the candidate points on the table dividing line to obtain fusion features; wherein, the candidate points are the vertices of the candidate cells corresponding to the table dividing line; A decoding module, configured to use the decoder to perform decoding processing on the fusion features to obtain the predicted structural information of the candidate points; A construction module, configured to construct a total loss function based on the predicted reference position and the predicted structural information; An adjustment module, configured to use the total loss function to adjust the model parameters of the encoder and the model parameters of the decoder; The table dividing line includes: a row dividing line and a column dividing line; The reference features include: row reference features and column reference features; The query features include: row query features and column query features; The fusion features include: row fusion features and column fusion features; The fusion module is further configured to: Perform an addition process on the row reference features and the row query features to obtain the row fusion features; Perform an addition process on the column reference features and the column query features to obtain the column fusion features; The decoder includes: a row decoder and a column decoder; The candidate points include: a first candidate point on the row dividing line and a second candidate point on the column dividing line; The decoding module is further configured to: Use the row decoder to perform row decoding processing on the row fusion features to obtain the predicted structural information of the first candidate point, and the predicted structural information of the first candidate point includes: the ordinate of the first candidate point, and the downward connection relationship; Using the column decoder, perform column decoding processing on the column fusion feature to obtain the predicted structure information of the second candidate point, where the predicted structure information of the second candidate point includes: the abscissa of the second candidate point, and the right connection relationship.

14. The apparatus according to claim 13, wherein the encoder includes: an image feature extraction network and an image feature enhancement network; the encoding module is further configured to: Use the image feature extraction network to perform image feature extraction processing on the table image to obtain initial image features of multiple sizes; Use the image feature enhancement network to perform enhancement processing on the initial image features of multiple sizes to obtain the target image features.

15. The apparatus according to claim 13, wherein the table structure recognition model further includes: a row regression network and a column regression network; the table dividing line includes: a row dividing line and a column dividing line; the determining module is further configured to: Use the row regression network to perform row regression processing on the target image features to determine the predicted reference position of the row dividing line; Use the column regression network to perform column regression processing on the target image features to determine the predicted reference position of the column dividing line.

16. The device according to any one of claims 13-15, wherein, The constructing module is further configured to: Construct a first loss function based on the predicted reference position and the true reference position of the table reference line; Construct a second loss function based on the predicted structure information and the true structure information of the candidate point; Construct the total loss function based on the first loss function and the second loss function.

17. An electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 1-8.

18. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-8.

19. A computer program product, comprising a computer program, where the computer program implements the method according to any one of claims 1-8 when executed by a processor.

Citation Information

Patent Citations

  • Table image recognition method and device, computer equipment and storage medium

    CN113283355A

  • Table detection method and device based on semantic segmentation, equipment and medium

    CN113449620A