Table identification method
By converting tables into image formats and combining semantic context analysis and multimodal hybrid networks, the problems of low accuracy of complex table recognition and high consumption of computing resources in the prior art are solved, and efficient and flexible table recognition and visual analysis are achieved.
Patent Information
- Application Number
- CN202510441230.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-07-22
AI Technical Summary
When handling complex table structures, existing table recognition technology has low recognition accuracy, high computing resources consumption, poor adaptability, and difficult to meet the real-time and flexibility needs of finance, medical and other fields.
By converting tables into image formats, combining semantic context analysis, cross-cell entity recognition and dependency syntax analysis, a multi-modal hybrid network is built, visual and text features are integrated, and a lightweight network architecture is adopted to reduce computational costs and generate visual table semantic feature images.
It significantly improves the recognition accuracy of complex tables, improves adaptability and flexibility to different fields, reduces calculation costs, and meets real-time processing needs.
Smart Images

Figure CN120356233A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for table recognition. Background Art
[0002] With the popularization of digital office, the automatic recognition and analysis of table data have become crucial in fields such as finance, healthcare, and education. However, existing table recognition technologies have many limitations when dealing with complex tables, as follows:
[0003] 1. Limitations of traditional OCR technology: Although optical character recognition (OCR) technology can effectively extract text content in tables, its recognition accuracy drops significantly when facing complex table structures (such as merged cells and slanted headers). OCR technology usually treats tables as plain text, ignoring the two-dimensional structure information of tables, resulting in poor performance in parsing table logical relationships.
[0004] 2. Deficiencies of the rule template method: The table recognition method based on rule templates relies on predefined table structure patterns. However, the table structures in actual application scenarios vary greatly, and the rule template method is difficult to adapt to diverse table formats. Moreover, when facing new types of tables, it is necessary to frequently adjust the rules manually, lacking flexibility and scalability.
[0005] 3. Limitations of deep learning methods: Existing deep learning table recognition methods mainly focus on single-modal feature extraction, such as only relying on visual features or text features. However, table data has multi-modal characteristics, and single-modal features are difficult to comprehensively capture the semantic and structural information of tables. In addition, existing methods lack an effective dynamic weighting mechanism when dealing with cross-cell entity associations and table logical structure parsing, resulting in insufficient parsing ability for complex tables.
[0006] 4. Computational cost and efficiency issues: Some high-precision deep learning table recognition models require huge computational resources for training and inference, which limits their applications in resource-constrained environments (such as mobile devices and edge computing scenarios). At the same time, existing methods are computationally inefficient when dealing with large-scale table data and are difficult to meet the requirements of application scenarios with high real-time requirements.
[0007] In summary, existing table recognition technologies have obvious deficiencies in aspects such as complex table structure parsing, multi-modal feature fusion, domain adaptability, and computational efficiency. Therefore, the present invention aims to propose a novel table recognition method, which can improve the accuracy, adaptability, and efficiency of table recognition by fusing multi-modal features, dynamically weighted neighborhood information, and deep learning technologies, so as to meet the requirements of efficient recognition and analysis of complex table data in practical applications. Summary of the Invention
[0008] The objective of the present invention is to provide a novel table recognition method, which converts table data into images and utilizes multi-adjacent relationship integration and deep learning techniques to improve the recognition accuracy and adaptability while reducing the computational cost.
[0009] The technical solutions adopted by the present invention are as follows:
[0010] A table recognition method, comprising the following steps:
[0011] 1) Convert the table into an image format, with each cell corresponding to a pixel block;
[0012] 2) Conduct semantic context analysis, extract the semantic context vectors of the cell entity content, and dynamically calculate the context weight scores of each cell and its neighboring cells;
[0013] 3) Based on the entity linking function, identify the entity content scattered in multiple cells, and form an entity link vector across cells through weighted integration;
[0014] 4) Combine the cell type features, location features, and semantic context relationship features, calculate the dependency syntax scores, and construct the logical structure relationship of the cells;
[0015] 5) Normalize the analysis results of steps 2)-4) and map them to the color space to generate a color or grayscale image reflecting the semantic features of the table;
[0016] 6) Construct a multi-modal hybrid network integrating a comprehensive vision branch, a text branch, and cross-attention, and train it using a multi-task loss function combining semantic classification loss and structural similarity loss;
[0017] 7) Use the trained multi-modal hybrid network to perform semantic classification and structural analysis on the table image, and generate the semantic labels and structural segmentation masks of the table;
[0018] 8) Locate the table header and content areas through image segmentation technology, and optimize the structural analysis results in combination with semantic features;
[0019] 9) Post-process the analysis results, including morphological operations and connected component analysis, to improve the recognition accuracy of the table structure;
[0020] 10) Output the semantic classification results, structural segmentation results, and visualization images of the table.
[0021] Furthermore, step 2) is specifically as follows:
[0022] Obtain the cell content word vector V through BERT semantic encoding i,j , and the formula is:
[0023]
[0024] where: w i,j is the cell text content in the i-th row and j-th column of the table;
[0025] The text content w of each cell in the table i,j After preprocessing including lowercasing and adding special markers, it is input into the BERT model. The BERT model outputs the mean vector of the last hidden layer, generating a 768-dimensional word vector V i,j ; For the word vector V i,j Perform L2 normalization to obtain the normalized word vector
[0026]
[0027] Furthermore, in the step 2), the context weight score αk between the cell and its neighboring cells is calculated through a multi-layer perceptron MLP, and the formula is:
[0028] αk = softmax(MLP([V i,j ; V k ))
[0029] where, [Vi,j; Vk] represents the concatenated vector of the word vector V i,j of the central cell and the word vector V k of the neighboring cell;
[0030] After calculating for the word vector of each cell, start calculating the context vector synthesis value, V k is the k-th cell, and the context vector synthesis formula is as follows:
[0031]
[0032] C i,j is the context vector synthesis value of the central cell (i,j);
[0033] V i,j is the word vector of the central cell (i,j), obtained from the BERT model;
[0034] α i,j is the self-attention weight of the central cell, output by the MLP, and calculated through the corresponding formula;
[0035] N(i,j) is the neighborhood range of the central cell (i,j);
[0036] αk is the weight score of the neighboring cell k, calculated from the dynamic attention weight.
[0037] sk is the cosine similarity score of the neighborhood cell k, and the calculation formula is: sk = cos_sim(V i,j , Vk)
[0038] sigmoid(sk): Compress the similarity score to (0, 1);
[0039] V k is the word vector of the neighborhood cell k, which is obtained by the BERT model.
[0040] Furthermore, the specific content of step 3) is: Calculate the probability distribution yk of the entity category through the context vector synthesis value Ci,j and the entity category projection matrix Went, and the formula is:
[0041] P(y k |C i,j ) = softmax(W ent C i,j + b ent )
[0042] bent: Trainable bias vector;
[0043] Went ∈ R 768×K : Entity category projection matrix;
[0044] The entity link strength calculation formula is:
[0045]
[0046] ε i,j,k represents the entity link strength between the cell (i, j) and the cell k;
[0047] represents the indicator function, which is 1 when the entity type of the cell k is the same as that of the cell (i, j), otherwise it is 0;
[0048] cos_sim(C i,j , C k ) represents calculating the cosine similarity between the context vectors of the cell (i, j) and the cell k.
[0049] Furthermore, in step 4), the formula for the dependency syntax score is:
[0050]
[0051] Where: represents that the cell k belongs to the neighborhood range of the cell (i, j), that is, consider the cells adjacent or related to the cell (i, j) in the table structure;
[0052] φpos(i,k): Represents the position feature function, which is used to measure the positional relationship between cell (i,j) and cell k;
[0053] φnum(w k ): Represents the numerical feature function, which is used to measure whether the content of cell k is of numerical type.
[0054] Furthermore, in step 6), the multi-task loss function is:
[0055] Ltotal = 0.6Lsem + 0.4Lstruct
[0056] where Lsem is the semantic classification loss and Lstruct is the structural similarity loss.
[0057] Furthermore, in step 7), the improved instance segmentation algorithm is based on Mask R-CNN as the basic framework for improvement, and the improvement points are:
[0058] (1) Replace the original ResNet-101 with the lightweight EfficientNet-B4;
[0059] (2) Introduce deformable convolution in the original ROI alignment module.
[0060] The beneficial effects of the present invention are:
[0061] 1. By combining semantic context analysis, cross-cell entity recognition, and dependency syntactic analysis, the present invention can effectively analyze the logical relevance of complex tables. Different from traditional OCR technologies, this method not only focuses on the text content but also deeply analyzes the two-dimensional structure information of the table, thus performing excellently in dealing with complex structures such as merged cells and diagonal headers.
[0062] 2. Aiming at the limitation of single-modal feature extraction in existing deep learning methods, the present invention innovatively integrates visual and text features. By constructing a multi-modal hybrid network and introducing a cross-attention mechanism, deep interaction between visual and text features is achieved, comprehensively capturing the semantic and structural information of the table, and significantly improving the recognition accuracy of complex tables.
[0063] 3. The present invention designs customized models for different fields. By integrating NLP (BERT) and computer vision (ResNet, Mask R-CNN) technologies, it can quickly adapt to specific table formats in fields such as finance and healthcare. This domain adaptability avoids the cumbersome process of frequently manually adjusting rules, greatly improving the flexibility and scalability of table recognition.
[0064] 4. The present invention maps semantic features such as context vectors and dependency syntactic scores into a color space to generate a visual image reflecting the content and structure of the table. This visualization method enhances the interpretability of semantic information, facilitates users to intuitively understand table data, and provides strong support for subsequent data analysis and decision-making. Description of the Drawings
[0065] Figure 1 To generate a grayscale image reflecting the semantic features of the table. Detailed Implementation Manner
[0066] The present invention will be further described below with reference to the accompanying drawings.
[0067] A table recognition method includes the following steps:
[0068] 1) Convert the table into an image format, with each cell corresponding to a pixel block;
[0069] 2) Perform semantic context analysis, extract the semantic context vectors of the cell entity content, and dynamically calculate the context weight scores of each cell and its neighboring cells;
[0070] 3) Based on the entity linking function, identify the entity content scattered in multiple cells, and form a cross-cell entity link vector through weighted integration;
[0071] 4) Combine the cell type features, position features, and semantic context relationship features, calculate the dependency syntactic scores, and construct the logical structure relationship of the cells;
[0072] 5) Normalize the analysis results in steps 2)-4) and map them to the color space to generate a color or grayscale image reflecting the semantic features of the table;
[0073] 6) Construct a multi-modal hybrid network integrating a visual branch, a text branch, and cross-attention, and train it using a multi-task loss function combining semantic classification loss and structural similarity loss;
[0074] 7) Use the trained multi-modal hybrid network to perform semantic classification and structural analysis on the table image, and generate the semantic labels and structural segmentation masks of the table;
[0075] 8) Locate the table header and content areas through image segmentation technology, and optimize the structural analysis results in combination with semantic features;
[0076] 9) Post-process the analysis results, including morphological operations and connected component analysis, to improve the recognition accuracy of the table structure;
[0077] 10) Output the semantic classification results, structural segmentation results, and visualization images of the table.
[0078] The technical solution of the present invention will be further described below through a certain business data table.
[0079]
[0080]
[0081] Table Image Conversion
[0082] Convert the table into an image format, with each cell corresponding to a pixel block. The image is composed of pixel blocks, and each element of the table is regarded as a pixel block.
[0083] Semantic Context Analysis of Key Value (K-V) Relationships
[0084] The key value (K-V) relationship obtains specific domain semantic analysis through the comprehensive calculation of three values: semantic context analysis, cross-cell entity recognition (entity link vector), and table-specific dependency syntactic analysis.
[0085] Since the expression of the same cell is different in different calculation steps, by using T i,j to represent the original cell position, and w i,j is the cell text content of cell T i,j in the i-th row and j-th column. In each calculation step, the definition of the calculation formula is used, and the association with the original cell definition is marked at the same time. The neighborhood range is a value that can be freely set. In the present invention, the 4-neighborhood relationship is used as an illustrative example. In actual applications, 8-neighborhood or more can be used.
[0086] 2.1 Semantic Context Analysis
[0087] 2.1.1) Obtain the cell content word vector V through BERT semantic encoding ij , and the formula is:
[0088]
[0089] Where: w i,j is the cell text content in the i-th row and j-th column of the table;
[0090] The present invention adopts the BERT model, and the input text is processed by the BERT model as follows:
[0091] Preprocessing of the input text w i,j : Lowercasing, adding special tokens ([CLS] and [SEP])
[0092] Output: Take the mean vector of the last hidden layer
[0093] R 768: The embedding code has 768 dimensions and is subjected to L2 normalization.
[0094] The text content w of each cell in the table i,j After preprocessing including lowercasing and adding special tokens, it is input into the BERT model. The BERT model outputs the mean vector of the last hidden state to generate a 768-dimensional word vector V. i,j ; For the word vector V i,j Perform L2 normalization to obtain the normalized word vector.
[0095]
[0096] 2.1.2) Dynamic attention weight calculation
[0097] αk = softmax(MLP([V i,j ; V k ))
[0098] Calculate the context weight score αk of the cell and its neighboring cells through the multi-layer perceptron MLP. Among them, [Vi,j; Vk] represents the concatenated vector of the word vector V of the central cell i,j and the word vector V of the neighboring cell k .
[0099] The multi-layer neurons of MLP perform non-linear transformation on the data to achieve classification or regression tasks.
[0100] softmax: Normalize all cells (including itself) in the neighborhood.
[0101] 2.1.3) Context vector synthesis
[0102] After calculating the word vector of each cell, start calculating the context vector synthesis value. Vk is the kth cell. In actual calculation, the cell subscript needs to be brought in to indicate which specific cell it is.
[0103]
[0104] C i,j is the context vector synthesis value of the central cell (i,j);
[0105] V i,j is the word vector of the central cell (i,j), obtained from the BERT model;
[0106] α i,j is the self-attention weight of the central cell, output by the MLP and calculated through the corresponding formula;
[0107] N(i,j) is the neighborhood range of the central cell (i,j);
[0108] αk is the weight score of the neighborhood cell k, which is calculated from the dynamic attention weight.
[0109] sk is the cosine similarity score of the neighborhood cell k, and the calculation formula is: sk = cos_sim(V i,j , Vk)
[0110] sigmoid(sk): Compresses the similarity score to (0,1);
[0111] V k is the word vector of the neighborhood cell k, which is obtained from the BERT model.
[0112] Example calculation
[0113] (a) Example of cosine similarity calculation:
[0114] sk = cos_sim(Vi,j, Vk): Calculates the cosine similarity score between this cell and other cells:
[0115]
[0116] (b) Example of similarity compression:
[0117]
[0118] (c) Context vector synthesis
[0119] Assume the central cell is T 2,3 , and its neighborhood cells include T 1,3 , T 2,2 , T 3,3 , T 2,4 , then:
[0120] α = [0.15 (itself), 0.18 (T1,3), 0.24 (T2,2), 0.35 (T3,3), 0.08 (T2,4)]
[0121] The contribution calculation of the neighbor cell T 1,3 = "2021" is as follows:
[0122] 0.18 × 0.70 × [0.3, 0.5]
[0123] Integrate the contributions of all neighborhood cells to calculate the context vector C 2,3 :
[0124] C2,3
[0125] = 0.15 × [0.2, 0.6] + 0.18 × 0.70 × [0.3, 0.5] + 0.24 × 0.66 × [0.8, 0.1] + 0.35 × 0.71 × [0.25, 0.55] + 0.08 × 0.60 × [0.4, 0.3]
[0126] = [0.03, 0.09] + [0.0378, 0.063] + [0.1267, 0.0158] + [0.0621, 0.1369] + [0.0192, 0.0144]
[0127] = [0.2758, 0.3191]
[0128] 2.2 Cross-cell entity recognition module
[0129] 2.2.1) Entity type annotation
[0130] Calculate the probability distribution yk of the entity category through the context vector synthesis value Ci,j and the entity category projection matrix Went. The formula is:
[0131] P(y k |C i,j ) = softmax(W ent C i,j + b ent )
[0132] Ci,j: 768-dimensional context vector from the semantic context module.
[0133] Went ∈ R 768×K : Entity category projection matrix (K is the number of categories). Map the context vector to the entity category space through the entity category projection matrix Went.
[0134] bent: Trainable bias vector.
[0135] yk: Probability distribution of the k-th entity category.
[0136] Use the softmax function to calculate the probability distribution of each entity category.
[0137] Example calculation
[0138] (a) Context vector:
[0139] C 2,3 = [0.2758, 0.3191]
[0140] (b) Entity category projection matrix:
[0141]
[0142] (c) Bias vector:
[0143]
[0144] (d) Process value calculation:
[0145] Process value = [0.2758 * 0.1 + 0.3191 * 0.3 +...,...] = [1.2, -0.5,...]
[0146] (e) Probability distribution calculation:
[0147] P(y|C i,j ) = softmax(process value)
[0148] Calculated as:
[0149] P(y|C i,j ) = 0.92.
[0150] 2.2.2) Entity link strength
[0151] The formula for calculating the entity link strength is:
[0152]
[0153] ε i,j,k represents the entity link strength between cell (i, j) and cell k;
[0154] represents the indicator function, which is 1 when the entity type of cell k is the same as that of cell (i, j), otherwise 0;
[0155] cos_sim(C i,j , C k ) represents the cosine similarity between the context vectors of cell (i, j) and cell k.
[0156] Similarity threshold: τ = 0.6 (only keep the links where E > τ)
[0157] Example calculation
[0158] (a) Content of neighboring cells:
[0159] The content of neighboring cell T 3,3 is "20,000", which is of the same numerical type as the central cell.
[0160] (b) Cosine similarity calculation:
[0161] cos_sim([0.2758, 0.3191], [0.25, 0.55]) = 0.89
[0162] (c) Entity Link Strength Calculation:
[0163] Use the indicator function I(·) to check if the entity types are the same, where I(·) = 1.
[0164] Calculate the entity link strength E 2,3,3,3 :
[0165] E 2,3,3,3 = I(·) × 0.89 = 1 × 0.89 = 0.89
[0166] Subscript Explanation
[0167] Subscripts in the formula: In the formula, the subscript of E i,j,k is 3, indicating the link strength between the cell in the i-th row, j-th column and the k-th cell.
[0168] Subscripts in actual calculation: In actual calculation, the subscript is 4 digits. For example, E 2,3,3,3 , indicating the link strength between the cell in the 2nd row, 3rd column and the cell in the 3rd row, 3rd column.
[0169] Unity: The formula and the actual calculation formula are essentially unified. The k in the formula represents the k-th cell, and specific coordinates need to be substituted in actual calculation.
[0170] 2.3 Table Dependency Syntactic Analysis Module
[0171] 2.3.1) Structure Feature Definition (see Table 1)
[0172]
[0173] Table 1
[0174] Example Calculation (cell T 2,3 )
[0175] (a) Position Decay Factor:
[0176] Δi = 2 - 1 = 1,
[0177] Δj = 3 - 1 = 2
[0178]
[0179] (b) Numerical Correlation Degree:
[0180] The number of numeric characters is 6, and the total number of characters is 7
[0181]
[0182] (c) Hierarchical Depth:
[0183] No nesting, the nesting level is 0
[0184] φ depth = log(1 + 0) = 0
[0185] 2.3.2) Dependency Score Calculation
[0186]
[0187] Calculation Conditions
[0188] Direction Sensitivity: Only consider neighbor cells in the same column or row.
[0189] Header Suppression: Automatically suppress the influence of header cells through the position attenuation factor.
[0190] Wherein: Indicates that cell k belongs to the neighborhood range of cell (i, j), that is, consider cells adjacent or related to cell (i, j) in the table structure;
[0191] φpos(i,k): Represents the position feature function, used to measure the position relationship between cell (i, j) and cell k;
[0192] φnum(w k ): Represents the numerical feature function, used to measure whether the content of cell k is of numerical type.
[0193] Example Calculation
[0194] (1) Contribution of Neighbor Cell T 3,3 :
[0195] Position Attenuation Factor φ pos = 0.33
[0196] Numerical Association Degree φ num = 1.71
[0197] Contribution Value: 0.33 × 1.71 = 0.56
[0198] (b) Dependency Score Calculation:
[0199] D 2,3 = 0.56 +... (Contributions of Other Neighbors)
[0200] 2.4. Color Mapping and Image Generation
[0201] 2.4.1) Channel Mapping Rules (see Table 2)
[0202]
[0203] Table 2
[0204] Example Calculation:
[0205] (a) Semantic context C 2,3 :
[0206] Calculate the L2 norm:
[0207] ||C 2,3 ||2 = 0.2758 2 +0.3191 2 = 0.42
[0208] Normalize and map to the R channel:
[0209]
[0210] (b) Entity linking ε:
[0211] Calculate the Sigmoid value:
[0212]
[0213] Map to the G channel:
[0214] G = 255 × 0.70 ≈ 181
[0215] (c) Dependency parsing D 2,3 :
[0216] Calculate the normalized value:
[0217]
[0218] Final RGB value:
[0219] RGB = (113, 181, 70)
[0220] 5. Deep learning model training
[0221] 1. Model architecture design
[0222] Adopt a multi-modal hybrid network, combining visual and text features:
[0223] Visual branch: Extract local image features (such as cell boundaries, color distributions) based on the improved ResNet-50.
[0224] Text branch: Use the pre-trained BERT model (fine-tuned) to generate semantic vectors of cell text.
[0225] Feature fusion module: Fuse visual and text features through the cross-attention mechanism to output joint feature vectors.
[0226] 2. Training dataset construction
[0227] Data source:
[0228] Public dataset: ICDAR 2019 Table Recognition Competition dataset (including complex table structure annotations).
[0229] Self-built domain dataset: Financial annual report tables (annotated fields include "Revenue", "Profit", etc.), medical report tables (annotated fields such as "Patient ID", "Diagnosis").
[0230] Data augmentation: Rotate the table images (±5°), add noise (Gaussian noise σ = 0.1), and perform color perturbation (±10% in the HSV space) to improve the robustness of the model.
[0231] 3. Loss function design
[0232] The model is trained using a multi-task loss function:
[0233] Semantic classification loss (Lsem): Cross-entropy loss, used to predict cell semantic labels (such as "numeric", "text", "header").
[0234] Structural similarity loss (Lstruct): Based on the Dice coefficient, optimize the overlap between the table region segmentation result and the ground truth annotation.
[0235] Joint loss: Ltotal = 0.6Lsem + 0.4Lstruct, balance the task priorities through hyperparameters.
[0236] 4. Training strategy
[0237] Transfer learning: Pre-train ResNet-50 on ImageNet, and pre-train the BERT model on WikiText-103.
[0238] Two-stage fine-tuning:
[0239] Freeze the visual branch, only train the text branch and the fusion module (learning rate 1e-4, Adam optimizer).
[0240] Jointly train all network parameters (learning rate 5e-5, introduce gradient clipping to prevent overfitting).
[0241] 5. Performance evaluation
[0242] Evaluation metrics:
[0243] Semantic classification accuracy (Accuracy), structural segmentation mIoU (Mean Intersection over Union).
[0244] End-to-end table recognition F1-score (combining text extraction and structure restoration).
[0245] Comparative experiment:
[0246] This method achieves an F1-score of 92.3% on the ICDAR test set, a 13.8% improvement over the traditional OCR + rule method (78.5%).
[0247] Image segmentation and structure recognition:
[0248] In this step, the table header and content areas are located through image segmentation technology, and semantic features are combined to optimize structure parsing. The specific process is as follows:
[0249] 1. Selection of image segmentation algorithm
[0250] Use Mask R-CNN as the basic framework and make the following improvements:
[0251] Feature extractor: Replace ResNet-101 with lightweight EfficientNet-B4 to improve the inference speed.
[0252] ROI alignment module: Replace the ordinary convolution in ROI Align to enhance the adaptability to irregular table boundaries.
[0253] 2. Segmentation process
[0254] Input preprocessing: Concatenate the semantic mapping image generated in step 3 (e.g., pixel value 205 corresponds to the "Revenue" field) with the original table image to form a dual-channel input.
[0255] Region proposal generation: Generate candidate regions through RPN (Region Proposal Network) and filter candidate boxes with a score higher than 0.7.
[0256] In the generated table image, the color of each cell combines three data dimensions:
[0257] Semantic intensity (red): Reflects the semantic density of the cell content.
[0258] Entity association (green): Reflects the entity association strength between the cell and other cells.
[0259] Structure dependence (blue): Reflects the structural features and dependence relationships of the cell.
[0260] The color mapping process is reversible, meaning that the original semantic feature data can be recovered from the generated color image. This reversibility ensures data integrity and recoverability.
[0261] Example illustration (e.g. Figure 1)
[0262] Dark red cells: indicate that the text at this position has a high semantic density (e.g., the keyword "profit").
[0263] Bright green cells: indicate that the value has a strong entity association with other cells (e.g., cross-sheet linkage calculation).
[0264] Blue-tone cells: indicate that the cell has significant structural features (e.g., table headers or nested cells).
[0265] 3. Structure determination rules
[0266] Semantic vector clustering: Extract BERT semantic vectors from the segmented regions and use the DBSCAN clustering algorithm:
[0267] Table header region: The clustering center corresponds to texts with high semantic weights such as "Financial Data" and "2021".
[0268] Content region: The clustering center is numerical fields (e.g., "100,000", "20%").
[0269] Location feature assistance: If the table header is in the first row or the first column, the segmentation result is preferentially corrected according to the location rule.
[0270] 4. Post-processing optimization
[0271] Morphological operation: Perform a closing operation (kernel size 3×3) on the segmentation mask to fill small holes and broken edges.
[0272] Connected component analysis: Merge isolated regions with an area less than 10 pixels.
[0273] 5. Performance verification
[0274] Test results:
[0275] Table header segmentation accuracy (IoU): 89.7%, an improvement of 17.3% compared to the traditional threshold segmentation method (72.4%).
[0276] Structure restoration accuracy: 95.2% for financial tables, and it drops to 87.6% for medical tables due to complex merged cells.
[0277] Analysis of failure cases: For merged cells (e.g., "Company A" across rows), it is necessary to further introduce a graph neural network (GNN) to optimize the topological relationship parsing.
[0278] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, several improvements can be made without departing from the principle of the present invention, and these improvements should also be regarded as the protection scope of the present invention.
Claims
1. A table recognition method, characterized in that: It includes the following steps: 1) Convert the table into an image format, with each cell corresponding to a pixel block; 2) Perform semantic context analysis, extract the semantic context vector of the cell entity content, and dynamically calculate the context weight score between each cell and its neighboring cells; 3) Based on the entity link function, identify the entity content scattered in multiple cells, and form an entity link vector across cells through weighted integration; 4) Combine the cell type features, location features, and semantic context relationship features to calculate the dependency syntax score and construct the logical structure relationship of the cells; 5) Normalize the analysis results of steps 2)-4) and map them to the color space to generate a color or grayscale image reflecting the semantic features of the table; 6) Construct a multi-modal hybrid network integrating the visual branch, text branch, and cross-attention, and train it using a multi-task loss function combining semantic classification loss and structural similarity loss; 7) Use the trained multi-modal hybrid network to perform semantic classification and structure parsing on the table image, and generate the semantic label and structure segmentation mask of the table; 8) Locate the table header and content area through image segmentation technology, and optimize the structure parsing result in combination with semantic features; 9) Post-process the parsing result, including morphological operations and connected component analysis, to improve the recognition accuracy of the table structure; 10) Output the semantic classification result, structure segmentation result, and visualization image of the table.
2. The table recognition method according to claim 1, wherein: Specifically, step 2) is as follows: Obtain the cell content word vector V through BERT semantic encoding i,j , and the formula is: where: w i,j is the cell text content in the i-th row and j-th column of the table; The text content w of each cell in the table i,j After preprocessing including lowercasing and adding special markers, it is input into the BERT model. The BERT model outputs the mean vector of the last layer's hidden states to generate a 768-dimensional word vector V i,j ; For the word vector V i,j Perform L2 normalization to obtain the normalized word vector 3. The table recognition method according to claim 2, wherein: In step 2), the context weight score αk between the cell and its neighboring cells is calculated through a multi-layer perceptron MLP, and the formula is: αk = softmax(MLP([V i,j ; V k )) Among them, [Vi,j; Vk] represents the concatenated vector of the word vector V of the central cell i,j and the word vector V k of the neighborhood cell; After calculating the word vectors of each cell, start calculating the synthetic value of the context vector, V k is the k-th cell, and the context vector synthesis formula is as follows: C i,j is the synthetic value of the context vector for the central cell (i, j); V i,j is the word vector of the central cell (i, j), obtained from the BERT model; α i,j is the self-attention weight of the central cell, output by the MLP and calculated through the corresponding formula; N(i,j) is the neighborhood range of the central cell (i,j); αk is the weight score of the neighboring cell k, which is calculated from the dynamic attention weight. $s_k$ is the cosine similarity score of neighboring cell $k$, and the calculation formula is: $s_k = \cos\_sim(V i,j , V_k) sigmoid(sk): Compress the similarity score to (0,1); V k is the word vector of the neighborhood cell k, which is obtained by the BERT model.
4. The table recognition method according to claim 3, wherein: Specifically, step 3) is as follows: Calculate the probability distribution yk of the entity category through the context vector synthesis value Ci,j and the entity category projection matrix Went, and the formula is: P(y k |C i,j ) = softmax(W ent C i,j + b ent ) bent: Trainable bias vector; Went∈R 768×K : Entity category projection matrix; The entity link strength calculation formula is: ε i,j,k represents the entity link strength between cell (i, j) and cell k; Indicates an indicator function that is 1 when the entity type of cell k is the same as the entity type of cell (i, j), and 0 otherwise; cos_sim(C i,j , C k ) represents calculating the cosine similarity between the context vectors of cell (i, j) and cell k.
5. The table recognition method according to claim 4, wherein: In step 4), the formula for the dependency syntax score is: Wherein: indicates that cell k belongs to the neighborhood range of cell (i, j), that is, cells adjacent or related to cell (i, j) in the table structure are considered; φpos(i,k): Represents the position feature function, which is used to measure the position relationship between the cell (i,j) and the cell k; φnum(w k ): Represents a numerical feature function used to measure whether the content of cell k is of numerical type.
6. The table recognition method according to claim 1, characterized in that: In step 6), the multi-task loss function is: Ltotal = 0.6Lsem + 0.4Lstruct where Lsem is the semantic classification loss and Lstruct is the structural similarity loss.
7. The table recognition method according to claim 1, wherein: In step 7), the improved instance segmentation algorithm is based on Mask R-CNN as the basic framework for improvement, and the improvement points are: (1) Replace the original ResNet-101 with the lightweight EfficientNet-B4; (2) Introduce deformable convolution into the original ROI alignment module.