Complex header identification method based on KV relation and image mapping

By converting the table into a two-dimensional matrix and mapping it into a color image, and using the image segmentation algorithm to identify complex table headers, the problem of low recognition accuracy in the prior art is solved, and the calculation complexity and the recognition accuracy are reduced.

CN120356234APending Publication Date: 2025-07-22HANGZHOU DIANZI UNIVERSTIY INFORMATION ENG SCHOOL
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510322789.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

In the prior art, when identifying the headers of complex tables, especially when multiple rows and multiple list headers, the recognition accuracy is not high and the calculation cost is high.

Method used

Convert the table into a two-dimensional matrix, calculate the KV relationship probability of each cell with cells in its eight neighborhoods, and map the probability to a color image, use the image segmentation algorithm to extract the table header area, and post-process it in combination with domain knowledge and custom rules to improve accuracy.

Benefits of technology

It reduces the computational complexity and improves the accuracy of complex table header recognition. It is suitable for various complex table forms and has high practical value.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356234A_ABST
    Figure CN120356234A_ABST
Patent Text Reader

Abstract

The invention discloses a complex header identification method based on a KV relation and image mapping, which comprises the following steps: converting a table to be processed into a two-dimensional matrix, and extracting text features from each cell in the table; the KV relation probability between each cell and cells in eight neighborhoods of the cell is calculated; mapping the calculated probability to a color space to generate a color image with a plurality of color areas; extracting a header color region from the color image by adopting an image segmentation algorithm; and in combination with domain knowledge and self-defined rules, the recognition result is post-processed to improve the accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for identifying complex table headers based on KV relationships and image mapping. Background Art

[0002] In digital office work and data mining, tabular data is ubiquitous in various documents. Traditional table header identification methods often rely on fixed position assumptions or KV relationship classification based on simple rules, but when faced with tables with complex structures (such as multi-row and multi-column headers), the identification accuracy is not high. In addition, although existing deep learning methods are effective, the training and inference costs are relatively high. The present invention aims to provide an efficient and general table header identification algorithm, which decomposes the table into KV relationship data and uses probability mapping and image segmentation to identify the table header area, thereby reducing the computational complexity and improving the identification accuracy. Summary of the Invention

[0003] The present invention provides a method for identifying complex table headers based on KV relationships and image mapping to solve the problems existing in the above-mentioned prior art. The present invention automatically identifies the table header area for various complex tables (including tables without headers, horizontal headers, vertical headers, multi-row headers, multi-column headers, etc.), solves the problems of complex table header positioning and classification, decomposes the table into K-V relationships, and generates an image through probability mapping to simplify the segmentation process.

[0004] The technical solutions adopted by the present invention are as follows:

[0005] A method for identifying complex table headers based on KV relationships and image mapping, comprising the following steps:

[0006] 1) Convert the table to be processed into a two-dimensional matrix, and extract text features for each cell in the table;

[0007] 2) Calculate the KV relationship probability between each cell and the cells in its eight-neighborhood, and then obtain the average KV relationship probability of the adjacent area of the cell;

[0008] 3) Map the calculated probability to the color space to generate a color image with multiple color regions;

[0009] 4) Use an image segmentation algorithm on the color image to extract the table header color region;

[0010] 5) Combine domain knowledge and custom rules to post-process the recognition result to improve the accuracy.

[0011] Further, convert the input table to be processed into an m×n matrix T, where each cell T (i,j) contains text data; use OCR and natural language processing technologies to convert the text data into a feature vector x ij∈R d 。

[0012] Furthermore, for each cell T (i,j) , its eight-neighborhood N (i,j) is defined. For each cell T(k, l) ∈ N(i, j) within the neighborhood, the pairwise KV relationship probability between each cell T (i,j) and its adjacent cells is calculated as follows:

[0013] p ij,kl = σ(w T · [x ij , x kl + b)

[0014] Where: is the Sigmoid function; [x ij , x kl is the vector obtained by concatenating the feature vectors of the two cells; w ∈ R 2d is the weight vector, and b is the bias;

[0015] Calculate the regional average KV relationship probability P ij between each cell and the cells within its eight-neighborhood:

[0016]

[0017] Furthermore, after obtaining the KV relationship probability P ij , a linear mapping is adopted:

[0018] R ij = 255 × P ij

[0019] G ij= 255 × (1 - P ij )

[0020] B ij= 128

[0021] Then each cell (i, j) is mapped to an RGB color: C(P ij ) = (R ij , G ij , B ij ).

[0022] The present invention has the following beneficial effects:

[0023] The present invention converts the input table into a two-dimensional matrix, calculates the KV relationship probability between each cell and its neighborhood, maps the probability values to a color image, and then uses image segmentation technology to identify the header area. By adaptively processing complex table structures (including multi-row and multi-column headers), the present invention not only reduces the computational complexity but also improves the recognition accuracy. The present invention can achieve good results in various complex table recognition tasks and has high practical value. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 is a flowchart of the present invention.

[0025] Figure 2 performs probability mapping on the example table of the present invention to obtain the color map corresponding to the table. DETAILED DESCRIPTION OF THE INVENTION

[0026] The present invention will be further described below with reference to the accompanying drawings.

[0027] As Figure 1 , a method for recognizing complex table headers based on KV relationships and image mapping according to the present invention includes the following steps:

[0028] 1) Convert the table to be processed into a two-dimensional matrix, and extract text features for each cell in the table;

[0029] 2) Calculate the KV relationship probability between each cell and the cells within its eight-neighborhood pairwise, and then obtain the average KV relationship probability of the adjacent area of the cell;

[0030] 3) Map the calculated probability to the color space to generate a color image with multiple color regions;

[0031] 4) Use an image segmentation algorithm on the color image to extract the header color region;

[0032] 5) Combine domain knowledge and custom rules to post-process the recognition results to improve the accuracy.

[0033] The specific solutions for implementing the above steps are as follows:

[0034] 1) Table preprocessing and feature extraction

[0035] Convert the input table to be processed into an m×n matrix T, where each cell T (i,j) contains text data; use OCR and natural language processing technologies to convert the text data into a feature vector x ij ∈R d .

[0036] 2) KV relationship probability calculation

[0037] For each cell T (i,j) , define its eight-neighborhood N (i,j) . For each cell T(k, l) ∈ N(i, j) in the neighborhood, calculate the pairwise KV relationship probability between each cell T (i,j) and its adjacent cells. The formula is as follows:

[0038] p ij,kl = σ(w T · [x ij , x kl + b)

[0039] Where: is the Sigmoid function; [x ij , x kl is the vector obtained by concatenating the feature vectors of the two cells; w ∈ R 2d is the weight vector, and b is the bias;

[0040] Calculate the regional average KV relationship probability P between each cell and the cells in its eight-neighborhood ij :

[0041]

[0042] The pairwise KV relationship between cells is prone to extreme value errors, while the regional average probability of KV based on the 8-neighborhood can compensate for the extreme value errors brought by the pairwise KV relationship in the neighborhood.

[0043] The method for obtaining the weight vector w in the above formula function is as follows:

[0044] Training data construction: Construct training samples from a large number of actual tables. Each sample consists of the concatenated feature vectors of a pair of cells and the corresponding KV relationship label (1 indicates the existence of a relationship, 0 indicates the non-existence).

[0045] Model architecture design: Design a fully connected neural network. Assume that the feature vector extracted from each cell is d. When the feature vectors of two cells are concatenated, the dimension of the input vector is 2d. The output is the predicted probability, indicating whether there is a KV relationship between the two cells.

[0046] Optimization algorithm: Use the gradient descent or Adam optimization algorithm to update w and b through backpropagation.

[0047] Pre-training and fine-tuning: Use a pre-trained NLP model (such as BERT, RoBERTa) as a feature extractor, and then fine-tune on specific domain data to obtain the optimal w parameter.

[0048] The above process ensures that w can learn effective KV relationship features from the data, thereby improving the recognition accuracy.

[0049] 3) Probability mapping to color values

[0050] Define the color mapping function C: [0, 1] → R 3 , such as a linear mapping:

[0051] R ij = 255 × P ij

[0052] G ij= 255 × (1 - P ij )

[0053] B ij= 128

[0054] Then each cell (i, j) is mapped to an RGB color: C(P ij ) = (R ij , G ij , B ij ).

[0055] 4) Image generation and segmentation

[0056] Rearrange the color values of all cells to form a color image III. Subsequently, use an image segmentation algorithm (such as a deep learning segmentation model based on U-Net or traditional threshold segmentation) to segment III and extract the header area. Post-processing steps (such as morphological operations) are used to smooth the boundaries, merge small regions, etc. to optimize the final recognition effect.

[0057] Example illustration

[0058] There is the following complex table (including multi-row and multi-column headers):

[0059] Sales data Regional sales Unit A |100 |10 Unit B |150 |12 Inventory data Warehouse inventory Unit A 200 8 Unit B 250 9 Market data Channel sales Unit C |180 |11

[0060] Suppose the calculated w = [1, -0.5, 0.3, 0.2, -0.1, 0.4] and b = 0.

[0061] Row 1: "Sales data", "Regional sales", "Unit A", "100", "$10";

[0062] Row 2: "Blank", "Blank", "Unit B", "150", "$12";

[0063] Row 3: "Inventory data", "Warehouse inventory", "Unit A", "200", "$8";

[0064] Row 4: "Blank", "Blank", "Unit B", "250", "$9";

[0065] Line 5: "Market Data", "Channel Sales", "Unit C", "180", "$11".

[0066] The corresponding feature vectors extracted through OCR and NLP (approximate values are taken in the example) are as follows:

[0067] Row 1:

[0068] Taking the cell T(1,1) "Sales Data" in the first row and first column and its neighborhood T(1,2) as an example, calculate:

[0069] p (1,1)(1,2) , = σ(w T [x 1,1 ,x 1,2 ), and then take the average of the neighborhood probabilities to obtain P 1,1 ≈0.734.

[0070] Then according to:

[0071] R ij = 255 × P ij ,

[0072] G ij = 255 × (1 - P ij ),

[0073] B ij = 128

[0074] Calculate the mapped color to be (187, 68, 128).

[0075] After the entire table is mapped, a color image is generated. Areas with darker red colors (higher P ij ) may indicate the table header area (such as Figure 2 ). After image segmentation, the system automatically extracts the table header area.

[0076] The present invention has the following advantages:

[0077] Reduce complexity: Only calculate the KV relationship probability, avoid processing the entire table data, and reduce the computational burden.

[0078] Strong adaptability: Applicable to various complex table forms such as tables without headers, horizontal, vertical, multi-row, or multi-column headers.

[0079] Intuitive visualization: Generate a color image through color mapping, which is convenient for debugging and subsequent image segmentation.

[0080] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, several improvements can be made without departing from the principle of the present invention, and these improvements should also be regarded as the protection scope of the present invention.

Claims

1. A method for identifying complex table headers based on KV relationships and image mapping, characterized in that: It includes the following steps: 1) Convert the table to be processed into a two-dimensional matrix, and extract text features for each cell in the table; 2) Calculate the KV relationship probability between each cell and the cells within its eight-neighborhood pairwise, and then obtain the average KV relationship probability of the adjacent area of the cell; 3) Map the calculated probability to the color space to generate a color image with multiple color regions; 4) Use an image segmentation algorithm for the color image to extract the header color region; 5) Combine domain knowledge and custom rules to post-process the recognition results to improve the accuracy.

2. The complex table header recognition method based on the KV relationship and image mapping according to claim 1, characterized in that: Convert the input table to be processed into an \(m\times n\) matrix \(T\), where each cell \(T\) (i,j) contains text data; use OCR and natural language processing technologies to convert the text data into a feature vector \(\mathbf{x}\) ij \(\in\mathbb{R}\) d .

3. The complex table header recognition method based on the KV relationship and image mapping according to claim 2, wherein: For each cell T (i,j) , define its eight-neighborhood N (i,j) . For each cell T(k, l) ∈ N(i, j) in the neighborhood, calculate the pairwise KV relationship probability between each cell T (i,j) and its adjacent cells. The formula is as follows: p ij,kl = σ(w T · [x ij , x kl + b) Wherein: is the Sigmoid function; [x ij , x kl is the vector obtained by concatenating two cell feature vectors; w ∈ R 2d is the weight vector, and b is the bias; Calculate the regional average KV relationship probability P between each cell and the cells within its eight-neighborhood ij :

4. The complex table header recognition method based on KV relationship and image mapping according to claim 3, characterized in that: Obtain the KV relationship probability P ij After that, adopt linear mapping: R ij = 255 × P ij G ij= 255×(1 - P ij ) B ij= 128 Then each cell (i, j) is mapped to an RGB color: C(P ij ) = (R ij , G ij , B ij ).