PDF document intelligent identification and content extraction method based on deep learning
Through deep learning methods, combined with table detection, cell merging and structure verification networks, the problems of low accuracy and poor robustness in complex table recognition and structured data extraction are solved, and efficient and automated table content extraction is achieved.
Patent Information
- Application Number
- CN202511309941.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-15
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2045-09-15
AI Technical Summary
Existing technologies have difficulty in efficiently identifying and extracting structured data from complex tables, especially when table lines are missing or the formats are diverse, resulting in low recognition accuracy, poor robustness, and inefficient reliance on manual operations.
A deep learning-based method is adopted, which utilizes table detection network, cell merging network and structure verification network. Through multi-level visual feature extraction, row and column structure analysis and cross-row and cross-column merging prediction, table area recognition, structure repair and content extraction are achieved.
It significantly improves the accuracy of table area detection and content extraction, realizes a fully automated process, adapts to different table styles and backgrounds, reduces image quality requirements, and improves work efficiency and recognition accuracy.
Smart Images

Figure CN120808373A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical fields of artificial intelligence, deep learning, computer vision and document image processing, in particular to a PDF document intelligent recognition and content extraction method based on deep learning. BACKGROUND
[0002] PDF documents are widely used in information storage and transmission due to their cross-platform and fixed format advantages. However, the tables in PDF documents, especially complex tables in the form of scanned documents or images, are often difficult to extract structured data directly. Traditional methods rely on heuristic rules or image edge detection, and have strong dependence on the integrity of table lines, consistency of styles, cell merging and splitting rules, and text direction in the processing process.
[0003] When facing actual situations such as missing table lines, complex structure and various formats, the recognition accuracy is low and the robustness is poor. This not only limits the degree of automation, but also forces a large amount of table information to rely on manual input, which is not only inefficient but also prone to errors, seriously restricting the efficient circulation and application of structured data in business processes.
[0004] At present, there is no effective solution to the problems in the related art. SUMMARY
[0005] In view of the problems in the related art, the present application proposes a PDF document intelligent recognition and content extraction method based on deep learning to overcome the above technical problems existing in the prior art.
[0006] To this end, the specific technical solutions adopted by the present application are as follows: A PDF document intelligent recognition and content extraction method based on deep learning, comprising: using a table detection network to recognize the table area of an input PDF whole-page image to obtain the positioning table area of each table in the PDF whole-page image; performing row and column structure analysis on the positioning table area, and using a cell merging network to predict the segmentation line of the row and column structure to obtain a basic grid structure; based on the basic grid structure, using the cell merging network to predict the merging of adjacent cells to obtain cells with cross-row or cross-column structure; using a structure verification network to detect and repair the consistency of the cells with cross-row or cross-column structure to obtain a repaired table structure; performing text recognition on each logical cell in the repaired table structure, and binding the row and column position information corresponding to each logical cell to obtain table content that can be output in a preset structured format.
[0007] Further, the table detection network is used to identify the table area of the input PDF whole-page image, and the positioning table area of each table in the PDF whole-page image is obtained, including: The table detection network is used to process the input PDF whole-page image, and multi-level visual features of the PDF whole-page image are extracted and fused through a backbone network, a neck network and a head network to obtain a feature map in the PDF whole-page image; Based on the feature map, the table boundary box of the PDF whole-page image is predicted to obtain the boundary box of the candidate table area as a preliminary estimate of the table position; Based on the smooth adaptive intersection over union loss function, the boundary box of the candidate table area is optimized to obtain the target table positioning box; Based on the target table positioning box, the region cutting processing of the PDF whole-page image is performed to obtain the positioning table area of each table in the PDF whole-page image.
[0008] Further, based on the feature map, the table boundary box of the PDF whole-page image is predicted to obtain the boundary box of the candidate table area as a preliminary estimate of the table position, including: A scale confidence prediction branch is introduced in the table detection network output, and the scale confidence is jointly trained using the smooth adaptive intersection over union loss function to obtain a scale confidence score; Based on the scale confidence score, a dynamic threshold function is constructed, and the detection results of different types of tables are screened; The dynamic threshold function is used to adaptively judge the predicted boundary box to obtain the boundary box of the candidate table area as a preliminary estimate of the table position.
[0009] Further, based on the smooth adaptive intersection over union loss function, the boundary box of the candidate table area is optimized to obtain the target table positioning box, including: The distance between the predicted box and the real box center point is punished using the smooth adaptive intersection over union loss function to obtain an initial optimization direction; Based on the initial optimization direction, the distance between the predicted box and the real box corner point is adjusted to obtain an optimized prediction box that conforms to the actual boundary of the table; The optimized prediction box is comprehensively measured using the smooth adaptive intersection over union loss function to obtain the target table positioning box.
[0010] Further, the row and column structure inside the positioning table area is analyzed, and a cell merging network is used to predict the segmentation line of the row and column structure to obtain a basic grid structure, including: A convolutional neural network is used to extract visual features of the table area image to obtain a two-dimensional feature map containing table row and column layout information; performing pooling operations in horizontal and vertical directions on the two-dimensional feature map to obtain a one-dimensional feature sequence representing row and column information; performing global context analysis on the one-dimensional feature sequence representing row and column information using a plurality of preset structure encoders to obtain a structure representation sequence containing global information; based on the structure representation sequence, performing binary classification prediction on each position to obtain logical row and column division line positions, and implementing preliminary construction of the basic grid structure.
[0011] Further, based on the structure representation sequence, performing binary classification prediction on each position to obtain logical row and column division line positions, and implementing preliminary construction of the basic grid structure includes: based on the obtained structure representation sequence, performing binary classification prediction on each position to obtain a probability sequence representing division line confidence; using the probability sequence of division line confidence to construct a cost graph with division points as nodes and probability as the basis, and obtaining a global cost structure; using a preset structure priori knowledge, defining a path transition cost function to obtain a complete path scoring mechanism; based on the constructed cost graph and path scoring mechanism, using a dynamic programming algorithm to solve all candidate paths to obtain a minimum total cost path; using the minimum total cost path to globally optimize the preliminary division line to obtain a logical row and column division line position with continuous structure and reasonable distribution, and implementing preliminary construction of the basic grid structure.
[0012] Further, based on the basic grid structure, using a cell merging network to perform merging prediction on adjacent cells to obtain cells with cross-row or cross-column structure includes: using a convolutional neural network to extract visual features of each cell, and combining normalized coordinates in the convolutional neural network and text embedding to obtain a multi-dimensional feature vector; based on the multi-dimensional feature vector, arranging all cells in row and column order into a sequence to obtain a cell feature sequence input to the cell merging network; using a self-attention mechanism of the cell merging network encoder to process the cell feature sequence, calculating the correlation between any two cells, and learning non-local merging logic; based on the feature representation output by the cell merging network encoder, performing merging prediction on each cell and its right or lower cell through a classification head to obtain cells with cross-row or cross-column structure.
[0013] Further, using a structure verification network to perform consistency detection and repair on the cells with cross-row or cross-column structure to obtain a repaired table structure includes: The combined cell structure is checked by using a structure checking network to obtain a cell combination situation with potential logical conflicts; Based on the logical conflict area output by the structure checking network, an editing action is performed on the logical conflict area to obtain a target structure rationality reward; The target structure rationality reward is used to optimize the strategy of all editing actions to obtain a repaired table structure with logical continuity and accurate cross-row and cross-column information.
[0014] Further, text recognition is performed on each logical cell in the repaired table structure, and the row and column position information corresponding to each logical cell is bound to obtain table content that can be output in a preset structured format, including: Each logical cell image slice in the repaired table structure is classified to obtain a category label; The category label is used to call the corresponding optical character recognition strategy for the logical cell image to obtain target extracted logical cell text content information; Based on the logical cell text content information and the repaired table structure, and by binding the row and column positions and merging information of each logical cell, table content that can be output in a preset structured format is obtained.
[0015] Further, the category label is used to call the corresponding optical character recognition strategy for the logical cell image to obtain target extracted logical cell text content information, including: Based on the final table structure, optical character recognition is performed on each logical cell to obtain preliminary logical cell text content; The preliminary logical cell text content is used to extract the logical cells across columns in the final table structure to trigger a consistency checking mechanism; Based on a semantic weighted edit distance algorithm, the header logical cell and the content of the sub-logical cells below it are compared for consistency to obtain a semantic consistency score result; The semantic consistency score result is used to judge and correct the recognized text to obtain target extracted logical cell text content information.
[0016] The beneficial effects of the present application are: 1、The present application combines the feature learning ability of the deep learning cell merging network to significantly improve the accuracy of table region detection, structure analysis and content extraction, especially in complex tables with incomplete lines and merged / detached cells, showing stronger robustness. Thus, it can effectively process various forms of PDF tables such as scanned documents and images, adapt to different table styles, fonts and backgrounds, reduce the requirements for input image quality, and ensure high-precision table recognition and content extraction.
[0017] 2、The application realizes a fully automatic process from table recognition to content extraction, greatly reduces manual intervention, and improves work efficiency. Thus, it is suitable for financial statements, contracts, invoices, questionnaires and other PDF documents containing complex tables. Its wide applicability and automation characteristics not only provide great convenience for data analysis and information management, but also greatly improve the intelligent level of industry workflow, reduce the error and time cost of manual operation.
[0018] 3、The application converts the cell merging problem into a classification task based on global context, uses the self-attention mechanism of Transformer for intelligent reasoning, and gets rid of the dependence on physical lines. Thus, the cell merging network can consider global information such as table layout, column division line, table header content, etc., to accurately determine the position of the merged cell. Through this strategy of first fixed structure and then classification reading, fine content recognition can be realized, greatly improving the efficiency and accuracy of table data extraction. BRIEF DESCRIPTION OF DRAWINGS
[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed in the embodiments will be briefly introduced as follows. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor on the basis of these drawings.
[0020] Figure 1 is a flowchart of a deep learning-based PDF document intelligent recognition and content extraction method according to an embodiment of the present application. DETAILED DESCRIPTION
[0021] In order to further illustrate the embodiments, the present application provides drawings which are part of the disclosure of the present application, mainly used to illustrate the embodiments, and can explain the operating principle of the embodiments in conjunction with the related description of the specification. With reference to these contents, those skilled in the art should understand other possible embodiments and advantages of the present application.
[0022] According to an embodiment of the present application, a deep learning-based PDF document intelligent recognition and content extraction method is provided.
[0023] The present application will be further described in conjunction with the drawings and specific embodiments, as shown in Figure 1 The deep learning-based PDF document intelligent recognition and content extraction method according to an embodiment of the present application comprises: S1, using a table detection network (YOLOv8 plus SAIoU Loss) to recognize the table area of the input PDF whole page image, and obtaining the positioning table area of each table in the PDF whole page image; Specifically, high-precision table positioning: through a specially optimized target detection network (i.e., table detection network), all the positions of the tables in the input whole-page image are accurately identified and framed.
[0024] Specifically, table detection network (YOLOv8 plus SAIoU Loss): YOLOv8 is selected as the basic framework, but its innovation lies in not using the conventional regression loss function, but using the SAIoU Loss loss function for training. This enables the network to achieve intelligent optimization of fast movement first and fine tuning later when positioning the table boundary, with much higher accuracy than conventional methods.
[0025] Specifically, YOLOv8 is used as the basic target detector. YOLOv8 is a typical one-stage detector, and its structure includes three parts: 1) Backbone (backbone network): responsible for extracting visual features at different levels from the input image.
[0026] 2) Neck (neck network): such as FPN (Feature Pyramid Network), responsible for fusing multi-scale features output by the backbone network to better detect targets of different sizes.
[0027] 3) Head (head network): performs final classification and bounding box regression prediction on the fused feature map.
[0028] It needs to be explained that the conventional target detection cell merging network (such as standard YOLO) usually uses general loss functions such as SmoothL1, GIoU, DIoU for bounding box regression during training. These loss functions are effective, but for tables with variable shapes and large aspect ratios, they have the problems of slow convergence and insensitivity to boundary accuracy. For example, GIoU can only provide effective movement gradient when the predicted box has no overlap with the true box, which is very common in the early stages of training, resulting in low training efficiency. The present invention does not use the above general loss function, but introduces a smooth adaptive intersection over union loss function (i.e., SAIoU Loss) to guide network training.
[0029] SAIoU Loss is used in the training process of the table detection network, and its working principle is to dynamically and in stages optimize the positioning of the bounding box, achieving intelligent training of fast positioning first and fine adjustment later, with the following specific steps: Early training stage (when the predicted box is far from the true box): At this stage, SAIoU Loss focuses on calculating and penalizing the center point distance between the predicted box and the true box.
[0030] This metric provides a clear and smooth gradient for the cell merging network, guiding the prediction box to quickly converge to the true table area, solving the problem of gradient disappearance when the two boxes do not overlap.
[0031] In the later stage of training (when the prediction box and the true box have good overlap): At this stage, the focus of SAIoU Loss shifts to minimizing the distance between the corresponding corner points of the two boxes.
[0032] Through fine adjustment of the corner points, the size and shape of the prediction box are fine-tuned to better fit the actual boundaries of the table, which is crucial for preventing rows and columns from being incorrectly cut.
[0033] Throughout the entire training process: Overlap (IoU) as the most basic metric is always part of the loss function, focusing on the overlap area between the prediction box and the true table area.
[0034] By adaptively combining and adjusting the focus of these three geometric metrics during training, SAIoU Loss guides the cell merging network to efficiently and accurately complete the table positioning task.
[0035] The working principle of SAIoU Loss: This loss function skillfully combines three different geometric metrics and can adaptively adjust the focus during training.
[0036] Overlap (IoU): Focuses on the overlap area between the prediction box and the true table area, which is the most basic metric.
[0037] Center point distance: In the early stage of training, when the prediction box is far from the true box, this item provides a clear and smooth gradient for the cell merging network by punishing the distance between the center points of the two boxes, guiding the prediction box to quickly converge to the correct position, solving the problem of gradient disappearance when the two boxes do not overlap.
[0038] Corner point distance: In the later stage of training, when the prediction box and the true box have good overlap, this item fine-tunes the size and shape of the prediction box by minimizing the distance between the corresponding corner points, making it better fit the actual boundaries of the table. This is crucial for ensuring that the rows and columns of the table are not incorrectly cut.
[0039] S2, analyze the row and column structure inside the positioned table area, and use the cell merging network to predict the segmentation line of the row and column structure to obtain the basic grid structure. Specifically, row and column segmentation: for each table region located, a Transformer-based cell merging network is used to intelligently predict the positions of all logical row and column segmentation lines, thereby dividing the table into a basic grid.
[0040] S3, based on the basic grid structure, the adjacent cells are merged by using the cell merging network to obtain cells with cross-row or cross-column structure; Specifically, cell merging: again using a Transformer cell merging network (i.e., a Transformer-based cell merging network), all cells in the basic grid are globally analyzed to predict which adjacent cells should be logically merged into a cross-row or cross-column cell.
[0041] S4, using a structure verification network to detect and repair the consistency of the cell with the cross-row or cross-column structure, to obtain the repaired table structure; Specifically, structure verification: in order to ensure the absolute logical correctness of the final structure, an intelligent verification agent based on reinforcement learning will review and fine-tune the merged structure, and automatically repair potential logical conflicts.
[0042] S5, text recognition is performed on each logical cell in the repaired table structure, and the row and column position information of each logical cell is bound, to obtain the table content that can be output as a preset structured format (i.e., HTML structured format).
[0043] Specifically, content extraction and structured output: on the completely determined table structure, fine-grained content recognition is performed on each logical cell, and the recognized text and its corresponding row and column position, merging information are bound, and finally output as an HTML structured format. Output HTML or JSON files containing complete structure and accurate content, input a PDF document.
[0044] In this optional embodiment, a table detection network is used to identify the table region of the input PDF whole-page image, and the located table region of each table in the PDF whole-page image includes: The table detection network is used to process the input PDF whole-page image, and the multi-level visual features of the PDF whole-page image are extracted and fused by the backbone network, neck network and head network to obtain the feature map in the PDF whole-page image. Based on the feature map, the table boundary box prediction of the PDF whole-page image is performed to obtain the boundary box of the candidate table region as a preliminary estimate of the table position. Based on the smooth adaptive intersection over union loss function (SAIoU Loss loss function), the boundary box of the candidate table region is optimized to obtain a target table positioning box. Based on the obtained target table positioning box, the region cutting processing is performed on the PDF whole page image to obtain the positioning table region of each table in the PDF whole page image.
[0045] Specifically, a document picture is received as input, and the YOLOv8 network processes the whole page image, extracts and fuses features through its backbone network (Backbone), neck network (Neck) and head network (Head). In the training process, the SAIoU Loss loss function will guide the network to perform intelligent optimization of fast movement first and fine tuning later, and finally output accurate boundary boxes of all table positions in the head network.
[0046] In this optional embodiment, based on the feature map, the table boundary box of the PDF whole page image is predicted to obtain the boundary box of the candidate table region as the preliminary estimation of the table position, including: A scale confidence prediction branch is introduced in the table detection network output, and the scale confidence is jointly trained using the smooth adaptive intersection over union loss function to obtain a scale confidence score; Based on the obtained scale confidence and classification confidence, a dynamic threshold function is constructed, and the detection results of different types of tables are screened; The adaptive judgment of the dynamic threshold function is used to obtain the boundary box of the candidate table region as the preliminary estimation of the table position.
[0047] Specifically, the present application introduces a scale-confidence adaptive thresholding (SCAT) mechanism coupled with network training to enhance the stability and accuracy of table detection of different scales and different clarity. The specific implementation is as follows: 1) Network output enhancement: modify the head (Head) of the table detection network (YOLOv8) to output an additional scale confidence (Scale Confidence Score) while predicting the bounding box (Bounding Box) and classification confidence (Classification Confidence). The score is used to represent the network's confidence in the scale and clarity of the target (i.e. table) in the current prediction box, and the learning of this value is guided by SAIoU Loss.
[0048] Adaptive threshold decision: In the inference stage, instead of using a fixed threshold (e.g. 0.5) to filter the detection results as in traditional methods, the invention uses a dynamic threshold function: .
[0049] wherein, S scale represents the scale confidence of the network output; C class represents the classification confidence; The core idea of this dynamic threshold function is: when the network predicts a high scale confidence (i.e. the network believes that this is a large size, clear feature table), the final judgment threshold will be automatically raised. This can effectively filter out non-table noise in the background and prevent false detection in large table areas, achieving the best of the best. When the network predicts a low scale confidence (i.e. the network believes that this is a small size or blurred table), the judgment threshold will be automatically lowered. This can save real tables with low classification confidence due to unclear features, greatly improving the recall rate for small targets and low-quality images.
[0050] In this optional embodiment, based on the smooth adaptive intersection over union loss function, the bounding box of the candidate table region is optimized to obtain the target table positioning box, including: The smooth adaptive intersection over union loss function is used to punish the distance between the predicted box and the real box center point to obtain an initial optimization direction; Based on the obtained initial optimization direction, the distance between the predicted box and the real box corner point is adjusted to obtain an optimized predicted box that conforms to the actual boundary of the table; The smooth adaptive intersection over union loss function is used to comprehensively measure the optimized predicted box to obtain the target table positioning box.
[0051] In this optional embodiment, the internal positioning table region is analyzed for row and column structure, and a cell merging network is used to predict the row and column structure segmentation line to obtain a basic grid structure, including: The convolutional neural network is used to extract visual features of the table region image to obtain a two-dimensional feature map containing table row and column layout information; The two-dimensional feature map is subjected to horizontal and vertical pooling operations to obtain a one-dimensional feature sequence representing row and column information; A number of preset structure encoders (i.e. Transformer encoders) are used to analyze the global context of the row and column one-dimensional feature sequence to obtain a structure representation sequence containing global information; Based on the structure representation sequence, each position is predicted for binary classification to obtain the logical row and column segmentation line position, achieving the construction of the preliminary basic grid structure.
[0052] Specifically, each table region image is input into a row-column segmentation network (CNN plus Transformer). The row-column segmentation network first extracts a visual feature map through CNN (i.e., ResNet), then converts the feature map into two one-dimensional feature sequences (one representing rows and one representing columns), and finally predicts the positions of all logical row and column segmentation lines through a fully connected layer.
[0053] Specifically, the row-column segmentation network (CNN plus Transformer): The core of this network is to convert two-dimensional table image features into one-dimensional feature sequences through compression (i.e., Pooling), and then input them into a Transformer encoder. Unlike traditional CNNs that can only see local pixels, the global attention mechanism of the Transformer allows the cell merging network to see the entire table layout when determining whether a certain place is a segmentation line, thus making more accurate judgments on complex cases such as wireless tables and broken tables.
[0054] It should be explained that a standard convolutional neural network (i.e., ResNet) is used as the backbone network, which inputs the detected table image and outputs a two-dimensional visual feature map (i.e., Feature Map), cleverly converting the two-dimensional table recognition task into two independent one-dimensional sequence labeling problems.
[0055] Specifically, the implementation of two-dimensional to one-dimensional: a CNN (ResNet) is used to extract a two-dimensional visual feature map (i.e., Feature Map) from the table image. Then, the pooling operation is used to convert it into two independent one-dimensional sequences.
[0056] For column segmentation: the two-dimensional feature map is globally averaged pooled in the horizontal direction to flatten it into a one-dimensional vertical feature sequence. Each element of this sequence condenses all the visual information of a vertical slice in the table image.
[0057] For row segmentation: similarly, the pooling is performed in the vertical direction to obtain a one-dimensional horizontal feature sequence, each element of which represents the information of a horizontal slice.
[0058] Through the above implementation, the complex two-dimensional structure recognition problem is cleverly decomposed into two relatively simple one-dimensional sequence labeling problems.
[0059] Specifically, the pooling method: when processing column segmentation, global average pooling is used. When row segmentation is used, the same pooling method is used. Row segmentation uses the same pooling mechanism as column segmentation, that is, global average pooling is performed in the vertical direction.
[0060] Specifically, the role of the Transformer encoder: after inputting the two one-dimensional feature sequences into the Transformer encoder, the encoder processes them using its core multi-head self-attention mechanism. The specific role is: each element in the sequence (representing an image slice) can simultaneously pay attention to all other elements in the sequence when calculating its new representation, and different attention weights are assigned according to the correlation. The purpose of this process is to give the cell merging network a global receptive field, so that the cell merging network can understand the context of each slice based on the layout of the entire table (such as table headers, borders, alignment, and other global information), rather than relying solely on local pixel information.
[0061] Specifically, the prediction method: after the Transformer encoder finishes processing, its output is connected to a simple fully connected layer (i.e., the classification head). The classification head performs binary classification prediction on each position in the sequence. The basis for determining whether a position is a segmentation point is the feature vector output by the Transformer encoder, which contains global context information. The document explicitly states that because the cell merging network has global reasoning capability, it considers global information such as alignment on the left side of the table, the border on the right side, and the layout of the table header when determining whether a certain position is a segmentation line. This global context-based judgment allows the invention to be independent of physical lines and exhibits high robustness to wireless tables and broken tables.
[0062] In this optional embodiment, based on the structure representation sequence, binary classification prediction is performed on each position to obtain the logical row and column segmentation line position, and a preliminary basic grid structure is constructed, including: Based on the obtained structure representation sequence, binary classification prediction is performed on each position to obtain a probability sequence representing the segmentation line confidence; Using the probability sequence of the segmentation line confidence, a cost graph is constructed with the segmentation points as nodes and the probability as the basis, and a global cost structure is obtained; Using the pre-set structure prior knowledge, a path transition cost function is defined to obtain a complete path scoring mechanism; Based on the constructed cost graph and path scoring mechanism, a dynamic programming algorithm is used to solve all candidate paths to obtain the path with the minimum total cost; The total cost minimum path is used to globally optimize the preliminary segmentation line to obtain a logically continuous and reasonably distributed logical row and column segmentation line position, and to realize the construction of a preliminary basic grid structure.
[0063] It should be explained that after the row and column segmentation based on the Transformer, the application introduces a global context guided dynamic programming (Global-Context Guided Dynamic Programming, GCG-DP) algorithm to optimize the segmentation line sequence output by the Transformer to ensure the absolute logical continuity of the segmentation line in an extremely irregular table, and the specific implementation is as follows: 1) Obtain a global context probability sequence: first, according to the original scheme, process the table image by using the row and column segmentation network (CNN plus Transformer) to output a probability sequence representing whether each row and column position is a segmentation line. This sequence contains the global understanding of the table layout by the cell merging network.
[0064] 2) Build a segmentation path cost graph: convert the above probability sequence into a cost graph. Each node in the graph represents a potential segmentation point, and the cost is inversely proportional to the probability output by the Transformer (the higher the probability, the lower the cost).
[0065] 3) Dynamic programming to solve the optimal path: apply a dynamic programming algorithm to find a path with the lowest total cost in the cost graph. The core of this process is to design a transition cost function (TransitionCost Function) that includes structural prior knowledge.
[0066] 4) Continuity reward: give a low transition cost or reward to a continuous and smooth segmentation line path to encourage the continuity of the line.
[0067] 5) Jump penalty: impose a high transition cost on paths that appear to be logically discontinuous, jumping or sharp, thereby suppressing discontinuous segmentation results in decision-making.
[0068] 6) Distance constraint: according to the average row and column width of the table, paths with too small or too large distances between two segmentation lines are penalized to ensure the reasonableness of the segmentation.
[0069] In this optional embodiment, based on the basic grid structure, the adjacent cells are merged by using the cell merging network to obtain a cell with a cross-row or cross-column structure, including: The visual features of each cell are extracted by using a convolutional neural network, and the normalized coordinates in the convolutional neural network and the text embedding are combined to obtain a multi-dimensional feature vector. Based on the multi-dimensional feature vector, all the cells are arranged in row and column order as a sequence to obtain the cell feature sequence input to the cell merging network; The self-attention mechanism of the cell merging network (Transformer) encoder is used to process the cell feature sequence, calculate the correlation degree between any two cells, and learn the non-local merging logic. Based on the feature representation output by the cell merging network encoder, the classification head is used to predict the merging of each cell with its right or lower cell, and the cell with cross-row or cross-column structure is obtained.
[0070] Specifically, based on the segmented basic grid, visual, position and preliminary text features are extracted for each cell, and the feature vectors of all cells are input as a sequence to the Transformer encoder. The cell merging network calculates the correlation degree between any two cells through the self-attention mechanism, and finally predicts which adjacent cells need to be logically merged by the classification head.
[0071] Specifically, the cell merging network (Transformer): This network inputs the features (visual, position, text) of all basic grid cells as a sequence into the Transformer encoder. The self-attention mechanism is used to calculate the correlation degree between any two cells, so as to learn complex merging rules, for example, the cell merging network can understand that two cells with similar content, consistent background color and in the same column should be merged.
[0072] It needs to be explained that the multi-dimensional features (visual, position, and preliminary recognized semantics) of all basic cells are input as a sequence into the cell merging network, and the self-attention mechanism of the Transformer calculates the correlation degree between any two cells, so as to learn complex merging rules, for example, the cell merging network can understand that two cells with similar content, consistent background color and in the same column should be merged, and a special classification head will predict whether each cell needs to be merged with its right or lower cell. After completing the row and column segmentation, a basic grid layout is obtained.
[0073] Specifically, for each segmented basic grid cell, a multi-dimensional feature vector is extracted, which is composed of three parts: 1) Visual features: From the ResNet backbone network, the corresponding region features are extracted according to the coordinates of the cell.
[0074] 2) Position feature: The normalized row and column coordinates of the cell in the grid.
[0075] 3) Semantic features: Through a lightweight OCR engine, the cell content is preliminarily identified, and the text is converted into a word embedding vector.
[0076] Specifically, the merging relationship prediction: a Transformer encoder cell merging network is used to predict whether adjacent grid cells should be merged. The feature vectors of all basic grid cells in the table are arranged in a sequence in the order from top to bottom and from left to right as the input of the Transformer. Similar to the row and column segmentation, the self-attention mechanism of the cell merging network calculates the correlation score between any two grid cells. At the output end of the Transformer, a classification head is designed. For each grid cell, the classification head predicts whether it needs to be merged with the cell to its right and whether it needs to be merged with the cell below it.
[0077] In this optional embodiment, a structure verification network (i.e., a Reinforcement Learning Agent) is used to detect and repair the consistency of the cells with cross-row or cross-column structures, and the repaired table structure includes: The structure verification network is used to check the merged cell structure, and the cell combination with potential logical conflicts is obtained. Based on the logical conflict area output by the structure verification network, editing actions are performed on the logical conflict area to obtain a target structure rationality reward. The target structure rationality reward is used to optimize the strategy of all editing actions, and a repaired table structure with logical continuity and accurate cross-row and cross-column information is obtained.
[0078] Specifically, the structure verification network: this is an intelligent agent based on reinforcement learning. Its core is a small policy network (i.e., Policy Network, usually MLP) that learns how to correct the table structure through a series of editing actions (such as merging and splitting) to obtain the highest structure rationality reward. This is a dynamic, self-learning correction process, rather than an immutable rule.
[0079] In this optional embodiment, text recognition is performed on each logical cell in the repaired table structure, and the row and column position information corresponding to each logical cell is bound, and the table content that can be output in a preset structured format includes: Each logical cell image slice in the repaired table structure is classified to obtain a category label. The category label is used to call the corresponding optical character recognition strategy (i.e., OCR strategy) for recognition to obtain the target extracted logical cell text content information. Based on the logical cell text content information and the repaired table structure, and the row and column positions and merging information of each logical cell are bound, to obtain the table content which can be output as a pre-set structured format.
[0080] Specifically, on the final table structure after verification, first use a small CNN to classify the content of each logical cell image slice (such as pure text, pure numbers, blank, etc.). Then, according to the classification results, call the corresponding OCR strategy for fine content recognition. The recognized text is bound with its corresponding row and column position, merging information (i.e. rowspan, colspan), and output as a structured file in HTML or JSON format.
[0081] Specifically, on the fully determined table structure, a small convolutional neural network (MobileNet) will be used first to quickly classify each final logical cell image slice into predefined categories such as pure text, pure numbers, mixed text and graphics, or blank. According to the classification results, different recognition strategies are adopted. For example, for areas classified as pure numbers, a number-optimized OCR cell merging network is called. This way, computing resources are concentrated where they are most needed, achieving adaptive and fine-grained recognition, avoiding false recognition of image noise as characters, and improving efficiency. After fine-grained recognition of each cell, the final OCR result is bound with its logical position (row number, column number) and merging information after RL verification, generating standard, machine-readable HTML file structured data.
[0082] In this optional embodiment, the logical cell image is called with a category label to invoke the corresponding optical character recognition strategy for recognition, and the target extracted logical cell text content information includes: Based on the final table structure, perform optical character recognition on each logical cell to obtain the preliminary logical cell text content; Use the preliminary logical cell text content to extract the logical cells across columns in the final table structure, triggering a consistency verification mechanism; Based on the semantic weighted edit distance algorithm (SWED algorithm), compare the contents of the table header logical cell and its sub-logical cells below for consistency, and obtain the semantic consistency score result; Use the semantic consistency score result to judge and correct the recognized text to obtain the target extracted logical cell text content information.
[0083] It needs to be explained that the application introduces a cross-cell content verification mechanism based on semantic-weighted edit distance (SWED) for logical verification of the recognition results of merged cells, thereby improving the content extraction accuracy under complex table headers. The specific implementation is as follows: 1) Preliminary content extraction: complete table structure analysis and content OCR of each logical cell.
[0084] 2) Trigger verification mechanism: automatically identify merged cells in the structure (especially cross-tables as table headers).
[0085] 3) Cross-cell content verification: For a table header cell spanning multiple columns, the contents of multiple sub-column cells below it are logically consistent with the contents of the table header cell. This verification is done by the SWED algorithm.
[0086] 4) Semantic weighting: Traditional edit distance algorithms treat all characters equally. The SWED algorithm classifies characters by semantics (e.g. Chinese characters, English letters, numbers, punctuation) before calculating the distance. When comparing, the replacement between different categories of characters will result in a very high cost. For example, matching the Chinese character gold in the table header with the number 8 in the sub-column content will result in a much higher cost than matching the amount with the amount.
[0087] 5) Consistency score: By calculating the SWED score between the table header text and the first row text of each sub-column, the system can determine whether the table header content recognized by OCR is consistent with the data below it in terms of semantic categories. For example, the name table header should be text, and the quantity table header should be numbers.
[0088] In order to facilitate the understanding of the above technical solutions of the application, the following will explain the PDF document intelligent recognition and content extraction based on deep learning in the actual process of the application.
[0089] I. Input file: Scan PDF standard file as input image.
[0090] II. Table detection: First, accurately identify all table regions in the table through the deep learning target detection cell merging network, and mark them with a bounding box.
[0091] III. Structure Analysis: For each identified table region, further analyze its internal structure. For example, there might be a total cell that combines data from multiple months, or a detail cell that splits a single item into multiple sub-items. The cell merging network of the present application can accurately identify these merged or split cells and restore the true logical structure of the table. Even if the table lines are incomplete or there are interfering lines, the cell merging network can accurately determine the boundaries of rows and columns through learned features.
[0092] IV. Content Extraction: For each structured cell, call the OCR engine to recognize the text and numerical content within.
[0093] V. Data Output: Finally, output the recognized structured data in HTML format, which contains the row, column information of the table, the coordinates, content of each cell, and its belonging logical row and column, and even can mark the cross-row and cross-column information of merged cells.
[0094] In summary, with the above technical solutions of the present application, the present application realizes a fully automated process from table recognition to content extraction, greatly reducing manual intervention and improving work efficiency. It is suitable for financial statements, contracts, invoices, questionnaires and other PDF documents containing complex tables. Its wide applicability and automation characteristics not only provide great convenience for data analysis and information management, but also greatly improve the intelligent level of industry workflow, reduce errors and time cost of manual operation. The present application converts the cell merging problem into a classification task based on global context, uses the self-attention mechanism of Transformer for intelligent reasoning, and gets rid of the dependence on physical lines. Thus, the cell merging network can consider global information such as table layout, column division lines, table header content, etc., to accurately determine the position of merged cells. Through this strategy of first structure and then classification, fine content recognition can be achieved, greatly improving the efficiency and accuracy of table data extraction.
[0095] The above only describes the preferred embodiments of the present application and is not intended to limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.
Claims
1. A method for intelligent recognition and content extraction of PDF documents based on deep learning, characterized in that: include: Use the table detection network to identify the table area of the input PDF full-page image and obtain the localized table area of each table in the PDF full-page image; Perform row and column structure analysis on the positioning table area, and use the cell merging network to predict the dividing lines of the row and column structure to obtain the basic grid structure; Based on the basic grid structure, the cell merging network is used to merge and predict adjacent cells to obtain cells with cross-row or cross-column structures; Use the structure verification network to perform consistency detection and repair on cells with cross-row or cross-column structures to obtain the repaired table structure; Text recognition is performed on each logical cell in the repaired table structure, and the row and column position information corresponding to each logical cell is bound to obtain table content that can be output in a preset structured format.
2. The method for intelligent identification and content extraction of PDF documents based on deep learning according to claim 1, characterized in that: The method of using a table detection network to identify a table area on an input PDF full-page image to obtain a localized table area of each table in the PDF full-page image includes: The table detection network is used to process the input PDF full-page image. The backbone network, neck network and head network are used to extract and fuse the multi-level visual features of the PDF full-page image to obtain the feature map of the PDF full-page image. Based on the feature map, the table bounding box is predicted for the entire PDF page image to obtain the bounding box of the candidate table area as a preliminary estimate of the table location; Based on the smooth adaptive intersection-over-union loss function, the bounding box of the candidate table area is optimized to obtain the target table positioning box; Based on the target table positioning frame, the PDF full-page image is cropped to obtain the positioning table area of each table in the PDF full-page image.
3. The method for intelligent identification and content extraction of PDF documents based on deep learning according to claim 2 is characterized in that: The method of performing table bounding box prediction on the PDF full page image based on the feature map to obtain the bounding box of the candidate table area as a preliminary estimate of the table location includes: A scale confidence prediction branch is introduced into the output of the table detection network, and the scale confidence is jointly trained using a smoothed adaptive intersection-over-union loss function to obtain a scale confidence score. Based on the scale confidence score, a dynamic threshold function is constructed and the detection results of different types of tables are screened; The dynamic threshold function is used to adaptively judge the predicted bounding box to obtain the bounding box of the candidate table area as a preliminary estimate of the table location.
4. The method for intelligent identification and content extraction of PDF documents based on deep learning according to claim 2, characterized in that: The optimization of the bounding box of the candidate table area based on the smooth adaptive intersection-over-union loss function to obtain the target table positioning box includes: The smooth adaptive intersection-over-union loss function is used to penalize the distance between the center point of the predicted box and the true box to obtain the initial optimization direction; Based on the initial optimization direction, the distance between the corner points of the predicted box and the real box is adjusted to obtain an optimized predicted box that meets the actual boundaries of the table; The smooth adaptive intersection-over-union loss function is used to comprehensively measure the optimized prediction box to obtain the target table positioning box.
5. The method for intelligent identification and content extraction of PDF documents based on deep learning according to claim 1, characterized in that: The row and column structure analysis is performed on the interior of the positioning table area, and the dividing lines of the row and column structure are predicted using the cell merging network to obtain the basic grid structure, including: Use convolutional neural networks to extract visual features from table area images and obtain a two-dimensional feature map containing table row and column layout information; Perform pooling operations in the horizontal and vertical directions on the two-dimensional feature map to obtain a one-dimensional feature sequence representing row and column information; Using several preset structural encoders to perform global context analysis on the row and column one-dimensional feature sequences, a structural representation sequence containing global information is obtained; Based on the structure representation sequence, a binary classification prediction is performed on each position to obtain the logical row and column dividing line position, thereby realizing the construction of a preliminary basic grid structure.
6. The method for intelligent identification and content extraction of PDF documents based on deep learning according to claim 5, characterized in that: The structure representation sequence is based on which a binary classification prediction is performed on each position to obtain the logical row and column dividing line position, thereby realizing the construction of a preliminary basic grid structure. The steps include: Based on the obtained structure representation sequence, a binary classification prediction is performed on each position to obtain a probability sequence representing the confidence of the segmentation line; The probability sequence of the segmentation line confidence is used to construct a cost graph with segmentation points as nodes and probability as the basis to obtain the global cost structure; Using the preset structural prior knowledge, we define the path transfer cost function and obtain a complete path scoring mechanism; Based on the constructed cost graph and path scoring mechanism, a dynamic programming algorithm is used to solve all candidate paths and obtain the path with the minimum total cost; The preliminary segmentation line is globally optimized using the total cost minimum path to obtain the logical row and column segmentation line positions with continuous structure and reasonable distribution, thus realizing the construction of the preliminary basic grid structure.
7. The method for intelligent identification and content extraction of PDF documents based on deep learning according to claim 1, characterized in that: The method of using a cell merging network to merge and predict adjacent cells based on the basic grid structure to obtain cells with a cross-row or cross-column structure includes: Use convolutional neural networks to extract the visual features of each cell and combine the normalized coordinates and text embedding in the convolutional neural network to obtain a multi-dimensional feature vector; Based on the multi-dimensional feature vector, all cells are arranged into a sequence in row and column order to obtain the cell feature sequence input to the cell merging network; The self-attention mechanism of the cell merging network encoder is used to process the cell feature sequence, calculate the correlation between any two cells, and learn the non-local merging logic; Based on the feature representation output by the cell merging network encoder, the classification head is used to merge and predict each cell with the cell to its right or below to obtain a cell with a cross-row or cross-column structure.
8. The method for intelligent identification and content extraction of PDF documents based on deep learning according to claim 1, characterized in that: The method of using the structure check network to perform consistency detection and repair on cells with cross-row or cross-column structures to obtain a repaired table structure includes: The structure verification network is used to check the structure of the merged cells to obtain the cell combinations with potential logical conflicts; Based on the logical conflict area output by the structure verification network, perform editing actions on the logical conflict area and obtain the target structure rationality reward; The target structure rationality reward is used to optimize the strategies of all editing actions, and a repaired table structure with logical continuity and accurate cross-row and cross-column information is obtained.
9. The method for intelligent identification and content extraction of PDF documents based on deep learning according to claim 1, characterized in that: The text recognition is performed on each logical cell in the repaired table structure, and the row and column position information corresponding to each logical cell is bound to obtain table content that can be output in a preset structured format, including: Classify each logical cell image slice in the repaired table structure to obtain a category label; Use the category label to call the corresponding optical character recognition strategy to identify the logical cell image and obtain the target extracted logical cell text content information; Based on the text content information of the logical cells and the repaired table structure, and by binding the row and column positions and merging information to each logical cell, the table content that can be output in a preset structured format is obtained.
10. The method for intelligent identification and content extraction of PDF documents based on deep learning according to claim 9, characterized in that: The method of using the category label to call the corresponding optical character recognition strategy to identify the logical cell image and obtain the target extracted logical cell text content information includes: Based on the final table structure, perform optical character recognition on each logical cell to obtain preliminary logical cell text content; Using the preliminary logical cell text content, extract the cross-column logical cells in the final table structure to trigger the consistency check mechanism; Based on the semantic weighted edit distance algorithm, the consistency of the header logical cell and its underlying sub-logical cells is compared to obtain the semantic consistency score result; The semantic consistency scoring results are used to judge and correct the recognized text to obtain the target extracted logical cell text content information.
Citation Information
Patent Citations
System and method for extracting tables for PDF document
CN110516208A
Method for obtaining house layout information and network model training method and device
CN111340938A
End-to-end table restoration method
CN115563936A
Table structure identification
CN117558015A
Knowledge graph and rule constraint combined data intelligent analysis method and system
CN118606440A
Cited By
Complex table recognition method and system
CN121600537A