Information Extraction Method, Apparatus, Device, and Storage Medium

By using the preset network model to segment the table image and identify text lines, the problem of text information extraction in complex tables is solved, and the information extraction effect with high accuracy and efficiency is achieved.

CN114120345BActive Publication Date: 2025-05-30CHINA MOBILE COMM LTD RES INST +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010902717.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-09-01
Publication Date
2025-05-30
Estimated Expiration
2040-09-01

AI Technical Summary

Technical Problem

Accurate extraction of text information in complex tables has become a key issue, and it is difficult for the prior art to effectively handle tables of multiple data storage methods and types.

Method used

The preset network model is used to segment and locate the table images, determine the text lines in the cells, and form table information by identifying the text. The specific steps include: using the first network model to segment the table area, determining text lines in combination with the second network model, using the third network model to perform text recognition, and finally building the table structure and information.

Benefits of technology

It realizes the accurate extraction of text information in complex tables, can effectively process tabular data of different types and structures, and improves the accuracy and efficiency of information extraction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114120345B_ABST
    Figure CN114120345B_ABST
Patent Text Reader

Abstract

The present invention discloses an information extraction method, apparatus, device and storage medium. Among them, the method includes: collecting a table image; using a preset first network model to segment and locate the table area in the table image to obtain at least two cells; for each of the at least two cells, combining a preset second network model to determine the text lines in the corresponding cell; using a preset third network model to respectively recognize the text lines in the at least two cells to obtain recognized text; determining the table structure corresponding to the at least two cells, and forming table information by using the table structure and the recognized text.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing, and in particular, to a method, apparatus, device, and storage medium for information extraction. Background Art

[0002] With the explosion of data in the network, how to extract information from a large amount of data has become crucial. In practical applications, a large amount of data can be stored in a table. As the data stored in the table increases, the structure of the table becomes more and more complex. As the structure of the table becomes more and more complex, the ways of storing data in the table can be various, and the types of data stored in the table are also various. Therefore, how to accurately extract text information from the table has become a key problem. Summary of the Invention

[0003] In view of this, embodiments of the present invention are expected to provide a method, apparatus, device, and storage medium for information extraction.

[0004] The technical solution of the embodiments of the present invention is implemented as follows:

[0005] At least one embodiment of the present invention provides an information extraction method, the method including:

[0006] Collect a table image;

[0007] Use a preset first network model to segment and locate the table area in the table image to obtain at least two cells;

[0008] For each of the at least two cells, combine a preset second network model to determine the text lines in the corresponding cell;

[0009] Use a preset third network model to respectively recognize the text lines in the at least two cells to obtain recognized text;

[0010] Determine the table structure corresponding to the at least two cells, and use the table structure and the recognized text to form table information.

[0011] In addition, according to at least one embodiment of the present invention, the using a preset first network model to segment and locate the table area in the table image to obtain at least two cells includes:

[0012] Use the table image as the input of the preset first network model, perform mapping from input to output on the table image to obtain the feature map and feature map information of the table area in the table image; the feature map information represents the line type corresponding to each feature point in the table area;

[0013] Using the feature map information, determine the coordinates of multiple feature points corresponding to at least two types of line segments from the feature map; and form at least two cells using the coordinates of the multiple feature points.

[0014] Select at least two cells that meet the first preset condition from the at least two formed cells.

[0015] In addition, according to at least one embodiment of the present invention, the selecting at least two cells that meet the first preset condition from the at least two formed cells includes:

[0016] For each of the at least two cells, determine whether the height of the corresponding cell is less than or equal to the height threshold and whether the length of the corresponding cell is less than or equal to the length threshold.

[0017] When it is determined that the height of the corresponding cell is less than or equal to the height threshold and the length of the corresponding cell is less than or equal to the length threshold, discard the corresponding cell.

[0018] Use the remaining at least two cells among the at least two cells as the at least two cells that meet the first preset condition.

[0019] In addition, according to at least one embodiment of the present invention, the determining the text line in the corresponding cell for each of the at least two cells in combination with a preset second network model includes:

[0020] For each of the at least two cells, determine at least two first text boxes included in the corresponding cell in combination with a preset second network model.

[0021] Select at least two second text boxes that meet the second preset condition from the at least two first text boxes.

[0022] Stitch the text within the at least two second text boxes to obtain the text line.

[0023] In addition, according to at least one embodiment of the present invention, the selecting at least two second text boxes that meet the second preset condition from the at least two first text boxes includes:

[0024] Perform horizontal sorting on the at least two first text boxes within the corresponding cell to obtain the sorted at least two first text boxes.

[0025] For the i-th first text box among the sorted at least two first text boxes, search for the j-th text box whose overlapping height with the i-th text box meets the second preset condition in the horizontal positive direction; and search for the k-th text box whose overlapping height with the j-th text box meets the second preset condition in the horizontal reverse direction.

[0026] Determine a first horizontal distance between the first i text boxes and the j-th text box; and determine a second horizontal distance between the j-th text box and the k-th text box;

[0027] When the first horizontal distance is greater than or equal to the second horizontal distance, splice the text in at least two text boxes between the i-th text box and the j-th text box to obtain a text line.

[0028] In addition, according to at least one embodiment of the present invention, searching for the j-th text box whose coincidence height with the i-th text box satisfies a second preset condition includes:

[0029] Search for at least one second text box whose horizontal distance from the i-th text box is greater than or equal to a distance threshold;

[0030] Calculate the coincidence height between the at least one second text box and the i-th text box to obtain at least one coincidence height;

[0031] Use the second text box corresponding to the maximum coincidence height among the at least one coincidence height as the j-th text box that satisfies the second preset condition.

[0032] In addition, according to at least one embodiment of the present invention, determining the table structure corresponding to the at least two cells includes:

[0033] Determine at least two first cells located in a reference row and at least two second cells located in a reference column from the at least two cells;

[0034] Determine a plurality of cells having a subordinate relationship with the at least two first cells; and determine a plurality of cells having a subordinate relationship with the at least two second cells;

[0035] Based on the determined plurality of cells having a subordinate relationship, construct a tree structure;

[0036] Use the tree structure as the table structure of the at least two cells.

[0037] In addition, according to at least one embodiment of the present invention, constructing a tree structure based on the determined plurality of cells having a subordinate relationship includes:

[0038] Use the plurality of cells having a subordinate relationship with the at least two first cells to construct a tree structure in a first direction;

[0039] Use the plurality of cells having a subordinate relationship with the at least two second cells to construct a tree structure in a second direction;

[0040] Wherein, the first direction is different from the second direction.

[0041] At least one embodiment of the present invention provides an information extraction device, including:

[0042] An acquisition unit for acquiring a table image;

[0043] A first processing unit for segmenting and positioning a table area in the table image by using a preset first network model to obtain at least two cells;

[0044] A second processing unit for determining a text line in each of the at least two cells in combination with a preset second network model;

[0045] A third processing unit for respectively identifying the text lines in the at least two cells by using a preset third network model to obtain recognized text;

[0046] A fourth processing unit for determining a table structure corresponding to the at least two cells and forming table information by using the table structure and the recognized text.

[0047] In addition, according to at least one embodiment of the present invention, the first processing unit is specifically configured to:

[0048] Take the table image as an input of a preset first network model, perform a mapping from input to output on the table image to obtain a feature map and feature map information of the table area in the table image; the feature map information represents the line segment type corresponding to each feature point in the table area;

[0049] Use the feature map information to determine the coordinates of multiple feature points corresponding to at least two line segment types from the feature map; and form at least two cells by using the coordinates of the multiple feature points;

[0050] Select at least two cells that meet a first preset condition from the formed at least two cells.

[0051] In addition, according to at least one embodiment of the present invention, the first processing unit is specifically configured to:

[0052] For each of the at least two cells, determine whether the height of the corresponding cell is less than or equal to a height threshold and whether the length of the corresponding cell is less than or equal to a length threshold;

[0053] When it is determined that the height of the corresponding cell is less than or equal to the height threshold and the length of the corresponding cell is less than or equal to the length threshold, discard the corresponding cell;

[0054] Use the remaining at least two cells among the at least two cells as the at least two cells that meet the first preset condition.

[0055] In addition, according to at least one embodiment of the present invention, the second processing unit is specifically configured to:

[0056] For each of the at least two cells, in combination with a preset second network model, determine at least two first text boxes included in the corresponding cell;

[0057] Select at least two second text boxes that meet the second preset condition from the at least two first text boxes;

[0058] Concatenate the texts in the at least two second text boxes to obtain a text line.

[0059] In addition, according to at least one embodiment of the present invention, the second processing unit is specifically configured to:

[0060] Perform horizontal sorting on at least two first text boxes within the corresponding cell to obtain at least two sorted first text boxes;

[0061] For the i-th first text box among the at least two sorted first text boxes, search in the horizontal positive direction for the j-th text box whose overlapping height with the i-th text box meets the second preset condition; and search in the horizontal reverse direction for the k-th text box whose overlapping height with the j-th text box meets the second preset condition;

[0062] Determine the first horizontal distance between the i-th text box and the j-th text box; and determine the second horizontal distance between the j-th text box and the k-th text box;

[0063] When the first horizontal distance is greater than or equal to the second horizontal distance, concatenate the texts in at least two text boxes between the i-th text box and the j-th text box to obtain a text line.

[0064] In addition, according to at least one embodiment of the present invention, the second processing unit is specifically configured to:

[0065] Search for at least one second text box whose horizontal distance from the i-th text box is greater than or equal to a distance threshold;

[0066] Calculate the overlapping height between the at least one second text box and the i-th text box to obtain at least one overlapping height;

[0067] Use the second text box corresponding to the maximum overlapping height among the at least one overlapping height as the j-th text box that meets the second preset condition.

[0068] In addition, according to at least one embodiment of the present invention, the fourth processing unit is specifically configured to:

[0069] Determine at least two first cells located in a reference row and at least two second cells located in a reference column from the at least two cells;

[0070] Determine a plurality of cells having a subordinate relationship with the at least two first cells; and determine a plurality of cells having a subordinate relationship with the at least two second cells;

[0071] Construct a tree structure based on the determined plurality of cells having a subordinate relationship;

[0072] Use the tree structure as the table structure of the at least two cells.

[0073] In addition, according to at least one embodiment of the present invention, the fourth processing unit is specifically configured to:

[0074] Construct a tree structure in a first direction by using a plurality of cells having a subordinate relationship with the at least two first cells;

[0075] Construct a tree structure in a second direction by using a plurality of cells having a subordinate relationship with the at least two second cells;

[0076] Wherein, the first direction is different from the second direction.

[0077] At least one embodiment of the present invention provides an electronic device, including:

[0078] A communication interface for collecting a table image;

[0079] A processor for using a preset first network model to segment and locate a table area in the table image to obtain at least two cells; for each cell in the at least two cells, combining a preset second network model to determine a text line in the corresponding cell; and using a preset third network model to respectively identify the text lines in the at least two cells to obtain identified text; determining a table structure corresponding to the at least two cells, and forming table information by using the table structure and the identified text.

[0080] At least one embodiment of the present invention provides an electronic device, including a processor and a memory for storing a computer program that can run on the processor,

[0081] Wherein, when the processor is used to run the computer program, it executes the steps of any of the above methods.

[0082] At least one embodiment of the present invention provides a storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of any of the above methods.

[0083] The information extraction method, device, equipment and storage medium provided by the embodiments of the present invention collect a table image; use a preset first network model to segment and locate the table area in the table image to obtain at least two cells; for each of the at least two cells, combine a preset second network model to determine the text lines in the corresponding cell; use a preset third network model to respectively recognize the text lines in the at least two cells to obtain recognized text; determine the table structure corresponding to the at least two cells, and use the table structure and the recognized text to form table information. By adopting the technical solution of the embodiments of the present invention, the table structure formed by at least two cells in the table area of the table image is extracted, and the text line information corresponding to each of the at least two cells is extracted. In this way, table information is formed based on the table structure and the text line information. Since the table information can be formed by combining the structure and text information of the table, the table information can be accurately extracted. BRIEF DESCRIPTION OF THE DRAWINGS

[0084] Figure 1 is a schematic flowchart of the implementation of the information extraction method according to the embodiments of the present invention;

[0085] Figure 2 is a schematic diagram of the feature map corresponding to the table area in the table image according to the embodiments of the present invention;

[0086] Figure 3 is a schematic diagram of the cells in the table area of the table image according to the embodiments of the present invention;

[0087] Figure 4 is a schematic flowchart of the implementation of determining at least two cells in the table image according to the embodiments of the present invention;

[0088] Figure 5 is a schematic diagram of the text box in the cell of the table image according to the embodiments of the present invention;

[0089] Figure 6 is a schematic flowchart of the implementation of splicing the text in the cell of the table image to obtain a text line according to the embodiments of the present invention;

[0090] Figure 7 is a schematic diagram of the table structure in the table image according to the embodiments of the present invention Figure 1 ;

[0091] Figure 8 is a schematic diagram of the table structure in the table image according to the embodiments of the present invention Figure 2 ;

[0092] Figure 9 is a schematic flowchart of the implementation of determining the table structure corresponding to at least two cells in the table image according to the embodiments of the present invention;

[0093] Figure 10It is a schematic diagram of the composition structure of the information extraction device according to an embodiment of the present invention;

[0094] Figure 11 It is a schematic diagram of the composition structure of the electronic device according to an embodiment of the present invention; Detailed implementation manners

[0095] Before introducing the technical solution of the embodiment of the present invention, the related technologies will be described first.

[0096] In the related technologies, the discriminator network D-Net of the semantic segmentation network UNet and the generative adversarial network GAN can be used to perform layout and text line analysis on the PDF using distance relationships and perform text recognition, so that the recognized text can carry structural information and restore the original layout structure of the PDF. However, the defect is that it is not applicable enough to table images such as various bills, because the structural information of the text in the table not only exists in the paragraphs adjacent to each other in the overall layout, but also various structural information exists in the paragraphs far apart. Since the table can be regarded as a two-dimensional matrix layout, it is not possible to simply use distance relationships to analyze the layout. In the related technologies, one of the Faster RCNN model, CTPN model, SegLink model, and EAST model can be used to detect text lines for business licenses, and the DenseNet+CTC text recognition model is used to recognize the detected text lines to obtain the recognized text. However, the defect is that there is a lack of analysis of the layout. Therefore, for complex license and certificate tables, model training needs to be carried out again.

[0097] Based on this, in various embodiments of the present invention, a table image is collected; a preset first network model is used to segment and locate the table area in the table image to obtain at least two cells; for each of the at least two cells, in combination with a preset second network model, the text lines in the corresponding cell are determined; a preset third network model is used to recognize the text lines in the at least two cells respectively to obtain the recognized text; the table structure corresponding to the at least two cells is determined, and table information is formed by using the table structure and the recognized text.

[0098] The present invention will be further described in detail below with reference to the drawings and embodiments.

[0099] The embodiment of the present invention provides an information extraction method, as Figure 1 shown, the method includes:

[0100] Step 101: Collect a table image;

[0101] Step 102: Use a preset first network model to segment and locate the table area in the table image to obtain at least two cells;

[0102] Step 103: For each of the at least two cells, in combination with a preset second network model, determine the text lines in the corresponding cell;

[0103] Step 104: Use a preset third network model to respectively recognize the text lines in the at least two cells to obtain recognized texts;

[0104] Step 105: Determine the table structure corresponding to the at least two cells, and use the table structure and the recognized texts to form table information.

[0105] Here, in Step 101, the table image may refer to an image containing a table. In practical applications, the table image can be obtained by photographing a document containing a table. Specifically, it can be table images corresponding to various bills, such as airplane tickets, train tickets, and so on.

[0106] Here, in Step 102, the first network model may specifically be a U-Net. The segmentation and positioning of the table area in the table image may refer to determining the coordinate points of the table area in the table image that belong to at least two types of line segments, and determining at least two cells based on the multiple coordinate points.

[0107] Here, in Step 103, the second network model may specifically be a Connectionist Text Proposal Network (CTPN). By using the second network model to determine the text lines in the corresponding cell, it can avoid the problem that the distance between the characters in the same row in a cell is too far and is misrecognized as not being in the same text line and thus cannot be merged, and also avoid the problem that the distance between the characters in different rows in a cell is too close and is misrecognized as not being in the same text line and thus cannot be merged.

[0108] Here, in Step 104, the third network model may specifically be a Convolutional Recurrent Neural Network (CRNN), which is an end-to-end text recognition network. The recognized texts may include Chinese characters, numbers, letters, and so on.

[0109] Here, in Step 105, using the table structure and the recognized texts to form table information can list the text information with dependencies in the table, and subsequent data analysis and so on can be performed using the text information with dependencies.

[0110] In practical applications, a cell can be formed by a closed area composed of 2 straight line segments and 2 vertical line segments, or can be formed by a closed area composed of 2 straight line segments, 2 vertical line segments and 1 oblique line segment; since the closed area is composed of 4 vertices, therefore, the table area in the table image can be segmented by positioning the vertices of the cell to obtain multiple cells.

[0111] Based on this, in one embodiment, the using a preset first network model to segment and locate the table area in the table image to obtain at least two cells includes:

[0112] Taking the table image as the input of the preset first network model, performing mapping from input to output on the table image to obtain the feature map and feature map information of the table area in the table image; the feature map information characterizes the line segment type corresponding to each feature point in the table area;

[0113] Using the feature map information, determining the coordinates of multiple feature points corresponding to at least two line segment types from the feature map; and forming at least two cells by using the coordinates of the multiple feature points;

[0114] Selecting at least two cells that meet the first preset condition from the at least two formed cells.

[0115] Here, the first network model may refer to a U-Net model; the line segment type may refer to a straight line segment, a vertical line segment, an oblique line segment, etc.

[0116] Here, the process of training the U-Net model may include: using various collected table pictures, using the method of manual annotation to mark the areas of the inner and outer borders in the table, and taking the borders belonging to the horizontal line segments as one category, the borders belonging to the vertical line segments as one category, and the borders belonging to the oblique line segments or other line segments as one category to obtain three categories, thereby obtaining the border structure of the table. The obtained border structure of the table can be used as the initial training set, and a trained U-Net model can be obtained by performing data augmentation methods such as adding noise to the training set.

[0117] Here, using the trained U-Net to segment and locate each area of the input table image. For example, inputting a table image with a width and height of W×H into the U-Net model, the U-Net model outputs a feature map with a size of W_out×H_out, and the feature map information is 3×W_out×H_out; among them, according to the feature map information, the line segment type corresponding to each feature point in the feature map can be determined, that is, the category probability of each feature point in the feature map belonging to the horizontal line, vertical line, oblique line or other line segments of the cell. In practical applications, when establishing the coordinate system of the feature map, the lower left edge of the table image can be used as the coordinate origin.

[0118] For example, as Figure 2 shown, assume that the feature map information includes 7 feature points, denoted as feature point 1, feature point 2, feature point 3, feature point 4, feature point 5, feature point 6, and feature point 7. The line segment types corresponding to feature point 1 are a straight line segment and a vertical line segment, that is, feature point 1 belongs to the intersection of the straight line segment and the vertical line segment; the line segment types corresponding to feature point 2 are a straight line segment and a vertical line segment, that is, feature point 2 belongs to the intersection of the straight line segment and the vertical line segment; the line segment types corresponding to feature point 3 are a straight line segment and a vertical line segment, that is, feature point 3 belongs to the intersection of the straight line segment and the vertical line segment; the line segment types corresponding to feature point 4, feature point 5, feature point 6, and feature point 7 are all straight line segments, that is, feature point 4, feature point 5, feature point 6, and feature point 7 belong to the points on the straight line segment. In this way, the two-dimensional coordinates of feature point 1, feature point 2, and feature point 3 can be used as the vertices of the cell, and combined with other coordinate points in the feature map to form a closed area to obtain a cell. The two-dimensional coordinates can refer to the x-axis and y-axis coordinates.

[0119] Here, if the collected table image is tilted, the tilt angle of the table image needs to be corrected. Among them, the tilt angle can refer to the tilt angle between the arrangement of the text in the table image and the reference; the reference refers to the arrangement order of the text in the table image being horizontally arranged from left to right.

[0120] Here, the methods for rotating the table image to correct the tilt angle include the following three:

[0121] The first method is to use the adaptive threshold binarization technique in the traditional open-source OpenCV, that is, first binarize the background and text of the entire table image; then rotate according to the gradient of the binarized table image to correct the tilt angle;

[0122] The second method is to use the classification method of the CNN network in deep learning, that is, regard the images with different text tilt angles as different categories, such as eight categories of 0, 45, 90, 135, 180, 225, 270, and 315 degrees, etc., train an 8-category classification network with a small network such as MobileNet, and perform rotation tilt correction according to the classification result;

[0123] The third method is to first collect images of the reference horizontal angles of different categories of charts, use the SIFT or SURF operator in OpenCV to match the charts between the images, and perform rotation tilt correction on the images according to the matching result.

[0124] It should be noted that if the background of the table image is not particularly complex, the first method can be used; if the inclination angle of the table image is relatively fixed, the second method can be adopted; otherwise, the third method is used.

[0125] In practical applications, considering that a cell can have one or multiple borders. For example, a cell includes an inner border and an outer border. To avoid recognizing the enclosed area formed by the inner border as a cell, the enclosed area formed by the outer border can be recognized and used as the cell.

[0126] Based on this, in one embodiment, selecting at least two cells that meet the first preset condition from the at least two formed cells includes:

[0127] For each cell in the at least two cells, determine whether the height of the corresponding cell is less than or equal to the height threshold and whether the length of the corresponding cell is less than or equal to the length threshold;

[0128] When it is determined that the height of the corresponding cell is less than or equal to the height threshold and the length of the corresponding cell is less than or equal to the length threshold, discard the corresponding cell;

[0129] Use the remaining at least two cells in the at least two cells as the at least two cells that meet the first preset condition.

[0130] Here, according to the intersection relationship of horizontal line segments, vertical line segments, and oblique line segments, the coordinates of the cell vertices are obtained by dividing the table area in the table image, and after constructing the cells based on the vertex coordinates, if the width and height of the cell are too small, the cell is discarded. For example, the cell is formed by the inner border, or the border of the cell cannot cover the text.

[0131] For example, as Figure 3 shown, feature point 1, feature point 2, and feature point 3 construct a cell, denoted as cell 1, and feature point 4, feature point 5, and feature point 6 construct a cell, denoted as cell 2. Assuming the height threshold is 8 pixels and the length threshold is 8 pixels, the height of cell 1 is 7 pixels and the width is 7 pixels, and the height of cell 2 is 9 pixels and the width is 9 pixels. Then cell 1 is a cell formed by the inner border, cell 2 is a cell formed by the outer border, discard cell 1, and use cell 2 as the cell that meets the preset condition.

[0132] Here, when screening the cells after determining the cell positions, prior knowledge can also be used to exclude cells with heights and widths smaller than the minimum text box anchor used in text detection and recognition.

[0133] In one example, asFigure 4 As shown, a process for determining at least two cells in a table image is described, including:

[0134] Step 401: Collect the table image; use the table image as the input of a preset first network model, perform mapping from input to output on the table image, and obtain the feature map and feature map information of the table area in the table image;

[0135] Wherein, the feature map information characterizes the line segment type corresponding to each feature point in the table area.

[0136] Step 402: Use the feature map information to determine the coordinates of multiple feature points corresponding to at least two line segment types from the feature map; and use the coordinates of the multiple feature points to form at least two cells.

[0137] Step 403: For each of the at least two cells, determine whether the height of the corresponding cell is less than or equal to the height threshold and whether the length of the corresponding cell is less than or equal to the length threshold; when it is determined that the height of the corresponding cell is less than or equal to the height threshold and the length of the corresponding cell is less than or equal to the length threshold, execute Step 404;

[0138] Step 404: Discard the corresponding cell; and use the remaining at least two cells among the at least two cells as at least two cells that meet the first preset condition.

[0139] Here, determining at least two cells in the table image has the following advantages:

[0140] (1) Use a preset first network model to segment and locate each area of the table in the table image to extract the vertex information of the table border, and construct cells based on the extracted vertex information. Specifically, when using a preset first network model such as U-Net to extract each area of the table in the image, horizontal lines, vertical lines, diagonal lines or other line segments are classified as one type respectively. In this way, the image feature information extracted by U-Net contains 3 types of line segments, and the vertex positions of the cell borders are determined according to the closed areas formed by these three types of image line segments; among them, the intersection points of different types of line segments are the vertices of the cells.

[0141] (2) For cells with multiple borders, by extracting the edge information corresponding to the inner and outer borders of the table, the cells formed by the inner borders are excluded, and then the text information of the cells formed by the outer borders is extracted, improving the information extraction efficiency. In addition, after determining the cell positions, when screening the cells, prior knowledge can be used to ensure that the height and width of the cells cannot be less than the size of the smallest text detection text box anchor, that is, cells with a height and width less than the smallest text box anchor will be excluded.

[0142] (3) Using a preset first network model such as the U-Net model can better locate the position of the basic unit in the table, that is, the cell, which provides a reference for subsequent text detection and word recognition for semantic analysis such as better sentence segmentation.

[0143] In actual application, in order to accurately extract all the text in a cell, multiple text boxes with different heights and widths can be used to align with the text in the table area of the table image first. Then, the text boxes that are not within a single cell are excluded. In this way, multiple text boxes included in a cell are obtained. Considering that the sizes of multiple text boxes included in a cell can be different and their distribution positions can be different, in order to avoid merging text that is not in the same row into a text line, multiple text boxes that can make multiple texts in the same row can be selected from multiple text boxes, and the texts in the selected multiple text boxes are spliced to obtain a text line.

[0144] Based on this, in one embodiment, for each of the at least two cells, in combination with a preset second network model, determining the text line in the corresponding cell includes:

[0145] For each of the at least two cells, in combination with a preset second network model, determining at least two first text boxes included in the corresponding cell;

[0146] Selecting at least two second text boxes that meet the second preset condition from the at least two first text boxes;

[0147] Splicing the texts in the at least two second text boxes to obtain a text line.

[0148] Here, the preset second network model may refer to the CTPN network model. The text line means that the positions of multiple texts are in the same row.

[0149] Here, considering that the texts in at least two first text boxes included in the corresponding cell may not correspond to a single text line and may correspond to multiple text lines, in this case, at least two second text boxes that meet the second preset condition can be selected from the at least two first text boxes; the texts in the at least two second text boxes are spliced to obtain a text line.

[0150] Here, taking the case where the text in the table image is arranged horizontally as an example, first, a large number of table images are collected. Using the k-means algorithm, the sizes of all texts in each table image are counted, and the sizes of the text boxes are determined according to the sizes of the texts. For example, the text box is represented by anchor. The horizontal width of the anchor is 8 pixels, and the vertical height is: 8 pixels, 11 pixels, 16 pixels, 23 pixels, 33 pixels, 48 pixels, 68 pixels, 97 pixels, 139 pixels, 198 pixels, a total of 10 anchors. Then, the softmax classification of the RPN network is used to determine the text in the table image. Finally, using Boundingbox regression, the central coordinate y and height of the text are determined, and combined with the sizes of the preset 10 text boxes, through regression calculation, the text boxes are aligned with the text, so as to determine multiple text boxes.

[0151] Here, after aligning the text boxes with the text, in order to determine the text boxes within a cell, the following constraint conditions can be used to eliminate the text boxes that are not within the same cell from the multiple text boxes aligned with the text. The constraint conditions specifically include:

[0152] The first constraint condition, that is, the length of the text line does not exceed the width of the area of a single cell in the table, that is, the terminator is within the cell;

[0153] The second constraint condition, that is, the text boxes with a large interval between single characters within the same cell are also merged;

[0154] The third constraint condition, that is, the text lines with detected cross-cell dividing lines are filtered. Here, the judgment of cross-cell is based on whether the center point coordinates of the rectangle frame of the detected text line exceed the cell boundary coordinates as whether the text line crosses the cell;

[0155] The fourth constraint condition, that is, the cells where the height of the text line covers the entire height of the cell are deleted. That is, the lowest y coordinate of the text line is less than the lowest y coordinate of the cell and the highest y coordinate of the text line is greater than the highest y coordinate of the cell.

[0156] It should be noted that through the above four constraint conditions, it can be ensured that the subsequent obtained text lines are covered within a cell, and there will be no problem of crossing cells, nor will there be a problem of missing content in the same cell.

[0157] In practical applications, considering that the multiple texts in the text line can be far apart or close to each other, therefore, for each text box among the multiple text boxes of a cell, multiple text boxes that are far away from the corresponding text box can be determined. In this way, it is possible to determine all the texts included in a text line with the highest probability.

[0158] Based on this, in one embodiment, selecting at least two second text boxes that meet the second preset condition from the at least two first text boxes includes:

[0159] Performing horizontal sorting on the at least two first text boxes in the corresponding cell to obtain the sorted at least two first text boxes;

[0160] For the i-th first text box among the sorted at least two first text boxes, search for the j-th text box whose overlapping height with the i-th text box meets the second preset condition in the horizontal positive direction; and search for the k-th text box whose overlapping height with the j-th text box meets the second preset condition in the horizontal reverse direction;

[0161] Determine the first horizontal distance between the i-th text box and the j-th text box; and determine the second horizontal distance between the j-th text box and the k-th text box;

[0162] When the first horizontal distance is greater than or equal to the second horizontal distance, splice the texts in the at least two text boxes between the i-th text box and the j-th text box to obtain a text line.

[0163] Here, performing horizontal sorting on the at least two first text boxes in the corresponding cell may refer to performing horizontal sorting according to the center coordinates of the at least two first text boxes.

[0164] In practical applications, in order to accurately extract a text line, for each text box among multiple text boxes in a cell, when determining multiple text boxes that are relatively far from the corresponding text box, select the text box with the most overlapping areas among the text boxes that are relatively far away.

[0165] Based on this, in one embodiment, searching for the j-th text box whose overlapping height with the i-th text box meets the second preset condition includes:

[0166] Search for at least one second text box whose horizontal distance from the i-th text box is greater than or equal to the distance threshold;

[0167] Calculate the overlapping height between the at least one second text box and the i-th text box to obtain at least one overlapping height;

[0168] Take the second text box corresponding to the maximum overlapping height among the at least one overlapping height as the j-th text box that meets the second preset condition.

[0169] Here, the CTPN algorithm, which is a text line construction algorithm, can be used to splice the texts in the cell to obtain a text line. The specific implementation process may include:

[0170] Step 1: Horizontally sort multiple text boxes (anchor boxes) within a cell according to the x-axis coordinate;

[0171] Step 2: For each anchor box, perform a forward search, that is, along the positive horizontal direction, find a series of candidate anchors within the same cell that have the largest possible horizontal distance from the anchor_i; from the candidate anchors, select the anchors whose vertical overlap height with anchor_i is overlap > 0.7 to obtain multiple anchors; select the anchor with the largest overlap height from the multiple anchors and denote it as anchor_j. It should be noted that when there are multiple anchors with the largest overlap height, select the anchor with the farthest horizontal distance from anchor_i and denote it as anchor_j. Then, along the negative horizontal direction, find a series of candidate anchors within the same cell that have the largest possible horizontal distance from the anchor_j; from the candidate anchors, select the anchors whose vertical overlap height with anchor_j is overlap > 0.7 to obtain multiple anchors; select the anchor with the largest overlap height from the multiple anchors and denote it as anchor_k. It should be noted that when there are multiple anchors with the largest overlap height, select the anchor with the farthest horizontal distance from anchor_j and denote it as anchor_k.

[0172] Step 3: Denote the distance from text box anchor_i to text box anchor_j as score_ij, and the distance from text box anchor_j to text box anchor_k as score_jk; compare score_ij and score_jk. If score_ij > score_jk, then i and j form a longest connection, and set Graph(i,j) = True, that is, text box anchor_i and text box anchor_j are connected; otherwise, it means that i and j are not the longest connection, that is, this connection must be included in another longer connection.

[0173] Step 4: Merge text lines based on whether the nodes corresponding to the text boxes in the comprehensive feature graph (Graph) are connected and whether text box anchor_i and text box anchor_j are in the same cell.

[0174] For example, such as Figure 5As shown, there are 5 first text boxes in cell 1, represented by text box 1, text box 2, text box 3, text box 4, and text box 5. The 5 first text boxes are sorted horizontally. The text boxes 1, 2, and 3 at the same horizontal position are grouped together, and the text boxes 4 and 5 at the same horizontal position are grouped together. Assume the distance threshold is 1 mm. The horizontal distance between text box 1 and text box 2 is 1 mm, and the horizontal distance between text box 1 and text box 3 is 2 mm. Then, in the horizontal positive direction, for text box 1, text box 2 and text box 3 are searched. Assume the coincidence degree between text box 3 and text box 1 is the largest. Then, in the horizontal reverse direction, for text box 3, text box 1 is searched. Since the horizontal distance between text box 1 and text box 3 is equal to the horizontal distance from text box 3 to text box 1, therefore, the text in text box 1, text box 2, and text box 3 is spliced to obtain a text line, such as "I love China".

[0175] In an example, as Figure 6 shown, the process of splicing the text in the cell of the table image to obtain a text line is described, including:

[0176] Step 601: Horizontally sort at least two first text boxes in the cell to obtain at least two sorted first text boxes;

[0177] Step 602: For the i-th first text box among the at least two sorted first text boxes, search for the j-th text box whose coincidence height with the i-th text box meets the second preset condition in the horizontal positive direction; and search for the k-th text box whose coincidence height with the j-th text box meets the second preset condition in the horizontal reverse direction;

[0178] Step 603: Determine the first horizontal distance between the i-th text box and the j-th text box; and determine the second horizontal distance between the j-th text box and the k-th text box;

[0179] Step 604: When the first horizontal distance is greater than or equal to the second horizontal distance, splice the text in at least two text boxes between the i-th text box and the j-th text box to obtain a text line.

[0180] Here, when splicing the text in the cell of the table image to obtain a text line, the following advantages are available:

[0181] (1) When splicing the text in the cell into a text line, search and connect the longest text boxes in the horizontal positive direction and the horizontal reverse direction, so as to extract the complete text line according to the longest connected text boxes. It can avoid the problem that the texts cannot be merged due to the long distance between the texts in the related technology, accurately extract the text line information, and ensure that the content of the text line is not missing.

[0182] (2) The CTPN text detection algorithm is improved. That is, without conflict between the text detection range and the position of the cell, the text box with the longest connection is searched in the horizontal positive and horizontal reverse directions, thus avoiding the problems of line breaking, interruption, omission, or excessive length of the detected text line in the related art.

[0183] (3) The improved CTPN is used to detect text in the whole table picture. That is, the table image is input into the CTPN network, and the spatial and sequence feature vectors of the table image are obtained by using the feature extraction networks CNN, BiLSTM network, and FC convolution of the CTPN network; the obtained spatial and sequence feature vectors are input into the RPN network in Faster-RCNN to align the text in the table image with the preset text box; and the text box with the longest connection is searched in the horizontal positive and horizontal reverse directions, and the text contained in the text box with the longest connection is spliced to obtain the text line.

[0184] In actual application, considering that text information can be stored in the table in the way of attribute name and attribute value, that is, there is a specific dependency relationship between the texts in the table. For example, for an air ticket table, the attribute name can be: departure place, and the attribute value can be: Beijing.

[0185] Based on this, in one embodiment, the at least two first cells located in the reference row and the at least two second cells located in the reference column are determined from the at least two cells;

[0186] A plurality of cells having a subordinate relationship with the at least two first cells are determined; and a plurality of cells having a subordinate relationship with the at least two second cells are determined;

[0187] Based on the determined plurality of cells having a subordinate relationship, a tree structure is constructed;

[0188] The tree structure is used as the table structure of the at least two cells.

[0189] Here, in the table structure, each node corresponds to each cell in the table, and the attribute information of each node is the text information in each cell.

[0190] In actual application, in order to accurately extract the text information in the table, a tree structure with a subordinate relationship can be constructed for multiple cells in the table. In this way, the corresponding text information can be stored in the corresponding nodes in the tree structure subsequently, improving the accuracy of text information extraction.

[0191] Based on this, in one embodiment, the constructing a tree structure based on the determined plurality of cells having a subordinate relationship includes:

[0192] Construct a tree structure in the first direction by using multiple cells that have a subordinate relationship with the at least two first cells;

[0193] Construct a tree structure in the second direction by using multiple cells that have a subordinate relationship with the at least two second cells;

[0194] Wherein, the first direction is different from the second direction.

[0195] Here, according to the characteristic point coordinates of the table area in the table image, the leftmost column and the uppermost row in the table can be used as the reference column and the reference row, and a multi-way tree structure with a subordinate relationship can be established from left to right and from top to bottom.

[0196] Table 1 is a schematic diagram of the table structure. As shown in Table 1, the cells located in the reference row are the cells in the uppermost row, that is, a - b - c, and the cells located in the reference column are the cells in the leftmost row, that is, a - d. As shown in Table 1, the cells in the uppermost row are a - b - c, which are the first-level child nodes of two trees respectively. Then, from left to right, the parent-child relationship of the tree nodes is established according to the inclusion relationship between the coordinates of the previous cells, and a tree can be obtained, as Figure 7 shown; similarly, a tree from top to bottom can also be obtained, as Figure 8 shown. In the tree structure, if the number of forks of the child nodes of any node is greater than 1, it can be considered that the content of the child nodes has a subordinate relationship with the parent node. If the child node has only one, such as Figure 7 a → b → c under the tree from left to right in, they have a parallel relationship. It should be noted that when constructing the parent-child relationship, for the left-to-right subtree, the height of the child node must be completely included by the parent node to form a parent-child node relationship; when constructing the top-to-bottom subtree, the width of the child node must be completely included by the parent node to form a parent-child node relationship.

[0197]

[0198] Table 1

[0199] In an example, as Figure 9 shown, the process of determining the table structure corresponding to at least two cells in the table image is described, including:

[0200] Step 901: Determine at least two first cells located in the reference row and at least two second cells located in the reference column from the at least two cells;

[0201] Step 902: Construct a tree structure in the first direction by using multiple cells that have a subordinate relationship with the at least two first cells;

[0202] Step 903: Construct a tree structure in the second direction by using multiple cells that have a subordinate relationship with the at least two second cells;

[0203] wherein, the first direction is different from the second direction.

[0204] Here, after determining the table structure corresponding to at least two cells in the table image, the recognized text obtained by recognizing the text line can be stored in the node corresponding to the table structure.

[0205] Here, after obtaining the recognized text by recognizing the text line, the recognized text can also be retrieved from a preset dictionary. When the recognized text is retrieved, the recognized text is stored in the corresponding node of the table structure.

[0206] For common tables, some prior phrase dictionaries are established. For example, for invoice tables, the following can be established: "supplier", "purchaser", "seller", "invoice code", etc. When a cell contains a single text line, if the recognized text obtained by recognizing the text line of the cell is not in the dictionary, the recognized text is not stored in the table structure; when a cell contains multiple text lines, if the recognized text obtained by recognizing the multiple text lines of the cell is not in the dictionary, the recognized text is not stored in the table structure, that is, the one or more text lines distributed within the same cell are stored only when they belong to the phrase in the dictionary.

[0207] Here, determining the table structure corresponding to at least two cells in the table image has the following advantages:

[0208] (1) The structures of the two trees constructed in the first direction and the second direction can save the text information of the entire table. Since there is a subordinate relationship between the nodes in the tree structure, the text information stored in the nodes also has a subordinate relationship, so that the information with a dependency relationship can be accurately extracted, and the character recognition and content extraction of the entire table can be completed.

[0209] (2) When storing the text information corresponding to a cell by using the table structure, if the text line information corresponding to the cell is in the preset dictionary, the text line information is stored in the corresponding node of the table structure; if the text line information corresponding to the cell is not in the preset dictionary, the text line information is not stored in the corresponding node of the table structure. In this way, the accuracy of the extracted table information is ensured.

[0210] Adopting the technical solution of the embodiment of the present invention, the table structure formed by at least two cells in the table area of the table image is extracted, and the text line information corresponding to the at least two cells is extracted. In this way, table information is formed based on the table structure and the text line information. Since the table information can be formed by combining the structure and text information of the table, the table information can be accurately extracted.

[0211] To implement the information extraction method of the embodiments of the present invention, the embodiments of the present invention further provide an information extraction device, which is set on a terminal. Figure 10 It is a schematic diagram of the composition structure of the information extraction device according to the embodiments of the present invention; as Figure 10 shown, the device includes:

[0212] An acquisition unit 101, configured to acquire a table image;

[0213] A first processing unit 102, configured to use a preset first network model to segment and locate a table area in the table image to obtain at least two cells;

[0214] A second processing unit 103, configured to, for each of the at least two cells, determine a text line in the corresponding cell in combination with a preset second network model;

[0215] A third processing unit 104, configured to use a preset third network model to respectively identify the text lines in the at least two cells to obtain recognized text;

[0216] A fourth processing unit 105, configured to determine a table structure corresponding to the at least two cells, and form table information by using the table structure and the recognized text.

[0217] In one embodiment, the first processing unit 102 is specifically configured to:

[0218] Take the table image as the input of a preset first network model, perform a mapping from the input to the output of the table image, and obtain a feature map and feature map information of the table area in the table image; the feature map information represents the line type corresponding to each feature point in the table area;

[0219] Use the feature map information to determine the coordinates of multiple feature points corresponding to at least two line types from the feature map; and use the coordinates of the multiple feature points to form at least two cells;

[0220] Select at least two cells that meet a first preset condition from the at least two formed cells.

[0221] In one embodiment, the first processing unit 102 is specifically configured to:

[0222] For each of the at least two cells, determine whether the height of the corresponding cell is less than or equal to a height threshold and whether the length of the corresponding cell is less than or equal to a length threshold;

[0223] When it is determined that the height of the corresponding cell is less than or equal to the height threshold and the length of the corresponding cell is less than or equal to the length threshold, discard the corresponding cell;

[0224] Use the remaining at least two cells among the at least two cells as the at least two cells that meet the first preset condition.

[0225] In addition, according to at least one embodiment of the present invention, the second processing unit is specifically configured to:

[0226] For each cell among the at least two cells, in combination with a preset second network model, determine at least two first text boxes included in the corresponding cell;

[0227] Select at least two second text boxes that meet the second preset condition from the at least two first text boxes;

[0228] Concatenate the texts in the at least two second text boxes to obtain a text line.

[0229] In one embodiment, the second processing unit 103 is specifically configured to:

[0230] Perform horizontal sorting on at least two first text boxes in the corresponding cell to obtain at least two sorted first text boxes;

[0231] For the i-th first text box among the at least two sorted first text boxes, search for the j-th text box whose overlapping height with the i-th text box meets the second preset condition in the horizontal positive direction; and search for the k-th text box whose overlapping height with the j-th text box meets the second preset condition in the horizontal reverse direction;

[0232] Determine the first horizontal distance between the first i-th text box and the j-th text box; and determine the second horizontal distance between the j-th text box and the k-th text box;

[0233] When the first horizontal distance is greater than or equal to the second horizontal distance, concatenate the texts in at least two text boxes between the i-th text box and the j-th text box to obtain a text line.

[0234] In one embodiment, the second processing unit 103 is specifically configured to:

[0235] Search for at least one second text box whose horizontal distance from the i-th text box is greater than or equal to the distance threshold;

[0236] Calculate the overlapping height between the at least one second text box and the i-th text box to obtain at least one overlapping height;

[0237] Use the second text box corresponding to the maximum coincidence height among the at least one coincidence height as the j-th text box that meets the second preset condition.

[0238] In one embodiment, the fourth processing unit 105 is specifically configured to:

[0239] Determine at least two first cells located in the reference row and at least two second cells located in the reference column from the at least two cells;

[0240] Determine multiple cells having a subordinate relationship with the at least two first cells; and determine multiple cells having a subordinate relationship with the at least two second cells;

[0241] Construct a tree structure based on the determined multiple cells having a subordinate relationship;

[0242] Use the tree structure as the table structure of the at least two cells.

[0243] In one embodiment, the fourth processing unit 105 is specifically configured to:

[0244] Use the multiple cells having a subordinate relationship with the at least two first cells to construct a tree structure in the first direction;

[0245] Use the multiple cells having a subordinate relationship with the at least two second cells to construct a tree structure in the second direction;

[0246] Wherein, the first direction is different from the second direction.

[0247] In practical applications, the acquisition unit 101 can be implemented by a communication interface in the information extraction device; the first processing unit 102, the second processing unit 103, the third processing unit 104, and the fourth processing unit 105 are implemented by a processor in the information extraction device in combination with the communication interface.

[0248] It should be noted that: when the information extraction device provided in the above embodiment performs information extraction, only the above division of each program module is used for illustration. In practical applications, the above processing can be allocated to different program modules according to needs, that is, the internal structure of the device is divided into different program modules to complete all or part of the above-described processing. In addition, the information extraction device provided in the above embodiment and the information extraction method embodiment belong to the same concept, and the specific implementation process is detailed in the method embodiment, which will not be repeated here.

[0249] The embodiment of the present invention also provides an electronic device, as Figure 11 shown, including:

[0250] A communication interface 111 capable of interacting with other devices;

[0251] The processor 112, connected to the communication interface 111, is configured to execute the methods provided by one or more of the above technical solutions on the intelligent device side when running a computer program. The computer program is stored on the memory 113.

[0252] It should be noted that: For the specific processing procedures of the processor 112 and the communication interface 111, please refer to the method embodiments for details and will not be elaborated here.

[0253] Of course, in actual applications, the various components in the electronic device 110 are coupled together through the bus system 114. It can be understood that the bus system 114 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 114 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, in Figure 11 all the various buses are labeled as the bus system 114.

[0254] The memory 113 in the embodiments of the present application is used to store various types of data to support the operation of the terminal 110. Examples of these data include: any computer program for operating on the electronic device 110.

[0255] The methods disclosed in the embodiments of the present application above can be applied to the processor 112 or implemented by the processor 112. The processor 112 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above methods can be completed by the integrated logic circuit in the hardware of the processor 112 or instructions in software form. The above-mentioned processor 112 may be a general-purpose processor, a digital data processor (DSP, Digital Signal Processor), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor 112 can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or any conventional processor, etc. Combining the steps of the methods disclosed in the embodiments of the present application, it can be directly embodied as being executed by a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software module may be located in a storage medium, and this storage medium is located in the memory 113. The processor 112 reads the information in the memory 113 and combines its hardware to complete the steps of the foregoing methods.

[0256] In an exemplary embodiment, the electronic device 110 may be implemented by one or more application specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontroller units (MCUs), microprocessors, or other electronic components, and is used to execute the foregoing method.

[0257] It can be understood that the memory (memory 103) in the embodiments of the present application can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM, Read Only Memory), a programmable read-only memory (PROM, Programmable Read-Only Memory), an erasable programmable read-only memory (EPROM, Erasable Programmable Read-Only Memory), an electrically erasable programmable read-only memory (EEPROM, Electrically Erasable Programmable Read-Only Memory), a ferromagnetic random access memory (FRAM, ferromagnetic random access memory), a flash memory (Flash Memory), a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM, Compact Disc Read-Only Memory); the magnetic surface memory can be a disk memory or a tape memory. The volatile memory can be a random access memory (RAM, Random Access Memory), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as a static random access memory (SRAM, Static Random Access Memory), a synchronous static random access memory (SSRAM, Synchronous Static Random Access Memory), a dynamic random access memory (DRAM, Dynamic Random Access Memory), a synchronous dynamic random access memory (SDRAM, Synchronous Dynamic Random Access Memory), a double data rate synchronous dynamic random access memory (DDR SDRAM, Double Data Rate Synchronous Dynamic Random Access Memory), an enhanced synchronous dynamic random access memory (ESDRAM, Enhanced Synchronous Dynamic Random Access Memory), a synchronous link dynamic random access memory (SLDRAM, SyncLink Dynamic Random Access Memory), a direct rambus random access memory (DRRAM, Direct Rambus Random Access Memory).The memories described in the embodiments of the present application are intended to include, but are not limited to, these and any other suitable types of memories.

[0258] In an exemplary embodiment, the embodiments of the present invention further provide a storage medium, namely a computer storage medium, specifically a computer-readable storage medium, such as a memory 113 storing a computer program, and the above computer program can be executed by a processor 112 of a control server 110 to complete the steps described in the foregoing control server-side method. The computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, Flash Memory, magnetic surface memory, optical disc, or CD-ROM.

[0259] It should be noted that: "first", "second", etc. are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence.

[0260] In addition, the technical solutions described in the embodiments of the present invention can be combined arbitrarily without conflict.

[0261] The above is only a preferred embodiment of the present invention and is not intended to limit the protection scope of the present invention.

Claims

1. An information extraction method, characterized in that, the method includes: Collecting a table image; Using a preset first network model to segment and locate the table area in the table image to obtain at least two cells; For each of the at least two cells, combining a preset second network model to determine the text lines in the corresponding cell; Using a preset third network model to separately identify the text lines in the at least two cells to obtain identified text; Determining the table structure corresponding to the at least two cells, and using the table structure and the identified text to form table information; wherein, for each of the at least two cells, combining a preset second network model to determine the text lines in the corresponding cell includes: For each of the at least two cells, combining a preset second network model to determine at least two first text boxes included in the corresponding cell; Selecting at least two second text boxes that meet a second preset condition from the at least two first text boxes; Concatenating the text in the at least two second text boxes to obtain a text line; wherein, selecting at least two second text boxes that meet a second preset condition from the at least two first text boxes, and concatenating the text in the at least two second text boxes to obtain a text line includes: Performing horizontal sorting on the at least two first text boxes within the corresponding cell to obtain the sorted at least two first text boxes; For the i-th first text box among the sorted at least two first text boxes, searching in the horizontal positive direction for the j-th text box whose overlapping height with the i-th text box meets the second preset condition; and searching in the horizontal reverse direction for the k-th text box whose overlapping height with the j-th text box meets the second preset condition; Determining the first horizontal distance between the i-th text box and the j-th text box; and determining the second horizontal distance between the j-th text box and the k-th text box; When the first horizontal distance is greater than or equal to the second horizontal distance, concatenating the text in at least two text boxes between the i-th text box and the j-th text box to obtain a text line; The searching for the j-th text box whose overlapping height with the i-th text box meets the second preset condition includes: Searching for at least one second text box whose horizontal distance from the i-th text box is greater than or equal to a distance threshold; Calculating the overlapping height between the at least one second text box and the i-th text box to obtain at least one overlapping height; Taking the second text box corresponding to the maximum overlapping height among the at least one overlapping height as the j-th text box that meets the second preset condition.

2. The method according to claim 1, characterized in that, the using a preset first network model to segment and locate the table area in the table image to obtain at least two cells includes: Taking the table image as the input of a preset first network model, performing mapping from the input to the output of the table image to obtain the feature map and feature map information of the table area in the table image; the feature map information characterizes the line segment type corresponding to each feature point in the table area; Using the feature map information, determine the coordinates of multiple feature points corresponding to at least two line types from the feature map; and form at least two cells using the coordinates of the multiple feature points; Select at least two cells that meet the first preset condition from the at least two formed cells to obtain at least two cells; Among them, the selecting at least two cells that meet the first preset condition from the at least two formed cells includes: For each cell among the at least two cells, determine whether the height of the corresponding cell is less than or equal to the height threshold and whether the length of the corresponding cell is less than or equal to the length threshold; When it is determined that the height of the corresponding cell is less than or equal to the height threshold and the length of the corresponding cell is less than or equal to the length threshold, discard the corresponding cell; Use the remaining at least two cells among the at least two cells as the at least two cells that meet the first preset condition.

3. The method according to claim 1, wherein, the determining the table structure corresponding to the at least two cells includes: Determine at least two first cells located in the reference row and at least two second cells located in the reference column from the at least two cells; Determine multiple cells having a subordinate relationship with the at least two first cells; and determine multiple cells having a subordinate relationship with the at least two second cells; Based on the determined multiple cells having a subordinate relationship, construct a tree structure; Use the tree structure as the table structure of the at least two cells.

4. The method according to claim 3, wherein, the constructing a tree structure based on the determined multiple cells having a subordinate relationship includes: Use the multiple cells having a subordinate relationship with the at least two first cells to construct a tree structure in the first direction; Use the multiple cells having a subordinate relationship with the at least two second cells to construct a tree structure in the second direction; wherein, the first direction is different from the second direction.

5. An information extraction device, wherein, includes: An acquisition unit for acquiring a table image; A first processing unit for segmenting and positioning the table area in the table image using a preset first network model to obtain at least two cells; A second processing unit for, for each cell among the at least two cells, determining the text lines in the corresponding cell in combination with a preset second network model; A third processing unit for respectively identifying the text lines in the at least two cells using a preset third network model to obtain the recognized text; A fourth processing unit for determining the table structure corresponding to the at least two cells and forming table information using the table structure and the recognized text; Among them, the second processing unit is specifically used for: For each cell among the at least two cells, determining at least two first text boxes included in the corresponding cell in combination with a preset second network model; Selecting at least two second text boxes that meet the second preset condition from the at least two first text boxes; Stitching the text in the at least two second text boxes to obtain the text line; Among them, the second processing unit is further configured to: Perform horizontal sorting on at least two first text boxes in the corresponding cell to obtain at least two sorted first text boxes; For the i-th first text box among the at least two sorted first text boxes, search for the j-th text box whose overlapping height with the i-th text box satisfies a second preset condition in the horizontal positive direction; and search for the k-th text box whose overlapping height with the j-th text box satisfies the second preset condition in the horizontal reverse direction; Determine the first horizontal distance between the i-th text box and the j-th text box; and determine the second horizontal distance between the j-th text box and the k-th text box; When the first horizontal distance is greater than or equal to the second horizontal distance, splice the texts in at least two text boxes between the i-th text box and the j-th text box to obtain a text line; The searching for the j-th text box whose overlapping height with the i-th text box satisfies the second preset condition includes: searching for at least one second text box whose horizontal distance from the i-th text box is greater than or equal to a distance threshold; calculating the overlapping height between the at least one second text box and the i-th text box to obtain at least one overlapping height; and using the second text box corresponding to the maximum overlapping height among the at least one overlapping height as the j-th text box that satisfies the second preset condition.

6. An electronic device Characterized in that It includes: A communication interface for collecting a table image; A processor for segmenting and positioning a table area in the table image by using a preset first network model to obtain at least two cells; For each of the at least two cells, combining a preset second network model to determine text lines in the corresponding cell; and using a preset third network model to respectively recognize the text lines in the at least two cells to obtain recognized texts; Determine the table structure corresponding to the at least two cells, and form table information by using the table structure and the recognized texts; Among them, the processor is specifically configured to: For each of the at least two cells, combine a preset second network model to determine at least two first text boxes included in the corresponding cell; Select at least two second text boxes that satisfy a second preset condition from the at least two first text boxes; Splice the texts in the at least two second text boxes to obtain a text line; Among them, the processor is further configured to: Perform horizontal sorting on at least two first text boxes in the corresponding cell to obtain at least two sorted first text boxes; For the i-th first text box among the at least two sorted first text boxes, search for the j-th text box whose overlapping height with the i-th text box satisfies a second preset condition in the horizontal positive direction; and search for the k-th text box whose overlapping height with the j-th text box satisfies the second preset condition in the horizontal reverse direction; Determine the first horizontal distance between the i-th text box and the j-th text box; and determine the second horizontal distance between the j-th text box and the k-th text box; When the first horizontal distance is greater than or equal to the second horizontal distance, splice the texts in at least two text boxes between the i-th text box and the j-th text box to obtain a text line; Searching for the j-th text box whose coincidence height with the i-th text box satisfies a second preset condition includes: searching for at least one second text box whose horizontal distance from the i-th text box is greater than or equal to a distance threshold; calculating the coincidence height between the at least one second text box and the i-th text box to obtain at least one coincidence height; and using the second text box corresponding to the maximum coincidence height among the at least one coincidence height as the j-th text box that satisfies the second preset condition.

7. An electronic device Characterized in that It includes a processor and a memory for storing a computer program that can run on the processor, wherein, when the processor is used to run the computer program, it executes the steps of the method according to any one of claims 1 to 4.

8. A storage medium, on which a computer program is stored, Characterized in that When the computer program is executed by a processor, it realizes the steps of the method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Document Field Detection And Parsing

    US20170351913A1

  • Determining functional and descriptive elements of application images for intelligent screen automation

    US20190294641A1