Table information extraction method, computer program product and information extraction system
By analyzing the cell area shape and dependency relationship of table images, combining deep learning and OCR technology to distinguish merged cells from ordinary cells, the problem of low efficiency in extracting complex table information in existing technologies is solved, and efficient and accurate information extraction and automated processing are achieved.
Patent Information
- Application Number
- CN202510702282.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-09-12
AI Technical Summary
When processing tables with complex structures, especially when merging cells, the existing technology has low information extraction efficiency, resulting in inaccurate data extraction.
By obtaining the shape and dependency of the cell area of the table image, distinguishing merged cells from ordinary cells, using HRNet, U-NET, GAN and other models for image segmentation and boundary recognition, combining OCR technology to extract text content, and using decision tree, SVM and other models to analyze semantic relevance to ensure complete information extraction.
It improves the efficiency and accuracy of information extraction, avoids data fragmentation, ensures the information integrity and logical coherence in complex tables, adapts to tables of different formats and design styles, reduces manual intervention, and improves the level of automated processing.
Smart Images

Figure CN120635920A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of document parsing, and in particular to a table information extraction method, table information extraction device, computer program product, and information extraction system. Background Art
[0002] In the digital age, annual reports and financial statements, as a core component of annual information disclosure by financial institutions and businesses, provide a detailed record of key information such as a company's financial status, operating results, and risk management. This information is an indispensable source of data for investors, analysts, regulators, and the public, assessing corporate health, guiding investment decisions, and conducting market oversight. With the increasing complexity of global financial markets and rising demands for transparency, the centralized analysis of large volumes of annual reports and financial statements has become increasingly crucial.
[0003] Initially, the processing of documents such as bank annual reports was highly manual. Data extraction, verification, and analysis were time-consuming, labor-intensive, and error-prone. With the development of computer technology, basic text processing software began to be used for document storage and simple retrieval.
[0004] However, when processing a table with a complex structure, especially a table with merged cells, the information extraction method is not accurate, resulting in low efficiency of the extracted information. Summary of the Invention
[0005] The main purpose of this application is to provide a table information extraction method, table information extraction device, computer program product and information extraction system, so as to at least solve the problem of low efficiency of information extracted from tables in the prior art.
[0006] In order to achieve the above-mentioned purpose, according to one aspect of the present application, a method for extracting information from a table is provided, comprising: obtaining an image to be identified, wherein the image to be identified is an image of a table; extracting relevant information of a cell area in the image to be identified, wherein the relevant information includes shape and / or dependency relationship, wherein the cell area is an area having a single rectangular frame, the shape is a geometric shape between a first boundary line of the cell area and a second boundary line of an adjacent cell area of the cell area having an intersection, the first boundary line is a boundary line composed of pixel points of the boundary of the cell area detected, and the dependency relationship is whether there is semantic content between the first text of the cell area and the second text of the adjacent cell area. Correlation; determining the type of the cell area according to the relevant information, wherein the type includes merged cells and ordinary cells, the merged cells include multiple cells, and the ordinary cells include one cell; in a case where the type is the ordinary cell, extracting the information of the cell area; in a case where the type is the merged cell, extracting the information of all the cell areas in the coverage area, wherein the coverage area is the cell area including the merged cell and the cell area having the semantic correlation with the merged cell; sending the extracted information to the security monitoring system of the financial institution, so that the security monitoring system performs data monitoring based on the information of the cell area.
[0007] Optionally, after obtaining the image to be identified, the method further includes: obtaining a first recognition model to be trained, wherein the first recognition model to be trained is one of an HRNet model, a U-NET model, and a GAN model; forming the images to be identified into a first training set, and using the first training set and corresponding region labels to train the first recognition model to be trained to obtain a cell recognition model, wherein the region label is the region where the cell is located in the image of the first training set; inputting the image to be identified into the cell recognition model to obtain the cell region corresponding to the image to be identified.
[0008] Optionally, extracting relevant information of the cell area in the image to be identified includes: using edge detection technology to perform boundary identification on the cell area to obtain multiple current boundary lines; using edge detection technology to perform boundary identification on the adjacent cell area to obtain multiple adjacent boundary lines; extracting the current boundary line and the adjacent boundary line with intersection points to obtain the first boundary line and the second boundary line; obtaining a second recognition model to be trained, wherein the second recognition model to be trained is one of a decision tree model, an SVM model, and a random forest model; combining the first boundary line and the second boundary line into a second training set, and using the second training set and the corresponding shape label to train the second recognition model to be trained to obtain a shape recognition model, wherein the shape label is the geometric shape between the boundary lines in the second training set; inputting the first boundary line and the second boundary line into the shape recognition model to obtain shape recognition results corresponding to the first boundary line and the second boundary line.
[0009] Optionally, extracting relevant information of the cell area in the image to be identified includes: obtaining a third recognition model to be trained, wherein the third recognition model to be trained is one of an LSTM model, a word embedding model, and an NTN model; adding a Transformer module before the output layer of the third recognition model to be trained, so that the output of the Transformer module is used as the input of the output layer to obtain a text recognition model; using OCR extraction technology to extract the text content of the cell area to obtain the first text; using OCR extraction technology to extract the text content of the adjacent cell area to obtain the second text; combining the first text and the second text into a third training set, using the third training set and corresponding relationship labels to train the text recognition model to obtain a relationship recognition model, wherein the relationship label is a semantic correlation relationship between the first text and the second text; inputting the first text and the second text into the relationship recognition model to obtain the dependency relationship corresponding to the first text and the second text.
[0010] Optionally, determining the type of the cell area based on the relevant information includes: when the relevant information includes the shape and the shape is a T-shape, determining that the type of the cell area is the merged cell; when the relevant information includes the shape and the shape is not a T-shape, determining that the type of the cell area is the ordinary cell.
[0011] Optionally, determining the type of the cell area based on the relevant information includes: determining that the type of the cell area is the merged cell when the relevant information includes the dependency relationship and there is semantic correlation between the first text of the cell area and the second text of the adjacent cell area; and determining that the type of the cell area is the ordinary cell when the relevant information includes the dependency relationship and there is no semantic correlation between the first text of the cell area and the second text of the adjacent cell area.
[0012] Optionally, the type of the cell area is determined based on the relevant information, including: determining that the type of the cell area is the merged cell when at least one of the following conditions is satisfied: the shape is a T-shape and there is a semantic correlation between the first text in the cell area and the second text in the adjacent cell area; and determining that the type of the cell area is the ordinary cell when the shape is not a T-shape and there is no semantic correlation between the first text in the cell area and the second text in the adjacent cell area.
[0013] According to another aspect of the present application, a table information extraction device is provided, comprising: a first acquisition unit, for acquiring an image to be identified, wherein the image to be identified is an image of a table; a first extraction unit, for extracting relevant information of a cell area in the image to be identified, wherein the relevant information includes a shape and / or a dependency relationship, wherein the cell area is an area having a single rectangular frame, the shape is a geometric shape between a first boundary line of the cell area and a second boundary line of an adjacent cell area of the cell area having an intersection, the first boundary line is a boundary line composed of pixel points of the boundary of the cell area detected, and the dependency relationship is whether there is semantic correlation between the first text of the cell area and the second text of the adjacent cell area; a determination unit, for determining the information of the cell area based on the relevant information. The information determines the type of the cell area based on the information, wherein the type includes merged cells and ordinary cells, the merged cells include multiple cells, and the ordinary cells include one cell; a second extraction unit is used to extract the information of all the cell areas based on the types of all the cell areas by adopting OCR extraction technology; a third extraction unit is used to extract the information of all the cell areas in the coverage area when the type is the merged cell, wherein the coverage area is the cell area including the merged cell and the cell area having the semantic relevance with the merged cell; a sending unit is used to send the extracted information to the security monitoring system of the financial institution, so that the security monitoring system performs data monitoring based on the information of the cell area.
[0014] According to another aspect of the present application, a computer program product is provided, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements the steps of any one of the table information extraction methods.
[0015] According to another aspect of the present application, an information extraction system is provided, comprising: one or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include an information extraction method for executing any one of the described tables.
[0016] By applying the technical solution of the present application, an analysis is first performed based on the shape and / or dependency relationship to distinguish whether the cell area is a merged cell or an ordinary cell. In this way, the type of cell area can be accurately identified. Different types of cell areas have different ways of extracting information. For ordinary cells, direct extraction is sufficient. For merged cells, not only the information of the merged cells is extracted, but also the information of the cells associated with the merged cells is extracted. In this way, the merged cells and the cells associated with the merged cells can be treated as a whole area, that is, the covering area, and information of these associated cells can be extracted as a whole. This ensures that the information related to the merged cells is integrated together, the data is retained, and data fragmentation is avoided, thereby improving the efficiency of the information extraction method. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The drawings that constitute part of this application are used to provide a further understanding of this application. The illustrative embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation on this application. In the drawings:
[0018] Figure 1 A hardware structure block diagram of a mobile terminal for executing a table information extraction method provided in an embodiment of the present application is shown;
[0019] Figure 2 A schematic diagram of a process for extracting information from a table according to an embodiment of the present application is shown;
[0020] Figure 3 Another schematic diagram of the process of extracting table information is shown;
[0021] Figure 4 A schematic diagram showing the intersection shape of the boundary line between two adjacent cells;
[0022] Figure 5 A schematic diagram showing the coverage area is shown;
[0023] Figure 6 A structural block diagram of a table information extraction device provided according to an embodiment of the present application is shown.
[0024] The above drawings include the following reference numerals:
[0025] 102. Processor; 104. Memory; 106. Transmission device; 108. Input / output device. DETAILED DESCRIPTION
[0026] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in this application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0027] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.
[0028] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present application described here. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0029] For ease of description, some nouns or terms involved in the embodiments of the present application are explained below:
[0030] Structured Data: Structured data refers to data with a clear predefined format and organization. This data follows a consistent pattern or architecture when stored and processed.
[0031] Unstructured Data: Unstructured data refers to data that does not follow a fixed format or predefined pattern. It has no unified organizational form and internal relationships, and is difficult to effectively store, manage, and process using traditional relational database management systems.
[0032] OCR (Optical Character Recognition): OCR refers to the application process of optical character recognition technology, which can automatically recognize printed or handwritten text visible to the human eye and convert it into a digital text format that can be read and edited by a computer.
[0033] As introduced in the background technology, the efficiency of extracting information from tables in the prior art is low. To solve the above problem, the embodiments of the present application provide a table information extraction method, a table information extraction device, a computer program product and an information extraction system.
[0034] The technical solutions in the embodiments of the present invention will be described clearly and completely below with reference to the accompanying drawings in the embodiments of the present invention.
[0035] The method embodiments provided in the embodiments of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Taking running on a mobile terminal as an example, Figure 1 FIG. 1 is a hardware structure block diagram of a mobile terminal for a table information extraction method according to an embodiment of the present invention. Figure 1 As shown, the mobile terminal may include one or more ( Figure 1 Only one is shown) a processor 102 (the processor 102 may include but is not limited to a microprocessor MCU or a programmable logic device FPGA and other processing devices) and a memory 104 for storing data, wherein the mobile terminal may also include a transmission device 106 and an input and output device 108 for communication functions. It will be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the mobile terminal. Figure 1 More or fewer components than shown, or with Figure 1 Different configurations shown.
[0036] The memory 104 can be used to store computer programs, such as software programs and modules of application software, such as the computer program corresponding to the device information display method in the embodiment of the present invention. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, implementing the above-mentioned method. The memory 104 may include a high-speed random access memory and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory 104 may further include a memory remotely located relative to the processor 102, and these remote memories can be connected to the mobile terminal via a network. Examples of the above-mentioned networks include but are not limited to the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof. The transmission device 106 is used to receive or send data via a network. Specific examples of the above-mentioned network may include a wireless network provided by the mobile terminal's communication provider. In one example, the transmission device 106 includes a network adapter (Network Interface Controller, abbreviated as NIC), which can be connected to other network devices via a base station to communicate with the Internet. In one example, the transmission device 106 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0037] In this embodiment, a method for extracting information from a table running on a mobile terminal, a computer terminal or a similar computing device is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0038] Figure 2 FIG. 1 is a flow chart of a table information extraction method according to an embodiment of the present application. Figure 2 As shown, the method includes the following steps:
[0039] Step S201, obtaining an image to be recognized, wherein the image to be recognized is an image of a table;
[0040] Specifically, PDF reports or scanned documents received from financial institutions or enterprises can be converted into grayscale images and binarized to facilitate subsequent table recognition and cell analysis to obtain images to be recognized.
[0041] Step S202: extracting relevant information of a cell region in the image to be recognized, wherein the relevant information includes shape and / or dependency relationship, wherein the cell region is a region having a single rectangular frame, the shape is a geometric shape between a first boundary line of the cell region and a second boundary line of an adjacent cell region having an intersection, the first boundary line is a boundary line composed of pixel points detected to obtain a boundary of the cell region, and the dependency relationship is whether there is semantic correlation between a first text in the cell region and a second text in the adjacent cell region;
[0042] Specifically, the shape analysis of cells involves identifying the boundaries of each cell and the geometric shapes between the boundary lines. The "first boundary line" mentioned here refers to the boundary line of a specific cell area obtained by detection, which is usually composed of pixels identified by image processing technology. The "second boundary line" refers to the boundary line of the cell area adjacent to the cell. There will be an "intersection point" where these two sets of boundary lines intersect or cross.
[0043] Dependency judgment focuses on the semantic relationship between cell contents (i.e., "first text" and "second text"). It is a higher-level analysis that goes beyond simple shape recognition. It examines the text in adjacent cells in the same row or column to analyze whether there is semantic relevance between these texts, such as whether there is logical subordination, parallelism, or causal relationship. For example, a cell may contain the title "Operating Income", while the cell below it lists specific numerical values. There is an obvious semantic dependency between the two.
[0044] Step S203, determining the type of the cell area according to the relevant information, wherein the types include merged cells and ordinary cells, the merged cells include multiple cells, and the ordinary cells include one cell;
[0045] Specifically, after obtaining the cell shape and dependency information, the cell area is classified to determine whether it is a normal cell or a merged cell.
[0046] Step S204, when the cell type is the common cell, extracting information of the cell area;
[0047] Specifically, for the area of common cells, the text information can be directly extracted using OCR technology.
[0048] Step S205: When the type is the merged cell, extract information of all the cell areas in a covered area, wherein the covered area is the cell area including the merged cell and the cell areas having the semantic relevance with the merged cell;
[0049] Specifically, for merged cells, the entire area they cover is identified, including other cell areas with semantic relevance. Considering that merged cells may contain information from multiple cells, covering the entire relevant area ensures that all involved data is correctly extracted and integrated, thus avoiding information omissions.
[0050] Step S206: sending the extracted information to the security monitoring system of the financial institution, so that the security monitoring system performs data monitoring based on the information in the cell area.
[0051] Specifically, the processed text information, including data from both regular and merged cells, is transmitted in a structured format to the financial institution's security monitoring system. This can be achieved through an API interface, ensuring data security and confidentiality. Based on the extracted information, the security monitoring system monitors data integrity, logic, and compliance, such as verifying the consistency of financial data and checking for unusual transactions.
[0052] In the existing scheme, table information is extracted directly, and the existence of merged cells in the table is not taken into account at all, which will result in the extracted information having no logical coherence. Through this embodiment, analysis is first performed based on the shape and / or dependency relationship to distinguish whether the cell area is a merged cell or an ordinary cell. In this way, the type of cell area can be accurately identified. Different types of cell areas have different ways of extracting information. For ordinary cells, direct extraction is sufficient. For merged cells, not only the information of the merged cells is extracted, but also the information of the cells associated with the merged cells is extracted. In this way, the merged cells and the cells associated with the merged cells can be regarded as a whole area, that is, the coverage area, and information of these associated cells can be extracted as a whole. This ensures that the information related to the merged cells is integrated together, the data is retained, and data fragmentation is avoided, thereby improving the efficiency of the information extraction method.
[0053] Specifically, the security monitoring system can identify unusual patterns in current reports, such as unusual transaction amounts, unusual times or dates, and financial ratios that contradict historical data. It can also check the logical consistency between financial statements, such as whether the data in the income statement and balance sheet match each other. It can also automatically generate reports summarizing the results of data monitoring and issue alerts for any potential issues identified.
[0054] Based on monitoring results, the security monitoring system automatically generates risk assessment reports, including risk levels, predicted trends, and recommended countermeasures. These reports use charts and key indicators to help senior management and the risk management team quickly understand the risk situation.
[0055] Specifically, existing solutions often cannot accurately identify the associations and hierarchical relationships between cells when dealing with highly complex or unconventional merged cell layouts, resulting in omissions or misinterpretations in data extraction. Faced with annual changes in formats like annual reports or the diversity of formats between different banks, many technologies (such as hard-coded rules or static templates) are difficult to adapt automatically and require frequent manual adjustments and maintenance. The complexity of merged cells makes it difficult to ensure the continuity and correctness of the logical relationships between data points, and existing solutions perform poorly in maintaining data integrity. Solutions that rely on a lot of manual intervention or computationally intensive algorithms make them inefficient when processing large amounts of annual report data and cannot meet the business needs of rapid data analysis.
[0056] Specifically, the purpose of this solution is to achieve efficient and accurate extraction of complex merged cell table data in bank annual reports, while enhancing the system's adaptability and automation level, ensuring the accurate extraction of complex table data to improve recognition accuracy, adapting to tables of different formats and design styles to enhance adaptability and flexibility, reducing manual intervention, and achieving rapid processing of large-scale data to improve processing efficiency and automation level.
[0057] Specifically, existing solutions often lead to data extraction errors or information omissions due to algorithm limitations when identifying and reconstructing complex merged cell structures. This solution significantly improves the accuracy and completeness of table recognition through innovative design, ensuring the accuracy of extracted data.
[0058] Specifically, to address the diversity of table formats and designs in documents, existing solutions often rely on fixed templates or rules, resulting in poor adaptability. This solution designs a highly adaptive parsing solution that can intelligently identify and adapt to tables of different styles, layouts, and formats, eliminating the need for frequent manual adjustments.
[0059] Specifically, existing solutions are inefficient and lack automation when processing large-scale data, requiring extensive manual intervention. One of the goals of this solution is to build an efficient automated processing platform that significantly reduces manual intervention, accelerates the data extraction process, and is suitable for the rapid processing of large-scale datasets.
[0060] Specifically, the process is as follows Figure 3 As shown, it includes 1. input and preprocessing, 2. precise positioning of advanced table areas, 3. deep analysis of table structure and template adaptation, 4. data extraction and intelligent enhancement processing, 5. large-scale data parallel processing and optimization, and 6. user interaction and intelligent feedback system, which are introduced below.
[0061] For 1. Input and preprocessing, including:
[0062] 1.1. Document Receipt and Security Check:
[0063] Receive files, perform virus scans, and verify file integrity. Before files are entered into the system for processing, they are first received. This step involves receiving user-uploaded files, whether transmitted over the network or uploaded locally. As files are received, they undergo security checks, including virus scanning and file integrity verification, to prevent malware intrusion or data corruption, ensuring system security and data reliability.
[0064] 1.2. File type identification and format conversion:
[0065] Use machine learning models to identify file types and select appropriate conversion tools (such as Aspose Libraries) for format conversion. File type identification is to determine whether the file is a PDF, Word document, Excel spreadsheet, or other type by analyzing the file's features and metadata. For table information extraction tasks, documents that usually need to be processed are in PDF or scanned image formats, which may not directly support deep learning model processing. Therefore, in this step, machine learning models or other recognition technologies will be used to automatically determine the type of input file, and based on the recognition results, appropriate tools (such as Aspose Libraries) will be selected to convert the file into a format that the model can operate, such as an image or a specific document format, to facilitate subsequent table recognition and information extraction.
[0066] 1.3. Metadata extraction and compliance checking:
[0067] Extract document metadata and check for copyright and privacy policy compliance. Metadata is data about data and can include information such as the file's creator, creation date, modification history, and copyright information. Extracting metadata provides additional context for subsequent data processing, helping the system better understand and process the file content. Compliance checks assess whether documents comply with legal and regulatory requirements, such as copyright laws and privacy policies, to ensure that data processing does not infringe copyrights or leak sensitive information, thereby protecting the rights and interests of businesses and users.
[0068] 1.4. Document structure analysis and stratification:
[0069] Natural language processing (NLP) techniques are applied to identify chapter titles, paragraphs, and other information, constructing a document structure. Image processing techniques are then used to convert the PDF document into an analyzable image format, performing operations such as binarization, noise removal, and deskew. Document structure analysis involves breaking down the document into its various components, such as chapter titles, paragraphs, lists, and tables, to construct the document structure. During this step, natural language processing (NLP) techniques are used to identify the text structure within the document. For example, chapter titles and body paragraphs are identified by analyzing text features such as formatting, spacing, and font size. Furthermore, for non-text content, such as tables, image processing techniques are applied to convert the PDF document into binary images (i.e., black and white) to simplify image analysis. During this conversion process, preprocessing operations such as noise removal and deskew are performed to improve the accuracy of table region recognition. Binarization reduces pixel values in an image to two states (0 or 1, representing black and white). Noise removal removes irrelevant or cluttered pixels that could interfere with analysis. Deskew ensures that the text is level, facilitating subsequent text line recognition.
[0070] In the specific implementation process, after obtaining the image to be identified, the above method also includes the following steps: obtaining a first recognition model to be trained, wherein the above first recognition model to be trained is one of the HRNet model, the U-NET model, and the GAN model; forming the above images to be identified into a first training set, and using the above first training set and the corresponding region labels to train the above first recognition model to be trained to obtain a cell recognition model, wherein the above region labels are the regions where the cells are located in the images of the above first training set; inputting the above image to be identified into the above cell recognition model to obtain the above cell region corresponding to the above image to be identified.
[0071] In this solution, the first recognition model to be trained is suitable for image segmentation. A large number of labeled first training sets are used to train the first recognition model to be trained, and a trained cell recognition model can be obtained. The cell recognition model can automatically segment all cell areas, so that the cell areas can be identified more accurately.
[0072] Specifically, the U-NET model can be selected as the first recognition model to be trained. With its excellent image segmentation capabilities, U-NET can efficiently identify the area of cells, whether they are ordinary cells or merged cells, and can provide detailed information. The cell area information extracted from a large number of annotated images is used as the first training set, and is used together with the corresponding area labels to train the U-NET model. During the training process, the model can be iterated multiple times until the model's performance on the test set reaches a predetermined accuracy threshold of 95%, thereby ensuring that the model can stably recognize various table layouts.
[0073] Deep learning models have demonstrated powerful capabilities in image recognition and segmentation. HRNet, U-NET, and GAN models (especially their variants, such as specific architectures for image segmentation) each have their own advantages in cell recognition tasks.
[0074] HRNet (High-Resolution Network) is a convolutional neural network designed to preserve high-resolution features. It enhances the resolution and details of features by exchanging information between parallel high-resolution streams and low-resolution streams. When identifying cell regions, HRNet's advantage lies in its ability to process subtle image structures, such as cell regions, even if the boundaries of these regions may be very thin or discontinuous.
[0075] U-NET is a convolutional neural network specifically designed for image segmentation. It has a "U"-shaped architecture consisting of an encoder (for capturing image features) and a decoder (for generating segmentation maps). The U-NET model passes high-resolution feature maps between the encoder and decoder through "skip connections", which enables the model to retain more detailed information when segmenting cell areas.
[0076] GANs (Generative Adversarial Networks) can be used for image-to-image translation tasks, for example, converting an original image into an image with cell region segmentation. The cGANs model consists of a generator and a discriminator. The generator is responsible for generating segmentation maps, while the discriminator evaluates the similarity between the generated maps and the true labels.
[0077] Specifically, this solution uses the U-Net architecture as its foundation, a convolutional neural network specifically suited for image segmentation tasks. The network outputs a probability map of whether each pixel belongs to a border, content, or blank area, enabling precise identification of table cell regions.
[0078] Specifically, a customized loss function can be used, combining boundary IoU (Intersection over Union) loss and weighted cross-entropy loss, to ensure that the model accurately segments while also correctly classifying cell content and boundaries. Loss function is a core concept in the training process of machine learning and deep learning models. It is used to measure the difference between the model's predicted results and its actual results, thereby guiding the model to update its parameters to reduce this difference. Customizing the loss function means designing or adjusting the form of the loss function based on the requirements of a specific task to make it more consistent with the goals of model training.
[0079] The IoU (intersection over union) loss function is:
[0080]
[0081] in, is the predicted box coordinate, is the real frame coordinate.
[0082]
[0083] |B p ∪B t |=|B p |+|B t |-|B p ∩B t |.
[0084] The weighted cross entropy loss is:
[0085]
[0086] N represents the total number of categories, y l ∈{0,1} is the one-hot encoding of the true label, p l ∈(0,1) is the predicted probability, is the category weight coefficient.
[0087] Boundary Intersection over Union (IoU) loss is a loss function specifically designed for object detection and segmentation tasks. IoU stands for Intersection over Union (IoU), also known as "intersection over union ratio" in Chinese. In table recognition scenarios, the key concern is whether the model can accurately identify cell boundaries. The IoU loss function evaluates the accuracy of boundary recognition by calculating the overlap ratio between the model's predicted bounding box and the actual bounding box. A higher IoU indicates a better match between the predicted and true boundaries.
[0088] Specifically, the IoU loss function is calculated by first finding the intersection area of the predicted boundary region and the true boundary region, and then dividing it by the union area of these two boundary regions. During training, the model attempts to maximize this ratio to optimize the accuracy of boundary detection.
[0089] Categorical cross-entropy loss (CCL) is primarily used for classification tasks in supervised learning. It measures the difference between the model's predicted probability distribution for a sample's class and the actual label probability distribution. In table recognition, CCL can be used to determine the accuracy of a model's prediction of whether a cell is content, a border, or a blank area. Simply put, if the model predicts a cell border when it is actually content, the CCL will be large, prompting the model to adjust its parameters to reduce this prediction error.
[0090] Combining boundary IoU loss and classification cross-entropy loss is to optimize two aspects of performance simultaneously during model training: on the one hand, IoU loss focuses on improving the accuracy of boundary detection; on the other hand, cross-entropy loss focuses on improving the accuracy of cell content classification. This joint optimization strategy ensures that the model can not only accurately "see" where the cell is (i.e. boundary recognition), but also understand what is inside the cell (i.e. content classification).
[0091] In table recognition tasks, simply identifying cell boundaries is not enough; the cell contents (such as numbers and text) must also be correctly classified. By using these two loss functions simultaneously, the model can achieve a good balance between segmentation and classification tasks, accurately determining the cell boundaries while effectively identifying different types of information within the cells, thereby improving the accuracy and completeness of the entire table information extraction.
[0092] In addition, in order to improve the generalization ability of the model, various data augmentation techniques such as rotation, scaling, and shearing can be used to simulate tables of different formats and design styles. Data augmentation can be applied to all models in this solution.
[0093] Based on the pre-trained model, fine-tuning is performed using a small amount of annotated complex tabular data, and a continuous learning mechanism is enabled to allow the model to self-optimize in newly encountered documents.
[0094] Fine-tuning is the process of applying a pre-trained model to a specific task. Even if a pre-trained model already has good general performance, further optimization may be required to achieve ideal performance in certain specific fields or for special problems. In this case, although the pre-trained table recognition model can handle general table structures, its performance may not be satisfactory for special tables with complex merged cells. Therefore, by fine-tuning the model using a small amount of manually annotated complex table data, the model's performance when handling merged cell tables can be specifically improved, allowing the model to more accurately recognize these complex structures and understand the logical relationships between cells.
[0095] Continuous learning, sometimes also called lifelong learning, is a method that enables machine learning models to continuously learn from new data and update their knowledge base. When working with complex tables, enabling continuous learning means that the model doesn't stop learning after fine-tuning. Instead, each time the model encounters a new document, especially one with a novel structure or design, it automatically analyzes its characteristics, learns from them, and optimizes its algorithm to better adapt to future tables.
[0096] In some embodiments, extracting relevant information of the cell area in the above-mentioned image to be identified can be specifically achieved through the following steps: using edge detection technology to perform boundary identification on the above-mentioned cell area to obtain multiple current boundary lines; using edge detection technology to perform boundary identification on the above-mentioned adjacent cell area to obtain multiple adjacent boundary lines; extracting the above-mentioned current boundary line and the above-mentioned adjacent boundary line with intersection points to obtain the above-mentioned first boundary line and the above-mentioned second boundary line; obtaining a second recognition model to be trained, wherein the above-mentioned second recognition model to be trained is one of a decision tree model, an SVM model, and a random forest model; combining the above-mentioned first boundary line and the above-mentioned second boundary line into a second training set, and using the above-mentioned second training set and the corresponding shape label to train the above-mentioned second recognition model to be trained to obtain a shape recognition model, wherein the above-mentioned shape label is the geometric shape between the boundary lines in the above-mentioned second training set; inputting the above-mentioned first boundary line and the above-mentioned second boundary line into the above-mentioned shape recognition model to obtain shape recognition results corresponding to the above-mentioned first boundary line and the above-mentioned second boundary line.
[0097] In this solution, edge detection technology can effectively identify the boundary lines of cells. By forming the first boundary line and the second boundary line into a second training set for training, a shape recognition model can be obtained, and then the shape can be more accurately recognized through the shape recognition model.
[0098] Specifically, the edge detection technology may be a Canny edge detection algorithm.
[0099] The Canny edge detection algorithm is applied to the cell region and adjacent cell regions. This algorithm identifies edges in the image based on gradient direction and magnitude, resulting in multiple current boundary lines and multiple adjacent boundary lines. The Canny algorithm is highly effective in detecting fine boundary lines, providing reliable boundary information even in poor document quality. From the current boundary line and adjacent boundary lines, edges with intersections are selected; these intersections are key to identifying cell segmentation and merging cells.
[0100] The random forest model can be selected as the second recognition model to be trained because it can handle multiple classification tasks and is very effective for the recognition and classification of boundary line shapes. The first boundary line and the second boundary line extracted from the intersection analysis are used to form a second training set, and shape labels (such as rectangle, trapezoid, etc.) are annotated for them. Through multiple iterative training, a shape recognition model is obtained, which can recognize and classify different shapes based on the geometric features of the boundary line. The first boundary line and the second boundary line extracted from the image to be recognized are input into the trained shape recognition model, and the model outputs the recognition result of the boundary line shape.
[0101] 2. Advanced table area precise positioning, including:
[0102] 2.1. Multimodal feature extraction:
[0103] Document images are grayscaled and binarized, combined with feature extraction methods such as HOG and LBP. Multimodal feature extraction is a multi-level analysis of document images, aiming to capture features of table areas from multiple perspectives. This involves converting document images into grayscale and binary images, a common image preprocessing step that facilitates subsequent feature detection and segmentation. Grayscaling reduces color information, making image processing more efficient; binarization simplifies pixels in grayscale images to black or white, helping to highlight boundaries and text.
[0104] Next, we combine feature extraction methods like HOG (Histogram of Oriented Gradients) and LBP (Local Binary Patterns) to further analyze the texture and orientation information in the image. HOG captures the edges and outlines of objects in an image, while LBP focuses on changes in brightness patterns in local areas of the image. These methods enable more detailed identification of table features, such as cell boundaries and text distribution.
[0105] 2.2. Iterative training of deep learning models:
[0106] Use object detection models such as YOLOv5 to locate tables, and customize training sets to improve accuracy. Iterative training of deep learning models utilizes object detection models such as YOLOv5 (You Only Look Once version 5) to accurately locate table areas within documents. YOLOv5 is an advanced real-time object detection model that analyzes the entire image at once, identifying and framing multiple objects. It is ideal for table detection because it considers the relationship between a table and its surroundings, rather than examining each cell in isolation.
[0107] During the model training phase, a custom training set containing images labeled with tables is used to teach the model to recognize tables of varying styles and structures. Through multiple rounds of iterative training (i.e., "fine-tuning"), the deep learning model gradually improves its accuracy in locating table areas until it reaches the desired performance level.
[0108] 2.3. Table structure and style analysis:
[0109] Identify the border line style (solid line, dashed line), color, and infer the complexity of the table structure. For example, use Canny edge detection or Hough transform to find continuous and parallel lines, which often represent the rows and columns of the table. This step focuses on identifying and understanding the internal structure and appearance style of the table, such as the type of border line (solid line, dashed line), color, continuity and parallelism of the lines, which are key factors in evaluating the complexity and layout of the table. The Canny edge detection algorithm can find edge information in the image, including the boundary lines of the table; while the Hough transform is used to find and identify straight lines, especially those that form the rows and columns of the table.
[0110] Algorithms based on connected component analysis: The binary image is segmented into connected regions, and based on features such as area and shape, the regions that are likely to be tables are determined. The connected component analysis algorithm analyzes a binary image, segmenting it into several connected regions (i.e., regions consisting of pixels of the same color). It then determines which regions are most likely to be tables based on features such as shape, size, and position. This method can help identify the actual extent of a table, even when the table's boundaries are unclear or obscured.
[0111] 2.4. Distributed Image Processing:
[0112] Leverage GPU clusters to accelerate OCR image processing, improving processing speed and quality. Distributed image processing leverages the parallel computing power of GPU (graphics processing unit) clusters to accelerate image processing tasks, particularly OCR (Optical Character Recognition). Because deep learning models, particularly convolutional neural networks, consume significant computing resources when processing images, using GPU clusters can significantly shorten image processing and model inference time, improving overall processing speed and quality.
[0113] In a distributed environment, image processing tasks are split up and run simultaneously across multiple GPU nodes, with each node responsible for processing a portion of the image. This approach not only speeds up processing but also improves OCR accuracy through the high-performance computing power of the GPU, which is particularly important for processing large-scale or high-resolution document images.
[0114] Specifically, fast scanning algorithms (such as contour tracing or edge detection techniques) can be used to pre-process documents to quickly identify the boundaries of table areas and their general structure, providing a basis for subsequent analysis. After identifying the type, table elements such as row height, column width, and font style can also be identified. Using K-means or other clustering algorithms, different types of cells and merged cell patterns can be counted.
[0115] In addition, adaptive templates can be constructed. Based on clustering results, a template framework can be dynamically constructed, including cell size, merging rules, and other factors, to form a preliminary parsing template. Genetic algorithms or gradient descent methods can be introduced to iteratively optimize the template based on its performance on actual data, adjusting parameters such as merging rules, column widths, and row heights until the optimal fit is achieved.
[0116] In the specific implementation process, the relevant information of the cell area in the above-mentioned image to be recognized is extracted, which can be achieved through the following steps: obtaining a third recognition model to be trained, wherein the above-mentioned third recognition model to be trained is one of the LSTM model, the word embedding model, and the NTN model; adding a Transformer module before the output layer of the above-mentioned third recognition model to be trained, so that the output of the above-mentioned Transformer module is used as the input of the above-mentioned output layer to obtain a text recognition model; using OCR extraction technology to extract the text content of the above-mentioned cell area to obtain the above-mentioned first text; using OCR extraction technology to extract the text content of the above-mentioned adjacent cell area to obtain the above-mentioned second text; combining the above-mentioned first text and the above-mentioned second text into a third training set, using the above-mentioned third training set and the corresponding relationship label to train the above-mentioned text recognition model to obtain a relationship recognition model, wherein the above-mentioned relationship label is the semantic correlation relationship between the above-mentioned first text and the above-mentioned second text; inputting the above-mentioned first text and the above-mentioned second text into the above-mentioned relationship recognition model to obtain the above-mentioned dependency relationship corresponding to the above-mentioned first text and the above-mentioned second text.
[0117] In this solution, the third recognition model to be trained can be used to process text sequence data, and the introduction of the Transformer enhances its ability to capture long-range dependencies through the self-attention mechanism. This means that the model can not only understand the text information within a single cell, but also intelligently analyze the logical relationships between texts, thereby accurately determining whether there are dependencies between texts.
[0118] Specifically, the LSTM model is a good choice because it can process sequential data and is very effective at understanding contextual information in text sequences. However, the LSTM model's limitations become apparent when dealing with long-range dependencies and text relationships in complex merged cells. Adding a Transformer module before the output layer of the LSTM model uses its self-attention mechanism to capture long-range dependencies between elements in a text sequence, enhancing the model's understanding capabilities.
[0119] Using advanced OCR technology, the model extracts text content from a cell region to produce the first text. Simultaneously, it extracts text from an adjacent cell region to produce the second text. A third training set is constructed, including the first and second texts, along with the semantic relationships between them, as relationship labels. Through training, the model learns the dependencies between different texts and intelligently classifies and predicts them based on these relationships, ultimately producing a relationship recognition model. The first and second texts are then fed into the trained relationship recognition model, which then outputs the recognition results for the dependency relationships between the texts.
[0120] Specifically, the network architecture of this solution is as follows: a Transformer layer is connected before the output layer of the third recognition model to be trained. This layer captures the contextual dependencies and long-range relationships of merged cells in a table. Through a self-attention mechanism, the logical connections between cells are learned, which subsequently assists in identifying merged cells that span multiple rows or columns.
[0121] In addition, the solution of this application also includes a personalized push service that collects user behavior data, including but not limited to browsing history, search keywords, and data access frequency. Browsing history: records the pages, documents, or data sets that users have visited, especially the table information they have viewed repeatedly. Search keywords: tracks the search terms used by users, which often reflect the specific needs or interests of users. Data access frequency: analyzes the frequency with which users access specific data or documents. Data with high frequency of access usually indicates the user's key areas of focus.
[0122] Use clustering algorithms (such as K-means, DBSCAN, or hierarchical clustering) to segment the user population into different market segments or role types based on user behavior data. For example, users can be divided into roles such as "investor," "analyst," and "regulator," each of which may have different data preferences.
[0123] Apply sequential pattern mining or time series analysis to identify patterns in user behavior as they browse, search, or interact with data. For example, an analyst may tend to look at overview data first before diving into the details, while an investor may jump directly to the profit and loss statement or balance sheet.
[0124] Integrate the above analysis results to build a detailed user profile, including the user's background information, role, preferred data types, key data areas of interest, common operations, etc. Based on the user profile, apply machine learning (such as collaborative filtering, content-based recommendation, or hybrid recommendation algorithms) to predict the types of data and specific data items that the user may be interested in in the future, thus achieving personalized recommendations.
[0125] By integrating user role analysis with customized information push services, we achieve highly personalized reporting data and efficient information delivery. We also provide precise recommendations tailored to the needs of professionals, such as investors' focus on financial health and analysts' need for in-depth data analysis.
[0126] In addition, after extracting the information, the solution of this application also includes: identifying the language of the information, obtaining a translation request (a request to translate the text into a target language), and translating the information into the specified target language based on the translation request. The translation tool can be SDL Trados, MemoQ, DeepL, etc. The solution performs an automatic translation operation to translate the information into the user-specified target language while maintaining the format consistency of the numerical information, ensuring the accuracy of the numerical expression before and after translation.
[0127] During the translation process, the numerical formats such as currency symbols, decimal points and thousands separators are automatically adjusted so that the translated information conforms to the habits of the target language, avoiding information distortion caused by differences in understanding of numerical formats and enhancing the global interoperability of data.
[0128] For example, it extracts information about common cells and merged cells in financial reports and translates the extracted text information into the target language. Meanwhile, the numerical consistency calibration algorithm ensures the correct conversion of numerical formats, such as converting the English "$1,234,567.89" into the Chinese "1,234,567.89USD" or the Japanese "1,234,567.89ドル", thus maintaining the integrity and accuracy of the data.
[0129] Specifically, the emergence of optical character recognition (OCR) technology marks a major step forward in automated document processing. OCR technology can convert text in paper or electronic documents into editable digital text, greatly improving the efficiency of document digitization. However, for tables with complex structures and merged cells, the recognition accuracy and structural restoration capabilities of OCR technology are limited. With the popularization of structured data formats such as XML and JSON, and the development of data mining and analysis technologies, people have begun to explore how to more effectively extract structured data from documents such as annual reports. Rule-based parsers and XML parsing technologies are used to extract data from documents, but they still face challenges when processing complex merged cells.
[0130] In one solution, the application of OCR technology in complex merged cell processing mainly includes the following steps:
[0131] Preprocessing: Perform denoising, binarization, tilt correction and other preprocessing operations on the annual report document to prepare for subsequent OCR recognition.
[0132] Text Recognition: Uses OCR technology to recognize text in documents, including individual characters, words, and sentences. For tables, first extract the text within the table area one by one.
[0133] Table structure analysis: This step attempts to identify table boundaries, rows and columns, and the presence of merged cells. This step typically relies on image processing techniques such as edge detection and row and column segmentation algorithms.
[0134] Merged Cell Processing: Existing technologies use relatively basic methods to identify merged cells, such as comparing the content or location information of adjacent cells to infer the extent of the merged cell. However, this step is often limited by the complexity of the merged cells and the accuracy of OCR recognition.
[0135] Data integration and output: Integrate the identified and parsed data into a structured format (such as CSV or database) for further analysis.
[0136] In some embodiments, the type of the above-mentioned cell area is determined based on the above-mentioned relevant information, which can be specifically achieved through the following steps: when the above-mentioned relevant information includes the above-mentioned shape and the above-mentioned shape is a T-shape, the above-mentioned type of the above-mentioned cell area is determined to be the above-mentioned merged cell; when the above-mentioned relevant information includes the above-mentioned shape and the above-mentioned shape is not a T-shape, the above-mentioned type of the above-mentioned cell area is determined to be the above-mentioned ordinary cell.
[0137] In this solution, for tables with complex step sizes, T-shaped cells are often an obvious sign of merged cells. The cell type can be determined by the set shape, thereby further improving the recognition accuracy of the cell type.
[0138] Specifically, the boundary line between two adjacent cells is as follows: Figure 4 As shown, Figure 4 The intersection of the bold boundary lines is a T-shape. The coverage area is as follows Figure 5 As shown, Figure 5 In , the text content of the merged cell is the unit, and the text contents of the subordinate cells are all sub-units. Therefore, the area where the merged cell and the cells under the merged cell are located is the coverage area.
[0139] Specifically, a cell area has multiple boundary lines, and an adjacent cell area has multiple boundary lines. It is very likely that there are multiple first boundary lines, and there are multiple second boundary lines in the adjacent cell area. As long as the cell area and the adjacent cell area have any two first boundary lines and second boundary lines with an intersection in the shape of a T, the cell area is a merged cell.
[0140] During the specific implementation process, the type of the above-mentioned cell area is determined according to the above-mentioned relevant information, which can be achieved through the following steps: when the above-mentioned relevant information includes the above-mentioned dependency relationship and there is semantic correlation between the above-mentioned first text of the above-mentioned cell area and the above-mentioned second text of the above-mentioned adjacent cell area, the above-mentioned type of the above-mentioned cell area is determined to be the above-mentioned merged cell; when the above-mentioned relevant information includes the above-mentioned dependency relationship and there is no semantic correlation between the above-mentioned first text of the above-mentioned cell area and the above-mentioned second text of the above-mentioned adjacent cell area, the above-mentioned type of the above-mentioned cell area is determined to be the above-mentioned ordinary cell.
[0141] In this scheme, the information in the merged cells usually has a close semantic relationship with the surrounding cells. This relationship reflects the data ownership and logical connection between cells, that is, whether there is semantic relevance between the texts. By determining whether there is semantic relevance between the texts, the type of cell can be further accurately determined.
[0142] Specifically, a relationship recognition model is used to perform dependency analysis on the first text in a cell and the second text in an adjacent cell. The model outputs the dependency strength between the texts to determine their semantic relevance. Based on the results of the dependency analysis, if there is a significant semantic relevance between the first and second texts (e.g., the relationship strength exceeds 0.85), the cell region is determined to be a merged cell. Conversely, if there is no semantic relevance between them (e.g., the relationship strength is less than 0.15), the cell region is determined to be a normal cell.
[0143] For example, a financial analysis company is processing a complex bank annual report. The report contains a large amount of tabular data, some of which use merged cell designs.
[0144] Specifically, using smart fill and dependency analysis algorithms, the company first conducted a deep semantic understanding of the text within the cell and analyzed the dependency strength between the first and second texts through a relationship recognition model. When a cell pair with a dependency strength exceeding 0.85 was identified, the company marked the cell area as a merged cell, indicating that the information between the cell and the adjacent cells is closely related and may be part of a summary or common description. For cell pairs with a dependency strength below 0.15, they are treated as ordinary cells, which generally means that there is no direct logical connection between the cells.
[0145] Specifically, 3. Deep analysis of table structure and template adaptation includes:
[0146] 3.1. Fine-grained cell cutting:
[0147] Combining deep learning with traditional image processing technology, line segment detection and intersection analysis are implemented to determine the intersection of rows and columns, thereby dividing each cell. Fine-grained cell cutting refers to further subdividing the interior of the table on the basis of accurately identifying the table boundaries to determine the specific location and range of each cell. This step utilizes the combination of deep learning models and traditional image processing technology, where deep learning models (such as U-Net or other segmentation networks) are used to provide high-precision table area segmentation, while traditional image processing technology (such as line segment detection and intersection analysis) is responsible for locating the intersection of cell rows and columns to achieve precise cell cutting.
[0148] Line segment detection (e.g., through Canny or Hough transform) can help identify straight lines in a table, while intersection analysis determines the location of cells based on the intersection of these lines, thereby dividing the table into multiple independent cells. This fine-grained segmentation is crucial for accurately parsing cell contents.
[0149] 3.2. Advanced merged cell recognition:
[0150] A long short-term memory (LSTM) network is used to analyze text sequences and identify cross-row and cross-column merges. The text content is clustered based on its location distribution within the image to infer the location and size of cells. Advanced merged cell recognition aims to address the complexity of merged cells, where cells may span multiple rows or columns to form larger cells. This process uses a long short-term memory (LSTM) network to analyze the sequential characteristics of the text content within cells and cluster analysis based on their location distribution within the image to infer the extent of the merged cells.
[0151] LSTM is a neural network particularly well-suited for processing sequential data. It can remember long-term dependencies and is very effective for analyzing sequential patterns in text content. For example, it can identify repeated text at the beginning of a row or column, which often indicates merged cells. Furthermore, cluster analysis of cell positions helps the system understand the layout of merged cells. Based on the distribution of text in an image, it can infer the boundaries and size of merged cells, thereby accurately identifying complex merged cells.
[0152] 3.3. Template automatic generation and adaptive adjustment:
[0153] Based on the historical template library and the current table's features, a genetic algorithm is used to optimize template matching. Automatic template generation and adaptive adjustment are adaptive strategies for different table styles and structures. First, based on the historical template library (i.e., a collection of previously processed table styles and structures) and the specific features of the current table, an attempt is made to generate a preliminary table parsing template. This template includes information such as cell size, position, and merging rules to help the system understand the table's structure.
[0154] Adaptive adjustment optimizes the initial template using a genetic algorithm. This optimization method, inspired by natural selection and genetics, simulates evolutionary processes (such as selection, crossover, and mutation) to find the optimal solution. In this scenario, the genetic algorithm iteratively adjusts template parameters (such as cell size and merge ranges) to find the template that best matches the current table structure, thereby improving the accuracy and efficiency of information extraction.
[0155] 3.4. Dynamic header identification and mapping:
[0156] The word embedding model is used to understand the meaning of table headers and establish a mapping relationship between the headers and data fields. Dynamic header recognition and mapping refers to the system's ability to automatically identify table header information (i.e., column or row headers) and, based on the meaning of the header content, establish a mapping relationship between the header and the data fields in the table. This process utilizes word embedding models, such as Word2Vec or BERT, to understand the semantics of the header text and determine the relevance of the header to specific data fields.
[0157] Word embedding models convert text into vector representations, capturing similarities and semantic relationships between words. These models can parse the meaning of table headers and match them with the data within cells, automatically establishing connections between headers and data fields. This is crucial for subsequent data analysis and processing, ensuring that data is correctly classified and stored.
[0158] Specifically, data extraction and intelligent enhancement processing in 4. include:
[0159] Build a logical table structure based on cell positional relationships to extract key financial indicators and other information. For example, the merged cell recognition algorithm analyzes the blank space, text content, and row / column spans between adjacent cells to identify and record the scope and logical relationships of merged cells. Using the table title row, headers, and merged cells as clues, a tree-like data structure is constructed to represent the table's hierarchical relationships and data ownership, facilitating subsequent data reorganization. Structured table data can be stored in a database or exported to Excel or other formats. The data reorganization logic can follow the following pattern.
[0160] 4.1. Advanced Text Purification and Entity Recognition:
[0161] Apply pre-trained models such as BERT for deep text purification and identify entities such as dates and amounts. Advanced text purification refers to a more sophisticated text cleaning and normalization process supported by deep learning models. For example, text purification can be performed using pre-trained models such as BERT (Bidirectional Encoder Representations from Transformers). The BERT model can understand the context of the text, allowing it to more accurately identify and correct common spelling errors, grammatical issues, or formatting inconsistencies when processing text data, such as standardized currency symbols and date formats, to ensure the cleanliness and consistency of text data.
[0162] Named Entity Recognition (NER) automatically identifies and classifies specific types of entities in text, such as dates, amounts, company names, and product codes. In the context of financial reports and bank statements, entity recognition is particularly important because it helps the system identify key financial metrics and other relevant information, such as net profit, total liabilities, and transaction dates, which are the cornerstones of financial analysis and report interpretation.
[0163] 4.2. Semantic Understanding and Intelligent Filling:
[0164] Automatically fill in missing values based on semantic understanding of the table context, such as using knowledge graphs to complete information. Semantic understanding means that the system can understand the meaning behind the text, not just the literal content. In the context of table information extraction, this means that the system can parse the actual meaning of the data fields based on the context and structure of the table, such as one column may be about costs, while another column is about revenue. To further enhance the completeness and accuracy of the data, the technology also integrates a smart fill function that can automatically identify and fill in missing data. For example, if a cell is missing a value, but its column contains numerical information for other cells, the missing values can be reasonably guessed and filled in by analyzing the patterns of these values or using external data sources such as knowledge graphs, thereby providing a more comprehensive data set for data analysis.
[0165] 4.3. Data quality assessment and optimization suggestions:
[0166] Data quality is assessed through statistical testing, providing recommendations for data cleaning and structural optimization. Data quality assessment is the process of ensuring the reliability and applicability of extracted data. This includes checking data accuracy, completeness, consistency, and timeliness. The system automatically assesses the quality of extracted data through statistical testing and consistency checks, and based on the results, provides optimization recommendations, such as whether to retrain the model, adjust algorithm parameters, or add preprocessing steps, to further improve data extraction.
[0167] 4.4. Real-time Tuning of Machine Learning Models:
[0168] Implement online learning to adjust model parameters in real time based on user feedback and improve data extraction accuracy. Real-time tuning means that the system can instantly adjust the model's parameters and behavior based on user feedback or new data received during operation to adapt to changing conditions or improve performance. In the information extraction scenario, this means that every time the system processes a new batch of reports or encounters a change in the table structure, it will collect data on model performance and user corrections through online learning, automatically updating the model to improve its accuracy and efficiency in future tasks. This continuous self-adjustment capability is key to the system's ability to flexibly respond to diverse and ever-changing data inputs and is an important guarantee for improving user experience.
[0169] In some embodiments, the type of the above-mentioned cell area is determined based on the above-mentioned relevant information, which can be specifically achieved through the following steps: when the above-mentioned shape is a T-shape and at least one of the above-mentioned first text in the above-mentioned cell area and the above-mentioned second text in the above-mentioned adjacent cell area has semantic correlation, the above-mentioned type of the above-mentioned cell area is determined to be the above-mentioned merged cell; when the above-mentioned shape is not a T-shape and there is no semantic correlation between the above-mentioned first text in the above-mentioned cell area and the above-mentioned second text in the above-mentioned adjacent cell area, the above-mentioned type of the above-mentioned cell area is determined to be the above-mentioned ordinary cell.
[0170] In this solution, the cell type can be comprehensively judged based on the two criteria of whether the shape and text have semantic relevance. The judgment logic is more detailed, so the cell type can be determined more accurately.
[0171] Specifically, for 5. Large-scale data parallel processing and optimization, we adopted a microservices architecture for Docker containerized deployment. We leveraged Kubernetes to manage the cluster, automatically scaling capacity based on task queue lengths to achieve dynamic load balancing and resource scheduling. We also provided open tuning interfaces, supporting Apache Flink for real-time data stream processing and Spark for batch data analysis, as well as performance monitoring and log analysis.
[0172] Specifically, for 6. User interaction and intelligent feedback system, a multi-dimensional user interface design can be designed, customized interfaces can be designed for different roles, and personalized data dictionaries can be associated with roles such as data analysts and business users to meet the independent needs of different users.
[0173] Specifically, in existing solutions, data extraction technology often misreads or misses when processing merged cells, especially when faced with tables with complex designs and irregular layouts. This results in a significant reduction in the accuracy of data extraction and low recognition accuracy of merged cells. Due to the diversity of annual report formats, current technologies often require manual pre-definition of templates or format adjustments, which increases the workload, reduces processing efficiency, and may introduce human errors, resulting in poor adaptability and greater reliance on manual intervention. Existing methods are unable to effectively maintain the integrity and logical associations of data, especially in complex table structures, resulting in deviations in analysis results, poor data integrity, and a lack of logic. When processing large amounts of annual report data, limited by algorithm efficiency or computing resources, the processing speed of existing solutions may not be able to meet business needs for rapid response, resulting in low processing efficiency.
[0174] Specifically, the solution of this application significantly improves the accuracy and completeness of recognition through innovative design, ensures the accuracy of extracted data, and improves the recognition accuracy of merged cells. The highly adaptive parsing method can intelligently identify and adapt to tables of different styles, layouts, and formats without the need for frequent manual adjustments, thereby enhancing adaptability and flexibility. By building an efficient and automated processing platform, manual intervention is greatly reduced, the data extraction process is accelerated, and it is suitable for the rapid processing of large-scale data sets, improving processing efficiency and automation level.
[0175] Alternatively, you can split large tables into smaller chunks or divide the entire document set into independently processed tasks, with each task responsible for identifying and parsing a portion of the table. Using container orchestration tools like Kubernetes or Docker Swarm automatically allocates CPU, memory, and GPU resources based on task requirements, achieving load balancing.
[0176] In addition, GPU acceleration can be used for convolution operations and Transformer layers using libraries such as CUDA, significantly reducing training and inference time. A pipelined operation model for data preprocessing, model inference, and result integration can be established to reduce waiting time between tasks and improve overall throughput.
[0177] In summary, the solution of this application provides an improved merged cell identification and positioning technology that can accurately identify and segment the merged cell structure in complex annual report forms, overcoming the problem that the existing technology easily leads to incomplete or erroneous data extraction when processing such forms. Merged cell logic reconstruction and information extraction are realized, and a technical solution for processing merged cells in annual report forms is applied for protection, which specifically describes how to determine the relationship between the merged cell boundaries and internal information through image analysis, model recognition or other related technologies. The semantics of merged cells can be understood, and the semantic understanding of regular report names can be supported to distinguish detailed data from summary data and associate them.
[0178] In view of the above, the significant advantages brought by the solution of this application are as follows:
[0179] Compared to traditional OCR text recognition technology, the solution in this application greatly improves the accuracy and adaptability of table recognition by adopting advanced deep learning models such as U-Net and Transformer. In particular, it can accurately segment and identify the contents of each cell when faced with tables with complex layouts such as merged cells. This technological innovation fundamentally addresses the limitations of traditional methods in recognizing complex table structures, improves the accuracy of data extraction, and makes it possible to automatically process table data. This can significantly reduce the cost and error rate of manual intervention, especially in fields such as finance, scientific research, and big data analysis.
[0180] In addition, the integration of intelligent verification and logic repair modules has built a closed-loop process from data identification to quality control, ensuring the high reliability of data processing results. This module captures and corrects potential errors through a multi-dimensional verification mechanism, and combined with the semantic understanding capabilities of artificial intelligence, ensures that the data is not only correctly extracted, but also logically consistent and reasonable, which is crucial for subsequent data analysis and decision support. Compared with existing technologies, this solution adopts an innovative strategy that can accurately identify and segment complex merged cell structures in annual reports, thereby avoiding data omissions or misinterpretations, and greatly improving the integrity of information extraction. It meets the goal of efficiently and accurately extracting key data from massive annual reports and tables, greatly improving the overall information extraction speed and accuracy, and providing financial institutions, regulators and corporate decision makers with more reliable data support tools.
[0181] The embodiments of the present application also provide an information extraction device for a table. It should be noted that the information extraction device for a table in the embodiments of the present application can be used to execute the information extraction method for a table provided in the embodiments of the present application. The device is used to implement the above-mentioned embodiments and preferred implementation modes, and the details that have been described will not be repeated here. As used below, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceived.
[0182] The following introduces the table information extraction device provided in the embodiment of the present application.
[0183] Figure 6 This is a structural block diagram of a table information extraction device according to an embodiment of the present application. Figure 6 As shown, the device includes:
[0184] A first acquiring unit 10 is configured to acquire an image to be identified, wherein the image to be identified is an image of a table;
[0185] A first extraction unit 20 is configured to extract relevant information of a cell area in the image to be recognized, wherein the relevant information includes a shape and / or a dependency relationship, wherein the cell area is a region having a single rectangular frame, the shape is a geometric shape between a first boundary line of the cell area and a second boundary line of an adjacent cell area having an intersection, the first boundary line is a boundary line composed of pixel points detected to obtain a boundary of the cell area, and the dependency relationship is whether there is semantic correlation between a first text in the cell area and a second text in the adjacent cell area;
[0186] a determination unit 30, configured to determine the type of the cell area according to the relevant information, wherein the type includes a merged cell and a normal cell, the merged cell includes multiple cells, and the normal cell includes one cell;
[0187] A second extraction unit 40 is configured to extract information of the cell area when the cell type is the ordinary cell.
[0188] A third extraction unit 50 is configured to, when the type is the merged cell, extract information of all the cell areas in a covered area, wherein the covered area is the cell area including the merged cell and the cell areas having the semantic relevance with the merged cell;
[0189] The sending unit 60 is configured to send the extracted information to a security monitoring system of a financial institution, so that the security monitoring system performs data monitoring based on the information in the cell area.
[0190] In the existing scheme, table information is extracted directly, and the existence of merged cells in the table is not taken into account at all, which will result in the extracted information having no logical coherence. Through this embodiment, analysis is first performed based on the shape and / or dependency relationship to distinguish whether the cell area is a merged cell or an ordinary cell. In this way, the type of cell area can be accurately identified. Different types of cell areas have different ways of extracting information. For ordinary cells, direct extraction is sufficient. For merged cells, not only the information of the merged cells is extracted, but also the information of the cells associated with the merged cells is extracted. In this way, the merged cells and the cells associated with the merged cells can be regarded as a whole area, that is, the coverage area, and information of these associated cells can be extracted as a whole. This ensures that the information related to the merged cells is integrated together, the data is retained, and data fragmentation is avoided, thereby improving the efficiency of the information extraction method.
[0191] During the specific implementation process, the above-mentioned device also includes a second acquisition unit, a training unit and a processing unit. The second acquisition unit is used to obtain a first recognition model to be trained after obtaining the image to be recognized, wherein the above-mentioned first recognition model to be trained is one of the HRNet model, the U-NET model, and the GAN model; the training unit is used to group the above-mentioned images to be recognized into a first training set, and use the above-mentioned first training set and the corresponding area label to train the above-mentioned first recognition model to be trained to obtain a cell recognition model, wherein the above-mentioned area label is the area where the cell is located in the image of the above-mentioned first training set; the processing unit is used to input the above-mentioned image to be recognized into the above-mentioned cell recognition model to obtain the above-mentioned cell area corresponding to the above-mentioned image to be recognized.
[0192] In this solution, the first recognition model to be trained is suitable for image segmentation. A large number of labeled first training sets are used to train the first recognition model to be trained, and a trained cell recognition model can be obtained. The cell recognition model can automatically segment all cell areas, so that the cell areas can be identified more accurately.
[0193] In some embodiments, the first extraction unit includes a first recognition module, a second recognition module, a first extraction module, a first acquisition module, a first training module and a first processing module, wherein the first recognition module is used to use edge detection technology to perform boundary recognition on the above-mentioned cell area to obtain multiple current boundary lines; the second recognition module is used to use edge detection technology to perform boundary recognition on the above-mentioned adjacent cell area to obtain multiple adjacent boundary lines; the first extraction module is used to extract the above-mentioned current boundary line and the above-mentioned adjacent boundary line with intersection points to obtain the above-mentioned first boundary line and the above-mentioned second boundary line; the first acquisition module is used to obtain a second recognition model to be trained, wherein the above-mentioned second recognition model to be trained is one of a decision tree model, an SVM model, and a random forest model; the first training module is used to form the above-mentioned first boundary line and the above-mentioned second boundary line into a second training set, and use the above-mentioned second training set and corresponding shape labels to train the above-mentioned second recognition model to be trained to obtain a shape recognition model, wherein the above-mentioned shape label is the geometric shape between the boundary lines in the above-mentioned second training set; the first processing module is used to input the above-mentioned first boundary line and the above-mentioned second boundary line into the above-mentioned shape recognition model to obtain shape recognition results corresponding to the above-mentioned first boundary line and the above-mentioned second boundary line.
[0194] In this solution, edge detection technology can effectively identify the boundary lines of cells. By forming the first boundary line and the second boundary line into a second training set for training, a shape recognition model can be obtained, and then the shape can be more accurately recognized through the shape recognition model.
[0195] During the specific implementation process, the first extraction unit includes a second acquisition module, an addition module, a second extraction module, a third extraction module, a second training module and a second processing module. The second acquisition module is used to obtain a third recognition model to be trained, wherein the third recognition model to be trained is one of an LSTM model, a word embedding model and an NTN model; the addition module is used to add a Transformer module before the output layer of the third recognition model to be trained, so that the output of the Transformer module is used as the input of the output layer to obtain a text recognition model; the second extraction module is used to use OCR extraction technology to extract the text content of the above-mentioned cell area to obtain the above-mentioned first text; the third extraction module is used to use OCR extraction technology to extract the text content of the above-mentioned adjacent cell area to obtain the above-mentioned second text; the second training module is used to combine the above-mentioned first text and the above-mentioned second text into a third training set, and use the above-mentioned third training set and the corresponding relationship label to train the above-mentioned text recognition model to obtain a relationship recognition model, wherein the above-mentioned relationship label is the semantic correlation relationship between the above-mentioned first text and the above-mentioned second text; the second processing module is used to input the above-mentioned first text and the above-mentioned second text into the above-mentioned relationship recognition model to obtain the above-mentioned dependency relationship corresponding to the above-mentioned first text and the above-mentioned second text.
[0196] In this solution, the third recognition model to be trained can be used to process text sequence data, and the introduction of the Transformer enhances its ability to capture long-range dependencies through the self-attention mechanism. This means that the model can not only understand the text information within a single cell, but also intelligently analyze the logical relationships between texts, thereby accurately determining whether there are dependencies between texts.
[0197] In some embodiments, the determination unit includes a first determination module and a second determination module. The first determination module is used to determine that the type of the above-mentioned cell area is the above-mentioned merged cell when the above-mentioned relevant information includes the above-mentioned shape and the above-mentioned shape is a T-shape; the second determination module is used to determine that the type of the above-mentioned cell area is the above-mentioned ordinary cell when the above-mentioned relevant information includes the above-mentioned shape and the above-mentioned shape is not a T-shape.
[0198] In this solution, for tables with complex step sizes, T-shaped cells are often an obvious sign of merged cells. The cell type can be determined by the set shape, thereby further improving the recognition accuracy of the cell type.
[0199] During the specific implementation process, the determination unit includes a third determination module and a fourth determination module. The third determination module is used to determine that the type of the above-mentioned cell area is the above-mentioned merged cell when the above-mentioned relevant information includes the above-mentioned dependency relationship, and the above-mentioned first text of the above-mentioned cell area and the above-mentioned second text of the above-mentioned adjacent cell area have semantic correlation; the fourth determination module is used to determine that the above-mentioned type of the above-mentioned cell area is the above-mentioned ordinary cell when the above-mentioned relevant information includes the above-mentioned dependency relationship, and the above-mentioned first text of the above-mentioned cell area and the above-mentioned second text of the above-mentioned adjacent cell area do not have semantic correlation.
[0200] In this scheme, the information in the merged cells usually has a close semantic relationship with the surrounding cells. This relationship reflects the data ownership and logical connection between cells, that is, whether there is semantic relevance between the texts. By determining whether there is semantic relevance between the texts, the type of cell can be further accurately determined.
[0201] In some embodiments, the determination unit includes a fifth determination module and a sixth determination module. The fifth determination module is used to determine that the type of the above-mentioned cell area is the above-mentioned merged cell when at least one of the above-mentioned shape is a T-shape and there is semantic correlation between the above-mentioned first text in the above-mentioned cell area and the above-mentioned second text in the above-mentioned adjacent cell area is satisfied; the sixth determination module is used to determine that the type of the above-mentioned cell area is the above-mentioned ordinary cell when the above-mentioned shape is not a T-shape and there is no semantic correlation between the above-mentioned first text in the above-mentioned cell area and the above-mentioned second text in the above-mentioned adjacent cell area.
[0202] In this solution, the cell type can be comprehensively judged based on the two criteria of whether the shape and text have semantic relevance. The judgment logic is more detailed, so the cell type can be determined more accurately.
[0203] The above-mentioned table information extraction device includes a processor and a memory. The above-mentioned first acquisition unit, first extraction unit, determination unit, second extraction unit, third extraction unit, and sending unit are all stored as program units in the memory. The processor executes the above-mentioned program units stored in the memory to implement the corresponding functions. The above-mentioned modules are all located in the same processor; alternatively, the above-mentioned modules can be located in different processors in any combination.
[0204] The processor includes a kernel, which retrieves the corresponding program unit from the memory. One or more kernels can be set, and the problem of low efficiency of the information extracted from the table in the prior art can be solved by adjusting the kernel parameters.
[0205] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM, and the memory includes at least one memory chip.
[0206] An embodiment of the present invention provides a computer-readable storage medium, which includes a stored program. When the program is executed, the device where the computer-readable storage medium is located is controlled to execute the information extraction method of the table.
[0207] An embodiment of the present invention provides a processor, which is used to run a program, wherein the information extraction method of the table is executed when the program is run.
[0208] An embodiment of the present invention provides a device comprising a processor, a memory, and a program stored in the memory and executable on the processor. When the processor executes the program, it implements at least the steps of a table information extraction method. The device herein may be a server, a PC, a PAD, a mobile phone, or the like.
[0209] The present application also provides a computer program product, including a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the table information extraction method in each embodiment of the present application are implemented.
[0210] The present application also provides an information extraction system, comprising one or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include an information extraction method for executing any one of the above-mentioned tables.
[0211] Obviously, those skilled in the art will appreciate that the various modules or steps of the present invention described above can be implemented using a general-purpose computing device, can be centralized on a single computing device, or can be distributed across a network of multiple computing devices. They can be implemented using program code executable by the computing device, and thus, can be stored in a storage device and executed by the computing device. In some cases, the steps shown or described herein can be performed in a different order than that shown, or can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, the present invention is not limited to any particular combination of hardware and software.
[0212] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0213] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0214] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0215] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0216] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.
[0217] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.
[0218] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.
[0219] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.
[0220] The above description is merely a preferred embodiment of the present application and is not intended to limit the present application. Various modifications and variations are possible for those skilled in the art. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present application shall be included within the scope of protection of the present application.
Claims
1. A method for extracting information from a table, characterized in that: include: Acquire an image to be recognized, wherein the image to be recognized is an image of a table; Extracting relevant information of a cell area in the image to be recognized, wherein the relevant information includes a shape and / or a dependency relationship, wherein the cell area is an area having a single rectangular frame, the shape is a geometric shape between a first boundary line of the cell area and a second boundary line of an adjacent cell area having an intersection, the first boundary line is a boundary line composed of pixel points detected to obtain a boundary of the cell area, and the dependency relationship is whether there is semantic correlation between a first text in the cell area and a second text in the adjacent cell area; Determine the type of the cell area according to the relevant information, wherein the type includes a merged cell and a normal cell, the merged cell includes multiple cells, and the normal cell includes one cell; In the case where the type is the ordinary cell, extracting information of the cell area; In the case where the type is the merged cell, extracting information of all the cell areas in a covered area, wherein the covered area is the cell area including the merged cell and the cell areas having the semantic relevance with the merged cell; The extracted information is sent to a security monitoring system of a financial institution, so that the security monitoring system performs data monitoring based on the information in the cell area.
2. The method according to claim 1, characterized in that After acquiring the image to be recognized, the method further includes: Obtain a first recognition model to be trained, wherein the first recognition model to be trained is one of an HRNet model, a U-NET model, and a GAN model; The images to be recognized are grouped into a first training set, and the first to-be-trained recognition model is trained using the first training set and corresponding region labels to obtain a cell recognition model, wherein the region label is the region where the cell is located in the image of the first training set; The image to be identified is input into the cell recognition model to obtain the cell area corresponding to the image to be identified.
3. The method according to claim 1, characterized in that Extracting relevant information of a cell area in the image to be identified includes: Using edge detection technology to identify the boundaries of the cell area and obtain multiple current boundary lines; Using edge detection technology to identify the boundaries of the adjacent cell areas to obtain multiple adjacent boundary lines; Extracting the current boundary line and the adjacent boundary line having an intersection to obtain the first boundary line and the second boundary line; Obtaining a second recognition model to be trained, wherein the second recognition model to be trained is one of a decision tree model, an SVM model, and a random forest model; The first boundary line and the second boundary line are combined into a second training set, and the second to-be-trained recognition model is trained using the second training set and corresponding shape labels to obtain a shape recognition model, wherein the shape labels are geometric shapes between the boundary lines in the second training set; The first boundary line and the second boundary line are input into the shape recognition model to obtain shape recognition results corresponding to the first boundary line and the second boundary line.
4. The method according to claim 1, wherein Extracting relevant information of a cell area in the image to be identified includes: Obtain a third recognition model to be trained, wherein the third recognition model to be trained is one of an LSTM model, a word embedding model, and an NTN model; Adding a Transformer module before the output layer of the third recognition model to be trained, so that the output of the Transformer module is used as the input of the output layer to obtain a text recognition model; Using OCR extraction technology to extract text content of the cell area to obtain the first text; Using OCR extraction technology to extract text content of the adjacent cell area to obtain the second text; The first text and the second text are combined into a third training set, and the text recognition model is trained using the third training set and corresponding relationship labels to obtain a relationship recognition model, wherein the relationship label is a semantic correlation relationship between the first text and the second text; The first text and the second text are input into the relationship recognition model to obtain the dependency relationship between the first text and the second text.
5. The method according to claim 1, wherein Determining the type of the cell area according to the relevant information includes: In a case where the relevant information includes the shape and the shape is a T-shape, determining that the type of the cell area is the merged cell; In a case where the relevant information includes the shape and the shape is not a T-shape, it is determined that the type of the cell area is the ordinary cell.
6. The method according to claim 1, characterized in that Determining the type of the cell area according to the relevant information includes: In a case where the relevant information includes the dependency relationship and there is semantic relevance between the first text in the cell area and the second text in the adjacent cell area, determining that the type of the cell area is the merged cell; In a case where the relevant information includes the dependency relationship and there is no semantic correlation between the first text in the cell area and the second text in the adjacent cell area, it is determined that the type of the cell area is the ordinary cell.
7. The method according to claim 1, characterized in that Determining the type of the cell area according to the relevant information includes: If at least one of the following conditions is met: the shape is a T-shape; and the first text in the cell area and the second text in the adjacent cell area have semantic relevance, determining that the type of the cell area is the merged cell; When the shape is not a T-shape and there is no semantic correlation between the first text in the cell area and the second text in the adjacent cell area, the type of the cell area is determined to be the ordinary cell.
8. A device for extracting information from a table, characterized in that: include: A first acquiring unit is configured to acquire an image to be identified, wherein the image to be identified is an image of a table; a first extraction unit, configured to extract relevant information of a cell area in the image to be recognized, wherein the relevant information includes a shape and / or a dependency relationship, wherein the cell area is an area having a single rectangular frame, the shape is a geometric shape between a first boundary line of the cell area and a second boundary line of an adjacent cell area of the cell area having an intersection, the first boundary line is a boundary line composed of pixel points detected to obtain a boundary of the cell area, and the dependency relationship is whether there is semantic correlation between a first text in the cell area and a second text in the adjacent cell area; a determining unit, configured to determine a type of the cell area according to the relevant information, wherein the type includes a merged cell and a normal cell, the merged cell includes a plurality of cells, and the normal cell includes a single cell; a second extraction unit, configured to extract information of all the cell regions by using an OCR extraction technology based on the types of all the cell regions; a third extraction unit, configured to extract information of all the cell areas in a covered area when the type is the merged cell, wherein the covered area is the cell area including the merged cell and the cell areas having the semantic relevance with the merged cell; A sending unit is used to send the extracted information to a security monitoring system of a financial institution, so that the security monitoring system performs data monitoring based on the information in the cell area.
9. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the table information extraction method according to any one of claims 1 to 7 are implemented.
10. An information extraction system, characterized in that: include: One or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the one or more programs include a method for executing the information extraction method of the table according to any one of claims 1 to 7.