Money report key information extraction method and system

By combining optical character recognition technology and relative position relationship matching methods, the problems of table recognition errors and wrong rows and columns in financial report review are solved, achieving higher recognition accuracy and accuracy of extracting key information of financial report.

CN120198928AActive Publication Date: 2025-06-24ZHONGQI LIANXIN (BEIJING) TECH CO LTD
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202510675374.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-06-24
Estimated Expiration
2045-05-23

AI Technical Summary

Technical Problem

It is difficult for the existing technology to effectively deal with problems such as tilt, seal occlusion and picture distortion during the financial report review process, resulting in table identification errors and wrong rows and columns, affecting the accuracy of automated identification.

Method used

The recognized report pictures are initially identified through optical character recognition technology (OCR), identifying the table header and cell positions of the target data, matching the data with the table header based on the relative position relationship, correcting the problem of wrong rows and columns, and improving the recognition accuracy through steps such as character replacement and reversal judgment.

Benefits of technology

It has achieved effective handling of problems such as tilt, seal occlusion and picture distortion in the financial report, improved the accuracy of table recognition and the accuracy of extracting key information of financial report, and reduced the impact of wrong rows and columns and OCR recognition errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198928A_ABST
    Figure CN120198928A_ABST
Patent Text Reader

Abstract

The invention provides a financial report key information extraction method and system. The method comprises the steps of obtaining a report picture to be recognized, processing the report picture to be recognized into a first table by adopting optical character recognition, performing character recognition on cells in a first area, and performing first marking or second marking on a target data column based on a recognition result; traversing each column of the first table, judging whether the column is a character column or not based on the number of character cells and the number of digital cells of each column, dividing a second table into a plurality of areas based on the position of the character column, and determining the position relation between a first mark and a second mark for each area; and querying a target cell in the character column of each area, if the target cell is found, determining at least one digital column on one side of the column based on the position of the column where the target cell is located, matching the first mark or the second mark based on the position of the digital column, and determining the information of the target cell corresponding to the first mark or the second mark.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and in particular to a method and system for extracting key information from financial reports. Background Art

[0002] With the development of deep learning technology, financial report information extraction technology has been widely applied to the work of financial report auditing. Deep learning technology plays a crucial role in the field of table extraction, providing strong support for accurately and efficiently extracting table data from complex documents. In the table extraction task, deep learning models demonstrate excellent recognition and understanding capabilities. For the extraction of table content, deep learning models can combine optical character recognition (OCR) technology to accurately recognize and extract the text information in the table. Through training with large-scale datasets, the model can learn the text features under different fonts, font sizes, and layouts, and achieve high-precision extraction of various complex table contents. Deep learning technology also supports end-to-end table extraction, integrating steps such as table detection, structure recognition, and content extraction into a unified model, greatly simplifying the processing flow, improving the extraction efficiency, and providing strong guarantees for data processing and analysis in many fields such as finance, healthcare, and scientific research.

[0003] However, during the auditing process, there are often a large number of pictures such as customers' photos, scanned documents, and printed photos. Such pictures often contain tilted pictures, pictures with seals covering key words, and pictures with imperfect focus. The existing open-source models are not yet perfect in table structure detection, seal removal, and image distortion removal, etc., which are prone to causing situations of misaligned rows and columns, and are prone to causing recognition errors during the process of automatically identifying target information. Summary of the Invention

[0004] In view of this, embodiments of the present invention provide a method and system for extracting key information from financial reports to eliminate or improve one or more defects existing in the prior art.

[0005] One aspect of the present invention provides a method for extracting key information from financial reports. The steps of the method include: Obtain a report picture to be recognized, and use optical character recognition to process the report picture to be recognized into a first table, and perform character replacement on a first preset area in the first table to obtain a second table; Perform text recognition on the cells in the first area of the second table, and perform a first mark or a second mark on the target data column based on the recognition result; Traverse each column of the first table, determine whether the column is a character string based on the number of text cells and the number of numeric cells in each column, divide the second table into multiple regions based on the position of the character string, and determine the positional relationship between the first mark and the second mark for each region; Query the position of the target cell based on the text in the character string of each area. If the target cell is found in the area, determine at least one numeric column of the columns on one side of the column based on the position of the column where the target cell is located, and match the first marker or the second marker based on the position of the numeric column to determine the information of the first marker or the second marker corresponding to the target cell.

[0006] Adopting the above solution, this solution first performs a preliminary recognition on the report picture to be recognized including the report through optical character recognition technology (OCR). However, there may be certain problems of misaligned rows and columns in the preliminary results. This solution further recognizes the column position where the header of the target data is located and the row position where the target cell is located. Based on the relative position relationship between the actual data position on one side of the recognized target cell and the column position where the header of the target data is located, the data is matched with the header. When there are problems of misaligned rows and columns, the matching can also be performed according to the relative position relationship to complete the extraction of the data and ensure the extraction accuracy.

[0007] In some embodiments of the present invention, in the step of performing character replacement on the first preset area in the first table to obtain the second table, read the text in each cell in the first preset area. If the read text is the text to be replaced, replace the text with the preset replacement text.

[0008] In some embodiments of the present invention, in the step of performing character recognition on the cells in the first area of the second table and performing the first marker or the second marker on the target data column based on the recognition result, perform a first recognition on the cells in the first area of the second table to complete the preliminary marking, mark the cells as the first pre-marker, the second pre-marker or unmarked. For the cells marked with the first pre-marker and the second pre-marker, use the second recognition to perform a reverse determination on the cells marked with the first pre-marker and the second pre-marker.

[0009] In some embodiments of the present invention, in the step of performing a reverse determination on the cells marked with the first pre-marker and the second pre-marker using the second recognition: If the reverse determination of the cell marked with the first pre-marker fails, mark the cell with the first marker; if the reverse determination of the cell marked with the first pre-marker succeeds, mark the cell with the second marker; If the reverse determination of the cell marked with the second pre-marker fails, mark the cell with the second marker; if the reverse determination of the cell marked with the second pre-marker succeeds, mark the cell with the first marker.

[0010] In some embodiments of the present invention, the step of traversing each column of the first table and determining whether the column is a character string based on the number of text cells and the number of numeric cells in each column includes: For each cell in a column, make a determination one by one, determine the value of the text variable based on the number of text cells, and determine the value of the numeric variable based on the number of numeric cells; Compare the value of the text variable with the value of the numeric variable, and compare the value of the text variable with the text variable threshold to determine whether this column is a string column.

[0011] In some embodiments of the present invention, in the step of comparing the value of the text variable with the value of the numeric variable, and comparing the value of the text variable with the text variable threshold to determine whether this column is a string column, if the value of the text variable is greater than the value of the numeric variable and the value of the text variable is greater than the text variable threshold, then determine that this column is a string column.

[0012] In some embodiments of the present invention, in the step of dividing the second table into multiple regions based on the position of the string column, calculate from left to right based on the position of the string column. If a string column is recognized, calculate to the right. If another string column is recognized, take the columns between this string column and the other string column adjacent on the right as one region; if the table boundary is recognized, take the columns between this string column and the table boundary as one region.

[0013] In some embodiments of the present invention, in the step of determining the positional relationship between the first marker and the second marker for each region, calculate the distance between the cell of the first marker and the string columns on the left and right sides of this region or the table boundary, calculate the distance between the cell of the second marker and the string columns on the left and right sides of this region or the table boundary, and determine the positional relationship between the cell of the first marker and the cell of the second marker through the two distance values.

[0014] In some embodiments of the present invention, in the step of querying the position of the target cell based on the text in the string column of each region, if the target cell is found in the region, determine at least one numeric column of the columns on one side of the column where the target cell is located based on the position of the column where the target cell is located, and match the first marker or the second marker based on the position of the numeric column, the steps for determining the information of the first marker or the second marker corresponding to the target cell include: Match the target text in the string column of each region. If the target cell where the target text is located is found, determine at least one numeric column of the columns on one side of the column where the target cell is located based on the position of the column where the target cell is located; If there are two numeric columns on one side of the column where the target cell is located, based on the relative positions of the two numeric columns corresponding to the positional relationship between the cell of the first marker and the cell of the second marker, respectively determine the numeric columns corresponding to the first marker and the second marker, and determine the information of the first marker and the second marker corresponding to the row where the target cell is located in the two numeric columns.

[0015] In some embodiments of the present invention, if there is a numeric column on one side of the column where the target cell is located, then based on the horizontal position of the numeric column and the horizontal positions of the first-marked cell and the second-marked cell, it is determined whether the numeric column corresponds to the first mark or the second mark, and the information corresponding to the first mark or the second mark in the numeric column for the row where the target cell is located is determined.

[0016] In some embodiments of the present invention, if there is a numeric column on one side of the column where the target cell is located, then it is simultaneously determined whether there are at least three numeric columns in the area adjacent to the area where the target cell is located. If so, the numeric column closest to this area in the adjacent area is included in the calculation of this area, and the calculation is performed based on the existence of two numeric columns on one side of the column where the target cell is located.

[0017] In some embodiments of the present invention, in the step of determining at least one numeric column of the columns on one side of the column based on the position of the column where the target cell is located, the cells in a column are determined one by one. If the content of the cell is not 0 and the length of the digits except the decimal point is greater than a preset length, a true count is added to this column; if the content of the cell is 0 or not 0 but the length of the digits except the decimal point is not greater than the preset length, a false count is added to this column. The true count and the false count in a column are compared to determine whether this column is a numeric column.

[0018] A second aspect of the present invention also provides a financial report key information extraction system, which includes a computer device. The computer device includes a processor and a memory. Computer instructions are stored in the memory, and the processor is used to execute the computer instructions stored in the memory. When the computer instructions are executed by the processor, the system implements the steps implemented by the method described above.

[0019] A third aspect of the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps implemented by the aforementioned financial report key information extraction method.

[0020] The additional advantages, objectives, and features of the present invention will be partially described below and will become partially apparent to those of ordinary skill in the art after studying the following text, or can be learned from the practice of the present invention. The objectives and other advantages of the present invention can be pointed out and obtained specifically in the description and the accompanying drawings.

[0021] Those skilled in the art will understand that the objectives and advantages that can be achieved by the present invention are not limited to the above specific descriptions, and the above and other objectives that the present invention can achieve will be more clearly understood according to the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] The accompanying drawings described herein are used to provide a further understanding of the present invention, form a part of this application, and do not constitute a limitation to the present invention.

[0023] Figure 1 It is a schematic diagram of the first implementation manner of the method for extracting key information from the financial report of this solution; Figure 2 It is a schematic diagram of the second implementation manner of the method for extracting key information from the financial report of this solution; Figure 3 It is a schematic diagram of the third implementation manner of the method for extracting key information from the financial report of this solution; Figure 4 It is a schematic diagram of the implementation manner of step S300 of this solution; Figure 5 It is a page schematic diagram of the table processed by this solution. Detailed implementation manner

[0024] To make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in combination with the implementation manners and the accompanying drawings. Herein, the illustrative implementation manners of the present invention and their descriptions are used to explain the present invention, but do not constitute a limitation to the present invention.

[0025] Herein, it also needs to be noted that in order to avoid obscuring the present invention due to unnecessary details, only the structures and / or processing steps closely related to the solution according to the present invention are shown in the accompanying drawings, while other details less related to the present invention are omitted.

[0026] In the prior art, with the development of deep learning technology, the financial report information extraction technology has been widely applied to the work of financial report auditing. During the auditing process, a large number of pictures such as photos of customers, scanned documents, and printed photos are often involved. Such pictures often include tilted pictures, pictures with seals covering key words, and pictures with imperfect focus.

[0027] However, the existing open-source models are not perfect in terms of table structure detection, stamp removal, image distortion removal, etc. Therefore, after a series of steps, there will be exaggerated situations such as misaligned rows and columns, incorrect cell text recognition, and column names and column contents not in the same column in the generated Excel table of the financial report. Therefore, it is quite difficult to extract key or interesting information based on such an Excel table, and the extraction rates brought by different extraction rules are very different. It is worth mentioning that the OCR recognition effect has a relatively high accuracy for the values of many fields. Therefore, it is of great significance to extract key information from some irregularly arranged Excel tables. A higher extraction rate can directly provide data reference for enterprises and provide a basis for their decision-making. At present, there are not many cases of extracting key information from the Excel after financial report extraction on the market. Most of the existing ones have invested a lot in the early algorithms, such as solving the problem of misaligned rows and columns. However, few people use post-algorithms, that is, the ideas and algorithms in this article, to extract information based on the existing Excel table, to offset the problem of misaligned rows and columns and the OCR recognition problem brought by the previous model to a certain extent.

[0028] As Figure 1 and 5 shown, the present invention proposes a method for extracting key information from financial reports. The steps of this method include: Step S100, obtain the report picture to be recognized, and use optical character recognition to process the report picture to be recognized into a first table, and perform character replacement on a first preset area in the first table to obtain a second table; In the specific implementation process, the report in the report picture to be recognized can refer to the financial report of an enterprise, and the main body includes the income statement, cash flow statement and balance sheet. It is regularly released by an enterprise or organization to reflect its financial position, operating results and cash flow.

[0029] In the specific implementation process, using optical character recognition to process the report picture to be recognized into a first table, the optical character recognition (OCR) technology mainly converts the text in the image into a computer-readable text format. The working principle involves image processing technologies such as binarization, noise removal, and character segmentation to accurately recognize the text in the image. OCR is usually used to process unstructured image data such as scanned documents, text in photos or screenshots.

[0030] In the specific implementation process, in the step of using optical character recognition to process the report picture to be recognized into a first table, OCR is combined with a deep learning model. Specifically, the deep learning model can be NLP. The combination of OCR+NLP is applied in the field of information extraction, such as automatically extracting the required information from documents such as bills and contracts. OCR is responsible for converting the document image into text, and NLP is responsible for parsing these texts and extracting useful data.

[0031] In the specific implementation process, deep learning methods include the following techniques: (1) Artificial Neural Networks (ANN): including Multilayer Perceptrons (MLP), Convolutional Neural Networks (CNN), Recurrent Neural Networks (RNN), etc. ANN is the basis of deep learning and is used to simulate and learn complex non-linear relationships. (2) Convolution operation: The core operation in convolutional neural networks for extracting features from input images. The convolutional kernel (filter) slides over the image and performs feature extraction in local regions. (3) Pooling: An operation in convolutional neural networks for reducing the size of the feature map, extracting the main features, and reducing computational complexity, such as max pooling and average pooling. (4) Recurrent Neural Networks (RNN): Used to process sequential data (such as text, speech), with memory capabilities and able to consider information from previous time steps. (5) Long Short-Term Memory (LSTM) and Gated Recurrent Unit (GRU): Variants of RNN, designed to solve the long-term dependence problem faced by traditional RNNs and are particularly suitable for processing long sequence data. (6) Generative Adversarial Networks (GANs): An adversarial model composed of a generator and a discriminator, used to generate new data samples with a sense of reality, such as images, audio, etc. (7) Autoencoders: An unsupervised learning method for learning an effective representation of data, including sparse autoencoders, variational autoencoders, etc. (8) Deep Reinforcement Learning (DRL): Combining deep learning and reinforcement learning, used to solve tasks that require long-term decision-making and environmental interaction, such as game play optimization, robot control, etc. (9) Transfer Learning: Utilizing pre-trained models to be adjusted on new tasks to accelerate the training speed and improve performance. (10) Batch Normalization: A technique for accelerating the neural network training process and enhancing the model's generalization ability. (11) Optimization algorithms: such as Stochastic Gradient Descent (SGD) and its variants (such as Adam, RMSProp, etc.), used to adjust the weights and biases of the neural network to minimize the loss function.(12) Regularization techniques: such as Dropout, L1, and L2 regularization, are used to prevent overfitting in neural networks and improve generalization ability.

[0032] Step S200: Perform character recognition on the cells in the first area of the second table, and perform the first marking or the second marking on the target data column based on the recognition result. Step S300: Traverse each column of the first table, determine whether the column is a character string based on the number of text cells and the number of numeric cells in each column, divide the second table into multiple areas based on the position of the character string, and determine the positional relationship between the first marking and the second marking for each area. Step S400: Query the position of the target cell based on the text in the character string of each area. If the target cell is found in the area, determine at least one numeric column of the column on one side of the column where the target cell is located based on the position of the column where the target cell is located, and match the first marking or the second marking based on the position of the numeric column to determine the information of the first marking or the second marking corresponding to the target cell.

[0033] With the above solution, this solution first performs preliminary recognition on the to-be-recognized report picture including the report through the optical character recognition technology (OCR). However, there may be certain problems of misaligned rows and columns in the preliminary results. This solution further recognizes the column position where the header of the target data is located, and recognizes the row position where the target cell is located. Based on the relative positional relationship between the actual data position on one side of the recognized target cell and the column position where the header of the target data is located, match the data with the header. When there are problems of misaligned rows and columns, it can also be matched according to the relative positional relationship to complete the data extraction and ensure the extraction accuracy.

[0034] In some embodiments of the present invention, in the step of performing character replacement on the first preset area in the first table to obtain the second table, read the text in each cell in the first preset area. If the read text is the text to be replaced, replace the text with the preset replacement text.

[0035] In the specific implementation process, for the fields of the text in the cell such as 'year-end balance', 'beginning balance', 'year-end number', and 'beginning number'. If the core words observed in this solution are 'end' and 'beginning'. That is, if 'end' or 'beginning' appears in the field, it corresponds to 'year-end balance' and 'beginning balance'. Therefore, if OCR misrecognizes 'end' and 'beginning', it may have a greater impact; this solution needs to analyze the scenarios where 'end' and 'beginning' may be misrecognized by OCR. Specifically, the common misrecognitions of 'end', such as the text to be replaced of 'this', 'water', 'come', 'not', 'technique', and 'Zhu', can be replaced with the replacement text 'end'; the common misrecognitions of 'beginning', such as the text to be replaced of 'cut','repair', 'lining','shirt', and 'clothes', can be replaced with the replacement text 'beginning'.

[0036] Adopting the above solution can correct common recognition errors of recognized text.

[0037] As Figure 2 shown, in some embodiments of the present invention, in the step of performing text recognition on the cells in the first region of the second table and performing a first mark or a second mark on the target data column based on the recognition result, it includes step S210 of performing a first recognition on the cells in the first region of the second table to complete a preliminary mark, and marking the cells as a first pre-mark, a second pre-mark, or an unmarked. For the cells marked with the first pre-mark and the second pre-mark, in step S220, a second recognition is used to perform a reverse determination on the cells marked with the first pre-mark and the second pre-mark.

[0038] In some embodiments of the present invention, in the step of performing a first recognition on the cells in the first region of the second table to complete a preliminary mark and marking the cells as a first pre-mark, a second pre-mark, or an unmarked, since 'beginning' often corresponds to cells such as 'beginning of XX year', and 'end' often corresponds to cells such as 'end of XX year', the first pre-mark is a 'beginning' mark, the second pre-mark is an 'end' mark, and the unmarked is a mark marked as neither 'beginning' nor 'end'.

[0039] Through multi-sample data analysis, it is not feasible to identify only through 'end' and 'beginning' because sometimes due to problems such as picture quality issues (such as problems caused by out-of-focus blur) and problems caused by seal covering, this keyword may not be recognized. Therefore, a more stringent field recognition method, that is, combined recognition, is required.

[0040] For fields containing the word 'beginning', words such as 'last year' and 'previous year' may also appear in different companies. Therefore, we first screen for 'last', 'previous', and 'beginning'; then, we screen for 'beginning of period', and then for fields starting with 'last' and ending with 'amount', fields starting with 'previous' and ending with 'amount', fields starting with 'year' and ending with 'amount', fields starting with 'year' and ending with 'quantity', and fields starting with 'year' and ending with'remainder'. For fields containing the word 'end', in order to maintain the accuracy of recognition, the recognition of 'end of period' is supplemented. Subsequently, we screen for fields starting with 'period' and ending with 'amount', fields starting with 'period' and ending with'remainder', and fields starting with 'period' and ending with 'quantity'. The same processing is also done for the keyword fields to be extracted.

[0041] Specifically, the first mark corresponds to the beginning balance, and the second mark corresponds to the ending balance.

[0042] For example, since the end of the previous year is the beginning of this year, when 'end' is identified, it cannot be directly equated to the end of this year. If 'last year' or 'last year' is determined at the same time, the end of last year or the end of the previous year is equivalent to the beginning of this year. In this case, the original corresponding 'end' mark is reversed to the 'beginning' mark, so that the corresponding final first mark is the balance at the beginning of the year. The same applies when 'beginning' is identified.

[0043] like Figure 3 As shown, in some embodiments of the present invention, the step of using the second identification to perform a reversal determination on the first pre-marked and the second pre-marked cells includes: Step S221, if the reversal determination of the first pre-marked cell fails, the cell is first marked; if the reversal determination of the first pre-marked cell succeeds, the cell is second marked; Step S222: if the reversal determination of the second pre-marked cell fails, the cell is second-marked; if the reversal determination of the second pre-marked cell succeeds, the cell is first-marked.

[0044] Using the above scheme, this scheme needs to ensure that the first mark and the second mark correspond to the header phrase, such as the correspondence between "year-end balance". Therefore, this scheme first uses a single word for judgment, such as constructing a preliminary correspondence through the correspondence between the last word and the "year-end balance", and further corrects the original correspondence error through reversal judgment to ensure recognition accuracy.

[0045] like Figure 4 As shown, in some embodiments of the present invention, the step of traversing each column of the first table and determining whether the column is a text column based on the number of text cells and the number of digital cells in each column includes: Step S310, determining each cell of a column one by one, determining the text variable value based on the number of text cells, and determining the numeric variable value based on the number of numeric cells; Step S320, compare the text variable value with the digital variable value, and compare the text variable value with the text variable threshold to determine whether the column is a text column.

[0046] In the specific implementation process, the number of numeric cells is the numeric variable value, and the number of text cells is the text variable value.

[0047] In some embodiments of the present invention, in the step of comparing the text variable value with the digital variable value and comparing the text variable value with the text variable threshold to determine whether the column is a text column, if the text variable value is greater than the digital variable value and the text variable value is greater than the text variable threshold, then the column is determined to be a text column.

[0048] In the specific implementation process, the text variable threshold is 4. In the step of comparing the text variable value with the digital variable value, and comparing the text variable value with the text variable threshold to determine whether the column is a text column, if the text variable value is greater than the digital variable value, and the text variable value is greater than 4, then the column is determined to be a text column.

[0049] In some embodiments of the present invention, the step of dividing the second table into multiple areas based on the position of the text column includes, step S330, calculating from left to right based on the position of the text column, if a text column is recognized, then calculating to the right, if another text column is recognized, then the column between the text column and another adjacent text column on the right is regarded as one area; if a table boundary is recognized, then the column between the text column and the table boundary is regarded as one area.

[0050] In some embodiments of the present invention, the step of determining the positional relationship between the first mark and the second mark for each area includes, step S340, calculating the distance between the cell of the first mark and the text column on the left side and the text column on the right side of the area or the table boundary, calculating the distance between the cell of the second mark and the text column on the left side and the text column on the right side of the area or the table boundary, and determining the positional relationship between the cell of the first mark and the cell of the second mark by the two distance values.

[0051] In the specific implementation process, in the step of determining the positional relationship between the first marked cell and the second marked cell by using two distance values, both distance values ​​are horizontal distances, and the positional relationship between the first marked cell and the second marked cell is a horizontal positional relationship, that is, the first marked cell is on the right or left side of the second marked cell.

[0052] In some embodiments of the present invention, the position of the target cell is searched based on the text in the text column of each region, if the target cell is found in the region, at least one digital column of the column on one side of the column is determined based on the position of the column where the target cell is located, and the first mark or the second mark is matched based on the position of the digital column, and the step of determining the information of the target cell corresponding to the first mark or the second mark includes: Match the target text in the text column of each area, and if the target cell where the target text is located is matched, determine at least one numeric column on one side of the column based on the position of the column where the target cell is located; If there are two digital columns on one side of the column where the target cell is located, based on the relative positions of the two digital columns and the corresponding positional relationship between the cells marked with the first mark and the cells marked with the second mark, the digital columns corresponding to the first mark and the second mark are determined respectively, and the information corresponding to the first mark and the second mark in the two digital columns of the row where the target cell is located is determined.

[0053] Using the above scheme, if the positional relationship between the cell with the first mark and the cell with the second mark is that the cell with the first mark is on the left of the cell with the second mark, then based on the left-right positional relationship of the two number columns, the number column on the left corresponds to the first mark, and the number column on the right corresponds to the second mark; even if there is a certain misalignment between the two number columns, the correspondence can be completed based on the relative relationship, that is, for the columns required by this scheme, which are the year-end balance and the beginning-of-the-year balance, the above steps can ensure that the data in the two columns correspond, thereby improving the recognition accuracy.

[0054] In the specific implementation process, in the step of matching the target text in the text column of each area, the text column of each area is replaced. For example, if the target text is 'notes receivable', it is necessary to recognize that the 'ticket' in 'notes receivable' may be recognized as 'period', so 'payable period' is mapped to 'notes payable'.

[0055] In some embodiments of the present invention, if there is a numerical column on one side of the column where the target cell is located, the numerical column is determined to correspond to the first mark or the second mark based on the horizontal position of the numerical column and the horizontal position of the cell with the first mark and the cell with the second mark, and the information that the row where the target cell is located corresponds to the first mark or the second mark in the numerical column is determined.

[0056] Using the above scheme, even if there is only one digital column, the corresponding first marker or second marker is determined based on the distance between the digital column and the columns where the first marker and the second marker are located. Specifically, the columns where the first marker and the second marker with a closer distance are located are used as the corresponding first marker or second marker.

[0057] In some embodiments of the present invention, if there is a number column on one side of the column where the target cell is located, it is simultaneously determined whether there are at least three number columns in an area adjacent to the area where the target cell is located. If so, a number column close to the current area in the adjacent area is included in the calculation of the current area, and calculations are performed based on the existence of two number columns on one side of the column where the target cell is located.

[0058] With the above solution, in the case where there is only one digital column, there may be a deviation to other areas. Including adjacent areas in the calculation can further correct the deviation of the column.

[0059] In some embodiments of the present invention, if a region includes at least three digital columns, and at least one of the digital columns is included in the adjacent region for calculation, the remaining digital columns are calculated for this region. While solving the column offset, the calculation amount of the region with multiple digital columns is reduced, and the calculation accuracy of the region with multiple digital columns can also be guaranteed at the same time.

[0060] In some embodiments of the present invention, in the step of determining at least one numerical column of the columns on one side of a column based on the position of the column where the target cell is located, the cells in a column are determined one by one. If the content of a cell is not 0 and the length of the digits except the decimal point is greater than a preset length, a true count is added to the column; if the content of the cell is 0 or not 0, but the length of the digits except the decimal point is not greater than the preset length, a false count is added to the column. The true count and the false count in a column are compared to determine whether the column is a numerical column.

[0061] In the specific implementation process, the preset length is 6.

[0062] The beneficial effects of this solution include that the logic operation of this solution is purely rule-based and requires very limited computing resources; this solution can tolerate problems caused by misaligned columns and incorrect OCR recognition of column names within a certain range; this solution applies the rule experience brought by a large number of diverse data sets; the computer operation resource overhead of this solution is limited; this solution has strong flexibility and can quickly modify and optimize the solution for new situations to achieve key information extraction; for different key information, only the font mapping rule needs to be modified, and the extraction logic and method do not need to be modified.

[0063] In summary, this solution uses a post-processing algorithm to perform key information extraction on the Excel after financial report recognition. It makes up for the OCR problems caused by seals and the problems of misaligned rows and columns; for the Excel generated after financial report recognition with the influence of seals, this solution provides the logic and method for information extraction of the Excel with the influence of table structure; in the present invention, a solution is provided for solving the influence of seals on OCR recognition, especially for column names in certain positions; in the present invention, a certain degree of remedy is carried out for the misaligned column situation in the table structure detection caused by reasons such as image distortion and tilt.

[0064] An embodiment of the present invention also provides a financial report key information extraction system, which includes a computer device. The computer device includes a processor and a memory. Computer instructions are stored in the memory, and the processor is used to execute the computer instructions stored in the memory. When the computer instructions are executed by the processor, the system implements the steps implemented by the method described above.

[0065] An embodiment of the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps implemented by the aforementioned financial report key information extraction method. The computer-readable storage medium may be a tangible storage medium, such as a random access memory (RAM), memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, register, floppy disk, hard disk, removable storage disk, CD-ROM, or any other form of storage medium known in the technical field.

[0066] Those of ordinary skill in the art should understand that the various exemplary components, systems, and methods described in connection with the embodiments disclosed herein can be implemented in hardware, software, or a combination of both. Specifically, whether to implement in hardware or software depends on the specific application and design constraints of the technical solution. A person skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention. When implemented in hardware, it can be, for example, an electronic circuit, an application-specific integrated circuit (ASIC), appropriate firmware, a plug-in, a functional card, and so on. When implemented in software, the elements of the present invention are programs or code segments for performing the required tasks. The program or code segment can be stored in a machine-readable medium or transmitted through a data signal carried in a carrier wave over a transmission medium or a communication link.

[0067] It should be clear that the present invention is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present invention is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order between steps after understanding the spirit of the present invention.

[0068] In the present invention, the features described and / or illustrated for one embodiment can be used in the same way or in a similar way in one or more other embodiments, and / or combined with the features of other embodiments or replace the features of other embodiments.

[0069] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, various changes and variations can be made to the embodiments of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for extracting key information from financial reports, characterized in that, The steps of the method include: Obtain the report picture to be recognized, and use optical character recognition to process the report picture to be recognized into a first table. For the first preset area in the first table, perform character replacement to obtain a second table; Perform text recognition on the cells in the first area of the second table, and perform a first mark or a second mark on the target data column based on the recognition result; Traverse each column of the first table, determine whether the column is a character string based on the number of text cells and the number of numeric cells in each column, divide the second table into multiple regions based on the position of the character string, and determine the positional relationship between the first mark and the second mark for each region; Query the position of the target cell based on the text in the character string of each region. If the target cell is found in the region, determine at least one numeric column of the column on one side of the column where the target cell is located based on the position of the column where the target cell is located, and match the first mark or the second mark based on the position of the numeric column to determine the information of the first mark or the second mark corresponding to the target cell.

2. The method for extracting key information from financial reports according to claim 1, wherein In the step of performing character replacement on the first preset area in the first table to obtain a second table, read the text in each cell in the first preset area. If the read text is the text to be replaced, replace the text with the preset replacement text.

3. The method for extracting key information from financial reports according to claim 1, wherein In the step of performing text recognition on the cells in the first area of the second table and performing a first mark or a second mark on the target data column based on the recognition result, perform a first recognition on the cells in the first area of the second table to complete a preliminary mark, and mark the cells as a first pre-mark, a second pre-mark, or an unmarked. For the cells with the first pre-mark and the second pre-mark; Use a second recognition to perform a reverse determination on the cells with the first pre-mark and the second pre-mark. If the reverse determination on the cells with the first pre-mark fails, perform a first mark on the cells; If the reverse determination on the cells with the first pre-mark is successful, perform a second mark on the cells; If the reverse determination on the cells with the second pre-mark fails, perform a second mark on the cells; If the reverse determination on the cells with the second pre-mark is successful, perform a first mark on the cells.

4. The method for extracting key financial report information according to claim 1, wherein The step of traversing each column of the first table and determining whether the column is a character string based on the number of text cells and the number of numeric cells in each column includes: Perform a one-by-one determination on each cell in a column, determine the text variable value based on the number of text cells, and determine the numeric variable value based on the number of numeric cells; Compare the text variable value with the numeric variable value, and compare the text variable value with the text variable threshold to determine whether the column is a character string.

5. The method for extracting key financial report information according to claim 4, wherein In the step of comparing the text variable value with the numeric variable value and comparing the text variable value with the text variable threshold to determine whether the column is a character string, if the text variable value is greater than the numeric variable value and the text variable value is greater than the text variable threshold, determine that the column is a character string.

6. The method for extracting key financial report information according to claim 1, wherein In the step of dividing the second table into multiple regions based on the position of the character string, the position of the character string is calculated from left to right. If a character string is recognized, the calculation proceeds to the right. If another character string is recognized, the columns between this character string and the adjacent character string on the right are taken as one region. If the table boundary is recognized, the columns between this character string and the table boundary are taken as one region.

7. The method for extracting key financial report information according to claim 1, wherein In the step of determining the positional relationship between the first marker and the second marker for each region, calculate the distances between the cell of the first marker and the character string on the left side and the character string on the right side or the table boundary of this region, and calculate the distances between the cell of the second marker and the character string on the left side and the character string on the right side or the table boundary of this region. Determine the positional relationship between the cell of the first marker and the cell of the second marker through the two distance values.

8. The method for extracting key financial report information according to any one of claims 1 to 7, characterized in that, In the step of querying the position of the target cell based on the characters in the character string of each region, if the target cell is found in the region, determine at least one numerical column of the columns on one side of the column where the target cell is located based on the position of the column where the target cell is located, and match the first marker or the second marker based on the position of the numerical column, and determine the information of the target cell corresponding to the first marker or the second marker, which includes: Match the target text in the character string of each region. If the target cell where the target text is located is found, determine at least one numerical column of the columns on one side of the column where the target cell is located based on the position of the column where the target cell is located. If there are two numerical columns on one side of the column where the target cell is located, based on the relative positions of the two numerical columns corresponding to the positional relationship between the cell of the first marker and the cell of the second marker, respectively determine the numerical columns corresponding to the first marker and the second marker, and determine the information of the row where the target cell is located corresponding to the first marker and the second marker in the two numerical columns.

9. The method for extracting key financial report information according to claim 8, wherein In the step of determining at least one numerical column of the columns on one side of the column based on the position of the column where the target cell is located, judge each cell in a column one by one. If the content of the cell is not 0 and the length of the digits except the decimal point is greater than the preset length, add a true count to this column; if the content of the cell is 0 or not 0 but the length of the digits except the decimal point is not greater than the preset length, add a false count to this column. Compare the true count and the false count in a column to determine whether this column is a numerical column.

10. A key information extraction system for financial reports, characterized in that, The system includes a computer device, the computer device includes a processor and a memory, the memory stores computer instructions, and the processor is used to execute the computer instructions stored in the memory. When the computer instructions are executed by the processor, the system implements the steps implemented by the method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Report processing method and device

    CN113177551A

  • Financial statement processing method and device

    CN115457578A

  • Financial statement information extraction method and device based on OCR technology

    CN115546806A

  • Table recognition method and device, electronic equipment and storage medium

    CN116071768A

  • Financial statement reconstruction method and device, computer equipment and medium

    CN116311304A