Picture recognition technology-based financial report table data processing method

By adopting a method based on image recognition technology in the financial data processing system, combining natural language processing and deep learning image algorithms, the difficulties of traditional systems in table recognition, text recognition and error correction recognition are solved, and more efficient and accurate financial report form data processing is achieved.

CN120148060APending Publication Date: 2025-06-13WONDERAIS (SHENZHEN) TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510303118.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

Traditional financial data processing systems have problems of insufficient accuracy and timeliness when processing financial statement data, especially in terms of difficulty in form recognition, text recognition and error correction recognition, resulting in a large number of manual verification and repair.

Method used

The data processing method of financial statement forms based on image recognition technology is adopted, combined with natural language processing, traditional image algorithms and deep learning image algorithms, and the automation rate and accuracy of data processing are improved through steps such as image preprocessing, table detection, table recognition and data error correction.

Benefits of technology

It significantly improves the automation rate and accuracy of data processing of financial statements, reduces the need for manual processing, and reduces the average processing time from 30 minutes to less than 3 minutes, improving the timeliness of data processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120148060A_ABST
    Figure CN120148060A_ABST
Patent Text Reader

Abstract

The invention relates to a financial report table data processing method based on a picture recognition technology, and the method comprises the steps: obtaining a financial announcement, and carrying out the following operations on a picture in the financial announcement: carrying out the image preprocessing: carrying out the detection of a line edge through an FLD algorithm, and correcting an inclined picture; table detection: constructing a YOLO table positioning model to obtain position and direction information of the table in the picture, and performing correct rotation; table identification: detecting table lines, and structuring a table into cell picture data; constructing a CRNN + CTC text recognition model and performing recognition, and constructing a BERT model for table classification; data error correction: extracting a financial subject list, constructing a BiLSTM-CRF subject fault-tolerant model, and obtaining a preliminary correct standard subject list; screening is carried out according to a cosine similarity algorithm, and balance verification is carried out according to a financial balance formula; and data processing and warehousing. The problem that a financial data processing system is insufficient in accuracy and timeliness in financial report table data processing is solved, and the automation rate of picture data processing is increased.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a data processing technology, and particularly to a method for quickly and accurately processing picture financial report form data. Background Art

[0002] In the financial industry, the information provided in financial statements is an important basis for national economic management departments to improve macro financial management and for investors to make investment decisions. Therefore, the financial industry has extremely high requirements for the timeliness and accuracy of financial data. Among financial report information, a large part is carried by pictures. However, in traditional financial data processing systems, as long as there are situations such as blurred pictures, tilted content, and slightly complex layout structures, a large amount of incorrect data will be generated. The main reasons are as follows:

[0003] 1. Difficulty in table recognition: Table positioning and table type recognition are both technical difficulties therein;

[0004] 2. Difficulty in text recognition of pictures: Especially for blurred, distorted, and text information with seal interference, the error rate of such information recognition is very high;

[0005] 3. Difficulty in correcting recognized incorrect content: It is difficult for the system to know which content is correctly recognized and which is incorrectly recognized, and of course it cannot automatically correct errors.

[0006] It requires financial professionals to spend a lot of time verifying and repairing financial account names and financial values, and it is difficult to meet the accuracy and timeliness requirements of enterprises for processing financial report form data. Summary of the Invention

[0007] Aiming at the problems of insufficient accuracy and timeliness in processing financial report form data by traditional financial data processing systems, a method for processing financial report form data based on picture recognition technology is proposed. Based on the characteristics of financial data, methods such as natural language processing, traditional image algorithms, and deep learning image algorithms are combined to improve the automation rate of picture data processing.

[0008] The technical solution of the present invention is as follows:

[0009] A method for processing financial report form data based on picture recognition technology, which obtains a financial announcement and performs the following operations on the picture in the financial announcement:

[0010] Step 1: Picture preprocessing; The specific steps are as follows:

[0011] Step 1.1. Convert the picture into JPG format;

[0012] Step 1.2. Adjust the picture size to 2000*1500;

[0013] Step 1.3. Use the FLD algorithm to detect the edges of the lines, count the inclination angles of the lines with a length greater than 20 pixels, and correct the inclined pictures according to the average inclination angle;

[0014] Step 2: Table detection; The specific steps are as follows:

[0015] Step 2.1. Build a YOLO table localization model, extract the corrected pictures for table localization sample annotation, and divide them into a training set and a test set. Train the YOLO table localization model through the training set and detect it through the test set to make the YOLO table localization model achieve the target localization effect;

[0016] Step 2.2. Send the corrected pictures to be recognized into the trained YOLO table localization model to obtain the position and orientation information of the tables in the pictures;

[0017] Step 2.3. Rotate the table pictures according to the orientation information obtained in the previous step to ensure the correctness of the picture orientation;

[0018] Step 3: Table recognition; The specific steps are as follows:

[0019] Step 3.1. Use the FLD algorithm to detect the table lines and structure the table into cell picture data;

[0020] Step 3.2. Build a CRNN+CTC text recognition model, scale the height of the structured cell pictures to 32 pixels, then perform sample annotation on the scaled pictures, divide the annotated sample data into a training set and a test set, train the CRNN+CTC text recognition model through the training set, and detect the training effect through the test set to make the recognition effect of the CRNN+CTC text recognition model reach the target effect;

[0021] Step 3.3. Build a BERT table classification model. After the structured cell pictures are recognized by the CRNN+CTC text recognition model, extract all the financial subject information from them, perform sample annotation, and then divide the annotated samples into a training set and a test set for building the BERT table classification model. Train through the training set and test the training effect through the test set to make the BERT table classification model reach the expected effect;

[0022] Step 3.4. Scale the cell pictures to be recognized to pictures with a height of 32 pixels, send them into the CRNN+CTC text recognition model to be recognized as text, filter out the financial report subject text information, and then send it into the BERT table classification model for table classification to determine the type of financial table;

[0023] Step 4: Data error correction; The specific steps are as follows:

[0024] Step 4.1. Extract financial subjects from the classified tabular data;

[0025] Step 4.2. Construct a BiLSTM-CRF subject fault tolerance model, label the extracted financial subjects. The labeling method is to label subjects with different representations as standard financial subjects, and divide the standard financial subject samples into a training set and a test set. Train the BiLSTM-CRF subject fault tolerance model with the training set and test the training effect with the test set to make the BiLSTM-CRF subject fault tolerance model reach the expected effect;

[0026] Step 4.3. Send the subjects to be converted into the trained BiLSTM-CRF subject fault tolerance model to obtain a preliminary correct list of standard subjects;

[0027] Step 4.4. According to the cosine similarity algorithm, discard subjects with similarity < 0.8 and retain subjects with similarity ≥ 0.8;

[0028] Step 4.5. According to the financial balance formula, perform a balance check on the values of the corresponding subjects. If the balance check is passed, proceed to Step Five;

[0029] Step Five: Data processing and warehousing

[0030] Save the data that has passed the balance check into the database according to the database table structure.

[0031] Furthermore, in Step 3.1, structure the table into cell image data. For a wireless table, generate virtual table lines in combination with the text positioning result.

[0032] Furthermore, the ratio of the training set to the test set for the CRNN+CTC text recognition model, the BERT table classification model, and the BiLSTM-CRF subject fault tolerance model is 9:1.

[0033] The beneficial effects of the present invention are as follows:

[0034] Since processing picture financial data needs to overcome problems such as picture blurring, stamping, and tilting, it takes more than 30 minutes to process an average table. The new financial report processing method is based on the characteristics of financial data and combines methods such as natural language processing, traditional image algorithms, and deep learning image algorithms, and can accurately and efficiently complete the extraction and conversion of information in financial pictures.

[0035] Compared with the traditional financial report tabular data processing method based on OCR technology, this processing method will reach the following technical level:

[0036] 1. More automated data processing: The automation rate of financial report form data processing has increased from 20% to over 80%, greatly reducing the scenarios of manual processing of financial report form data;

[0037] 2. More accurate data processing: The accuracy of automated financial data processing has increased from 40% to over 90%, reducing the workload of manual review;

[0038] 3. More efficient data processing: The average processing time has been reduced from 30 minutes to within 3 minutes, and the timeliness of data processing has increased by more than 10 times, improving the manual processing speed.

[0039] Based on the technical solution proposed in the present invention, it has been put into use in the production environment. The automation rate of picture processing is >80%, and the timeliness of financial data processing has been improved from T+30 to T+0. Description of the Drawings

[0040] Figure 1 It is a flowchart of the method for processing financial report form data based on picture recognition technology of the present invention. Detailed Embodiment

[0041] The present invention will be described in detail below with reference to the drawings and specific embodiments. This embodiment is implemented on the premise of the technical solution of the present invention, and the detailed implementation manners and specific operation processes are given, but the protection scope of the present invention is not limited to the following embodiments.

[0042] A method for processing financial report form data based on picture recognition technology includes the following steps:

[0043] Step 1: Picture preprocessing; The specific steps are as follows:

[0044] Step 1.1. Convert the picture into JPG format; Use the PIL library in Python to directly convert the picture into JPG.

[0045] Step 1.2. Adjust the picture size to 2000*1500; Use the PIL library to directly adjust the picture size to 2000*1500.

[0046] Step 1.3. Use the FLD algorithm (see Note 1 for details) to detect the line edges; Use the image operation functions provided in the OpenCV library to convert the image into a grayscale image, then create an FLD detector object, and then use the FLD to detect lines with a length greater than 20 pixels. Calculate the average inclination angle of all the obtained lines, and correct the picture according to the inclination angle.

[0047] Note 1: The FLD algorithm (Fast Line Detection) is a method for quickly detecting straight lines in an image based on the improved Hough transform.

[0048] Step 2: Table Detection; The specific steps are as follows:

[0049] Step 2.1. Build a YOLO table localization model (see Note 2 for details):

[0050] First, annotate the position and orientation of the table in the corrected image. Divide the annotated samples into a training set and a test set in a ratio of 9:1. Use the training command of YOLO, input the training set data, model weights, configuration data, and test set data, and perform multiple rounds of iterative testing until a model with a correct rate meeting the business requirements is trained.

[0051] Note 2: The YOLO model is a single-stage object detection model and has been widely used in many fields.

[0052] Step 2.2. Table Position and Orientation Detection

[0053] After correcting the image to be recognized, send it into the trained YOLO table localization model to obtain the position and orientation information of the table in the image;

[0054] Step 2.3. Rotate the Image

[0055] According to the orientation information obtained in the previous step, use the OpenCV library to rotate the table image to ensure the correctness of the image orientation;

[0056] Step 3: Table Recognition; The specific steps are as follows:

[0057] Step 3.1. Structure the table into cell image data

[0058] Use the FLD algorithm to detect the table lines and structure the table into cell image data;

[0059] For a table without lines, it is necessary to combine the text localization results, generate virtual table lines outside the text, and then treat it as a table with lines for cell image cutting.

[0060] Step 3.2. Build a CRNN+CTC text recognition model (see Note 3 for details):

[0061] Scale the height of the structured cell images to 32 pixels, and annotate the scaled images. Divide the annotated samples into a training set and a test set in a ratio of 9:1. Train the CRNN+CTC text recognition model through the training set and detect the training effect through the test set to make the recognition effect of the CRNN+CTC text recognition model reach the target effect;

[0062] Note 3: The CRNN+CTC text recognition model is a deep learning model for text recognition. This model combines the image feature extraction ability of the convolutional neural network (CNN) and the sequence modeling ability of the recurrent neural network (RNN), and then uses the CTC loss function to solve the unsupervised sequence labeling problem. Its main working principle is as follows: First, extract graphic features through the CNN to generate a feature map, send the feature map into the Bi-LSTM (a deep learning architecture) layer to further extract sequence features, and finally decode through CTC to convert it into the final text sequence.

[0063] Step 3.3. Build a BERT table classification model (see Note 4 for details):

[0064] Use the CRNN+CTC text recognition model to recognize the structured cell images well, extract all the financial subject information in them, perform sample annotation, divide the table data into a training set and a test set according to a ratio of 9:1, train the BERT table classification model through the training set, and test the training effect through the test set to make the BERT table classification model achieve the expected effect;

[0065] Note 4: The BERT model is a pre-trained language representation model. Its core idea is to use a large amount of unlabeled text data for pre-training to learn general language representations that can capture deep language features, and is applicable to text classification, question answering systems, named entity recognition, etc.

[0066] Step 3.4. Identify the type of financial table:

[0067] Structure the image to be recognized into cell images with a height of 32 pixels, send them into the CRNN+CTC text recognition model to be recognized as text, filter out the financial report subject text information, and then send the recognized text into the BERT table classification model for table classification to determine the type of financial table;

[0068] Step Four: Data error correction; the specific steps are as follows:

[0069] Step 4.1. Extract financial subjects from the classified table data;

[0070] Step 4.2. Build a BiLSTM-CRF subject fault tolerance model (see Note 5):

[0071] Perform sample annotation on the extracted financial subjects. The annotation method is to label subjects with different representation methods as standard financial subjects, and divide the annotated data into a training set and a test set according to a ratio of 9:1. Train the BiLSTM-CRF subject fault tolerance model through the training set, and test the training effect through the test set to make the BiLSTM-CRF subject fault tolerance model achieve the expected effect;

[0072] Note 5: The BiLSTM-CRF model is a hybrid model that combines a recurrent neural network and a conditional random field, and is widely used in sequence labeling tasks in natural language processing.

[0073] Step 4.3. Send the list of subjects to be converted into the trained BiLSTM-CRF subject error tolerance model to obtain a preliminary correct standard subject list;

[0074] Step 4.4. According to the cosine similarity algorithm, discard the subjects with a similarity < 0.8, and retain the subjects with a similarity >= 0.8;

[0075] Step 4.5. According to the financial balance formula, perform a balance check on the values of the corresponding subjects. If the balance check passes, proceed to Step 5;

[0076] Step Five: Data processing and warehousing

[0077] Save the data that has passed the balance check into the warehouse according to the database table structure.

[0078] The above target effects, expected effects, etc. can be specifically controlled according to the actual situation.

[0079] The above-described embodiments only represent one implementation manner of the present invention, and its description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention patent shall be subject to the appended claims.

Claims

1. A method for processing financial report table data based on image recognition technology, characterized in that: Get the financial announcement and perform the following operations on the images in the financial announcement: Step 1: Image preprocessing; the specific steps are as follows: Step 1.

1. Convert the image to JPG format; Step 1.

2. Resize the image to 2000*1500; Step 1.

3. Use the FLD algorithm to detect line edges, count the tilt angles of lines longer than 20 pixels, and correct the tilted image based on the average tilt angle; Step 2: Table detection; the specific steps are as follows: Step 2.

1. Build a YOLO table positioning model, extract the corrected image to annotate the table positioning sample, and divide it into a training set and a test set. Train the YOLO table positioning model with the training set and test it with the test set to make the YOLO table positioning model achieve the target positioning effect; Step 2.

2. After correcting the image to be recognized, send it to the trained YOLO table positioning model to obtain the position and direction information of the table in the image; Step 2.

3. According to the direction information obtained in the previous step, rotate the table image to ensure the correct direction of the image; Step 3: Table recognition; the specific steps are as follows: Step 3.

1. Use the FLD algorithm to detect table lines and structure the table into cell image data; Step 3.

2. Build a CRNN+CTC text recognition model, scale the structured cell image height to 32 pixels, then annotate the scaled image sample, divide the annotated sample data into a training set and a test set, train the CRNN+CTC text recognition model with the training set, and test the training effect with the test set, so that the recognition effect of the CRNN+CTC text recognition model reaches the target effect; Step 3.

3. Build a BERT table classification model. After the structured cell image is recognized by the CRNN+CTC text recognition model, all financial account information is extracted, and sample annotation is performed. Then, the annotated samples are divided into a training set and a test set for building a BERT table classification model. The training set is used for training, and the training effect is tested by the test set, so that the BERT table classification model achieves the expected effect. Step 3.

4. Scale the cell image to be recognized to a height of 32 pixels, send it to the CRNN+CTC text recognition model to recognize it as text, filter out the text information of financial report items, and then send it to the BERT table classification model for table classification to determine the type of financial table; Step 4: Data error correction; the specific steps are as follows: Step 4.

1. Extract financial items from the classified table data; Step 4.

2. Build a BiLSTM-CRF subject fault-tolerant model and label the extracted financial subjects. The labeling method is to label subjects with different representation methods as standard financial subjects, and divide the standard financial subject samples into training sets and test sets. The BiLSTM-CRF subject fault-tolerant model is trained with the training set, and the training effect is tested with the test set, so that the BiLSTM-CRF subject fault-tolerant model can achieve the expected effect. Step 4.

3. Send the subjects to be converted into the trained BiLSTM-CRF subject error-tolerant model to obtain a preliminary correct standard subject list; Step 4.

4. According to the cosine similarity algorithm, the subjects with similarity < 0.8 are discarded, and the subjects with similarity >= 0.8 are retained; Step 4.

5. According to the financial balance formula, check the value of the corresponding account. If the balance check passes, proceed to step 5; Step 5: Data processing and storage The data that has passed the balance check will be saved in the database according to the database table structure.

2. The financial report table data processing method based on image recognition technology according to claim 1 is characterized in that: In step 3.1, the table is structured into cell image data, and for wireless tables, virtual table lines are generated in combination with the text positioning results.

3. The financial report table data processing method based on image recognition technology according to claim 1 is characterized in that: The ratio of training set to test set of CRNN+CTC text recognition model, BERT table classification model and BiLSTM-CRF subject fault tolerance model is 9:1.