Bank account bill automatic processing method and system based on OCR (Optical Character Recognition)

Through the OCR-based method, the watermarks and seals in bank statements are automatically identified and removed, and combined with deep learning and NLP technology, the automated processing and structured output of bank statements are realized, solving the problem of inefficient processing of bank statements and improving the accuracy and efficiency of data processing.

CN120296070APending Publication Date: 2025-07-11FUYANG NORMAL UNIVERSITY

Patent Information

Application Number
CN202510353218.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

Bank statements exist in the form of paper documents or electronic images, resulting in inefficient information extraction and poor accuracy, difficulty in achieving automated processing and data sorting, and unable to meet the real-time and efficient requirements of banking business.

Method used

Using an OCR-based method, the watermark and seal recognition model is constructed, the watermark and seal recognition module is used to remove watermarks and seals using multi-scale extrusion excitation attention module and adversarial network, and text recognition and table reconstruction are combined with deep learning and NLP technology to output structured Excel or JSON data.

Benefits of technology

It realizes the automated processing of bank statements, reduces manual operations, improves data processing efficiency and accuracy, supports structured output, and ensures data integrity and reliability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296070A_ABST
    Figure CN120296070A_ABST
Patent Text Reader

Abstract

The invention relates to an OCR (Optical Character Recognition)-based bank account bill automatic processing method. The method comprises the following steps: inputting bank account bill data in various formats; judging whether the document format is a picture format, and converting the document format into the picture format; constructing a watermark and seal identification model; removing the watermark and the seal on the picture; character recognition is carried out, and character content in the picture is extracted to form a text; reconstructing the text according to a table plate of the original document or picture; and adjusting and confirming the reconstructed table structure to obtain formatted table data, and exporting the formatted table data into Excel and JSON formats. Manual operation can be reduced, and the data processing efficiency is improved; the OCR and NLP technologies are combined, so that high-accuracy information extraction is realized; watermarks or seals can be automatically detected and removed, and the quality of text recognition is improved; format output of Excel, JSON and the like is supported, and subsequent analysis and storage are facilitated; the method has a verification function, and ensures the accuracy and integrity of final data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of image processing and optical character recognition, and in particular to an automated processing method and system for bank statement based on OCR. Background Art

[0002] After converting a paper document into an electronic image, it is difficult to edit and analyze due to its complex structure and background interference, which limits the value development. Although OCR technology has been mature in character recognition, it is still insufficient for complex documents. Document image reconstruction technology can convert images into editable data, and is widely used in fields such as bank statement analysis, file management, and automatic marking.

[0003] In the process of banking business and financial management, bank statements, as important financial vouchers, their accuracy and traceability are crucial for the operation, audit, and risk control of banks. However, many bank statement data still exist in the form of paper documents, electronic images, or scanned copies. Direct processing of these unstructured data formats faces many technical challenges and defects: First, paper and image-based statement data cannot be directly edited or automatically processed, resulting in low information extraction efficiency and heavy reliance on manual entry. Manual entry is not only time-consuming and laborious, but also prone to data inaccuracies due to human errors, affecting the reliability of subsequent analysis and decision-making; Second, the text in images or scanned copies may have problems such as blurring, tilting, occlusion, or background interference, increasing the recognition difficulty of optical character recognition (OCR) technology and further reducing the accuracy and efficiency of data extraction. In addition, unstructured data is difficult to seamlessly integrate with existing database systems, resulting in a complex and time-consuming data collation and classification process, unable to meet the requirements of banking business for real-time and high efficiency. Finally, due to the lack of intelligent data processing technology, banks are difficult to automate in aspects such as anomaly detection, trend analysis, and prediction of statement data, and still require a large amount of manual intervention, increasing the operating costs and the difficulty of risk management. In short, these defects seriously restrict the efficiency improvement of banks in digital transformation and the full excavation of data value. Summary of the Invention

[0004] To solve the problems of low parsing accuracy and low efficiency in the prior art, the primary object of the present invention is to provide an OCR-based automated processing method for bank statements that can effectively process input bank statement documents or pictures and output structured Excel or JSON data to achieve efficient and accurate automated parsing of bank statements.

[0005] To achieve the above object, the present invention adopts the following technical solutions: An OCR-based automated processing method for bank statements, the method includes the following steps in sequence:

[0006] (1) Input bank statement data in multiple formats, where the bank statement data includes documents and pictures;

[0007] (2) Automatically detect the format of the bank statement data and determine whether it is in picture format. If the determination result is yes, proceed to the next step; otherwise, convert the document format to picture format;

[0008] (3) Construct a watermark and seal recognition model to recognize watermarks and seals on the picture;

[0009] (4) Remove the watermarks and seals on the picture through an adversarial network to obtain a picture after removing the watermarks and seals;

[0010] (5) Perform optical character recognition on the picture after removing the watermarks and seals, and extract the text content in the picture to form a text;

[0011] (6) Reconstruct the text according to the table layout of the original document or picture to obtain a reconstructed table structure;

[0012] (7) Adjust and confirm the reconstructed table structure to obtain formatted table data;

[0013] (8) Export the formatted table data in Excel and JSON formats.

[0014] In step (3), the construction of the watermark and seal recognition model specifically means that the watermark and seal recognition model is obtained by taking the YOLOv9 model as the benchmark network and introducing a multi-scale squeeze-and-excitation attention module;

[0015] The multi-scale squeeze-and-excitation attention module is obtained by improving the EMA module: replace the global average pooling layer in the X direction, the global average pooling layer in the Y direction, and the two 2D global average pooling layers in the EMA module with an adaptive average pooling layer in the X direction, an adaptive average pooling layer in the Y direction, and two 2D adaptive average pooling layers respectively; connect a squeeze module in series after each of the two 2D adaptive average pooling layers; divide the input data into multiple groups, perform adaptive average pooling operations on different parts of the input data respectively, perform convolution operations using a 3×3 convolution kernel, splice the processing results of the adaptive average pooling layer in the X direction and the adaptive average pooling layer in the Y direction, then perform convolution operations using a 1×1 convolution kernel, re-weight the output features, then perform group normalization operations, perform adaptive average pooling operations, softmax function operations, and matrix multiplication operations on the features of the 1×1 and 3×3 scales, and then perform re-weighting, and finally output;

[0016] Connect the multi-scale squeeze-and-excitation attention module in series after the first GELAN module in the backbone network of the YOLOv9 model.

[0017] In step (3), the loss function of the watermark and seal recognition model adopts the Innerα-IoU loss function, and its formula is:

[0018]

[0019] In the formula, inter represents the intersection area of the predicted box and the ground truth box, and union represents the union area of the predicted box and the ground truth box. and b r represent the right boundary coordinates of the predicted box and the ground truth box respectively. and b l represent the left boundary coordinates of the predicted box and the ground truth box respectively. and b b represent the bottom boundary coordinates of the predicted box and the ground truth box respectively. and b t represent the top boundary coordinates of the predicted box and the ground truth box respectively. w gt and h gt represent the width and height of the predicted box respectively, and w and h represent the width and height of the ground truth box respectively. represents the abscissa value of the center point of the detection box. represents the ordinate value of the center point of the detection box. x c represents the abscissa value of the center point of the ground truth box, and y c represents the ordinate value of the center point of the ground truth box; ratio is a scaling factor with a value of 0.7; α is a power variable with a value of 3.0; eps is a constant with a value of 10 -7 .

[0020] In step (4), the adversarial network adopts the EraseNet adversarial network.

[0021] Step (6) specifically includes the following steps in sequence:

[0022] (6a) Scale the input image to obtain the scaled image;

[0023] (6b) Detect all text blocks in the scaled image, and obtain the coordinate information of the four corners of the text block and the text information;

[0024] (6c) Use the continuous local peak detection algorithm for the coordinate information to parse and obtain the table structure;

[0025] (6d) Perform cell positioning on the table structure, and reconstruct the obtained text information into the table structure to obtain the reconstructed table structure.

[0026] In step (6c), the continuous local peak detection specifically refers to: extracting the abscissa of the left border of the text block and the mean value of the abscissa of the left border, and taking the abscissa of the left border as the column left coordinate value; generating a column coordinate list from the column left coordinate values, using the dynamic local difference method to remove noise points from the column coordinate list, using the differential method to calculate the difference between adjacent points in the column coordinate list, obtaining a first-order difference sequence for detecting local peaks; the first-order difference represents the instantaneous change rate of the signal, traversing the first-order difference sequence, finding the position of the maximum value changing from negative to positive, i.e., the index value, and outputting the index value and the corresponding column left coordinate value;

[0027] Extracting the abscissa of the right border of the text block and the mean value of the abscissa of the right border as the column right coordinate value; extracting the abscissa of the right border of the text block and the mean value of the abscissa of the right border, and taking the abscissa of the right border as the column right coordinate value; generating a column coordinate list from the column right coordinate values, using the dynamic local difference method to remove noise points from the column coordinate list, using the differential method to calculate the difference between adjacent points in the column coordinate list, obtaining a first-order difference sequence for detecting local peaks; the first-order difference represents the instantaneous change rate of the signal, traversing the first-order difference sequence, finding the position of the maximum value changing from negative to positive, i.e., the index value, and outputting the index value and the corresponding column right coordinate value;

[0028] Taking the average value of the column right coordinate value of the first column and the column left coordinate value of the second column of the coordinate information as the final right coordinate of the first column, and traversing all column coordinates in this way to obtain the final column coordinates, and determining the bounding box of each cell through the final column coordinates to identify the content within the cell and obtain the table structure.

[0029] Another object of the present invention is to provide a system for an automated processing method of bank statement based on OCR, including:

[0030] An input module for receiving the input of bank statement data;

[0031] An image judgment module for determining whether the input bank statement data is in image format and converting the document format into a picture format;

[0032] A watermark and seal recognition model for detecting watermark and seal information in the picture;

[0033] A watermark and seal removal module using the EraseNet adversarial network for removing watermark and seal information;

[0034] A text recognition module for performing text recognition on the picture after removing the watermark and seal based on OCR technology, i.e., optical character recognition technology, to extract text, and the extracted text content forms a text;

[0035] A document structure reconstruction module, which is used to analyze the logical structure of a document and reconstruct the extracted text information in the form of a table of the original document or picture to obtain a reconstructed table structure; the original document is the document in the received bank statement data;

[0036] A verification module, which is used to adjust and confirm the reconstructed table structure to obtain formatted table data;

[0037] A formatted output module, which is used to export the formatted table data in Excel and JSON formats.

[0038] As can be seen from the above technical solutions, the beneficial effects of the present invention are as follows: First, it can perform automated processing: reduce manual operations and improve data processing efficiency; Second, it realizes high-precision recognition: combines deep learning OCR and NLP technologies to achieve high-accuracy information extraction; Third, it realizes intelligent watermark removal: can automatically detect and remove watermarks or seals to improve the quality of text recognition; Fourth, it performs structured output: supports output in formats such as Excel and JSON, facilitating subsequent analysis and storage; Fifth, it realizes intervention: has a verification function to ensure the accuracy and integrity of the final data. Description of the Drawings

[0039] Figure 1 is the method flow chart of the present invention;

[0040] Figure 2 is the structural schematic diagram of the watermark and seal recognition model in the present invention;

[0041] Figure 3 is the output schematic diagram of the present invention;

[0042] Figure 4 is the verification effect diagram of the present invention. Detailed Embodiments

[0043] As Figure 1 shown, an automated processing method for bank statements based on OCR, the method includes the following steps in sequence:

[0044] (1) Input bank statement data in multiple formats, and the bank statement data includes documents and pictures;

[0045] (2) Automatically detect the format of the bank statement data and determine whether it is in picture format. If the judgment result is yes, proceed to the next step; otherwise, convert the document format to picture format;

[0046] (3) Build a watermark and seal recognition model to recognize the watermarks and seals on the picture;

[0047] (4) Remove watermarks and seals on the image through an adversarial network to obtain the image after removing watermarks and seals;

[0048] (5) Perform text recognition on the image after removing watermarks and seals, and extract the text content in the image to form a text;

[0049] (6) Reconstruct the text according to the table layout of the original document or image to obtain the reconstructed table structure; Document reconstruction is a kind of NLP technology, namely natural language processing technology;

[0050] (7) Adjust and confirm the reconstructed table structure to obtain formatted table data to improve data integrity;

[0051] (8) Export the formatted table data into Excel and JSON formats to meet the requirements of different application scenarios.

[0052] As Figure 2 shown, in step (3), the construction of the watermark and seal recognition model specifically refers to: The watermark and seal recognition model is obtained by using the YOLOv9 model as the benchmark network and introducing a multi-scale squeeze-and-excitation attention module;

[0053] The multi-scale squeeze-and-excitation attention module is obtained by improving the EMA module: Replace the global average pooling layer in the X direction, the global average pooling layer in the Y direction, and the two 2D global average pooling layers in the EMA module with an adaptive average pooling layer in the X direction, an adaptive average pooling layer in the Y direction, and two 2D adaptive average pooling layers respectively; Connect a squeeze module in series after the two 2D adaptive average pooling layers; Divide the input data into multiple groups, that is, groups, and perform adaptive average pooling operations on different parts of the input data respectively, perform convolution operations using a 3×3 convolution kernel, splice the processing results of the adaptive average pooling layer in the X direction and the adaptive average pooling layer in the Y direction, then perform convolution operations using a 1×1 convolution kernel, re-weight the output features, that is, Re-weight, and then perform group normalization operations, that is, GroupNorm, and perform adaptive average pooling operations, softmax function operations, and matrix multiplication operations, that is, Matmul, on the features of the 1×1 and 3×3 scales again, re-weight the features again, and finally output;

[0054] Connect the multi-scale squeeze-and-excitation attention module after the first GELAN module in the backbone network of the YOLOv9 model. Introducing the multi-scale squeeze-and-excitation attention module can simultaneously process complex scenarios of different scale features and optimize the feature channel weights

[0055] In step (3), the loss function of the watermark and seal recognition model adopts the Innerα-IoU loss function, and its formula is:

[0056]

[0057]

[0058] Wherein, inter represents the intersection area between the predicted bounding box and the ground truth bounding box, and union represents the union area between the predicted bounding box and the ground truth bounding box. and b r respectively represent the right boundary coordinates of the predicted bounding box and the ground truth bounding box. and b l respectively represent the left boundary coordinates of the predicted bounding box and the ground truth bounding box. and b b respectively represent the bottom boundary coordinates of the predicted bounding box and the ground truth bounding box. and b t respectively represent the top boundary coordinates of the predicted bounding box and the ground truth bounding box. w gt and h gt respectively represent the width and height of the predicted bounding box, and w and h respectively represent the width and height of the ground truth bounding box. represents the abscissa value of the center point of the detection bounding box. represents the ordinate value of the center point of the detection bounding box. x c represents the abscissa value of the center point of the ground truth bounding box. y c represents the ordinate value of the center point of the ground truth bounding box; ratio is a scaling factor with a value of 0.7; α is a power variable with a value of 3.0; eps is a constant with a value of 10 -7 .

[0059] Introducing the Innerα-IoU loss function can suppress the influence of outliers on the model, accelerate the convergence speed, improve the detection accuracy and localization accuracy, and significantly enhance the feature expression ability of the network. The Innerα-IoU loss function focuses on the feature differences inside the bounding box. By suppressing the interference of irrelevant backgrounds on the loss calculation, reducing the influence of noise on the gradient update, enhancing the recognition effect of overlapping objects and small objects, and introducing a power variable α to dynamically adjust the contribution ratio of the localization and classification errors. When α is greater than 1, it will increase the loss and gradient of objects with a high intersection over union ratio, thereby improving the accuracy of bounding box regression, converging to the optimal solution faster, and improving the overall detection accuracy. Generally, good results can be obtained when α = 3.

[0060] In step (4), the adversarial network adopts the EraseNet adversarial network.

[0061] Step (6) specifically includes the following steps in sequence:

[0062] (6a) Scale the input image to obtain the scaled image.

[0063] (6b) Detect all text blocks in the scaled image, and obtain the coordinate information of the four corners of the text blocks and the text information;

[0064] (6c) For the coordinate information, use the continuous local peak detection algorithm to parse and obtain the table structure;

[0065] (6d) Perform cell localization on the table structure, reconstruct the obtained text information into the table structure, and obtain the reconstructed table structure.

[0066] In step (6c), the continuous local peak detection specifically refers to: extracting the abscissa of the left border of the text block and the mean value of the abscissa of the left border, and taking the abscissa of the left border as the column left coordinate value; generating a column coordinate list from the column left coordinate values, using the dynamic local difference method to remove noise points from the column coordinate list, using the differential method to calculate the difference between adjacent points in the column coordinate list, obtaining a first-order difference sequence for detecting local peaks; the first-order difference represents the instantaneous change rate of the signal, traversing the first-order difference sequence, finding the position of the maximum value of the change from negative to positive, that is, the index value, and outputting the index value and the corresponding column left coordinate value;

[0067] Extract the abscissa of the right border of the text block and the mean value of the abscissa of the right border, and take it as the column right coordinate value; extract the abscissa of the right border of the text block and the mean value of the abscissa of the right border, and take the abscissa of the right border as the column right coordinate value; generate a column coordinate list from the column right coordinate values, use the dynamic local difference method to remove noise points from the column coordinate list, use the differential method to calculate the difference between adjacent points in the column coordinate list, obtaining a first-order difference sequence for detecting local peaks; the first-order difference represents the instantaneous change rate of the signal, traversing the first-order difference sequence, finding the position of the maximum value of the change from negative to positive, that is, the index value, and outputting the index value and the corresponding column right coordinate value;

[0068] Take the average value of the column right coordinate value of the first column and the column left coordinate value of the second column of the coordinate information, and take this average value as the final column right coordinate of the first column. Traverse all column coordinates in this way to obtain the final column coordinates. Through the final column coordinates, determine the bounding box of each cell, identify the content within the cell, and obtain the table structure.

[0069] Another object of the present invention is to provide an OCR-based automated processing system for bank statement, including:

[0070] An input module for receiving the input of bank statement data;

[0071] An image judgment module for determining whether the input bank statement data is in image format and converting the document format into a picture format;

[0072] A watermark and seal recognition model for detecting watermark and seal information in the picture;

[0073] A watermark and seal removal module, which adopts the EraseNet adversarial network and is used to remove watermark and seal information;

[0074] A text recognition module, which is used to perform text recognition and extraction on the picture after removing the watermark and seal based on the OCR technology, i.e., optical character recognition technology, and the extracted text content forms a text;

[0075] A document structure reconstruction module, which is used to analyze the logical structure of the document and reconstruct the extracted text information according to the table format of the original document or picture to obtain a reconstructed table structure; the original document is the document in the received bank statement data;

[0076] A verification module, which is used to adjust and confirm the reconstructed table structure to obtain formatted table data;

[0077] A formatted output module, which is used to export the formatted table data in Excel and JSON formats.

[0078] Figure 3 The screenshot of the output Excel result after processing the Ping An Bank statement through the present invention.

[0079] As Figure 4 shown, the previous page and next page buttons can be used to view the text modification situations of other seal parts, Figure 4 and the verification area in it can be directly edited, modified by referring to the picture, and after verification is completed, confirm and save.

[0080] In summary, the present invention can perform automated processing: reduce manual operations and improve data processing efficiency; achieve high-precision recognition: combine deep learning OCR and NLP technologies to achieve high-accuracy information extraction; achieve intelligent watermark removal: can automatically detect and remove watermarks or seals to improve the quality of text recognition; perform structured output: support output in formats such as Excel and JSON, which is convenient for subsequent analysis and storage; achieve intervention: have a verification function to ensure the accuracy and integrity of the final data.

[0081] The above shows and describes the basic principle, main features and advantages of the present invention. Those skilled in the art of this industry should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification is only the principle of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.

Claims

1. An automated processing method for bank statement based on OCR, characterized in that: The method includes the following steps in sequence: (1) Input bank statement data in multiple formats, where the bank statement data includes documents and pictures; (2) Automatically detect the format of the bank statement data and determine whether it is in picture format. If the judgment result is yes, proceed to the next step; otherwise, convert the document format to picture format; (3) Construct a watermark and seal recognition model to recognize watermarks and seals on the picture; (4) Remove the watermarks and seals on the picture through an adversarial network to obtain a picture after removing the watermarks and seals; (5) Perform optical character recognition on the picture after removing the watermarks and seals, and extract the text content in the picture to form text; (6) Reconstruct the text according to the table layout of the original document or picture to obtain a reconstructed table structure; (7) Adjust and confirm the reconstructed table structure to obtain formatted table data; (8) Export the formatted table data in Excel and JSON formats.

2. The OCR-based automated bank statement processing method according to claim 1, wherein: In step (3), the construction of the watermark and seal recognition model specifically means that the watermark and seal recognition model is obtained by taking the YOLOv9 model as the baseline network and introducing a multi-scale squeeze-and-excitation attention module; The multi-scale squeeze-and-excitation attention module is obtained by improving the EMA module: replace the global average pooling layer in the X direction, the global average pooling layer in the Y direction, and the two 2D global average pooling layers in the EMA module with an adaptive average pooling layer in the X direction, an adaptive average pooling layer in the Y direction, and two 2D adaptive average pooling layers respectively; connect a squeeze module in series after each of the two 2D adaptive average pooling layers; divide the input data into multiple groups, perform adaptive average pooling operations on different parts of the input data respectively, perform convolution operations using a 3×3 convolution kernel, splice the processing results of the adaptive average pooling layer in the X direction and the adaptive average pooling layer in the Y direction, then perform convolution operations using a 1×1 convolution kernel, re-weight the output features, then perform group normalization operations, perform adaptive average pooling operations, softmax function operations, and matrix multiplication operations on the features of the two scales of 1×1 and 3×3, and then perform re-weighting, and finally output; Connect the multi-scale squeeze-and-excitation attention module after the first GELAN module in the backbone network of the YOLOv9 model.

3. The OCR-based automated processing method for bank statement according to claim 1, wherein: In step (3), the loss function of the watermark and seal recognition model uses the Innerα-IoU loss function, and its formula is: Wherein, inter represents the intersection area of the predicted box and the ground truth box, and union represents the union area of the predicted box and the ground truth box. and b r respectively represent the right boundary coordinates of the predicted box and the ground truth box. and b l respectively represent the left boundary coordinates of the predicted box and the ground truth box. and b b respectively represent the bottom boundary coordinates of the predicted box and the ground truth box. and b t respectively represent the top boundary coordinates of the predicted box and the ground truth box. w gt and h gt respectively represent the width and height of the predicted box, and w and h respectively represent the width and height of the ground truth box. represents the abscissa value of the center point of the detection box. represents the ordinate value of the center point of the detection box. x c represents the abscissa value of the center point of the ground truth box. y c represents the ordinate value of the center point of the ground truth box; ratio is a scaling factor with a value of 0.7; α is a power variable with a value of 3.0; eps is a constant with a value of 10 -7 .

4. The OCR-based automated processing method for bank statement according to claim 1, wherein: In step (4), the adversarial network uses the EraseNet adversarial network.

5. The OCR-based automated processing method for bank statement according to claim 1, wherein: Step (6) specifically includes the following steps in sequence: (6a) Scale the input picture to obtain a scaled picture; (6b) Detect all text blocks in the scaled picture, and obtain the coordinate information and text information of the four corners of the text blocks; (6c) Use the continuous local peak detection algorithm for the coordinate information to parse and obtain the table structure; (6d) Locate the cells in the table structure, and reconstruct the obtained text information into the table structure to obtain a reconstructed table structure.

6. The OCR-based automated processing method for bank statement according to claim 5, wherein: In step (6c), the continuous local peak detection specifically refers to: extracting the abscissa of the left border of the text block and the mean value of the abscissa of the left border, and taking the abscissa of the left border as the column left coordinate value; generating a column coordinate list from the column left coordinate value, using the dynamic local difference method to remove noise points from the column coordinate list, using the differential method to calculate the difference between adjacent points in the column coordinate list, obtaining a first-order difference sequence for detecting local peaks; the first-order difference represents the instantaneous change rate of the signal, traversing the first-order difference sequence, finding the position of the maximum value of the change from negative to positive, that is, the index value, and outputting the index value and the column left coordinate value corresponding to the index value; Extracting the abscissa of the right border of the text block and the mean value of the abscissa of the right border as the column right coordinate value; extracting the abscissa of the right border of the text block and the mean value of the abscissa of the right border, and taking the abscissa of the right border as the column right coordinate value; generating a column coordinate list from the column right coordinate value, using the dynamic local difference method to remove noise points from the column coordinate list, using the differential method to calculate the difference between adjacent points in the column coordinate list, obtaining a first-order difference sequence for detecting local peaks; the first-order difference represents the instantaneous change rate of the signal, traversing the first-order difference sequence, finding the position of the maximum value of the change from negative to positive, that is, the index value, and outputting the index value and the column right coordinate value corresponding to the index value; Taking the average value of the column right coordinate value of the first column and the column left coordinate value of the second column of the coordinate information as the final column right coordinate of the first column, and traversing all column coordinates in this way to obtain the final column coordinates, and determining the bounding box of each cell through the final column coordinates to identify the content within the cell and obtain the table structure.

7. A system for implementing the OCR-based automated processing method of bank statement according to any one of claims 1 to 6, characterized in that: Including: An input module for receiving the input of bank statement data; An image judgment module for determining whether the input bank statement data is in image format and converting the document format into a picture format; A watermark and seal recognition model for detecting watermark and seal information in the picture; A watermark and seal removal module using the EraseNet adversarial network for removing watermark and seal information; A text recognition module for performing text recognition and text extraction on the picture after removing the watermark and seal based on the OCR technology, i.e., optical character recognition technology, and forming text from the extracted text content; A document structure reconstruction module for analyzing the logical structure of the document and reconstructing the extracted text information according to the table layout of the original document or picture to obtain the reconstructed table structure; the original document is the document in the received bank statement data; A verification module for adjusting and confirming the reconstructed table structure to obtain formatted table data; A formatted output module for exporting the formatted table data into Excel and JSON formats.

Citation Information

Patent Citations

  • Tax declaration form identification method and device

    CN114445841A

  • Smoke and fire detection method based on super-resolution reconstruction and adaptive extrusion excitation

    CN115719463A

  • PDF scanning copy content identification method and device

    CN116311305A

  • Method for extracting building change area in double-time-phase remote sensing image based on twinborn mixed attention mechanism and multi-scale feature fusion

    CN118212532A

  • Substation equipment defect detection method and system based on improved YOLOv9

    CN118469986A

Cited By

  • Intelligent document detection method and system based on OCR and ES

    CN120954012A