System for automatically processing financial information by using image recognition technology
Through the adaptive binarization threshold calculation method, the binary processing problem of financial notes images under different paper colors and lighting conditions is solved, achieving higher recognition accuracy and completeness of financial information.
Patent Information
- Application Number
- CN202510222373.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-02-27
AI Technical Summary
When the prior art performs binary processing of the images of financial notes, it is difficult to adapt to different paper colors and uneven lighting, resulting in incomplete text and affecting the accuracy of financial information.
Adaptive binarization threshold calculation method is adopted, and binarization is performed by calculating the binarization threshold of each pixel point, combining the global threshold and the adjustment threshold, so as to adapt to different paper colors and lighting conditions.
It improves the binarization accuracy of financial notes images, reduces the incomplete text when the light is uneven, and improves the accuracy of identification of financial information.
Smart Images

Figure CN120107982A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of financial data processing, and in particular to a system for automatically processing financial information using image recognition technology. Background Art
[0002] It is a common financial information technology to obtain the financial information recorded on paper financial documents such as invoices through image recognition technology. After obtaining the image of the financial document, it is usually necessary to binarize the image to obtain a binary image, and then perform OCR recognition on the binary image to obtain the text information in the image. In the prior art, when binarizing the image of the financial document, it is usually based on a preset threshold or a global threshold. However, the preset threshold cannot be applied to financial documents with different paper colors, and the global threshold has a high probability of causing some pixels belonging to the text to be mistakenly divided as pixels belonging to the background when the lighting is uneven. This will cause the text in the image after binarization to be incomplete, affecting the accuracy of the financial information finally recognized. Summary of the invention
[0003] The purpose of the present invention is to disclose a system for automatically processing financial information using image recognition technology to solve the technical problems raised in the background technology.
[0004] In order to achieve the above object, the present invention provides the following technical solutions:
[0005] The present invention provides a system for automatically processing financial information using image recognition technology, comprising an image preprocessing module, the image preprocessing module comprising a grayscale unit, a noise reduction unit and a binarization unit;
[0006] The grayscale unit is used to grayscale the image of the financial bill to obtain a first image;
[0007] The noise reduction unit is used to reduce noise on the first image to obtain a second image;
[0008] The binarization unit is used to perform binarization processing on the second image to obtain a binarized image, including:
[0009] Calculate the binarization threshold of each pixel in the second image respectively;
[0010] Determine the grayscale value of the pixel point in the second image in the binary image based on the binary threshold;
[0011] Calculate the binarization threshold of each pixel in the second image respectively, including:
[0012] The first step is to obtain the i-th pixel p in the second image i, i∈[1,N], N is the total number of pixels in the second image;
[0013] The second step is to calculate p i The corresponding binarization threshold thre i , if p i The gray value is greater than thre i , then p i Save to set F;
[0014] The third step is to determine whether i is less than N. If so, add 1 to the value of i and go to the first step. If not, stop calculating.
[0015] thre i The calculation formula is:
[0016] thre i =k×To+(1-k)×Tp
[0017] k represents the adaptive weight, To represents the global threshold corresponding to the second image, and Tp represents the threshold value based on p i The calculated adjustment threshold.
[0018] Preferably, the image preprocessing module further includes a correction unit;
[0019] The correction unit is used to perform distortion correction on the binary image to obtain a corrected image.
[0020] Preferably, it also includes a text area detection module;
[0021] The text area detection module is used to identify the corrected image and obtain the area containing text in the corrected image.
[0022] Preferably, it also includes a text recognition module;
[0023] The text recognition module is used to recognize the area containing text and obtain the text in the area containing text.
[0024] Preferably, it also includes a content filling module;
[0025] The content filling module is used to fill the obtained text into the preset data template.
[0026] Preferably, it also includes an image acquisition module;
[0027] The image acquisition module is used to photograph financial bills and obtain images of the financial bills.
[0028] Preferably, performing noise reduction on the first image to obtain the second image includes:
[0029] The first image is denoised using a median filter algorithm or a bilateral filter algorithm to obtain a second image.
[0030] Preferably, determining the grayscale value of a pixel point in the second image in the binary image based on the binarization threshold comprises:
[0031] The size of the binary image is the same as that of the second image. For the pixel point cp in the jth row and mth column of the binary image j,m , j∈[1,J], m∈[1,M], J and M are the number of rows and columns of pixels in the binary image, cp j,m The gray value is determined as follows:
[0032] Get the pixel bp at the jth row and mth column in the second image j,m ;
[0033] If bp j,m The gray value is greater than bp j,m The corresponding binarization threshold, then cp j,m The grayscale value is 255, otherwise cp j,m The gray value is 0.
[0034] Preferably, the process of obtaining Tp includes:
[0035] In the second image, p i centered, with p i The pixel points whose distance is less than the adaptive distance S are stored in the set NB;
[0036] Use the following formula to calculate Tp:
[0037]
[0038] d u are pixels u and p i The distance between u Indicates the gradient value of pixel u, FL i Indicates p i The gradient value of FLm represents the maximum value of the gradient value of all pixels in NB, G u represents the gray value of pixel u, st represents the standardization of the variables in brackets, and λ is the threshold calculation weight.
[0039] Preferably, the calculation formula of the adaptive weight is:
[0040]
[0041] SU represents the set of pixels in the second image, NumSU represents the number of pixels in the second image, G vrepresents the gray value of pixel v in SU, aveGSU represents the mean gray value of pixel in SU, aveGu represents the mean gray value of pixel in NB, and NumNB represents the number of pixel in NB.
[0042] Beneficial effects:
[0043] In the process of recognizing the image of financial documents, the system of the present invention not only performs binarization processing based on the global threshold, but also considers the adjustment threshold calculated based on the pixel points in the binarization process. In this way, the binarization threshold calculated by the present invention is not only applicable to financial documents of different paper colors, but also because the adjustment threshold calculated based on the grayscale distribution of the pixel points is considered, it can reduce the probability of incomplete text in the binarized image when the lighting is uneven, thereby achieving more accurate recognition of the image of the financial document and improving the accuracy of the recognized financial information. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for describing the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.
[0045] Figure 1 A schematic diagram of a system for automatically processing financial information using image recognition technology according to the present invention. DETAILED DESCRIPTION
[0046] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0047] The present invention provides a system for automatically processing financial information using image recognition technology, comprising an image preprocessing module, the image preprocessing module comprising a grayscale unit, a noise reduction unit and a binarization unit;
[0048] The grayscale unit is used to grayscale the image of the financial bill to obtain a first image;
[0049] The noise reduction unit is used to reduce noise on the first image to obtain a second image;
[0050] The binarization unit is used to perform binarization processing on the second image to obtain a binarized image, including:
[0051] Calculate the binarization threshold of each pixel in the second image respectively;
[0052] Determine the grayscale value of the pixel point in the second image in the binary image based on the binary threshold;
[0053] Calculate the binarization threshold of each pixel in the second image respectively, including:
[0054] The first step is to obtain the i-th pixel p in the second image i , i∈[1,N], N is the total number of pixels in the second image;
[0055] The second step is to calculate p i The corresponding binarization threshold thre i , if p i The gray value is greater than thre i , then p i Save to set F;
[0056] The third step is to determine whether i is less than N. If so, add 1 to the value of i and go to the first step. If not, stop calculating.
[0057] thre i The calculation formula is:
[0058] thre i =k×To+(1-k)×Tp
[0059] k represents the adaptive weight, To represents the global threshold corresponding to the second image, and Tp represents the threshold value based on p i The calculated adjustment threshold.
[0060] In the above binarization process, the adjustment threshold calculated based on the pixel points is also considered. In this way, the binarization threshold calculated by the present invention is not only applicable to financial documents of different paper colors, but also because the adjustment threshold calculated based on the grayscale distribution of the pixel point position is considered, it can reduce the probability of incomplete text in the binarized image when the lighting is uneven, thereby achieving more accurate recognition of the image of the financial document.
[0061] The financial documents of the present invention include invoices, receipts, checks, delivery notes, etc.
[0062] In the present invention, the global threshold may be a threshold obtained by calculating the second image using the maximum inter-class variance method.
[0063] If the image has uneven lighting conditions, the global threshold may not be able to adapt to the brightness differences in different areas. For example, one part of the image may be brighter while another part is darker. Using a single threshold will result in bright or dark areas being incorrectly binarized.
[0064] For example, in a black text image on a white background, the brightness of the background will affect the choice of threshold. When using a global threshold, some dark background parts may be mistaken for the foreground, or some bright foreground areas may be mistaken for the background.
[0065] Therefore, the present invention improves the existing threshold calculation method and adds an adjustment threshold. The adjustment threshold is calculated based on the grayscale value around the pixel point. Therefore, even if the overall brightness of the image is uneven, the accuracy of the calculated threshold can be improved because the grayscale value around the pixel point is added in the threshold calculation process.
[0066] Preferably, the image preprocessing module further includes a correction unit;
[0067] The correction unit is used to perform distortion correction on the binary image to obtain a corrected image.
[0068] In the process of invoice image recognition to extract invoice information, distortion correction is a very important step, especially during the image acquisition process, image distortion may occur due to problems such as shooting angle, camera lens distortion or image deformation.
[0069] The necessity of distortion correction includes:
[0070] Ensure accuracy: Information in invoice images often requires the extraction of specific text, barcodes, or other identifiers. If there is distortion in the image, this information may be distorted, affecting the accuracy of the recognition system.
[0071] Alignment issues: When the invoice image is taken at an incorrect angle or the document is not completely parallel to the camera, it will cause perspective distortion between the parts of the image. Distortion correction can help restore the correct proportions of the image so that the text and graphics in the image can be properly aligned, which is convenient for subsequent optical character recognition (OCR) processing.
[0072] Improve OCR recognition rate: OCR (Optical Character Recognition) technology works best when processing clear, undistorted text. Distortion can cause difficulties in character recognition, so by correcting the image to make it more "flat", the recognition accuracy of the OCR system can be significantly improved.
[0073] Common distortion correction methods:
[0074] Perspective Transformation: If the image is taken from an oblique angle, the perspective transformation can "flatten" the image by mapping the four corners of the image to a front view.
[0075] Lens distortion correction: The distortion caused by the lens (such as barrel distortion or pincushion distortion) can be corrected by specific algorithms. The most common one is the distortion correction method based on camera calibration.
[0076] Image preprocessing: In some cases, image preprocessing (such as contrast enhancement, edge detection, etc.) can also help alleviate the problems caused by distortion.
[0077] Preferably, it also includes a text area detection module;
[0078] The text area detection module is used to identify the corrected image and obtain the area containing text in the corrected image.
[0079] Image segmentation techniques (such as connected domain analysis) can be used to extract areas that may contain text, or deep learning methods (such as convolutional neural networks, CNN) can be used for region detection to automatically identify text areas on bills.
[0080] Preferably, it also includes a text recognition module;
[0081] The text recognition module is used to recognize the area containing text and obtain the text in the area containing text.
[0082] Specifically, you can use OCR technology (such as Tesseract, Google Vision API, Microsoft OCR) to recognize text in an image. OCR technology recognizes characters in a text area and converts them into machine-readable text.
[0083] Preferably, it also includes a content filling module;
[0084] The content filling module is used to fill the obtained text into the preset data template.
[0085] Specifically, different bills generally have fixed formats. Therefore, corresponding data templates can be set in advance for different types of bills. After obtaining the area containing text, the data item corresponding to the text in the area can be determined according to the position of the area, so that the text in the area can be filled into the data template. There are multiple data items in the data template. For example, the data template for an invoice includes data items such as invoice number, invoice date, and amount. For the invoice date, it is generally in the upper right corner of the invoice. Of course, when performing image recognition, the coordinates can be used to more accurately locate the area where the text of the invoice date is located.
[0086] Preferably, it also includes an image acquisition module;
[0087] The image acquisition module is used to photograph financial bills and obtain images of the financial bills.
[0088] A scanner or a camera may be used to photograph the financial document to obtain the corresponding image.
[0089] Preferably, performing noise reduction on the first image to obtain the second image includes:
[0090] The first image is denoised using a median filter algorithm or a bilateral filter algorithm to obtain a second image.
[0091] Median filtering removes extreme values (i.e., noise points) by replacing each pixel value in the image with the median of its neighboring pixel values. This method retains the image edge information and is very helpful for text recognition.
[0092] Bilateral filtering smoothes an image by considering both the distance of spatial neighbors and the similarity of pixel values. It can remove noise from an image while maintaining edge clarity.
[0093] Preferably, determining the grayscale value of a pixel point in the second image in the binary image based on the binarization threshold comprises:
[0094] The size of the binary image is the same as that of the second image. For the pixel point cp in the jth row and mth column of the binary image j,m , j∈[1,J], m∈[1,M], J and M are the number of rows and columns of pixels in the binary image, cp j,m The gray value is determined as follows:
[0095] Get the pixel bp at the jth row and mth column in the second image j,m ;
[0096] If bp j,m The gray value is greater than bp j,m The corresponding binarization threshold, then cp j,m The grayscale value is 255, otherwise cp j,m The gray value is 0.
[0097] By judging by the threshold, the pixels in the text area and the pixels in the non-text area in the second image can present obviously different grayscale values in the binary image, which can improve the speed of subsequent text recognition.
[0098] Preferably, the process of obtaining Tp includes:
[0099] In the second image, p icentered, with p i The pixel points whose distance is less than the adaptive distance S are stored in the set NB;
[0100] Use the following formula to calculate Tp:
[0101]
[0102] d u are pixels u and p i The distance between u Indicates the gradient value of pixel u, FL i Indicates p i The gradient value of FLm represents the maximum value of the gradient value of all pixels in NB, G u represents the gray value of pixel u, st represents the standardization of the variables in brackets, and λ is the threshold calculation weight.
[0103] The Tp of the present invention is based on i The distance between the pixels is smaller than the distance calculated by the adaptive distance, and the adaptive distance can be adaptively changed as the contrast of the second image changes, and can effectively make a trade-off between the calculation efficiency of Tp and the effectiveness of Tp. If the value of S is too large, too many pixels will be involved in the calculation process of Tp, affecting the calculation efficiency of Tp. If the value of S is too small, when the contrast of the second image is low, the number of pixels involved in the calculation process of Tp may be insufficient, so that the calculated binarization threshold is too greatly affected by the global threshold, resulting in an increased probability that the pixels belonging to the text are mistakenly classified as the pixels belonging to the background.
[0104] The influence of the pixel points in the NB of the present invention on the Tp result increases with d u and FL u and FL i The absolute value of the difference between u The smaller the FL u and FL i The larger the absolute value of the difference between them, the greater the influence of the pixel u on the Tp result, so that Tp can comprehensively influence p from two perspectives. i The surrounding grayscale distribution is effectively represented so that the calculated binary threshold can be accurately included in p i The surrounding grayscale distribution improves the adaptability of the calculated binarization threshold.
[0105] The normalization process of the present invention can be performed based on the minimum-maximum normalization method.
[0106] The value range of the threshold calculation weight of the present invention may be [0.3, 0.7]. Furthermore, the threshold calculation weight of the present invention may be 0.5.
[0107] Preferably, the calculation formula of the adaptive weight is:
[0108]
[0109] SU represents the set of pixels in the second image, NumSU represents the number of pixels in the second image, G v represents the gray value of pixel v in SU, aveGSU represents the mean gray value of pixel in SU, aveGu represents the mean gray value of pixel in NB, and NumNB represents the number of pixel in NB.
[0110] The adaptive weight can change with the change of the contrast of the pixel points in SU. The greater the contrast of the pixel points in SU, that is, the larger the molecular part of the present invention, the greater the adaptive weight. In the process of calculating the binarization threshold, the greater the influence of To on the calculation result. In this way, when the contrast of the pixel points in SU is small and the binarization accuracy of the global threshold is low, the greater the influence of Tp on the result of the binarization threshold. Therefore, the adaptive weight can effectively perform adaptive control on the calculation result of the binarization threshold based on the contrast of the second image, so that the calculated binarization threshold is adapted to the actual contrast, and a more accurate binarization result is obtained.
[0111] Preferably, the calculation formula of the adaptive distance S is:
[0112]
[0113] NF means p i is the number of pixels belonging to F contained in a square area with a side length of 5 as the center, mr is the preset distance, and η is the distance determination weight.
[0114] The adaptive distance of the present invention is obtained by weighted calculation based on the number of pixel points belonging to F among the pixel points in the area of specified size and the contrast of the pixel points in SU. The greater the contrast of the pixel points in SU and the larger the value of NF, the smaller the value of S. In this way, the accuracy of the binarization result obtained based on the global threshold is higher. When the positive direction area contains more pixel points belonging to text, the value of S is smaller, and the number of pixels involved in the calculation process of Tp is smaller, which can not only improve the efficiency of obtaining the binarization threshold, but also further strengthen the edge of the text; when the contrast of the pixel points in SU is smaller and the value of NF is smaller, Tp is calculated based on more pixel points to obtain more accurate results.
[0115] The preset distance of the present invention may be 10. The distance determination weight may be 0.6.
[0116] The preferred embodiments of the present invention disclosed above are only used to help illustrate the present invention. The preferred embodiments do not describe all the details in detail, nor do they limit the invention to the specific implementation methods described. Obviously, many modifications and changes can be made according to the content of this specification. This specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the present invention, so that those skilled in the art can understand and use the present invention well. The present invention is limited only by the claims and their full scope and equivalents.
Claims
1. A system for automatically processing financial information using image recognition technology, characterized in that: It includes an image preprocessing module, which includes a grayscale unit, a noise reduction unit and a binarization unit; The grayscale unit is used to grayscale the image of the financial bill to obtain a first image; The noise reduction unit is used to reduce noise on the first image to obtain a second image; The binarization unit is used to perform binarization processing on the second image to obtain a binarized image, including: Calculate the binarization threshold of each pixel in the second image respectively; Determine the grayscale value of the pixel point in the second image in the binary image based on the binary threshold; Calculate the binarization threshold of each pixel in the second image respectively, including: The first step is to obtain the i-th pixel p in the second image i , i∈[1,N], N is the total number of pixels in the second image; The second step is to calculate p i The corresponding binarization threshold thre i , if p i The gray value is greater than thre i , then p i Save to set F; The third step is to determine whether i is less than N. If so, add 1 to the value of i and go to the first step. If not, stop calculating. thre i The calculation formula is: thre i =k×To+(1-k)×Tp k represents the adaptive weight, To represents the global threshold corresponding to the second image, and Tp represents the threshold value based on p i The calculated adjustment threshold.
2. A system for automatically processing financial information using image recognition technology according to claim 1, characterized in that: The image preprocessing module also includes a correction unit; The correction unit is used to perform distortion correction on the binary image to obtain a corrected image.
3. A system for automatically processing financial information using image recognition technology according to claim 2, characterized in that: It also includes a text region detection module; The text area detection module is used to identify the corrected image and obtain the area containing text in the corrected image.
4. A system for automatically processing financial information using image recognition technology according to claim 3, characterized in that: It also includes a text recognition module; The text recognition module is used to recognize the area containing text and obtain the text in the area containing text.
5. A system for automatically processing financial information using image recognition technology according to claim 4, characterized in that: It also includes a content filling module; The content filling module is used to fill the obtained text into the preset data template.
6. A system for automatically processing financial information using image recognition technology according to claim 1, characterized in that: Also includes an image acquisition module; The image acquisition module is used to photograph financial bills and obtain images of the financial bills.
7. A system for automatically processing financial information using image recognition technology according to claim 1, characterized in that: Denoising the first image to obtain a second image, comprising: The first image is denoised using a median filter algorithm or a bilateral filter algorithm to obtain a second image.
8. The system for automatically processing financial information using image recognition technology according to claim 1, characterized in that: Determining the grayscale value of a pixel point in the second image in the binary image based on the binary threshold value includes: The size of the binary image is the same as that of the second image. For the pixel point cp in the jth row and mth column of the binary image j,m , j∈[1,J], m∈[1,M], J and M are the number of rows and columns of pixels in the binary image, cp j,m The gray value is determined as follows: Get the pixel bp at the jth row and mth column in the second image j,m ; If bp j,m The gray value is greater than bp j,m The corresponding binarization threshold, then cp j,m The grayscale value is 255, otherwise cp j,m The gray value is 0.
9. A system for automatically processing financial information using image recognition technology according to claim 1, characterized in that: The process of obtaining Tp includes: In the second image, p i centered, with p i The pixel points whose distance is less than the adaptive distance S are stored in the set NB; Use the following formula to calculate Tp: d u are pixels u and p i The distance between u Indicates the gradient value of pixel u, FL i Indicates p i The gradient value of FLm represents the maximum value of the gradient value of all pixels in NB, G u represents the gray value of pixel u, st represents the standardization of the variables in brackets, and λ is the threshold calculation weight.
10. A system for automatically processing financial information using image recognition technology according to claim 9, characterized in that: The calculation formula of adaptive weight is: SU represents the set of pixels in the second image, NumSU represents the number of pixels in the second image, G v represents the gray value of pixel v in SU, aveGSU represents the mean gray value of pixel in SU, aveGu represents the mean gray value of pixel in NB, and NumNB represents the number of pixel in NB.
Citation Information
Patent Citations
File image binaryzation method
CN101021905A
Inferior-quality document image binaryzation method based on background estimation and energy minimization
CN107133929A
Intelligent electric meter chip image binarization processing method based on adaptive mixed threshold
CN111986222A
LED lamp wick defect detection method
CN115601367A
Carpal segmentation and recognition method and system, terminal and readable storage medium
US20190206052A1