An intelligent verification identification method, device and electronic equipment for financial statements

Through Hough spatial analysis and character error probability calculation, identification errors caused by sticking between characters and table lines in financial statements are solved, and the accuracy and efficiency of verification are improved.

CN120182987BActive Publication Date: 2025-07-18XIAN BAOKANG MEDICAL MANAGEMENT DATA TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510661221.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-22
Publication Date
2025-07-18
Estimated Expiration
2045-05-22

AI Technical Summary

Technical Problem

In the prior art, the accuracy of financial statement verification and identification is poor, especially when the text and table lines are stuck or blocked, character breakage and adhesion recognition errors occur frequently, resulting in confusion in number recognition and misjudgment of amounts.

Method used

By obtaining the scanned image of financial statements, using Hough space to analyze the edge line distribution to determine the text direction, adjust the image grayscale distribution, identify the text box, calculate the character error probability, and filter out the wrong characters.

Benefits of technology

It improves the accuracy and efficiency of financial statement verification, reduces character recognition errors, and ensures the accuracy of amount judgment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182987B_ABST
    Figure CN120182987B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of report character recognition, and particularly relates to an intelligent verification and identification method, device and electronic device for financial statements. The present invention obtains the starting probability of text regions in each layer in each direction; according to the change trend of the starting probability of text regions in all layers in different directions, a plurality of text recognition frames are obtained; according to the differences in different position dimensions between different text recognition frames and the character similarity, text recognition frame clustering clusters in each position dimension are obtained; according to the gray-scale distribution differences and character similarity between each text recognition frame and other text recognition frames within the same text recognition frame clustering cluster in all corresponding dimensions, the misrecognition probability of each character of each text recognition frame is obtained, and the incorrect characters are screened out to verify the financial statements. The present invention improves the efficiency and accuracy of financial statement verification by obtaining the misrecognition probability of each character in each text box.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of statement character recognition, and particularly to an intelligent verification and identification method, device and electronic device for financial statements. Background Art

[0002] As the core carrier of an enterprise's financial status, operating results and cash flow, financial statements play an important role in financial decision-making, risk assessment and audit supervision. Traditional methods for verifying financial statements rely on manual comparison and basic rule verification, which have limitations such as low efficiency, error-proneness, and difficulty in discovering deep-seated problems.

[0003] In the prior art, OCR is used to recognize financial statements. However, due to the complex format of the statements, numerous table lines, and the compact arrangement of financial data, there are prone to adhesion or occlusion between the text and the table lines, resulting in character breakage, incorrect adhesion recognition, etc., causing confusion in number recognition and misjudgment of amounts. The accuracy of financial statement verification and recognition is relatively poor. Summary of the Invention

[0004] In order to solve the technical problems of easy adhesion or occlusion between the text and the table lines and relatively poor accuracy of financial statement verification and recognition, the purpose of the present invention is to provide an intelligent verification and identification method, device and electronic device for financial statements. The specific technical solutions adopted are as follows:

[0005] The present invention provides an intelligent verification and identification method for financial statements, and the method includes:

[0006] Obtain a scanned image of the financial statement;

[0007] Obtain the edge lines in the scanned image, and based on the distribution of pixel points on different edge lines in the Hough space, obtain the text orientation; adjust the scanned image based on the text orientation, and based on the gray-scale distribution of pixel points in each layer in different directions of the adjusted scanned image, obtain the starting probability of the text area in each layer in each direction; based on the change trend of the starting probability of the text area in all layers in different directions, obtain multiple text recognition frames and corresponding characters;

[0008] Based on the differences in different position dimensions between different text recognition frames and character similarity, obtain the text recognition frame clustering clusters in each position dimension; based on the gray-scale distribution differences and character similarity between each text recognition frame and other text recognition frames within the same text recognition frame clustering cluster in all dimensions corresponding thereto, obtain the misrecognition probability of each character within each text recognition frame;

[0009] Based on the misrecognition probability of each character within different text recognition frames, screen out the incorrect characters and verify the financial statement.

[0010] Further, the method for obtaining the text direction includes:

[0011] Converting the pixel points on different edge lines to the Hough space through the polar coordinate mapping formula to obtain a curve composed of distances at different angles;

[0012] Obtaining the number of intersections of the curves at each angle in the Hough space, selecting the maximum value of the sequence composed of the number of intersections of the curves at all angles, and selecting the angle corresponding to the largest maximum value as the text direction.

[0013] Further, the method for obtaining the starting probability of the text region includes:

[0014] Obtaining the grayscale mean value of all pixel points on each layer in each direction as the pixel grayscale level in each direction on each layer;

[0015] Obtaining the mean difference value of the pixel grayscale levels of other layers between the front and back neighborhood ranges on each layer in each direction as the starting probability of the text region in each direction on each layer.

[0016] Further, the method for obtaining the text recognition frame and the corresponding characters includes:

[0017] Obtaining the extreme values of the sequence composed of the starting probabilities of the text regions in all layers in each direction, where the layer corresponding to the maximum value is used as the starting layer of the text region, and the layer corresponding to the minimum value is used as the ending layer of the text region;

[0018] For the row direction, the range between the adjacent starting layer and ending layer of the text region constitutes the text line region; based on the column direction within the text line region, the range between the adjacent starting layer and ending layer of the text region constitutes the text recognition frame;

[0019] Using an OCR engine to obtain the characters within the text recognition frame.

[0020] Further, the method for obtaining the text recognition frame clustering clusters includes:

[0021] The position dimension includes the starting horizontal position, ending horizontal position, starting vertical position, and ending vertical position;

[0022] According to the differences in each position dimension between different text recognition frames and the character similarity, obtaining the relative distance between different text recognition frames in each position dimension;

[0023] For each position dimension, performing hierarchical clustering on all text recognition frames according to the relative distance between different text recognition frames to obtain multiple text recognition frame clustering clusters.

[0024] Further, the method for obtaining the relative distance includes:

[0025] For each position dimension, obtain the differences in the position coordinates between different text recognition frames and normalize them as the first difference.

[0026] Obtain the DTW distance of the sequences composed of characters between different text recognition frames as the second difference.

[0027] Obtain the product of the first difference and the second difference as the relative distance between different text recognition frames.

[0028] Furthermore, the method for obtaining the misrecognition probability includes:

[0029] Perform DTW matching on the characters between different text recognition frames to obtain the matching characters of each character in other text recognition frames.

[0030] If each character in each text recognition frame is equal to the matching characters in other text recognition frames within the same text recognition frame clustering cluster under all corresponding dimensions, set the character difference to 0; otherwise, set it to the positive integer 1. Obtain the cosine angle of the sequences composed of the number of pixel points at all gray levels between different text recognition frames as the image difference.

[0031] For each character in each text recognition frame and all the matching characters in other text recognition frames within the same text recognition frame clustering cluster under all corresponding dimensions, obtain the ratio of the image difference to the character difference as the first ratio; obtain the ratio of the character difference to the image difference as the second ratio; obtain the sum of all the first ratios and the second ratios as the misrecognition probability of each character in each text recognition frame.

[0032] Furthermore, the method for obtaining the misspelled characters includes:

[0033] If the misrecognition probability of a character is greater than a preset misrecognition threshold, regard the corresponding character as a misspelled character.

[0034] The present invention also proposes an intelligent verification and identification device for financial statements, including an image acquisition module, a text recognition frame generation module, a character misrecognition calculation module, and a misspelled character screening and verification module:

[0035] Perform DTW matching on the characters between different text recognition frames to obtain the matching characters of each character in other text recognition frames.

[0036] If each character in each text recognition frame is equal to the matching characters in other text recognition frames within the same text recognition frame clustering cluster under all corresponding dimensions, set the character difference to 0; otherwise, set it to the positive integer 1. Obtain the cosine angle of the sequences composed of the number of pixel points at all gray levels between different text recognition frames as the image difference.

[0037] Obtain the ratio of the image difference to the character difference as the first ratio; obtain the ratio of the character difference to the image difference as the second ratio.

[0038] For each character in each text recognition box and all matching characters in other text recognition boxes within the same text recognition box clustering cluster under all corresponding dimensions, obtain the sum of all the first ratios and the second ratios as the misrecognition probability of each character in each text recognition box.

[0039] Further, the method for obtaining the incorrect characters includes:

[0040] If the misrecognition probability of a character is greater than a preset misrecognition threshold, take the corresponding character as an incorrect character.

[0041] The present invention proposes an intelligent verification and identification device for financial statements, including an image acquisition module, a text recognition box generation module, a character misrecognition calculation module, and an incorrect character screening and verification module:

[0042] Image acquisition module: Obtain a scanned image of the financial statement.

[0043] Text recognition box generation module: Obtain the edge lines in the scanned image, obtain the text orientation based on the distribution of pixel points on different edge lines in the Hough space; adjust the scanned image based on the text orientation, obtain the starting probability of the text area for each direction on each layer according to the gray level distribution of pixel points in different directions on the adjusted scanned image; obtain multiple text recognition boxes and corresponding characters according to the change trend of the starting probability of the text area for different directions on all layers.

[0044] Character misrecognition calculation module: Obtain the text recognition box clustering clusters for each position dimension according to the differences in different position dimensions between different text recognition boxes and character similarity; obtain the misrecognition probability of each character in each text recognition box according to the gray level distribution difference and character similarity between each text recognition box and other text recognition boxes within the same text recognition box clustering cluster under all corresponding dimensions.

[0045] Incorrect character screening and verification module: Screen out the incorrect characters according to the misrecognition probability of each character in different text recognition boxes, and verify the financial statement.

[0046] The present invention also proposes an electronic device, on which a program or instruction is stored, and when the program or instruction is executed by a processor, the steps of an intelligent verification and identification method for a financial statement as described above are implemented.

[0047] The present invention has the following beneficial effects:

[0048] In order to clearly outline the text and structure in an image, obtain multiple edge lines in the scanned image, obtain the text orientation based on the distribution of pixel points on different edge lines in the Hough space, which helps to correct the tilt angle of the text in the subsequent image; adjust the scanned image based on the text orientation, obtain the starting probability of the text area in each layer in each direction according to the gray-scale distribution of pixel points in different directions in the adjusted scanned image, and identify the possible positions of the text area; obtain multiple text recognition frames according to the change trend of the starting probability of the text area in all layers in different directions, and determine the boundary of the text area by analyzing the probability change trend; obtain the text recognition frame clustering clusters in each position dimension according to the differences in different position dimensions and character similarities between different text recognition frames, group the text recognition frames according to spatial position and character similarity, and reveal the structural characteristics of the text; obtain the misrecognition probability of each character in each text recognition frame according to the gray-scale distribution difference and character similarity between each text recognition frame and other text recognition frames in the same text recognition frame clustering cluster in all corresponding dimensions, which helps to improve the accuracy of the financial statement verification; screen out the incorrect characters and verify the financial statements. By obtaining the misrecognition probability of each character in each text box, the present invention improves the efficiency and accuracy of the financial statement verification. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] In order to more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0050] Figure 1 It is a flowchart of an intelligent verification and identification method for a financial statement provided by an embodiment of the present invention;

[0051] Figure 2 It is a flowchart of a method for obtaining a misrecognition probability provided by an embodiment of the present invention;

[0052] Figure 3 It is a structural block diagram of an intelligent verification and identification device for a financial statement provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0053] In order to further elaborate on the technical means and effects adopted by the present invention to achieve the intended invention purpose, the following specifically describes, in conjunction with the accompanying drawings and preferred embodiments, a method, device, and electronic device for intelligent verification and marking of financial statements according to the present invention, including their specific implementation manners, structures, features, and effects. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. In addition, specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.

[0054] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this invention belongs.

[0055] The following specifically describes the specific solutions of a method, device, and electronic device for intelligent verification and marking of financial statements provided by the present invention in conjunction with the accompanying drawings.

[0056] Please refer to Figure 1 , which shows a flowchart of a method for intelligent verification and marking of financial statements provided by an embodiment of the present invention. The specific method includes:

[0057] Step S1: Obtain a scanned image of the financial statement.

[0058] In an embodiment of the present invention, in order to ensure the marking accuracy of the financial statement and reduce digital recognition confusion and amount misjudgment caused by problems such as character breakage and adhesion recognition errors, it is necessary to collect the scanned image of the statement for analysis of incorrect characters; First, the financial statement includes a balance sheet, an income statement, a cash flow statement, and a statement of changes in owners' equity, and the text information is divided by tables; Based on a paper-based statement digital acquisition scanning device, the statement is collected to obtain a scanned image of the financial statement.

[0059] Step S2: Obtain the edge lines in the scanned image. According to the distribution of pixel points on different edge lines in the Hough space, obtain the text orientation; Adjust the scanned image based on the text orientation. According to the gray-scale distribution of pixel points in each layer in the adjusted scanned image in different directions, obtain the starting probability of the text area in each layer in each direction; According to the change trend of the starting probability of the text area in all layers in different directions, obtain the characters in multiple text recognition frames.

[0060] Edge lines are one of the important features in an image, which can clearly outline the object contours, shapes, and structures in the image. By detecting the edge lines, the image can be segmented into different regions or objects, and useful feature information can be extracted from the image;

[0061] Obtain the edge lines in the scanned image. It should be noted that, in some embodiments of the present invention, existing edge detection methods such as the Sobel operator and the Canny algorithm are used to obtain multiple edge lines in the scanned image; the specific means are well-known technical means to those skilled in the art and will not be elaborated here.

[0062] When the report is tilted during the scanning and shooting process, the text will also be tilted accordingly, affecting the recognition result of the final financial information; according to the distribution of pixel points on different edge lines in the Hough space, obtain the direction of the text.

[0063] Preferably, in an embodiment of the present invention, the method for obtaining the direction of the text includes:

[0064] Convert the pixel points on different edge lines to the Hough space through the polar coordinate mapping formula, and obtain a curve composed of distances at different angles;

[0065] Obtain the number of intersections of the curves at each angle in the Hough space, select the maximum value of the sequence composed of the number of intersections of the curves at all angles, and select the angle corresponding to the maximum maximum value as the direction of the text.

[0066] It should be noted that, in the embodiment of the present invention, the polar coordinate mapping formula is , where obtain the coordinates of the pixel point Substitute into the polar coordinate mapping formula to obtain different angles The distances when changing, , constitute the corresponding curve. The polar coordinate mapping method is a well-known technical means to those skilled in the art and will not be elaborated here.

[0067] Horizontal and vertical lines or backgrounds of different colors are used in financial statements to distinguish different data regions. Thicker or darker lines and changes in the background color will all affect the cutting of characters; adjust the scanned image based on the direction of the text. It should be noted that, in the embodiment of the present invention, the scanned image is rotated in the opposite direction based on the direction of the text for adjustment; according to the gray-scale distribution of pixel points in different directions on each layer of the adjusted scanned image, obtain the starting probability of the text region in each direction on each layer.

[0068] Preferably, in an embodiment of the present invention, the method for obtaining the starting probability of the text region includes:

[0069] Obtain the gray-scale mean value of all pixel points in each direction on each layer as the pixel gray-scale level in each direction on each layer;

[0070] Obtain the mean difference value of the pixel gray-scale levels of other layers between the front and rear neighborhood ranges in each direction on each layer as the starting probability of the text region in each direction on each layer.

[0071] It should be noted that in the embodiments of the present invention, the analysis of directions is based on the horizontal and vertical directions of the image, that is, the row direction and column direction of the text in the image.

[0072] It should be noted that in an embodiment of the present invention, the front neighborhood range is a range formed by a preset number of layers in front based on each layer, and the rear neighborhood range is a range formed by a preset number of layers behind based on each layer. Wherein, the preset number is 5. In other embodiments of the present invention, the size of the neighborhood range can be specifically set according to specific situations, and no limitation and elaboration will be made here.

[0073] By analyzing the starting probabilities of text regions in all layers in different directions, the possibility of text appearing at different positions in the image is reflected, and the position and direction of the text region are more accurately determined. According to the change trend of the starting probabilities of text regions in all layers in different directions, multiple text recognition frames and corresponding characters are obtained.

[0074] Preferably, in an embodiment of the present invention, the method for obtaining the text recognition frame and the corresponding character includes:

[0075] Obtain the extreme values in the sequence composed of the starting probabilities of text regions in all layers in each direction. The layer corresponding to the maximum value is used as the starting layer of the text region, and the layer corresponding to the minimum value is used as the ending layer of the text region.

[0076] For the row direction, the range between the adjacent starting layer and ending layer of the text region constitutes the text line region; based on the column direction within the text line region, the range between the adjacent starting layer and ending layer of the text region constitutes the text recognition frame;

[0077] Use an OCR engine to obtain the characters within the text recognition frame.

[0078] It should be noted that the recognized characters are manifested in various types such as letters, numbers, symbols, and Chinese characters. The OCR engine algorithm is a well-known technical means for those skilled in the art, and no elaboration will be made here.

[0079] Step S3: Obtain the text recognition frame clustering clusters for each position dimension according to the differences in different position dimensions between different text recognition frames and character similarities; obtain the misrecognition probability of each character within each text recognition frame according to the gray-scale distribution differences and character similarities between each text recognition frame and other text recognition frames within the same text recognition frame clustering cluster in all dimensions.

[0080] Considering text comparison in a similar format, there may be different alignment methods and position differences in the corresponding columns of data with the same format. By analyzing the differences in different position dimensions and character similarities, it helps to identify text recognition frames with similar features and improve the accuracy of recognition. According to the differences in different position dimensions and character similarities between different text recognition frames, a clustering cluster of text recognition frames for each position dimension is obtained.

[0081] Preferably, in an embodiment of the present invention, the method for obtaining the clustering cluster of text recognition frames includes:

[0082] According to the differences in each position dimension and character similarities between different text recognition frames, the relative distance between different text recognition frames in each position dimension is obtained;

[0083] Preferably, in an embodiment of the present invention, the method for obtaining the relative distance includes:

[0084] For each position dimension, obtain the difference in position coordinates between different text recognition frames and normalize it as the first difference;

[0085] Obtain the DTW distance of the character composition sequences between different text recognition frames as the second difference; obtain the product of the first difference and the second difference as the relative distance between different text recognition frames.

[0086] It should be noted that, in an embodiment of the present invention, the DTW distance between the same character types is 0, and the DTW distance between different character types is 1, and the DTW distance of the character composition sequences between different text recognition frames is calculated.

[0087] In an embodiment of the present invention, for any position dimension, the formula for the relative distance is expressed as:

[0088] ;

[0089] Wherein, represents the relative distance between the th text recognition frame and the th text recognition frame in the th position dimension; represents the position coordinate of the th text recognition frame in the th position dimension; represents the position coordinate of the th text recognition frame in the th position dimension; represents the th text recognition frame and the The DTW distance of the character composition sequence between text recognition frames is the second difference; represents a normalization function.

[0090] In the formula of relative distance, represents the th text recognition frame and the th text recognition frame at the difference between the position coordinates in the position dimension, and after normalization, the obtained first difference. The greater the first difference, the greater the difference in position between the th text recognition frame and the th text recognition frame in the position dimension, the greater the relative distance, and the smaller the possibility of showing similar properties; the greater the character similarity and the more consistent the character types, the greater the possibility of showing similarity and the smaller the relative distance.

[0091] For each position dimension, according to the relative distances between different text recognition frames, hierarchical clustering is performed on all text recognition frames to obtain multiple text recognition frame clustering clusters.

[0092] It should be noted that the position dimension includes the starting horizontal position, starting vertical position, ending horizontal position, and ending vertical position of the text recognition frame, that is, the horizontal and vertical positions of the upper diagonal vertices of the text recognition frame; hierarchical clustering combines similar text recognition frames together by recursively merging or differentiating the data, which helps to understand the content features of the text recognition frames. The specific means are well-known technical means to those skilled in the art and will not be elaborated here.

[0093] The difference in gray-scale distribution can reflect the similarity of the image regions between text recognition frames, and character similarity reflects the comparison of two characters themselves. By introducing the difference in gray-scale distribution and character similarity, the accuracy of the characters in the text recognition frame can be evaluated more comprehensively, and the incorrect characters can be recognized more effectively; according to the difference in gray-scale distribution between each text recognition frame and other text recognition frames within the same text recognition frame clustering cluster under all corresponding dimensions, as well as the character similarity, the incorrect recognition probability of each character in each text recognition frame is obtained.

[0094] Preferably, in an embodiment of the present invention, for the method of obtaining the incorrect recognition probability, please refer to Figure 2 , which shows a flowchart of a method for obtaining the incorrect recognition probability, including:

[0095] Step S201: Perform DTW matching on the characters between different text recognition frames to obtain the matching characters of each character in other text recognition frames.

[0096] DTW finds the optimal alignment path through dynamic programming, matches between different characters, and determines the matching characters of each character within other text recognition frames.

[0097] Step S202: If each character in each text recognition frame is equal to the matching characters in other text recognition frames within the same text recognition frame clustering cluster under all corresponding dimensions, set the character difference to 0; otherwise, set it to the positive integer 1. Obtain the cosine angle of the sequence composed of the number of pixel points at all gray levels between different text recognition frames as the image difference.

[0098] The sequence composed of the number of pixel points at all gray levels reflects the distribution of gray values in the image. The cosine angle can measure the correlation of gray distributions between two text recognition frames. The larger the cosine angle, the smaller the image similarity and the greater the image difference.

[0099] It should be noted that the cosine similarity between the sequences composed of the number of pixel points at all gray levels can be calculated, and the cosine angle is obtained through the inverse cosine function. The larger the cosine angle, the smaller the cosine similarity, that is, the smaller the image similarity and the greater the difference performance.

[0100] Step S203: For each character in each text recognition frame and all the matching characters in other text recognition frames within the same text recognition frame clustering cluster under all corresponding dimensions, obtain the ratio of the image difference to the character difference as the first ratio; obtain the ratio of the character difference to the image difference as the second ratio; obtain the sum of all the first ratios and the second ratios as the misrecognition probability of each character in each text recognition frame.

[0101] In an embodiment of the present invention, the formula for the misrecognition probability is expressed as:

[0102] ;

[0103] where represents the misrecognition probability of the th character in the th text recognition frame; represents the image difference between the th character in the th text recognition frame and the th matching character in the th other text recognition frame within the same text recognition frame clustering cluster under all corresponding dimensions; represents the character difference between the th character in the th text recognition frame and the th matching character in the th other text recognition frame within the same text recognition frame clustering cluster under all corresponding dimensions; Indicates the number of matching characters of each character in the text recognition box within other text recognition boxes; Indicates the number of other text recognition boxes within the same text recognition box clustering cluster corresponding to the nth text recognition box under all dimensions.

[0104] In the formula for the error recognition probability, adding 1 in is to avoid the denominator of the formula being 0 and the formula being meaningless; the smaller the image difference is, the more consistent the regional features of the text recognition box are. If the character difference is much larger than the image difference, the first ratio is smaller, the second ratio is larger, and the error recognition probability of each character is larger. The smaller the character difference, it indicates that the matching characters with other text recognition boxes within the same text recognition box clustering cluster are consistent, and the error recognition probability of each character is smaller; when the character difference is smaller and the characters are consistent, the larger the image difference, the more inconsistent the regional features of the text recognition box are, and the greater the error recognition probability.

[0105] Step S4: According to the error recognition probability of each character in different text recognition boxes, screen out the error characters and check the financial statements.

[0106] By analyzing the error recognition probability of each character, potential error characters can be identified, thereby improving the accuracy of the data.

[0107] Preferably, in an embodiment of the present invention, the method for obtaining error characters includes:

[0108] If the error recognition probability of a character is greater than a preset error threshold, the corresponding character is used as an error character.

[0109] It should be noted that the greater the error recognition probability, the greater the possibility that the character is incorrect. In an embodiment of the present invention, the size of the preset error threshold is 1.3; in other embodiments of the present invention, the size of the preset error threshold can be specifically set according to specific circumstances, and no limitation and elaboration are made here.

[0110] After obtaining the text recognition box with error characters, the user places the displayed area of the error text recognition box under the shooting device, re-recognizes the content of the report, and manually checks to determine whether a printing error has occurred to ensure the accuracy of the report data.

[0111] In summary, the present invention obtains the starting probability of the text area in each layer in each direction; obtains multiple text recognition frames according to the change trend of the starting probability of the text area in all layers in different directions; obtains the text recognition frame clustering clusters in each position dimension according to the position differences in different position dimensions between different text recognition frames and the character similarity; obtains the character misrecognition probability of each text recognition frame according to the gray-scale distribution difference and the character similarity between each text recognition frame and other text recognition frames within the same text recognition frame clustering cluster in all corresponding dimensions, screens out the misrecognized characters, and checks the financial statements. By obtaining the misrecognition probability of each character in each text box, the present invention improves the efficiency and accuracy of financial statement checking.

[0112] Based on the same application concept as the intelligent checking and marking method for a financial statement provided in the embodiments of the present application, this embodiment also proposes an intelligent checking and marking device for a financial statement, as Figure 3 shown. The device includes: an image acquisition module 301, a text recognition frame generation module 302, a character misrecognition calculation module 303, and a misrecognized character screening and checking module 304:

[0113] The image acquisition module 301: acquires a scanned image of the financial statement;

[0114] The text recognition frame generation module 302: obtains the edge lines in the scanned image, obtains the text orientation based on the distribution of pixel points on different edge lines in the Hough space; adjusts the scanned image based on the text orientation, obtains the starting probability of the text area in each layer in each direction according to the gray-scale distribution of pixel points in different directions in each layer of the adjusted scanned image; obtains multiple text recognition frames and corresponding characters according to the change trend of the starting probability of the text area in all layers in different directions;

[0115] The character misrecognition calculation module 303: obtains the text recognition frame clustering clusters in each position dimension according to the differences in different position dimensions between different text recognition frames and the character similarity; obtains the misrecognition probability of each character in each text recognition frame according to the gray-scale distribution difference and the character similarity between each text recognition frame and other text recognition frames within the same text recognition frame clustering cluster in all corresponding dimensions;

[0116] The misrecognized character screening and checking module 304: screens out the misrecognized characters according to the misrecognition probability of each character in different text recognition frames, and checks the financial statements.

[0117] It should be understood that the intelligent checking and marking device for a financial statement provided in this embodiment is used to execute the above-mentioned intelligent checking and marking method for a financial statement, and therefore has the same beneficial effects as the method adopted, run or implemented by the application program stored therein.

[0118] The present invention also provides an electronic device, on which a program or instruction is stored, and when the program or instruction is executed by a processor, the steps of an intelligent verification and identification method for a financial statement as described above are implemented.

[0119] It should be noted that: the above sequence of embodiments of the present invention is only for description and does not represent the superiority or inferiority of the embodiments. The processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0120] Each embodiment in this specification is described in a progressive manner, and the same or similar parts among the embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments.

Claims

1. An intelligent verification identification method for financial statements, characterized in that, The method includes: Obtaining a scanned image of a financial statement; Obtaining edge lines in the scanned image, obtaining the text orientation based on the distribution of pixel points on different edge lines in the Hough space; adjusting the scanned image based on the text orientation, obtaining the starting probability of the text area for each orientation on each layer according to the gray-scale distribution of pixel points in different orientations on each layer; obtaining multiple text recognition frames and corresponding characters according to the change trend of the starting probability of the text area for different orientations in all layers; Obtaining a text recognition frame clustering cluster for each position dimension according to the difference in different position dimensions between different text recognition frames and character similarity; obtaining the misrecognition probability of each character in each text recognition frame according to the gray-scale distribution difference and character similarity between each text recognition frame and other text recognition frames within the same text recognition frame clustering cluster under all dimensions corresponding to the text recognition frame; Screening out incorrect characters according to the misrecognition probability of each character in different text recognition frames and checking the financial statement; The method for obtaining the text recognition frame and corresponding characters includes: Obtaining the extreme values of the sequence composed of the starting probabilities of the text areas for each orientation in all layers, using the layer corresponding to the maximum value as the starting layer of the text area and the layer corresponding to the minimum value as the ending layer of the text area; For the row direction, the range between the adjacent starting layer and ending layer of the text area constitutes a text line area; based on the column direction within the text line area, the range between the adjacent starting layer and ending layer of the text area constitutes a text recognition frame; Using an OCR engine to obtain the characters within the text recognition frame; The method for obtaining the text recognition frame clustering cluster includes: Obtaining the relative distance between different text recognition frames in each position dimension according to the difference in each position dimension between different text recognition frames and character similarity; For each position dimension, performing hierarchical clustering on all text recognition frames according to the relative distance between different text recognition frames to obtain multiple text recognition frame clustering clusters.

2. The intelligent verification identification method for a financial statement according to claim 1, wherein The method for obtaining the text orientation includes: Converting the pixel points on different edge lines to the Hough space through a polar coordinate mapping formula to obtain a curve composed of distances at different angles; Obtaining the number of intersections of curves at each angle in the Hough space, selecting the maximum value of the sequence composed of the number of intersections of curves at all angles, and selecting the angle corresponding to the maximum maximum value as the text orientation.

3. The intelligent verification and identification method of a financial statement according to claim 1, characterized in that The method for obtaining the starting probability of the text area includes: Obtaining the gray-scale mean value of all pixel points on each layer for each orientation as the pixel gray-scale level for each orientation on each layer; Obtaining the mean difference value of the pixel gray-scale levels of other layers between the front and back neighborhood ranges for each orientation on each layer as the starting probability of the text area for each orientation on each layer.

4. The intelligent verification identification method for a financial statement according to claim 1, characterized in that, The method for obtaining the relative distance includes: For each position dimension, obtaining the difference in position coordinates between different text recognition frames and normalizing it as the first difference; Obtaining the DTW distance of the sequences composed of characters between different text recognition frames as the second difference; Obtaining the product of the first difference and the second difference as the relative distance between different text recognition frames.

5. The intelligent verification and identification method of a financial statement according to claim 1, characterized in that, The method for obtaining the error recognition probability includes: Performing DTW matching on the characters between different text recognition frames to obtain the matching characters of each character in other text recognition frames; If each character in each text recognition frame is equal to the matching characters in other text recognition frames within the same text recognition frame clustering cluster under all corresponding dimensions, the character difference is set to 0, otherwise it is set to the positive integer 1; obtaining the cosine angle of the sequence composed of the number of pixel points at all gray levels between different text recognition frames as the image difference; For each character in each text recognition frame and all the matching characters in other text recognition frames within the same text recognition frame clustering cluster under all corresponding dimensions, obtaining the ratio of the image difference to the character difference as the first ratio; obtaining the ratio of the character difference to the image difference as the second ratio; obtaining the sum of all the first ratios and the second ratios as the error recognition probability of each character in each text recognition frame.

6. The intelligent verification and identification method of a financial statement according to claim 1, characterized in that The method for obtaining the error characters includes: If there is a character whose error recognition probability is greater than the preset error threshold, the corresponding character is regarded as an error character.

7. An intelligent verification identification device for financial statements, characterized in that, It includes an image acquisition module, a text recognition frame generation module, a character error recognition calculation module, and an error character screening and verification module: Image acquisition module: Obtaining the scanned image of the financial statement; Text recognition frame generation module: Obtaining the edge lines in the scanned image, and based on the distribution of the pixel points on different edge lines in the Hough space, obtaining the text orientation; Adjusting the scanned image based on the text orientation, and based on the gray level distribution of the pixel points in each layer in different orientations of the adjusted scanned image, obtaining the starting probability of the text area in each layer in each orientation; Based on the change trend of the starting probability of the text area in all layers in different orientations, obtaining multiple text recognition frames and the corresponding characters; The method for obtaining the text recognition frames and the corresponding characters includes: Obtaining the extreme values of the sequence composed of the starting probability of the text area in all layers in each orientation, the layer corresponding to the maximum value is used as the starting layer of the text area, and the layer corresponding to the minimum value is used as the ending layer of the text area; For the row direction, the range between the adjacent starting layer and ending layer of the text area constitutes the text line area; Based on the column direction within the text line area, the range between the adjacent starting layer and ending layer of the text area constitutes the text recognition frame; Using an OCR engine to obtain the characters within the text recognition frame; Character error recognition calculation module: According to the differences in different position dimensions between different text recognition frames and the character similarity, obtaining the text recognition frame clustering cluster for each position dimension; According to the gray level distribution difference and the character similarity between each text recognition frame and other text recognition frames within the same text recognition frame clustering cluster under all corresponding dimensions, obtaining the error recognition probability of each character in each text recognition frame; The method for obtaining the text recognition frame clustering cluster includes: According to the differences in each position dimension between different text recognition frames and the character similarity, obtaining the relative distance between different text recognition frames in each position dimension; For each position dimension, hierarchical clustering is performed on all text recognition frames according to the relative distances between different text recognition frames, and multiple text recognition frame clustering clusters are obtained. Error character screening and verification module: According to the error recognition probability of each character in different text recognition frames, error characters are screened out to verify the financial statements.

8. An electronic device, characterized in that, The program or instruction is stored on the electronic device, and when the program or instruction is executed by the processor, the steps of an intelligent verification and identification method for a financial statement as described in any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Text image recognition method and device

    CN112507782A

  • Archive management system based on OCR image recognition

    CN117558005A