Intelligent checking identification method and device for financial statements and electronic equipment

By using Hough transformation and clustering technology in character recognition in financial statements, the text direction is determined and text recognition boxes are generated, and the probability of character error recognition is calculated, the problem of poor character recognition accuracy in financial statements is solved, and the verification efficiency and accuracy are improved.

CN120182987AActive Publication Date: 2025-06-20XIAN BAOKANG MEDICAL MANAGEMENT DATA TECH CO LTD +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510661221.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-22
Publication Date
2025-06-20
Estimated Expiration
2045-05-22

AI Technical Summary

Technical Problem

The character recognition accuracy of financial statements is poor, especially when there is prone to sticking or obstruction between text and table lines, resulting in character breakage and errors in sticking recognition.

Method used

By obtaining the scanned image of the financial statement, the edge lines in the scanned image are obtained, and the direction of the text is determined according to the distribution of pixel points on different edge lines in the Hough space, and the scanned image is adjusted to obtain the starting probability of the text area in each direction. Based on these probabilities, multiple text recognition boxes are generated, and by analyzing the position difference and character similarity between the text recognition boxes, clustering the recognition boxes, calculating the error recognition probability of each character, and finally filtering out the wrong characters.

Benefits of technology

It improves the efficiency and accuracy of financial statement verification, reduces identification errors caused by character breakage and adhesion, and enhances the accuracy and reliability of financial data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182987A_ABST
    Figure CN120182987A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of report character recognition, in particular to an intelligent checking identification method and device for financial statements and electronic equipment. The method comprises the following steps: obtaining a character region starting probability of each layer in each direction; obtaining a plurality of text recognition boxes according to the change trend of the starting probabilities of the text regions of all the layers in different directions; obtaining a text recognition box cluster under each position dimension according to the difference of different position dimensions among different text recognition boxes and character similarity; according to the gray level distribution difference between each text recognition box and other text recognition boxes in the same text recognition box cluster under all corresponding dimensions and the character similarity, the error recognition probability of each character of each text recognition box is obtained, error characters are screened out, and the financial statement is checked. The financial statement checking efficiency and accuracy are improved by obtaining the error recognition probability of each character in each textbox.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of statement character recognition, and particularly to an intelligent verification and identification method, device and electronic device for financial statements. Background Art

[0002] As the core carrier of a company's financial position, operating results and cash flow, financial statements play an important role in financial decision-making, risk assessment and audit supervision. Traditional methods for verifying financial statements rely on manual comparison and basic rule verification, which have limitations such as low efficiency, error-proneness, and difficulty in discovering deep-seated problems.

[0003] In the prior art, OCR is used to recognize financial statements. However, due to the complex format of the statements, numerous table lines, and the compact arrangement of financial data, there are prone to adhesion or occlusion between the text and the table lines, resulting in character breakage, incorrect adhesion recognition, etc., causing confusion in number recognition and misjudgment of amounts. The accuracy of financial statement verification and recognition is relatively poor. Summary of the Invention

[0004] In order to solve the technical problem of the prone adhesion or occlusion between the text and the table lines and the relatively poor accuracy of financial statement verification and recognition, the purpose of the present invention is to provide an intelligent verification and identification method, device and electronic device for financial statements, and the specific technical solutions adopted are as follows: The present invention proposes an intelligent verification and identification method for financial statements, and the method includes: Obtain a scanned image of the financial statement; Obtain the edge lines in the scanned image, and based on the distribution of pixel points on different edge lines in the Hough space, obtain the text direction; adjust the scanned image based on the text direction, and based on the gray-scale distribution of pixel points in each layer in different directions of the adjusted scanned image, obtain the starting probability of the text area in each direction in each layer; based on the change trend of the starting probability of the text area in all layers in different directions, obtain multiple text recognition frames and corresponding characters. Based on the differences in different position dimensions between different text recognition frames and character similarity, obtain the text recognition frame clustering clusters for each position dimension; based on the gray-scale distribution differences and character similarity between each text recognition frame and other text recognition frames within the same text recognition frame clustering cluster in all dimensions corresponding thereto, obtain the misrecognition probability of each character within each text recognition frame. Based on the misrecognition probability of each character within different text recognition frames, screen out the incorrect characters and verify the financial statement.

[0005] Further, the method for obtaining the text direction includes: Map the pixel points on different edge lines to the Hough space through the polar coordinate mapping formula to obtain a curve composed of distances at different angles; Obtain the number of intersections of the curves at each angle in the Hough space, select the maximum value of the sequence composed of the number of intersections of the curves at all angles, and select the angle corresponding to the maximum maximum value as the text writing direction.

[0006] Further, the method for obtaining the starting probability of the text area includes: Obtain the grayscale mean value of all pixel points on each layer in each direction as the pixel grayscale level of each direction on each layer; Obtain the mean difference of the pixel grayscale levels of other layers between the front and rear neighborhood ranges of each layer in each direction as the starting probability of the text area of each direction on each layer.

[0007] Further, the method for obtaining the text recognition frame and the corresponding characters includes: Obtain the extreme values of the sequence composed of the starting probabilities of the text areas of each direction on all layers. The layer corresponding to the maximum value is used as the starting layer of the text area, and the layer corresponding to the minimum value is used as the ending layer of the text area; For the row direction, the range between the adjacent starting layer and ending layer of the text area constitutes the text line area; based on the column direction within the text line area, the range between the adjacent starting layer and ending layer of the text area constitutes the text recognition frame; Use an OCR engine to obtain the characters within the text recognition frame.

[0008] Further, the method for obtaining the text recognition frame clustering clusters includes: The position dimension includes the starting horizontal position, the ending horizontal position, the starting vertical position, and the ending vertical position; According to the differences in each position dimension between different text recognition frames and the character similarity, obtain the relative distance between different text recognition frames in each position dimension; For each position dimension, perform hierarchical clustering on all text recognition frames according to the relative distance between different text recognition frames to obtain multiple text recognition frame clustering clusters.

[0009] Further, the method for obtaining the relative distance includes: For each position dimension, obtain the difference in the position coordinates between different text recognition frames and normalize it as the first difference; Obtain the DTW distance of the sequences composed of the characters between different text recognition frames as the second difference; Obtain the product of the first difference and the second difference as the relative distance between different text recognition frames.

[0010] Further, the method for obtaining the error recognition probability includes: Perform DTW matching on the characters between different text recognition frames to obtain the matching characters of each character in other text recognition frames; If each character in each text recognition frame is equal to the matching characters in other text recognition frames within the same text recognition frame clustering cluster under all corresponding dimensions, the character difference is set to 0; otherwise, it is set to the positive integer 1. Obtain the cosine angle of the sequence composed of the number of pixel points at all gray levels between different text recognition frames as the image difference; For each character in each text recognition frame and all the matching characters in other text recognition frames within the same text recognition frame clustering cluster under all corresponding dimensions, obtain the ratio of the image difference to the character difference as the first ratio; obtain the ratio of the character difference to the image difference as the second ratio; obtain the sum value of all the first ratios and the second ratios as the error recognition probability of each character in each text recognition frame.

[0011] Further, the method for obtaining the error character includes: If there is a character whose error recognition probability is greater than the preset error threshold, the corresponding character is taken as the error character.

[0012] The present invention also proposes an intelligent verification identification device for financial statements, including an image acquisition module, a text recognition frame generation module, a character error recognition calculation module, and an error character screening and verification module: Perform DTW matching on the characters between different text recognition frames to obtain the matching characters of each character in other text recognition frames; If each character in each text recognition frame is equal to the matching characters in other text recognition frames within the same text recognition frame clustering cluster under all corresponding dimensions, the character difference is set to 0; otherwise, it is set to the positive integer 1. Obtain the cosine angle of the sequence composed of the number of pixel points at all gray levels between different text recognition frames as the image difference; Obtain the ratio of the image difference to the character difference as the first ratio; obtain the ratio of the character difference to the image difference as the second ratio; For each character in each text recognition frame and all the matching characters in other text recognition frames within the same text recognition frame clustering cluster under all corresponding dimensions, obtain the sum value of all the first ratios and the second ratios as the error recognition probability of each character in each text recognition frame.

[0013] Further, the method for obtaining the error character includes: If there is a character whose error recognition probability is greater than the preset error threshold, the corresponding character is taken as the error character.

[0014] The present invention provides an intelligent verification and identification device for financial statements, including an image acquisition module, a text recognition frame generation module, a character error recognition and calculation module, and an error character screening and verification module: The image acquisition module: acquires a scanned image of the financial statement; The text recognition frame generation module: obtains the edge lines in the scanned image, and based on the distribution of pixel points on different edge lines in the Hough space, obtains the text orientation; adjusts the scanned image based on the text orientation, and based on the gray-scale distribution of pixel points in each layer in different orientations of the adjusted scanned image, obtains the starting probability of the text area in each layer in each orientation; based on the change trend of the starting probability of the text area in all layers in different orientations, obtains multiple text recognition frames and corresponding characters; The character error recognition and calculation module: based on the differences in different position dimensions between different text recognition frames and character similarity, obtains the text recognition frame clustering clusters in each position dimension; based on the gray-scale distribution differences and character similarity between each text recognition frame and other text recognition frames within the same text recognition frame clustering cluster in all dimensions corresponding to it, obtains the error recognition probability of each character in each text recognition frame; The error character screening and verification module: based on the error recognition probability of each character in different text recognition frames, screens out the error characters and verifies the financial statement.

[0015] The present invention also provides an electronic device, on which a program or instruction is stored, and when the program or instruction is executed by a processor, the steps of an intelligent verification and identification method for a financial statement as described above are implemented.

[0016] The present invention has the following beneficial effects: In order to clearly outline the text and structure in an image, obtain multiple edge lines in the scanned image, and obtain the text direction based on the distribution of pixel points on different edge lines in the Hough space, which helps to correct the tilt angle of the text in the image subsequently; adjust the scanned image based on the text direction, and obtain the starting probability of the text area in each layer in each direction according to the gray-scale distribution of pixel points in different directions in the adjusted scanned image, and identify the possible positions of the text area; obtain multiple text recognition frames according to the change trend of the starting probability of the text area in all layers in different directions. By analyzing the change trend of the probability, the boundary of the text area can be determined; obtain the clustering clusters of text recognition frames in each position dimension according to the differences in different position dimensions and character similarities between different text recognition frames, group the text recognition frames according to spatial position and character similarity, and reveal the structural characteristics of the text; obtain the misrecognition probability of each character in each text recognition frame according to the gray-scale distribution difference and character similarity between each text recognition frame and other text recognition frames in the same text recognition frame clustering cluster in all corresponding dimensions, which helps to improve the accuracy of financial statement verification; screen out the incorrect characters and verify the financial statements. By obtaining the misrecognition probability of each character in each text box, the present invention improves the efficiency and accuracy of financial statement verification. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0018] Figure 1 It is a flowchart of an intelligent verification and identification method for financial statements provided by an embodiment of the present invention; Figure 2 It is a flowchart of a method for obtaining misrecognition probability provided by an embodiment of the present invention; Figure 3 It is a structural block diagram of an intelligent verification and identification device for financial statements provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0019] In order to further elaborate on the technical means and effects adopted by the present invention to achieve the intended invention purpose, the following specifically describes, with reference to the accompanying drawings and preferred embodiments, a method, device, and electronic device for intelligent verification and marking of financial statements according to the present invention, including their specific implementation manners, structures, features, and effects. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. In addition, the specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.

[0020] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs.

[0021] The following specifically describes the specific solutions of a method, device, and electronic device for intelligent verification and marking of financial statements provided by the present invention with reference to the accompanying drawings.

[0022] Please refer to Figure 1 , which shows a flowchart of a method for intelligent verification and marking of financial statements provided by an embodiment of the present invention. The specific method includes: Step S1: Obtain a scanned image of the financial statement.

[0023] In an embodiment of the present invention, in order to ensure the marking accuracy of the financial statement and reduce the digital recognition confusion and amount misjudgment caused by problems such as character breakage and adhesion recognition errors, it is necessary to collect the scanned image of the statement for analysis of incorrect characters; First, the financial statement includes a balance sheet, an income statement, a cash flow statement, and a statement of changes in owners' equity, and the text information is divided by tables; Based on a digital acquisition and scanning device for paper statements, the statement is collected to obtain a scanned image of the financial statement.

[0024] Step S2: Obtain the edge lines in the scanned image. According to the distribution of pixel points on different edge lines in the Hough space, obtain the text orientation; Adjust the scanned image based on the text orientation. According to the gray level distribution of pixel points in each layer in the adjusted scanned image in different directions, obtain the starting probability of the text area in each layer in each direction; According to the change trend of the starting probability of the text area in all layers in different directions, obtain the characters in multiple text recognition frames.

[0025] Edge lines are one of the important features in an image, which can clearly outline the object contours, shapes, and structures in the image. By detecting edge lines, the image can be segmented into different regions or objects, and useful feature information can be extracted from the image; Obtain edge lines in the scanned image. It should be noted that in some embodiments of the present invention, existing edge detection methods such as the Sobel operator and the Canny algorithm are used to obtain multiple edge lines in the scanned image; the specific means are well known to those skilled in the art and will not be described in detail here.

[0026] When the report is tilted during the scanning process, the text will also be tilted, affecting the final recognition result of the financial information; the direction of the text is obtained based on the distribution of pixel points on different edge lines in the Hough space.

[0027] Preferably, in one embodiment of the present invention, the method for obtaining the direction of text includes: The pixel points on different edge lines are converted to Hough space through the polar coordinate mapping formula to obtain curves composed of distances at different angles; Get the number of curve intersections at each angle in the Hough space, select the maximum value of the sequence composed of the number of curve intersections at all angles, and select the angle corresponding to the largest maximum value as the text direction.

[0028] It should be noted that, in the embodiment of the present invention, the polar coordinate mapping formula is: , where the coordinates of the pixel points are obtained Substituting into the polar coordinate mapping formula, we get different angles Distance when changing , forming a corresponding curve, the polar coordinate mapping method is a technical means well known to those skilled in the art and will not be described in detail here.

[0029] Financial statements may use horizontal and vertical lines or backgrounds of different colors to distinguish different data areas. Thicker or darker lines and changes in background color will affect the cutting of characters. The scanned image is adjusted in direction based on the text. It should be noted that, in an embodiment of the present invention, the scanned image is rotated in the opposite direction for adjustment based on the direction of text travel. The grayscale distribution of pixels in different directions on each layer in the adjusted scanned image is used to obtain the starting probability of the text area in each direction on each layer.

[0030] Preferably, in one embodiment of the present invention, the method for obtaining the text region start probability includes: Obtain the grayscale mean of all pixels in each direction on each layer as the pixel grayscale level in each direction on each layer; The mean difference of the pixel grayscale levels of other layers between the front and back neighborhood ranges of each layer in each direction is obtained as the starting probability of the text area in each layer for each direction.

[0031] It should be noted that in the embodiments of the present invention, the analysis of directions is based on the horizontal and vertical directions of the image, that is, the row direction and column direction of the text in the image.

[0032] It should be noted that in an embodiment of the present invention, the front neighborhood range is a range composed of a preset number of previous layers based on each layer, and the rear neighborhood range is a range composed of a preset number of subsequent layers based on each layer. Among them, the preset number is 5. In other embodiments of the present invention, the size of the neighborhood range can be specifically set according to specific situations, and no limitation and elaboration are made here.

[0033] By analyzing the starting probabilities of text regions in all layers in different directions, the possibility of text appearing at different positions in the image is reflected, and the position and direction of the text region are more accurately determined. According to the change trend of the starting probabilities of text regions in all layers in different directions, multiple text recognition frames and corresponding characters are obtained.

[0034] Preferably, in an embodiment of the present invention, the method for obtaining the text recognition frame and the corresponding characters includes: Obtain the extreme values in the sequence composed of the starting probabilities of text regions in all layers in each direction. The layer corresponding to the maximum value is used as the starting layer of the text region, and the layer corresponding to the minimum value is used as the ending layer of the text region.

[0035] For the row direction, the range between the adjacent starting layer and ending layer of the text region forms a text line region; based on the column direction within the text line region, the range between the adjacent starting layer and ending layer of the text region forms a text recognition frame; Use an OCR engine to obtain the characters within the text recognition frame.

[0036] It should be noted that the recognized characters are in various types such as letters, numbers, symbols, and Chinese characters. The OCR engine algorithm is a well-known technical means for those skilled in the art, and no elaboration is made here.

[0037] Step S3: According to the differences in different position dimensions between different text recognition frames and the character similarity, obtain the text recognition frame clustering clusters for each position dimension; according to the gray-scale distribution differences and character similarity between each text recognition frame and other text recognition frames within the same text recognition frame clustering cluster in all dimensions, obtain the misrecognition probability of each character within each text recognition frame.

[0038] Considering text comparison in a similar format, there may be different alignment methods and position differences in the corresponding columns of data with the same format. By analyzing the differences in different position dimensions and character similarities, it helps to identify text recognition frames with similar features and improve the recognition accuracy. According to the differences in different position dimensions and character similarities between different text recognition frames, text recognition frame clustering clusters for each position dimension are obtained.

[0039] Preferably, in an embodiment of the present invention, the method for obtaining text recognition frame clustering clusters includes: According to the differences in each position dimension and character similarities between different text recognition frames, the relative distance between different text recognition frames in each position dimension is obtained; Preferably, in an embodiment of the present invention, the method for obtaining the relative distance includes: For each position dimension, obtain the difference in position coordinates between different text recognition frames and normalize it as the first difference; Obtain the DTW distance of the character composition sequences between different text recognition frames as the second difference; obtain the product of the first difference and the second difference as the relative distance between different text recognition frames.

[0040] It should be noted that in an embodiment of the present invention, the DTW distance between the same character types is 0, and the DTW distance between different character types is 1, and the DTW distance of the character composition sequences between different text recognition frames is calculated.

[0041] In an embodiment of the present invention, for any position dimension, the formula for the relative distance is expressed as: ; Wherein, represents the relative distance between the th text recognition frame and the th text recognition frame in the th position dimension; represents the position coordinate of the th text recognition frame in the th position dimension; represents the position coordinate of the th text recognition frame in the th position dimension; represents the DTW distance of the character composition sequences between the th text recognition frame and the th text recognition frame, that is, the second difference; represents the normalization function.

[0042] In the formula for the relative distance, represents the The difference between the position coordinates of the first text recognition box and the second text recognition box in the position dimension is normalized to obtain the first difference. The greater the first difference, the greater the difference in the position of the first text recognition box and the second text recognition box in the position dimension, the greater the relative distance, and the smaller the possibility of showing similar properties; the greater the character similarity and the more consistent the character types, the greater the possibility of showing similarity and the smaller the relative distance. For each position dimension, hierarchical clustering is performed on all text recognition boxes according to the relative distances between different text recognition boxes to obtain multiple text recognition box clustering clusters. It should be noted that the position dimension includes the starting horizontal position, starting vertical position, ending horizontal position, and ending vertical position of the text recognition box, that is, the horizontal and vertical positions of the upper diagonal vertices of the text recognition box; hierarchical clustering combines similar text recognition boxes together by recursively merging or differentiating the data, which helps to understand the content characteristics of the text recognition boxes. The specific means are well-known technical means to those skilled in the art and will not be elaborated here. The difference in gray-scale distribution can reflect the similarity of the image areas between text recognition boxes, and character similarity reflects the comparison of two characters themselves. By introducing the difference in gray-scale distribution and character similarity, the accuracy of the characters in the text recognition box can be evaluated more comprehensively, and incorrect characters can be recognized more effectively; according to the difference in gray-scale distribution and character similarity between each text recognition box and other text recognition boxes within the same text recognition box clustering cluster under all corresponding dimensions, the misrecognition probability of each character in each text recognition box is obtained.

[0043] Preferably, in an embodiment of the present invention, for the method of obtaining the misrecognition probability, please refer to

[0044] which shows a flowchart of a method for obtaining the misrecognition probability, including:

[0045] Step S201: Perform DTW matching on the characters between different text recognition boxes to obtain the matching characters of each character in other text recognition boxes.

[0046] DTW finds the optimal alignment path through dynamic programming, matches different characters, and determines the matching characters of each character in other text recognition boxes. Figure 2

[0047]

[0048] ​​​Step S202: If each character in each text recognition box is equal to the matching characters in other text recognition boxes within the same text recognition box clustering cluster under all corresponding dimensions, the character difference is set to 0; otherwise, it is set to the positive integer 1. Obtain the cosine angle of the sequence composed of the number of pixel points at all gray levels between different text recognition boxes as the image difference.

[0049] The sequence composed of the number of pixel points at all gray levels reflects the distribution of gray values in the image. The cosine angle can measure the distribution correlation of gray levels between two text recognition boxes. The larger the cosine angle, the smaller the image similarity and the greater the image difference.

[0050] It should be noted that the cosine similarity between the sequences composed of the number of pixel points at all gray levels can be calculated, and the cosine angle can be obtained through the inverse cosine function. The larger the cosine angle, the smaller the cosine similarity, that is, the smaller the image similarity and the greater the difference performance.

[0051] Step S203: For each character in each text recognition box and all matching characters in other text recognition boxes within the same text recognition box clustering cluster under all corresponding dimensions, obtain the ratio of the image difference to the character difference as the first ratio; obtain the ratio of the character difference to the image difference as the second ratio; obtain the sum of all the first ratios and the second ratios as the misrecognition probability of each character in each text recognition box.

[0052] In an embodiment of the present invention, the formula for the misrecognition probability is expressed as: ; Wherein, represents the misrecognition probability of the th character in the th text recognition box; represents the image difference between the th character in the th text recognition box and the th matching character in the th other text recognition box within the same text recognition box clustering cluster under all corresponding dimensions; represents the character difference between the th character in the th text recognition box and the th matching character in the th other text recognition box within the same text recognition box clustering cluster under all corresponding dimensions; represents the number of matching characters of each character in the text recognition box in other text recognition boxes; represents the The number of other text recognition frames within the same text recognition frame clustering cluster corresponding to each text recognition frame in all dimensions.

[0053] In the formula for the error recognition probability, Adding 1 is to avoid a zero denominator in the formula, making the formula meaningless; the image difference The smaller the image difference, the more consistent the regional features of the text recognition frame. If the character difference is much greater than the image difference, the first ratio is smaller, the second ratio is larger, and the error recognition probability of each character is greater. The smaller the character difference, it indicates that the matching characters are consistent with other text recognition frames within the same text recognition frame clustering cluster, and the error recognition probability of each character is smaller; when the character difference is small and the characters are consistent, the larger the image difference, the more inconsistent the regional features of the text recognition frame, and the greater the error recognition probability.

[0054] Step S4: According to the error recognition probability of each character in different text recognition frames, screen out the error characters and check the financial statements.

[0055] By analyzing the error recognition probability of each character, potential error characters can be identified, thereby improving the accuracy of the data.

[0056] Preferably, in an embodiment of the present invention, the method for obtaining error characters includes: If the error recognition probability of a character is greater than a preset error threshold, the corresponding character is taken as an error character.

[0057] It should be noted that the greater the error recognition probability, the greater the possibility that the character is incorrect. In an embodiment of the present invention, the size of the preset error threshold is 1.3; in other embodiments of the present invention, the size of the preset error threshold can be specifically set according to specific situations and will not be limited and elaborated here.

[0058] After obtaining the text recognition frame with error characters, the user places the displayed area of the error text recognition frame under the shooting device, re-recognizes the content of the report, and manually checks to determine whether a printing error has occurred to ensure the accuracy of the report data.

[0059] In summary, the present invention obtains the starting probability of the text area in each layer in each direction; obtains multiple text recognition frames according to the change trend of the starting probability of the text area in all layers in different directions; obtains the text recognition frame clustering clusters in each position dimension according to the position differences in different position dimensions between different text recognition frames and the character similarity; obtains the character misrecognition probability of each text recognition frame according to the gray distribution difference and the character similarity between each text recognition frame and other text recognition frames in the same text recognition frame clustering cluster in all corresponding dimensions, screens out the misrecognized characters, and checks the financial statements. By obtaining the misrecognition probability of each character in each text box, the present invention improves the efficiency and accuracy of financial statement checking.

[0060] Based on the same application concept as the intelligent checking and marking method for a financial statement provided in the embodiment of the present application, this embodiment also proposes an intelligent checking and marking device for a financial statement, as Figure 3 shown. The device includes: an image acquisition module 301, a text recognition frame generation module 302, a character misrecognition calculation module 303, and a misrecognized character screening and checking module 304: The image acquisition module 301: acquires a scanned image of the financial statement; The text recognition frame generation module 302: obtains the edge lines in the scanned image, obtains the text direction based on the distribution of pixel points on different edge lines in the Hough space; adjusts the scanned image based on the text direction, and obtains the starting probability of the text area in each layer in each direction according to the gray distribution of pixel points in different directions in each layer of the adjusted scanned image; obtains multiple text recognition frames and corresponding characters according to the change trend of the starting probability of the text area in all layers in different directions; The character misrecognition calculation module 303: obtains the text recognition frame clustering clusters in each position dimension according to the differences in different position dimensions between different text recognition frames and the character similarity; obtains the misrecognition probability of each character in each text recognition frame according to the gray distribution difference and the character similarity between each text recognition frame and other text recognition frames in the same text recognition frame clustering cluster in all corresponding dimensions; The misrecognized character screening and checking module 304: screens out the misrecognized characters according to the misrecognition probability of each character in different text recognition frames, and checks the financial statements.

[0061] It should be understood that the intelligent checking and marking device for a financial statement provided in this embodiment is used to execute the above-mentioned intelligent checking and marking method for a financial statement, and thus has the same beneficial effects as the method adopted, run or implemented by the application program stored therein.

[0062] The present invention also provides an electronic device, on which a program or instruction is stored, and when the program or instruction is executed by a processor, the steps of an intelligent verification and identification method for a financial statement as described above are implemented.

[0063] It should be noted that the above sequence of embodiments of the present invention is only for description and does not represent the superiority or inferiority of the embodiments. The processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0064] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments.

Claims

1. An intelligent verification and identification method for financial statements, characterized in that, The method includes: Obtain a scanned image of a financial statement; Obtain the edge lines in the scanned image. Based on the distribution of pixel points on different edge lines in the Hough space, obtain the text orientation. Adjust the scanned image based on the text orientation. According to the gray-scale distribution of pixel points in each layer in different orientations of the adjusted scanned image, obtain the starting probability of the text area in each orientation in each layer. According to the change trend of the starting probability of the text area in all layers in different orientations, obtain multiple text recognition frames and corresponding characters; According to the differences in different position dimensions between different text recognition frames and character similarity, obtain the text recognition frame clustering clusters in each position dimension. According to the gray-scale distribution differences and character similarity between each text recognition frame and other text recognition frames within the same text recognition frame clustering cluster in all dimensions, obtain the misrecognition probability of each character in each text recognition frame; According to the misrecognition probability of each character in different text recognition frames, screen out the incorrect characters and check the financial statement.

2. The intelligent verification and identification method for financial statements according to claim 1, characterized in that, The method for obtaining the text orientation includes: Convert the pixel points on different edge lines to the Hough space through the polar coordinate mapping formula to obtain a curve composed of distances at different angles; Obtain the number of intersections of curves at each angle in the Hough space, select the maximum value of the sequence composed of the number of intersections of curves at all angles, and select the angle corresponding to the maximum maximum value as the text orientation.

3. The intelligent verification and identification method for financial statements according to claim 1, characterized in that, The method for obtaining the starting probability of the text area includes: Obtain the gray-scale mean value of all pixel points in each layer in each orientation as the pixel gray-scale level in each orientation in each layer; Obtain the mean difference value of the pixel gray-scale levels in other layers between the front and rear neighborhood ranges in each layer in each orientation as the starting probability of the text area in each orientation in each layer.

4. The intelligent verification and identification method for financial statements according to claim 1, characterized in that, The method for obtaining the text recognition frame and corresponding characters includes: Obtain the extreme values of the sequence composed of the starting probabilities of the text areas in all layers in each orientation. The layer corresponding to the maximum value is used as the starting layer of the text area, and the layer corresponding to the minimum value is used as the ending layer of the text area; For the row direction, the range between the adjacent starting layer and ending layer of the text area forms the text line area. Based on the column direction within the text line area, the range between the adjacent starting layer and ending layer of the text area forms the text recognition frame; Use an OCR engine to obtain the characters within the text recognition frame.

5. The intelligent verification and identification method for financial statements according to claim 1, characterized in that, The method for obtaining the text recognition frame clustering clusters includes: According to the differences in each position dimension between different text recognition frames and character similarity, obtain the relative distance between different text recognition frames in each position dimension; For each position dimension, perform hierarchical clustering on all text recognition frames according to the relative distance between different text recognition frames to obtain multiple text recognition frame clustering clusters.

6. The intelligent verification and identification method for financial statements according to claim 5, characterized in that, The method for obtaining the relative distance includes: For each position dimension, obtain the difference in position coordinates between different text recognition frames and normalize it as the first difference; Obtain the DTW distance of the sequences composed of characters between different text recognition frames as the second difference; Obtain the product of the first difference and the second difference as the relative distance between different text recognition frames.

7. The intelligent verification and identification method for financial statements according to claim 1, characterized in that, The method for obtaining the error recognition probability includes: Performing DTW matching on the characters between different text recognition frames to obtain the matching characters of each character in other text recognition frames; If each character in each text recognition frame is equal to the matching characters in other text recognition frames within the same text recognition frame clustering cluster under all corresponding dimensions, the character difference is set to 0; otherwise, it is set to the positive integer 1. Obtain the cosine angle of the sequence composed of the number of pixel points at all gray levels between different text recognition frames as the image difference; For each character in each text recognition frame and all the matching characters in other text recognition frames within the same text recognition frame clustering cluster under all corresponding dimensions, obtain the ratio of the image difference to the character difference as the first ratio; obtain the ratio of the character difference to the image difference as the second ratio; obtain the sum value of all the first ratios and the second ratios as the error recognition probability of each character in each text recognition frame.

8. The intelligent verification and identification method for financial statements according to claim 1, characterized in that, The method for obtaining the error characters includes: If there is a character whose error recognition probability is greater than the preset error threshold, the corresponding character is taken as an error character.

9. An intelligent verification identification device for financial statements, characterized in that, It includes an image acquisition module, a text recognition frame generation module, a character error recognition calculation module, and an error character screening and verification module: Image acquisition module: Obtain the scanned image of the financial statement; Text recognition frame generation module: Obtain the edge lines in the scanned image, and based on the distribution of the pixel points on different edge lines in the Hough space, obtain the text direction. Adjust the scanned image based on the text direction, and based on the gray level distribution of the pixel points in each layer in the adjusted scanned image in different directions, obtain the starting probability of the text area in each layer in each direction; Based on the change trend of the starting probability of the text area in all layers in different directions, obtain multiple text recognition frames and the corresponding characters; Character error recognition calculation module: According to the differences in different position dimensions between different text recognition frames and the character similarity, obtain the text recognition frame clustering cluster for each position dimension; according to the gray level distribution difference and the character similarity between each text recognition frame and other text recognition frames within the same text recognition frame clustering cluster under all corresponding dimensions, obtain the error recognition probability of each character in each text recognition frame; Error character screening and verification module: Screen out the error characters according to the error recognition probability of each character in different text recognition frames, and verify the financial statement.

10. An electronic device, characterized in that, The program or instruction is stored on the electronic device, and when the program or instruction is executed by the processor, it implements the steps of the intelligent verification and marking method for a financial statement described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Text image recognition method and device

    CN112507782A

  • Archive management system based on OCR image recognition

    CN117558005A

  • Character image judgment method and device, computer equipment and storage medium

    CN119580293A

  • Method and apparatus for accounting management using OCR recognition based on artificial intelligence

    KR102507534B1