A financial ticket data processing method based on image recognition

The area division and adjustment of invoices is solved through image recognition technology, and the scanning inaccurate problem caused by overlapping multiple invoices is solved, achieving accurate acquisition and efficient archiving of invoice information.

CN116486423BActive Publication Date: 2025-07-11BAOTOU RARE EARTH PRODUCTS EXCHANGE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310446267.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-24
Publication Date
2025-07-11
Estimated Expiration
2043-04-24

AI Technical Summary

Technical Problem

When scanning financial notes in large batches, overlapping multiple invoices leads to inaccurate acquisition of scanned files, affecting identification and archiving.

Method used

The invoices are divided and filtered through image recognition technology, and the images of overlapping invoices are identified and adjusted to ensure that key areas such as table headers, tables and QR codes are aligned. The edge lines of the table area are determined by edge detection, and the image is adjusted by rotating and realizing accurate scanning.

Benefits of technology

Accurate image recognition and information extraction of overlapping invoices are realized, and the accuracy of invoice scanning and archiving efficiency are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116486423B_ABST
    Figure CN116486423B_ABST
Patent Text Reader

Abstract

The present invention discloses a financial ticket processing method based on image recognition. It mainly scans and recognizes a large number of invoice information through an invoice scanning device, facilitating the retrieval, sorting and management of invoices by staff. When the invoice scanning and recognition is incorrect, the scanned invoice image information is adjusted to align it with the scanning window and then scanned and recognized for entry.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of invoice management security, and relates to a financial ticket processing method based on image recognition. Background Art

[0002] With the development of the economy, invoices, as an indispensable financial instrument, in the prior art, in order to collect paper invoices filed in the warehouse, an intelligent processing method of mass scanning and recognition is often adopted (the main purposes of the intelligent archiving process of mass scanning and recognition of paper invoices are twofold. First, it is convenient to store the data of the truly scanned ticket evidence. Second, it is convenient to retrieve relevant ticket data at any time);

[0003] When performing mass scanning and recognition of financial tickets, invoices can be pushed and scanned one by one (i.e., pushed one by one) through supporting equipment. However, it is found that when performing mass invoice scanning, there may be a phenomenon where multiple invoices are pushed at once, resulting in overlapping of multiple invoices. This will cause errors in the scanned files and affect the normal image acquisition and extraction and recognition. Summary of the Invention

[0004] The present invention provides a financial ticket processing method based on image recognition and a readable storage medium, which solves the above technical problems proposed in the prior art.

[0005] The present invention provides a financial ticket processing method based on image recognition, including the following operating steps:

[0006] Obtain the text information of multiple sellers and the text information of multiple buyers, form a seller text set from the text information of multiple sellers; form a buyer text set from the text information of multiple buyers;

[0007] Divide the basic template invoice into regions according to its functions, and then obtain the divided sub-regions. The functions include the header, two-dimensional code, password area, and table area. And number the divided sub-regions in a preset order, and mark them as 1, 2,... k,... v in sequence, and then obtain the positions corresponding to each sub-region, and construct a set of positions of each sub-region W (W1, W2,... Wk,... Wv), where Wk represents the position where the k-th sub-region is located;

[0008] When performing a scanning process on the layout of the current entire invoice, and obtain the scanned information: The scanning of the layout of the entire invoice includes scanning the header, two-dimensional code, password area, and table area of the entire invoice, and then obtain the regional characterization information corresponding to each sub-region, and then construct a set of sub-region information H (H1, H2,... Hk,... Hv), where Hk represents the regional characterization information corresponding to the k-th sub-region;

[0009] Perform the scanning information screening and processing operation: obtain and then call the sub-regions where the purchaser and the seller are located in the table area of the current whole invoice, obtain the regional characterization information corresponding to the sub-regions of the purchaser and the seller, compare the regional characterization information corresponding to the sub-region of the purchaser with the purchaser text set, if there is a mismatch, determine that the sub-region is regarded as an abnormal sub-region, if it matches, it is regarded as a normal sub-region, and then filter the normal sub-region; compare the regional characterization information corresponding to the sub-region of the seller with the seller text set, if there is a mismatch, determine that the sub-region is regarded as an abnormal sub-region, if it matches, it is regarded as a normal sub-region, and then filter the normal sub-region; if the sub-region is an abnormal sub-region, mark the region as a marked region, and then process the image information in the current whole invoice scanning information corresponding to the marked region, confirm that it is aligned with the scanning window, and then scan and input the whole invoice information.

[0010] Specifically, processing the image information in the current whole invoice scanning information corresponding to the marked region specifically includes the following steps:

[0011] Segment the image information of the currently overlapping invoice scanned, and obtain the topmost invoice image as the target image; and use the horizontal line where the current scanning window is located as the preset horizontal line, and use the preset horizontal line as the reference line for the long edge line of the table.

[0012] Obtain the long edge line and short edge line of the table in the table area of the target image by scanning and recognizing the table edge, and initially judge whether it meets the first initial condition of the alignment position according to the long edge line of the table and the reference line of the long edge line of the table; the first initial condition refers to that the long edge line of the current table and the reference line of the long edge line of the table are in a parallel relationship.

[0013] If the long edge line of the scanned and recognized table is not parallel to the reference line of the long edge line of the table, adjust the target image, and perform scanning and recognition input on the adjusted target image that meets the second initial condition, including the following steps:

[0014] Determine the character boundaries (boundaries of characters) in the table area by the edge detection method.

[0015] Obtain the character boundaries of two consecutive current characters in the table area to determine the character spacing between the two consecutive characters, and determine whether the character spacing between the two consecutive characters is equal to the preset character spacing value. If so, consider the two consecutive characters as being on the same line, and further determine whether the semantics of multiple non - consecutive pairs of two characters in the same line have financial bill - related semantics; the semantics of the two consecutive characters are regarded as a phrase, and the semantics of multiple non - consecutive pairs of two characters in the same line are regarded as multiple non - consecutive phrases in the same line; (that is, the semantics of two consecutive characters form a phrase, and semantic judgment by multiple non - consecutive phrases in the same line is more accurate. The aforementioned single semantic judgment of two consecutive characters does have the situation of inaccurate semantic judgment. For example, the semantics of "China" has financial bill - related semantics, the semantics of "industry and commerce" has financial bill - related semantics, and the semantics of "country industry" does not have financial bill - related semantics. Therefore, semantic judgment by multiple non - consecutive phrases in the same line is more accurate).

[0016] If it is determined that multiple non - consecutive phrases in the same line have financial bill - related semantics, then determine any two characters in the same line with financial bill - related semantics as the first character and the second character in sequence, and further determine that the edge line of the long side of the table in the target table area of the target image parallel to the character extension vector direction from the first character to the second character is the first edge of the target table area of the target image;

[0017] Determine the perpendicular line to the first edge of the target table area of the target image as the second edge of the table, and rotate the first edge of the table and the second edge of the table clockwise (that is, the entire scanned target image is rotated), driving the target image to rotate clockwise for target image position adjustment;

[0018] At the same time, continuously detect whether the first edge of the target table area of the rotated target image is parallel to the preset horizontal line (that is, the reference line of the long - side edge line of the table). If so, consider that the table edge is in the aligned position, and then the top - most invoice as the target image is adjusted, and it is scanned, recognized, and entered;

[0019] Among them, the second initial condition means that the spacing between any two consecutive current characters in the same line is consistent with the preset character spacing value and multiple non - consecutive phrases in the same line meet the financial bill - related semantics.

[0020] Specifically, to determine that the semantics of two consecutive characters have financial bill - related semantics, it specifically includes the following steps:

[0021] Preset multiple sets of financial bill - related vocabulary and build a semantic library from multiple sets of financial bill - related vocabulary;

[0022] First, obtain the current two consecutive characters, and match them with the current semantic library. If the semantic match is successful, it is determined that the current two consecutive characters have the semantic association with financial bills; then obtain the corresponding set of financial bill-related words that match the current two consecutive characters.

[0023] It should be noted that if they are the same, it can be determined that the current characters are from the first character to the second character, and the long side edge line of the table parallel to the vector direction from the first character to the second character is determined as the first edge of the table.

[0024] Specifically, the character boundaries within the table area are determined by the edge detection method, which specifically includes the following steps:

[0025] Perform binary processing on the target image to obtain a black and white image;

[0026] Define the first edge and the second edge of the table area as the X-axis and Y-axis respectively. Based on the X and Y-axis directions as the basic directions, let the recognition point start from the X-axis coordinate of 0 and scan downward. The scanning width is 1px. If the RGB value of the current pixel point scanned is 0, record the X-axis coordinate A1 at this time, and then continue to scan downward to obtain the next X-axis coordinate A2;

[0027] The recognition point starts from A1 + 1 and scans vertically through each pixel point on the entire target image. If the RGB value of the current pixel point scanned is 255, record the X-axis coordinate B1 at this time; then continue to scan downward to obtain the next X-axis coordinate B2;

[0028] In the X interval (A1, (B - 1)1), the recognition point scans horizontally from the point where the Y coordinate is 0, and judges the R value of each point until the RGB value is equal to 0, then stops scanning and records the Y-axis coordinate C1 at this time; then continue to scan to the right to obtain the next Y-axis coordinate C2;

[0029] In the X interval (A1, (B - 1)1), the recognition point starts to scan horizontally from C1 + 1, and judges the R value of each pixel point. If the R value is equal to 255, stop scanning and record the Y-axis coordinate D1 at this time; then continue to scan to the right to obtain the next Y-axis coordinate D2;

[0030] Obtain the abscissa and ordinate of the horizontal and vertical four boundary points of the character, that is, Z1(A1, (B - 1)1, C1, (D - 1)1) is the character boundary.

[0031] Specifically, the calculation method of the distance between two consecutive characters is the difference between the boundary coordinates of the consecutive second character and the boundary coordinates of the first character. The calculation formula is:

[0032] Z2(A2, (B - 1)2, C2, (D - 1)2) - Z1(A1, (B - 1)1, C1, (D - 1)1)

[0033] Wherein, Z2 is the character boundary coordinate of the second character; Z1 is the first character boundary coordinate.

[0034] The present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the financial ticket processing method based on image recognition are implemented.

[0035] Compared with the prior art, the present invention has at least the following beneficial effects:

[0036] The present invention provides a financial ticket processing method and a readable storage medium based on image recognition. In the above-mentioned financial ticket processing method based on image recognition, the text information and image information in the whole invoice are recognized by scanning the whole invoice and stored in a database, which is convenient for staff to input the invoice code or other basic information so as to retrieve the relevant invoice scan data to obtain the bill evidence data. However, when scanning the whole invoice, if there is invoice overlap, especially in key sub-regions (such as the seller or buyer sub-regions in the form area), key recognition and standardized scanning are required. In this regard, the financial ticket processing method based on image recognition provided by the embodiments of the present invention screens the scanned invoice images, excludes the compliant scanned invoice images, and automatically processes the non-compliant invoice images to make the automatic scanning more convenient. The present invention processes the non-compliant invoice images, accurately obtains the invoice image information and compares it with the text information therein to make the processing accurate. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0038] Figure 1 It is the main implementation step flowchart of the financial ticket processing method based on image recognition provided by the embodiments of the present invention;

[0039] Figure 2 It is one of the flow schematic diagrams of the implementation manner for processing the image information in the scan information of the current overlapping invoice when the information in the detected sub-region in the financial ticket processing method based on image recognition provided by the embodiments of the present invention is incorrect;

[0040] Figure 3Schematic diagram of the invoice scanning window in the financial ticket processing method based on image recognition provided by the embodiments of the present invention;

[0041] Figure 4 Schematic diagram of multiple invoices overlapping and misaligned within the scanning window in the financial ticket processing method based on image recognition provided by the embodiments of the present invention;

[0042] Figure 5 Flowchart of one implementation manner of the implementation steps for processing the target image when the long-edge edge line of the table in the financial ticket processing method based on image recognition provided by the embodiments of the present invention is not parallel to the reference line of the long-edge edge line of the table;

[0043] Figure 6 Flowchart of one implementation manner of the implementation steps for determining that the semantics of two consecutive characters have financial bill-related semantics in the financial ticket processing method based on image recognition provided by the embodiments of the present invention;

[0044] Figure 7 Flowchart of the second implementation manner of processing the image information in the scanning information of the current overlapping invoices when the corresponding information in the detection sub-region is incorrect in the financial ticket processing method based on image recognition provided by the embodiments of the present invention;

[0045] Figure 8 Flowchart of the second implementation manner of the implementation steps for processing the target image when the long-edge edge line of the table in the financial ticket processing method based on image recognition provided by the embodiments of the present invention is not parallel to the reference line of the long-edge edge line of the table;

[0046] Figure 9 Flowchart of the second implementation manner of the implementation steps for determining that the semantics of two consecutive characters have financial bill-related semantics in the financial ticket processing method based on image recognition provided by the embodiments of the present invention;

[0047] Figure 10 Flowchart of the implementation steps for determining the character boundaries within the table area by the edge detection method in the financial ticket processing method based on image recognition provided by the embodiments of the present invention;

[0048] Figure 11 Simulation schematic diagram of determining the character boundaries within the table area by the edge detection method in the financial ticket processing method based on image recognition provided by the embodiments of the present invention. Detailed implementation manners

[0049] The above content is only an example and illustration of the concept of the present invention. Those skilled in the art of this technology can make various modifications or supplements to the described specific embodiments or use similar methods for substitution, as long as they do not deviate from the concept of the invention or exceed the scope defined by this claim book, they should all fall within the protection scope of the present invention.

[0050] Example 1

[0051] Please refer to Figure 1 As shown, Example 1 of the present invention provides a financial ticket processing method based on image recognition, and this method includes the following steps:

[0052] Step S1: In the initial state, first use the invoice tax control device to perform initialization collection on the purchasers and each seller of multiple invoices;

[0053] Step S2: Obtain the text information of multiple sellers and the text information of the purchaser, and form a seller text set with the text information of multiple sellers; form a purchaser text set with the text information of multiple purchasers; (The initialization collection of the purchasers and each seller of multiple invoices by using the invoice tax control device is directly obtained through the data interface method. There is another method, that is, the purchasers and each seller are pre-stored, and the common customer Party A information is manually input, etc., so as to obtain the text information of multiple sellers and the text information of the purchaser, and then construct the seller text set and the purchaser text set (hereinafter collectively referred to as the text set);

[0054] Step S3: Divide the basic template invoice into regions according to its functions, and then obtain each divided sub-region, where the functions include the header, two-dimensional code, password area, and table area, and number the divided sub-regions in a preset order, and mark them as 1, 2,... k,... v in sequence, and then obtain the positions corresponding to each sub-region, and construct a set W of positions of each sub-region (W1, W2,... Wk,... Wv), where Wk represents the position where the k-th sub-region is located;

[0055] Step S4: When performing a scanning processing operation on the layout of the current whole invoice, and obtain the scanned information: The scanning of the layout of the whole invoice includes scanning the header, two-dimensional code, password area, and table area of the whole invoice, and then obtain the region characterization information corresponding to each sub-region, and then construct a set H of information of each sub-region (H1, H2,... Hk,... Hv), where Hk represents the region characterization information corresponding to the k-th sub-region;

[0056] The said table area includes key information such as the purchaser and the seller. However, for the table area, the most crucial ones are the purchaser and the seller. At the same time, the purchaser and the seller correspond to their respective important information, namely the regional characterization information of the purchaser and the regional characterization information of the seller. Among them, the seller includes the seller's text information and the official seal corresponding to the seller; the purchaser includes the purchaser's text information, the product name or business item corresponding to the purchaser (the general term of the product name or business item is the product item), and the quantity corresponding to the purchaser's product name or business item. In other words, the purchaser includes the purchaser's text information, the product item corresponding to the purchaser, and the quantity of the purchaser's product item.

[0057] Step S5: Perform the scanning information screening and processing operation: Obtain and then call the sub-regions where the purchaser and the seller are located in the table area of the current entire invoice, obtain the regional characterization information corresponding to the sub-regions of the purchaser and the seller, compare the regional characterization information corresponding to the sub-region of the purchaser with the purchaser text set. If there is a mismatch, the sub-region is determined to be an abnormal sub-region. If it matches, it is regarded as a normal sub-region, and then the normal sub-region is filtered (if it matches, it means that the similarity between the regional characterization information corresponding to the sub-region of the purchaser and a certain purchaser information text in the purchaser text set is greater than the standard threshold); compare the regional characterization information corresponding to the sub-region of the seller with the seller text set. If there is a mismatch, the sub-region is determined to be an abnormal sub-region. If it matches, it is regarded as a normal sub-region, and then the normal sub-region is filtered (if it matches, it means that the similarity between the regional characterization information corresponding to the sub-region of the seller and a certain seller information text in the seller text set is greater than the standard threshold; or if it matches, it means that the regional characterization information corresponding to the sub-region is exactly the same as a certain seller text information in the seller text set); if the sub-region is an abnormal sub-region, mark this region as the marked region, process the image information in the scanning information of the current entire invoice corresponding to the marked region, confirm that it is aligned with the scanning window, and then scan and enter the information of the entire invoice.

[0058] In step S5, obtain the set of region characterization information corresponding to the sub-regions of the current entire invoice, and then call the sub-regions where the purchaser and the seller are located in the table area of the current entire invoice to obtain the region characterization information corresponding to the sub-regions of the purchaser and the seller. Compare the region characterization information corresponding to the sub-region of the purchaser with the purchaser text set. In the case of normal non-overlap, there should be a certain matching relationship between the set of region characterization information corresponding to the sub-regions of the currently scanned entire invoice and the text set stored in text form; if the information scanned in the sub-region (such as the seller text information and purchaser text information corresponding to the table area in the sub-region) is accurate, then regard this sub-region as a normal sub-region, and then filter this normal sub-region, and then filter the entire invoice (for some invoices, there may only be no occlusion of the table area, and other areas such as the password area are occluded or overlapped. Of course, the matching of the profit password area can also identify invoice overlap, but in the embodiments of the present invention, invoice overlap in non-table areas is not recognized or considered;

[0059] If the information corresponding to the sub-region is incorrect (that is, when scanning the invoice, due to the overlap of multiple invoices, the region characterization information corresponding to the sub-regions of the currently scanned entire invoice may not correspond to the text set at the corresponding position), then after processing the scanned information image, re-scan the entire invoice and perform the recognition and entry operation.

[0060] The above financial ticket processing method based on image recognition recognizes the text information and image information in the entire invoice through scanning the entire invoice and stores them in the database, which is convenient for staff to input the invoice code or some other basic information, so as to retrieve the relevant invoice scanning data to obtain the bill evidence data.

[0061] Analyzing the above technical solution, it can be seen that: A financial ticket processing method based on image recognition involved in the present invention divides the basic template invoice into regions according to its functions through an invoice tax control device or preset text information, and then obtains each divided sub-region, where the functions include a header, a two-dimensional code, a password area, and a table area, and numbers the divided sub-regions in a preset order, sequentially marked as 1, 2,... k,... v, and then obtains the positions corresponding to each sub-region, and constructs a set W (W1, W2,... Wk,... Wv) of the positions of each sub-region, where Wk represents the position where the k-th sub-region is located; when scanning, when performing a scanning process operation on the layout of the current entire invoice, and obtaining the scanned information: The scanning of the layout of the entire invoice includes scanning the header, two-dimensional code, password area, and table area of the entire invoice, and then obtaining the regional characterization information corresponding to each sub-region, and then constructing a set of information for each sub-region; Subsequently, it is screened and judged whether the regional characterization information corresponding to the scanned sub-region matches the pre-stored or called text set; if the regional characterization information corresponding to the sub-region is accurate, the sub-region is regarded as a normal sub-region, and then the normal sub-region is filtered and stored, and then the corresponding entire invoice is filtered and stored, that is, stored in the database;

[0062] If the regional characterization information corresponding to the sub-region is incorrect (when multiple invoices overlap during invoice scanning, resulting in the scanned entire invoice not corresponding), it is marked, and the image information in the current scanned information of the entire invoice corresponding to the marked area is reprocessed, and the entire invoice is scanned again and the recognition and entry operation is performed.

[0063] One of the specific implementation manners of processing the image information in the current scanned information of the entire invoice corresponding to the marked area;

[0064] As Figure 2 shown, in step S5, if the information corresponding to the sub-region is incorrect, after processing the scanned information image, the entire invoice is scanned again and the recognition and entry operation is performed (when it is detected that the information corresponding to the sub-region is incorrect, the current invoice is recognized as the current overlapping invoice, and the image information in the scanned information of the current overlapping invoice is processed), which specifically includes the following steps:

[0065] Step S51: Refer to Figure 4 , scan multiple invoices in the scanning window (at this time, multiple invoices overlap and are not placed correctly), divide the image information of the currently scanned overlapping invoice, and obtain an image of the topmost invoice as the target image; and use the horizontal line where the current scanning window is located as the preset horizontal line, and use the preset horizontal line as the reference line for the long edge of the table;

[0066] Refer to Figure 3, the outer frame is the current scanning window, and the inner frame is the scanned image of the entire invoice in the aligned state;

[0067] The horizontal line where the current scanning window is located is used as the preset horizontal line. What the current scanning window scans is a rectangular frame image. Explain the horizontal line where the current scanning window is located. For the current scanning window of the rectangle, a horizontal line is preset for the current overlapping invoice obtained by the window scanning range of the view window. It can also be understood that for the current overlapping invoice, a physical reference horizontal line is set from the left side to the right side with its outer border as the rectangle;

[0068] Step S52: Scan and identify the table edges in the table area of the target image to obtain the long table edge line and the short table edge line. According to the comparison line of the long table edge line and the long table edge line, initially judge whether it meets the first initial condition of the alignment position (that is, if the comparison line of the long table edge line and the long table edge line is parallel, it is initially determined that the table area of the target image may meet the alignment position condition, that is, when it is completely aligned or completely flipped 180 degrees, the comparison line of the long table edge line and the long table edge line is also parallel);

[0069] Step S53: If the scanned and identified long table edge line is not parallel to the comparison line of the long table edge line (the long table edge line is not parallel to the comparison line of the long table edge line), then adjust the target image, scan and identify and input it, see Figure 5 , including the following steps:

[0070] Step S531: Determine the boundaries of the characters (or simply characters) in the table area by the edge detection method;

[0071] Step S532: Judge whether the distance between two consecutive characters is equal to the preset character distance value. (If not, it can be initially determined that the two consecutive characters do not belong to the same line of characters, and they will be filtered out). If so (it can be determined that they are consecutive characters in the same line), then further judge whether the semantics of the two consecutive characters have the associated semantics of financial bills (if not, they will be filtered out);

[0072] Step S533: If it is judged that the semantics of two consecutive characters have the associated semantics of financial bills (preset associated semantics), then determine the above two consecutive characters with the associated semantics of financial bills as the first character and the second character in sequence, and further determine that the direction of the character extension vector from the first character to the second character (character extension vector direction) parallel to the long table edge line of the target table area in the target image is the first edge of the target table area of the target image;

[0073] Step S534: Determine the vertical line of the first edge of the target table area of the target image as the second edge of the table (i.e., the short edge line of the table), and rotate the first edge of the table and the second edge of the table clockwise (processed using an image rotation tool, that is, the entire scanned target image is rotated), driving the target image to rotate clockwise to adjust the position of the target image;

[0074] Step S535: At the same time, detect in real time whether the first edge of the target table area of the rotated target image is parallel to the preset horizontal line. If so, it is regarded that the table edge is in the aligned position. Furthermore, the topmost invoice as the target image is adjusted, and it is scanned, recognized, and entered.

[0075] Analyzing the above solution, it can be seen that in specific operations, when the corresponding information in the detected sub-region is incorrect, the current invoice is identified as the current overlapping invoice, and the image information in the scanned information of the current overlapping invoice is processed. The target image is obtained through image segmentation, and a preset parallel line is used as the reference line for the table edge line. Furthermore, the edge lines of the table area in the target image are scanned and recognized to obtain the long edge line and the short edge line, and it is determined whether the long edge line and the reference line of the table edge line meet the first initial condition, that is, the long edge line is parallel to the reference line of the table edge line; if not, it is regarded that the target image is not aligned with the scanning window, and the target image is rotated:

[0076] By detecting whether the distance between two consecutive characters is equal to the preset distance value, if so, it can be determined that these two consecutive characters are characters in the same row. Further, it is determined whether the semantics of multiple consecutive characters have semantics related to financial bills. If so, it can be determined that the long edge line of the table in the target image where the extension vector directions of any two characters of multiple consecutive characters are parallel is the first edge of the target table area of the target image, and the vertical line of the first edge of the table is determined as the second edge of the table.

[0077] At the same time, detect in real time whether the first edge line of the table area of the rotated target image is parallel to the preset horizontal line. If so, it is regarded that the table edge is in the aligned position. At this time, the information of the entire invoice can be completely recognized by scanning the image, and it is scanned, recognized, and entered.

[0078] Specifically, see Figure 6 , in step S533, determining that the semantics of two consecutive characters have semantics related to financial bills specifically includes the following steps:

[0079] Step S5331: Preset multiple sets of financial bill-related vocabulary and a semantic library constructed by multiple sets of financial bill-related vocabulary;

[0080] Step S5332: First, obtain the current two consecutive characters, match the current two consecutive characters with the current semantic library, and if the semantic match is successful, determine that the current two consecutive characters have the semantic association with financial bills; then obtain the corresponding set of financial bill-related words that match the current two consecutive characters.

[0081] It should be noted that if they are the same, it can be determined that the current character is from the first character to the second character, and then the long edge line of the table parallel to the vector direction from the first character to the second character is determined as the first edge of the table.

[0082] Analyzing the above solution, it can be seen that a financial ticket processing method based on image recognition provided by the present invention, through a preset semantic library constructed by presetting multiple sets of financial bill-related words and multiple sets of financial bill-related words, by matching multiple two consecutive characters obtained with the semantic library, if the match is successful, it is possible to obtain that the long edge line of the table parallel to the vector direction from any first character to the first character to the second character among multiple two consecutive characters is the first edge of the table.

[0083] The second specific implementation method of processing the image information in the current scanned information of the entire invoice corresponding to the marked area;

[0084] In this embodiment, for processing the image information in the current scanned information of the entire invoice corresponding to the marked area, refer to Figure 7 , specifically including the following steps:

[0085] Step S54: Segment the image information of the currently scanned overlapping invoice, obtain the topmost invoice image as the target image; and use the horizontal line where the current scanning window is located as the preset horizontal line, and use the preset horizontal line as the reference line for the long edge line of the table.

[0086] Step S55: Obtain the long edge line and short edge line of the table according to the table edge in the table area of the scanned and recognized target image, and initially judge whether it meets the first initial condition of the alignment position according to the long edge line of the table and the reference line of the long edge line of the table; the first initial condition refers to that the long edge line of the current table and the reference line of the long edge line of the table are in a parallel relationship.

[0087] Step S56: If the long edge line of the scanned and recognized table is not parallel to the reference line of the long edge line of the table, adjust the target image, and perform scanning and recognition input on the adjusted target image that meets the second initial condition, refer to Figure 8 , including the following steps:

[0088] Step S561: Determine the character boundary (the boundary of the character) in the table area through the edge detection method.

[0089] Step S562: Obtain the character boundaries of two consecutive current characters in the table area to determine the character spacing between the two consecutive characters, and determine whether the character spacing between the two consecutive characters is equal to a preset character spacing value. If so, consider the current two consecutive characters as being on the same line, and further determine whether the semantics of multiple non - consecutive pairs of consecutive characters within the same line have a financial bill - related semantics; the semantics of the above - mentioned two consecutive characters are regarded as a phrase, and the semantics of multiple non - consecutive pairs of consecutive characters within the same line are regarded as multiple non - consecutive phrases within the same line.

[0090] Step S563: If it is determined that multiple non - consecutive phrases within the same line have a financial bill - related semantics, determine any two characters within the same line with the financial bill - related semantics as the first character and the second character in sequence, and further determine that the edge line of the long side of the table in the target image parallel to the character extension vector direction from the first character to the second character is the first edge of the target table area of the target image.

[0091] Step S564: Determine the perpendicular line to the first edge of the target table area of the target image as the second edge of the table, and rotate the first edge of the table and the second edge of the table clockwise to drive the target image to rotate clockwise for target image position adjustment.

[0092] Step S565: At the same time, continuously detect whether the first edge of the target table area of the rotated target image is parallel to a preset horizontal line. If so, consider that the table edge is in the aligned position, and then the top - most invoice as the target image is adjusted, and it is scanned, recognized, and entered.

[0093] The second initial condition means that the spacing between any two consecutive current characters in the same line is consistent with the preset character spacing value and multiple non - consecutive phrases within the same line conform to the financial bill - related semantics.

[0094] During the execution of step S563 in the above - mentioned specific implementation, the determination that multiple non - consecutive phrases within the same line have a financial bill - related semantics refers to Figure 9 , and it specifically includes the following steps:

[0095] Step S5631: Preset multiple financial bill - related vocabulary sets and build a semantic library from the multiple financial bill - related vocabulary sets.

[0096] Step S5632: First, obtain multiple non - consecutive phrases within the same line, match the multiple non - consecutive phrases within the same line with the current semantic library. If the semantic match is successful, determine that the multiple non - consecutive phrases within the current same line have a financial bill - related semantics; then obtain the corresponding financial bill - related vocabulary set that matches the multiple non - consecutive phrases within the same line.

[0097] Specifically, as Figure 10 shown, in step S531 corresponding to the first specific implementation manner or step S561 of the second specific implementation manner, the boundaries of the characters within the table area (or simply referred to as characters) are determined by an edge detection method, including the following steps:

[0098] Step S5311: Binarize the target image (the binarization method here is common knowledge and will not be elaborated further), obtaining a black-and-white image (at this time, an image with an RGB value of 255 is white or an image with a value of 0 is black);

[0099] Step S5312: Define the first edge and the second edge of the table area as the X-axis and Y-axis respectively. Based on the X and Y-axis directions as the basic directions, let the recognition point start from the origin of the X and Y-axis intersection point and scan downward traversally, with a scanning traversal width of 1px. First, scan and traverse the character boundary of one of the two consecutive characters, Z1. If the RGB value of the currently scanned pixel point is 0, record the coordinate A1 at this time, thereby obtaining the coordinate A1 of the left boundary point of Z1 (the left boundary point (boundary pixel point) of Z1 is a complete coordinate A1). Then continue to scan downward traversally to obtain the pixel point coordinate A2 of the next character;

[0100] Step S5313: The recognition point starts from A1 + 1 and scans and traverses each pixel point on the entire target image vertically. If the RGB value of the currently scanned pixel point is 255, record the coordinate B1 at this time. Based on this coordinate B1, obtain the coordinate (B - 1)1 of the right boundary point of Z1 (that is, the abscissa of (B - 1)1 is the abscissa of the original coordinate point B1 - 1, and the ordinate is the ordinate value of the original coordinate point B1). Then continue to scan and traverse to the right along the X-axis direction to obtain the pixel point coordinate B2 of the next character;

[0101] Step S5314: In the M interval (A1, (B - 1)1), the recognition point scans vertically starting from the point where the Y coordinate is 0, and judges the R value of each point until the RGB value is equal to 0, then stops scanning and records the coordinate C1 at this time to obtain the coordinate C1 of the upper boundary point of Z1 (which is the coordinate C1 of the upper boundary point of Z1); then continue to scan and traverse downward along the Y-axis direction to obtain the pixel point coordinate C2 of the next character;

[0102] Step S5315: In the M interval (A1, (B - 1)1), the recognition point starts to scan vertically from the coordinate point C1 + 1 (the coordinate point C1 + 1 is obtained by keeping the X-axis value of the original coordinate point C1 unchanged and adding 1 to the Y-axis value of the original coordinate point C1 to get the Y-axis value of the coordinate point C1 + 1), and judges the R value of each pixel point. If the R value is equal to 255, stop scanning and record the coordinate D1 at this time (which is the coordinate (D - 1)1 of the lower boundary point of Z1);

[0103] Step S5316: According to the coordinates of the four boundary points of the current character Z1, namely the left boundary point, right boundary point, upper boundary point, and lower boundary point, calculate the character boundary of the current character Z1 based on the coordinates of the above four boundary points, and record it as Z1(A1, (B - 1)1, C1, (D - 1)1) as the character boundary;

[0104] Step S5317: Then repeat the above steps to obtain the coordinates of the character boundary points of another character Z2 in two consecutive characters, and record it as Z2(A2, (B - 1)2, C2, (D - 1)2) as the character boundary.

[0105] Analyzing the above solution, it can be seen that the first edge of the table area is defined as the X-axis, the second edge of the table area is defined as the Y-axis, starting from the intersection point of the X and Y axes as the origin, traverse each pixel point in turn, and obtain the coordinates of the boundary points of a character Z1 in two consecutive characters, that is, Z1(A1, (B - 1)1, C1, (D - 1)1). Further, repeat the coordinates of the Z1 boundary points to obtain the coordinates of the character boundary points of another character Z2 in two consecutive characters, that is, Z2(A2, (B - 1)2, C2, (D - 1)2).

[0106] Specifically, the calculation method for the distance between two consecutive characters is the difference between the boundary coordinates of the second consecutive character and the boundary coordinates of the first character. The calculation formula is:

[0107] Z2(A2, (B - 1)2, C2, (D - 1)2) - Z1(A1, (B - 1)1, C1, (D - 1)1);

[0108] In the formula, Z2 is the character boundary coordinate of the second character; Z1 is the character boundary coordinate of the first character.

[0109] In summary, the present invention obtains a large amount of invoice information through scanning, which is convenient for the staff to organize, manage, and retrieve a large number of invoices. And when the scanned image is inaccurate due to multiple invoices overlapping or even being rotated, resulting in inaccurate acquisition of invoice information, the topmost layer is obtained as the target image (i.e., the topmost invoice) by segmenting the scanned invoice image, and then the table lines in the table area of the target image are scanned to identify that the placement of the whole invoice is not aligned;

[0110] Furthermore, the consecutive characters in the table area are determined to be characters in the same row by the spacing between the boundary points of the consecutive characters. Then, it is determined that the consecutive characters within the same row of characters conform to the financial-related semantic information. Further, the edge line parallel to the straight line of any character vector direction within the current row of characters is determined as the first edge line (i.e., the long-side edge line). Then, the second edge line (i.e., the short-side edge line) that is connected to and perpendicular to the first edge line is determined. By rotating the first edge line and the second edge line (i.e., rotating the entire scanned target image), the target image is aligned with the scanning frame, and the invoice information is scanned and recognized for entry.

[0111] Embodiment 2

[0112] The present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the financial ticket processing method based on image recognition in Embodiment 1 are implemented.

[0113] The above content is only an example and illustration of the concept of the present invention. Those skilled in the art of the present technology can make various modifications or supplements to the described specific embodiments or use similar methods for substitution, as long as they do not deviate from the concept of the invention or exceed the scope defined by this claim book, they should fall within the protection scope of the present invention.

Claims

1. A financial ticket processing method based on image recognition, characterized in that: The method includes the following steps: Obtain the text information of multiple sellers and the text information of the buyer, form a seller text set from the text information of multiple sellers; form a buyer text set from the text information of multiple buyers; Divide the basic template invoice into regions according to its functions, and then obtain each divided sub-region, where the functions include the header, QR code, password area, and table area, and number the divided sub-regions in a preset order, and mark them as 1, 2,... k,... v in sequence, and then obtain the positions corresponding to each sub-region, and construct a set of sub-region positions W (W1, W2,... Wk,... Wv), where Wk represents the position where the kth sub-region is located; When performing a scanning process operation on the layout of the current entire invoice, and obtain the scanned information: The scanning of the layout of the entire invoice includes scanning the header, QR code, password area, and table area of the entire invoice, and then obtaining the region characterization information corresponding to each sub-region, and then constructing a set of sub-region information H (H1, H2,... Hk,... Hv), where Hk represents the region characterization information corresponding to the kth sub-region; Perform a scanning information screening process operation: Obtain and then call the sub-regions where the buyer and seller are located in the table area of the current entire invoice, obtain the region characterization information corresponding to the sub-regions of the buyer and seller, compare the region characterization information corresponding to the sub-region of the buyer with the buyer text set, if there is a mismatch, then determine that the sub-region is regarded as an abnormal sub-region, if it matches, it is regarded as a normal sub-region, and then filter the normal sub-region; compare the region characterization information corresponding to the sub-region of the seller with the seller text set, if there is a mismatch, then determine that the sub-region is regarded as an abnormal sub-region, if it matches, it is regarded as a normal sub-region, and then filter the normal sub-region; if the sub-region is an abnormal sub-region, then mark this region as a marked region, and then process the image information in the scanned information of the current entire invoice corresponding to the marked region, confirm that it is aligned with the scanning window, and then scan and input the information of the entire invoice.

2. The financial ticket processing method based on image recognition according to claim 1, characterized in that The processing of the image information in the scanned information of the current entire invoice corresponding to the marked region specifically includes the following steps: Segment the scanned image information of the current overlapping invoice, and obtain the topmost invoice image as the target image; and use the horizontal line where the current scanning window is located as the preset horizontal line, and use the preset horizontal line as the reference line for the long edge line of the table; Obtain the long edge line and short edge line of the table according to the table edge in the table area of the scanned and recognized target image, and initially judge whether it meets the first initial condition of the alignment position according to the long edge line of the table and the reference line of the long edge line of the table; The first initial condition refers to that the long edge line of the current table and the reference line of the long edge line of the table are in a parallel relationship; If the scanned and recognized long edge line of the table is not parallel to the reference line of the long edge line of the table, then adjust the target image, and perform scanning and recognition input on the adjusted target image that meets the second initial condition, including the following steps: Determine the character boundaries within the table area through edge detection method; Obtain the character boundaries of two consecutive current characters within the table area to determine the character spacing between the two consecutive characters, and determine whether the character spacing between the two consecutive characters is equal to the preset character spacing value. If so, further determine whether the semantics of the two consecutive characters have the associated semantics of financial bills; If it is determined that the semantics of the two consecutive characters have the associated semantics of financial bills, determine the two consecutive characters with the associated semantics of financial bills as the first character and the second character in sequence, and further determine that the table long-edge line of the target table area in the target image parallel to the character extension vector direction from the first character to the second character is the first edge of the target table area of the target image; Determine the perpendicular line to the first edge of the target table area of the target image as the second edge of the table, and rotate the first edge of the table and the second edge of the table clockwise to drive the target image to rotate clockwise for target image position adjustment; At the same time, continuously detect whether the first edge of the target table area of the rotated target image is parallel to the preset horizontal line. If so, it is regarded that the table edge is in the aligned position. Furthermore, the topmost invoice as the target image is adjusted and scanned and recognized and entered; The second initial condition refers to that the spacing between any two consecutive current characters in the same row is consistent with the preset character spacing value and the semantics of any two consecutive current characters in the same row conform to the associated semantics of financial bills.

3. The financial ticket processing method based on image recognition according to claim 2, wherein The determination that the semantics of two consecutive characters have the associated semantics of financial bills specifically includes the following steps: Preset multiple sets of financial bill-related vocabulary and construct a semantic library from the multiple sets of financial bill-related vocabulary; First, obtain the current two consecutive characters, match the current two consecutive characters with the current semantic library. If the semantic match is successful, it is determined that the current two consecutive characters have the associated semantics of financial bills; then obtain the corresponding set of financial bill-related vocabulary that matches the current two consecutive characters successfully.

4. The financial ticket processing method based on image recognition according to claim 3, wherein The processing of the image information in the current entire invoice scan information corresponding to the marked area specifically includes the following steps: Segment the image information of the currently scanned overlapping invoice to obtain an invoice image of the topmost layer as the target image; and use the horizontal line where the current scan window is located as the preset horizontal line, and use the preset horizontal line as the reference line for the table long-edge line; Obtain the table long-edge line and the table short-edge line according to the table edge of the table area in the scanned and recognized target image, and initially determine whether it meets the first initial condition of the aligned position according to the table long-edge line and the reference line of the table long-edge line; the first initial condition refers to that the current table long-edge line and the reference line of the table long-edge line are in a parallel relationship; If it is scanned and recognized that the table long-edge line is not parallel to the reference line of the table long-edge line, adjust the target image, and perform scanning and recognition and entry on the adjusted target image that meets the second initial condition, including the following steps: Determine the character boundaries within the table area through edge detection method; Obtain the character boundaries of two consecutive current characters in the table area to determine the character spacing between the two consecutive characters, and determine whether the character spacing between the two consecutive characters is equal to the preset character spacing value. If so, consider the current two consecutive characters as being on the same line, and further determine whether the semantics of multiple non - consecutive pairs of two characters within the same line have a financial bill - related semantics; The semantics of the above - mentioned two consecutive characters are regarded as a phrase, and the semantics of multiple non - consecutive pairs of two characters within the same line are regarded as multiple non - consecutive phrases within the same line; If it is determined that multiple non - consecutive phrases within the same line have a financial bill - related semantics, then determine any two characters within the same line having the financial bill - related semantics as the first character and the second character in sequence, and further determine that the long - side edge line of the target table area in the target image parallel to the character extension vector direction from the first character to the second character is the first edge of the target table area of the target image; Determine the perpendicular line to the first edge of the target table area of the target image as the second edge of the table, and rotate the first edge of the table and the second edge of the table clockwise to drive the target image to rotate clockwise for target image position adjustment; At the same time, real - time detect whether the first edge of the target table area of the rotated target image is parallel to the preset horizontal line. If so, consider that the table edge is in the aligned position, and then the top - most invoice as the target image is adjusted, and it is scanned, recognized, and entered; The second initial condition refers to that the spacing between any two consecutive current characters in the same line is consistent with the preset character spacing value and multiple non - consecutive phrases within the same line conform to the financial bill - related semantics.

5. The financial ticket processing method based on image recognition according to claim 4, wherein The determination that multiple non - consecutive phrases within the same line have a financial bill - related semantics specifically includes the following steps: Preset multiple financial bill - related vocabulary sets and build a semantic library from the multiple financial bill - related vocabulary sets; First, obtain multiple non - consecutive phrases within the same line, match the multiple non - consecutive phrases within the same line with the current semantic library. If the semantic match is successful, then determine that the multiple non - consecutive phrases within the current same line have a financial bill - related semantics; then obtain the corresponding financial bill - related vocabulary set that matches the multiple non - consecutive phrases within the same line successfully.

6. The financial ticket processing method based on image recognition according to claim 5, characterized in that The semantic library includes an enterprise information semantic library, a bank information semantic library, an industry service category information semantic library, and a product category information semantic library.

7. A method for processing financial tickets based on image recognition according to claim 2 or 4, characterized in that The determination of the character boundaries within the table area by the edge detection method specifically includes the following steps: Perform binarization processing on the target image to obtain a black - and - white image; Define the first edge of the table area and the second edge of the table area as the X - axis and the Y - axis respectively. Based on the X and Y - axis directions as the basic directions, make the recognition point start from the X - axis coordinate of 0 and scan downwards, and the scanning width is 1px; First, scan and traverse the character boundary of one of the two consecutive characters, Z1. If the RGB value of the currently scanned pixel point is 0, record the coordinate A1 at this time, and thus obtain the coordinate A1 of the left boundary point of Z1. Then continue to scan and traverse downward to obtain the pixel point coordinate A2 of the next character. The recognition point starts from A1 + 1 and vertically scans and traverses each pixel point on the entire target image. If the RGB value of the currently scanned pixel point is 255, record the coordinate B1 at this time. Based on the coordinate B1 at this time, obtain the coordinate of the right boundary point of Z1, which is (B - 1)1: that is, the abscissa of (B - 1)1 is the abscissa of the original coordinate point B1 minus 1, and the ordinate is the ordinate value of the original coordinate point B1. In the M interval (A1, (B - 1)1), the recognition point vertically scans starting from the point where the Y coordinate is 0, and judges the R value of each point until the RGB value is equal to 0, then stops scanning, and records the coordinate C1 at this time to obtain the coordinate C1 of the upper boundary point of Z1. In the M interval (A1, (B - 1)1), the recognition point starts vertically scanning from the coordinate point C1 + 1, and judges the R value of each pixel point. If the R value is equal to 255, then stops scanning and records the coordinate D1 at this time. According to the coordinates of the four boundary points, namely the left boundary point, the right boundary point, the upper boundary point, and the lower boundary point, of the current character Z1, calculate the character boundary of the current character Z1 based on the coordinates of the above four boundary points, and record it as Z1(A1, (B - 1)1, C1, (D - 1)1) as the character boundary. Then repeat the above steps to obtain the character boundary of the other character Z2 in the two consecutive characters, and record it as Z2(A2, (B - 1)2, C2, (D - 1)2) as the character boundary.

8. A financial ticket processing method based on image recognition according to claim 7, characterized in that, The character spacing is the character boundary coordinate of the second character minus the character boundary coordinate of the first character, and its calculation formula is: Z2(A2, (B - 1)2, C2, (D - 1)2) - Z1(A1, (B - 1)1, C1, (D - 1)1); In the formula, Z2 is the character boundary coordinate of the second character; Z1 is the character boundary coordinate of the first character.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Value-added tax invoice information extraction method

    CN110751136A

  • Financial account book intelligent classification management cloud platform based on big data and cloud computing

    CN113344686A