Bank electronic receipt identification method based on semantic analysis
By separating the color and structure layers of bank electronic receipts using a semantic analysis-based approach, and combining optical character recognition (OCR) technology with semantic recognition, the problems of model redundancy and low recognition accuracy in traditional systems are solved, achieving efficient and accurate bank electronic receipt recognition and storage.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-26
- Publication Date
- 2026-03-31
AI Technical Summary
Traditional bank electronic receipt recognition systems suffer from a large and redundant model architecture due to the diverse characteristics of bank electronic receipts, which affects operational efficiency and maintainability, and also results in low recognition accuracy.
A semantic analysis-based approach is adopted, which extracts and filters color to obtain text and structural layers. Combined with optical character recognition technology, text recognition and structural correction are performed. Structural data is used for semantic recognition and distributed storage to ensure the accuracy and immutability of the recognition results.
It improves the accuracy and efficiency of bank electronic receipt identification, reduces system complexity, ensures the accuracy and reliability of identification results, reduces manual intervention, and improves the work efficiency of auditors.
Smart Images

Figure CN121768024A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and in particular to a method for recognizing electronic bank receipts based on semantic analysis. Background Technology
[0002] A bank receipt is a written or electronic document issued by a bank to a customer after processing their deposit, withdrawal, or transfer transaction. It serves not only as proof of the transaction but also as an important source document for financial accounting. Traditional manual data entry methods for processing electronic bank deposit slips consume significant manpower and resources and have a high error rate.
[0003] In related technologies, the recognition of bank electronic receipts mainly relies on computer vision (CV) and optical character recognition (OCR) technologies. Specifically, this approach involves using deep learning to train various bank electronic receipt samples to build a corresponding accurate recognition model, thereby ensuring the accurate extraction and recognition of key information from bank electronic receipts.
[0004] Regarding the aforementioned technologies, due to the highly diverse nature of bank electronic receipts, those generated by different banks and in different business scenarios exhibit significant differences in layout, field format, and information content. This high degree of categorization necessitates that a complete recognition system must build and maintain a dedicated recognition model for each category of bank electronic receipts in practical applications. As the number of bank electronic receipt categories continues to increase, the number of recognition models integrated into the system also increases dramatically, ultimately leading to an exceptionally large and redundant model architecture for the entire recognition system. This severely impacts the system's operational efficiency and maintainability, indicating a need for improvement. Summary of the Invention
[0005] To reduce the bloat of the system model architecture, improve system operating efficiency, and ensure the accuracy of bank electronic receipt recognition, this application provides a bank electronic receipt recognition method based on semantic analysis.
[0006] This application provides a semantic analysis-based method for recognizing electronic bank receipts, employing the following technical solution:
[0007] A semantic analysis-based method for recognizing electronic bank receipts includes:
[0008] Obtain bank electronic receipts to be processed and record them as receipts to be processed. Extract and analyze the colors of the receipts to be processed to separate the color layers corresponding to different colors.
[0009] The color layer is filtered and analyzed to obtain the text layer for character recognition and the structure layer for structure recognition;
[0010] Based on optical character recognition technology, text layer is used to recognize characters and determine the text content of the text layer, which is recorded as the first text content.
[0011] Based on the text in the first text content, generate the corresponding text image, and compare the generated text image with the corresponding text in the receipt to be processed to determine the structural consistency of the text image.
[0012] If the structure of the generated text image is inconsistent with that of the receipt to be processed, the receipt to be processed is corrected according to the text image, and the corrected receipt to be processed is subjected to text recognition to obtain the second text content.
[0013] If the generated text image has the same structure as the receipt to be processed, the built-in structure table is corrected according to the structure layer to obtain the structure data, and the first text content or the second text content is classified and statistically analyzed according to the structure data to obtain the text to be verified.
[0014] Semantic recognition is performed on the text to be verified to determine the semantic coherence of the text. If the text to be verified does not have semantic coherence, the text to be verified is combined and judged based on the position information of the structural data and the text to be verified on the corresponding structure to determine whether the combination result has semantic coherence. The combination result with semantic coherence is then distributed and stored.
[0015] Preferably, the layer size of the receipt to be processed is determined to obtain the initial graphic data, and an initial graphic coordinate system is constructed based on the initial image data;
[0016] The color layers are compared with the initial graphic data to determine the pixel proportion of each color layer. The pixel proportion of each color layer is then compared with the built-in proportion threshold. Color layers with a pixel proportion less than the proportion threshold are retained to obtain the first filtering result.
[0017] Based on the initial graphic coordinate system, the continuity of pixels in each color layer in the first screening result is judged, and continuous pixels are recorded as a pixel set.
[0018] The length of the pixel set of each color layer in the first screening result is judged. If the length of the pixel set is greater than the built-in length threshold, the color layer corresponding to the location is determined to be a structural layer.
[0019] For color layers with pixel set lengths less than the built-in length threshold, position coordinates are selected, and the position coordinates of each pixel set are compared to determine whether the intervals between pixel sets are the same.
[0020] If the intervals between position coordinates are the same, then the color layer corresponding to that position coordinate is determined to be the text layer.
[0021] Preferably, based on the initial graphic coordinate system, the pixels of each color layer in the first screening result are marked with coordinates to obtain pixel coordinates;
[0022] The continuity of pixel coordinates is determined, and two pixels with adjacent coordinates are considered continuous. The pixel coordinates with continuous coordinates are counted to obtain the pixel set.
[0023] Based on the edge pixels of the pixel set, a rectangular region that fits the pixel set is constructed, and the length information of the rectangular region is used as the length information of the pixel set.
[0024] The coordinates corresponding to the center position of the rectangular region are recorded as the position coordinates of the pixel set.
[0025] Preferably, based on the text in the first text content, the corresponding generated text image is marked as a standard image;
[0026] Based on the generated text image, the corresponding text image in the receipt to be processed is selected to obtain the image to be matched;
[0027] The standard image and the image to be matched are matched proportionally to determine whether the image to be matched can overlap with the standard image.
[0028] If the image to be matched overlaps with the standard image, the text image is determined to have a consistent structure; otherwise, the text image is determined to have an inconsistent structure.
[0029] Preferably, if the standard image and the image to be matched do not overlap, the standard image is deformed based on the image to be matched until the standard image and the image to be matched overlap, and the image deformation process of the standard image is recorded to obtain the first deformation operation data.
[0030] Based on the purpose of each operation step in the first deformation operation data, the first deformation operation data is simplified to obtain the second deformation operation data.
[0031] Based on the second deformation operation data, the return receipt to be processed is adjusted in reverse to obtain the corrected return receipt.
[0032] Preferably, the built-in structure table is matched based on the initial graphic data to construct a first structure table corresponding to the initial graphic data;
[0033] Based on the structural layer, the position coordinates of the lines in the corresponding structural layer are determined, and the first structural table is matched according to the position coordinates of the lines to determine whether the position coordinates of the lines coincide with the cell border lines of the first structural table.
[0034] If the position coordinates of the structural layer lines coincide with the cell border lines of the first structural table, then the first structural table is divided according to the structural layer lines to obtain the divided cell blocks, and the divided cell blocks are merged to obtain the second structural table.
[0035] If the position coordinates of the structural layer lines do not coincide with the cell borders of the first structural table, the cells in the table are proportionally reduced according to the position coordinates of the structural layer lines until the position coordinates of the structural layer lines coincide with the cell borders of the proportionally reduced first structural table. Then, the adjusted first structural table is divided and cells are merged according to the structural layer lines to obtain the second structural table.
[0036] The table features of the second structure table are extracted to obtain the structure data.
[0037] Preferably, if the semantics of the text to be verified are not coherent, the text to be verified that is not semantically coherent is marked as the text to be analyzed, and the position coordinates of the text to be analyzed are marked to obtain the coordinate data to be analyzed.
[0038] The text to be analyzed is matched with the built-in return order keywords to determine the relevant keywords that can be matched for each character in the text to be analyzed, and these keywords are recorded as the keywords to be analyzed. The keywords to be analyzed include related words and related words.
[0039] Based on the structural data, the coordinate data of the related words in the keyword to be analyzed is obtained, and the surrounding text of the related words is selected according to the coordinate data to be analyzed to obtain the matching words;
[0040] The word to be matched is matched with the related words adjacent to the related words in the keyword to be analyzed. If the match is successful, the adjacent text at the corresponding position of the word to be matched is selected according to the positional relationship between the word to be matched and the related words. The selected text is then matched with the next related word in the keyword to be analyzed until the keyword to be analyzed is successfully matched.
[0041] Identify the keywords that have successfully matched and analyze them. Based on the keywords and their corresponding coordinate data, segment the text to be analyzed to obtain the segmented text.
[0042] The segmented text is combined and semantically recognized, and the combined results with semantic coherence are stored in a distributed manner.
[0043] In summary, this application includes at least one of the following beneficial technical effects:
[0044] 1. By extracting colors from the bank electronic receipts to be processed, they are divided into multiple color layers. This effectively reduces interference from other factors during the recognition of the receipt content, improving the accuracy of the recognition results. Optical character recognition (OCR) technology is used to recognize the separated text layers and construct standard text images based on the recognized characters. These standard text images are then compared with the corresponding images of the receipts to determine if there are any deformation issues, thus ensuring the accuracy of OCR. The structural data represented by the structural layers is used to classify and statistically analyze the accurately recognized text content, and semantic recognition is applied to the statistical results to further verify the correctness of the recognition and statistical results, further improving the accuracy of the bank electronic receipt recognition. Simultaneously, the accurately recognized bank electronic receipts are distributed and stored to ensure their immutability.
[0045] 2. Using the original layer size of the receipt to be processed, initial graphic data and initial graphic coordinates are constructed. Then, the color difference layer containing text and structural layers is filtered out based on the proportion of text color layers. The first filtering result is then filtered again using the initial graphic coordinate system. By comparing the length of the pixel set in each color layer, the structural layer is filtered out using the characteristics of the structural layer. By comparing the position coordinates of the pixel set in each color layer, the text layer is filtered out based on the interval between the position coordinates. This makes the recognition results more accurate when performing text recognition and structural judgment on the receipt to be processed based on the filtered results, thus improving the accuracy of bank electronic receipt recognition.
[0046] 3. By matching the text to be analyzed with built-in receipt keywords, the text of electronic bank receipts is further segmented based on its unique characteristics. Each character in the text is associated with receipt keywords to determine the possible keywords corresponding to each character. Then, based on the positional coordinates between characters, adjacent characters are selected, and the selected results are matched with the corresponding associated characters of the related keywords to determine the successfully matched keywords for each character. The text is then spatially segmented using the positions of the successfully matched keywords, ensuring that each text statement is in a separate cell. This guarantees the accuracy of the segmented text statements and improves the accuracy of electronic bank receipt recognition. Attached Figure Description
[0047] Figure 1 This is a flowchart illustrating the steps of the semantic analysis-based bank electronic receipt recognition method in this embodiment. Detailed Implementation
[0048] The following is in conjunction with the appendix Figure 1 This application will be described in further detail.
[0049] This application discloses a method for recognizing electronic bank receipts based on semantic analysis.
[0050] Example: Figure 1 As shown, the present invention provides a method for recognizing electronic bank receipts based on semantic analysis, comprising:
[0051] S1, obtain the bank electronic receipt to be processed and record it as a receipt to be processed, perform color extraction and statistics on the receipt to be processed, so as to separate the color layers corresponding to different colors;
[0052] S2, the color layer is filtered and analyzed to obtain the text layer for character recognition and the structure layer for structure recognition;
[0053] S3, based on optical character recognition technology, performs character recognition on the text layer to determine the text content of the text layer, which is recorded as the first text content;
[0054] S4. Based on the text in the first text content, generate the corresponding text image, and compare the generated text image with the corresponding text in the receipt to be processed to determine the structural consistency of the text image;
[0055] S5. If the structure of the generated text image is inconsistent with that of the receipt to be processed, the receipt to be processed is corrected according to the text image, and the corrected receipt to be processed is subjected to text recognition to obtain the second text content.
[0056] S6. If the generated text image is consistent with the structure of the receipt to be processed, the built-in structure table is corrected according to the structure layer to obtain the structure data, and the first text content or the second text content is classified and statistically analyzed according to the structure data to obtain the text to be verified.
[0057] S7 performs semantic recognition on the text to be verified to determine its semantic coherence. If the text lacks semantic coherence, it is combined with the corresponding text in the structure based on the location information of the structural data to determine whether the combined result has semantic coherence. Combined results with semantic coherence are then distributed and stored. If the text has semantic coherence, a new bank electronic receipt is constructed based on the semantically recognized text and structural data, and then distributed and stored. This also includes rule matching based on a built-in rule base on the stored bank electronic receipts, and risk alerts are issued for bank electronic receipts that trigger the rules in the rule base. By issuing alerts for risky bank electronic receipts, auditors can make targeted human judgments based on the risk alert signals, improving the efficiency of auditors in processing bank electronic receipts.
[0058] In this embodiment, color extraction is performed on the bank electronic receipts to be processed, dividing them into multiple color layers. This effectively reduces interference from other factors when recognizing the receipt content, improving the accuracy of the recognition results. Optical character recognition (OCR) technology is used to recognize the separated text layers, and standard text images are constructed from the recognized text. These standard text images are then compared with the corresponding text images on the receipts to determine if there are any deformation issues, thus ensuring the accuracy of OCR. Furthermore, the structural data represented by the structural layers is used to categorize and statistically analyze the accurately recognized text content, and semantic recognition is applied to the statistical results to further verify the correctness of the recognition and statistical results, further improving the accuracy of the bank electronic receipt recognition. Meanwhile, by distributing the accurately identified electronic bank receipts to ensure their immutability, the system performs rule matching on the accurately identified electronic bank receipts based on the built-in rule base, and issues risk alerts for electronic bank receipts that trigger the rules. This enables auditors to conduct precise analysis of electronic bank receipts based on risk alert signals, thereby improving the efficiency of auditors' work.
[0059] Exemplarily, when a bank electronic receipt to be recognized is obtained, since an electronic receipt mainly includes four parts: a seal, a receipt structure (explicit or implicit), bank information, and receipt content. Among them, the seal is mostly red, the receipt structure is light or dark, the bank information is a bank icon image, and the receipt content is mostly black. It can be seen that the colors of these four parts are all different. Therefore, color extraction can be performed on the bank electronic receipt to be processed, and pixel points of the same color are pieced together to form a color layer of a certain part. At this time, the color layer includes the white background area, the red seal area, the black receipt content area, the structural image of other colors, and the bank icon image of other colors.
[0060] By screening the color layer, the white background area is removed to avoid recognizing the red seal area and the bank icon image of other colors, and the recognition of the black receipt content area and the structural image of other colors is retained, so as to reduce the color layer to be analyzed and improve the processing efficiency of the system for data. At the same time, by performing color extraction on the receipt to be processed, interference from other factors to text recognition is reduced, and the recognition accuracy is improved.
[0061] After determining the text layer to be text-recognized and the structure layer to be structure-recognized, first, the text layer is recognized by the optical character recognition technology (OCR recognition technology) to determine the recognizable text content. When using the OCR recognition technology to recognize the text layer, the text layer can be reprocessed by methods such as decoloring, denoising, and rotation to ensure the accuracy of the recognition result of the OCR recognition technology.
[0062] When the text content in the text layer is recognized by using the OCR recognition technology, a text image is generated once for the text content. For example: if the recognized text contains the two characters "receipt", then separate images of "receipt" and "single" are generated to obtain the image representation of "receipt" and "single" under the default standard of the system, and then the "receipt" and "single" are compared with the characters at the corresponding positions on the receipt to be processed. Thus, according to the comparison result (whether the structure is consistent), it is determined whether there is a deformation situation in the receipt to be processed. If there is a deformation situation, there may be recognition errors and other situations when using the OCR recognition technology to recognize the deformed text layer. Therefore, the text layer needs to be corrected first, and then the corrected text layer is text-recognized, so as to further ensure the accuracy of the recognized text content, and thus ensure the correctness of reconstructing and storing the receipt to be processed.
[0063] After ensuring the correctness of the recognition results, it is also necessary to perform permutations and combinations on the recognized text to ensure the correct order between words. For example, for "hui" and "dan", the correct order is "huidan", while if the randomly combined order is "dan hui", at this time, "dan hui" may refer to the power system structure of the transmission line, which may cause misjudgment. Therefore, it is necessary to correctly sort the recognized text and perform text semantic recognition on the sorted text to be verified, so as to determine whether the composed text content is correct and whether it conforms to the expression scenario of bank receipts. Thus, the accuracy of the sorted text structure is achieved.
[0064] At the same time, after ensuring the correctness of the uploaded and stored bank electronic receipts, the stored bank electronic receipts can be matched with the built-in rule library to determine the riskiness of each bank electronic receipt, reducing the workload of auditors and improving work efficiency.
[0065] In step S2, the color layer is screened and analyzed to obtain a text layer for text recognition and a structure layer for structure recognition, including the following steps:
[0066] Synthesis of the above two aspects, the following is a more detailed description of the method for processing bank receipts:
[0067] S21, determine the layer size of the receipt to be processed to obtain initial graphic data, and construct an initial graphic coordinate system based on the initial image data;
[0068] S22, compare the color layer with the initial graphic data to determine the pixel proportion of each color layer, and compare the pixel proportion of each color layer with the built-in proportion threshold, and retain the color layer with a pixel proportion less than the proportion threshold to obtain the first screening result;
[0069] S23, based on the initial graphic coordinate system, judge the continuity of the pixel points in each color layer of the first screening result, and record the continuous pixel points as pixel sets;
[0070] S24, judge the length of the pixel sets in each color layer of the first screening result. If the length of the pixel set is greater than the built-in length threshold, it is determined that the color layer corresponding to this position area is the structure layer;
[0071] S25, select the position coordinates of the color layer with a pixel set length less than the built-in length threshold, and compare the position coordinates of each pixel set with each other to determine whether the intervals between the pixel sets are the same;
[0072] In this embodiment, initial graphic data and initial graphic coordinates are constructed using the original layer size of the receipt to be processed. Then, the color difference layer containing text and structural layers is selected based on the proportion of text color layers. The first selection result is then further filtered using the initial graphic coordinate system. By comparing the length of the pixel set in each color layer, the structural layer is selected based on its characteristics. By comparing the position coordinates of the pixel set in each color layer, the text layer is selected based on the interval between the position coordinates. This makes the recognition results more accurate when performing text recognition and structural judgment on the receipt to be processed based on the selected results, thus improving the accuracy of bank electronic receipt recognition.
[0073] For example, when filtering and analyzing the color layer to select the text layer containing the receipt content and the structure layer containing the receipt format, since the separated color layer is obtained based on the receipt to be processed, the initial graphic data and the initial graphic coordinate system corresponding to the initial graphic data are constructed based on the receipt to be processed.
[0074] Since most pixels in a bank electronic receipt are blank white pixels, and the text content's pixel proportion is less than a certain value, a threshold for this proportion is determined experimentally and used as a standard to judge the color layer. Assume the color layers separated from the bank electronic receipt include the text layer of the receipt content, the structural layer of the receipt format, the stamp layer, the background layer, and the icon layer containing the bank's logo; for example, printing an article on an A4 sheet with default margins, a two-character first-line indent, and 1.5 line spacing results in an approximate pixel occupancy of about 15%.
[0075] The order of pixels, in terms of quantity, is likely: background layer > text layer > stamp layer > icon layer > structure layer. Therefore, using a percentage threshold to filter color layers first effectively eliminates the background layer. Then, based on the initial graphic coordinate system, coordinate analysis is performed on the remaining text, stamp, icon, and structure layers. Because the text layer has a unique structure—character spacing between words, and the spacing within a line of text is uniform—this unique structure can be used to determine the remaining color layers. First, the pixel continuity of each color layer is determined using the initial graphic coordinate system, thus dividing a color layer into multiple discontinuous pixel sets. Then, the positional coordinate intervals of these discontinuous pixel sets are judged to determine if the color layer satisfies the special characteristics of the text structure, thus identifying the text layer. Furthermore, during the judgment of the text layer, the inclusion of Chinese characters and Arabic numerals, along with the unique structural characteristics of Chinese characters (vertical and horizontal structures), enhances the accuracy of the text layer judgment.
[0076] Meanwhile, since the structural layer is composed of lines used to distinguish the structure, the length of these lines is greater than that of general text content. By judging the length of consecutive pixel sets, it is possible to effectively determine whether a structural layer exists in the remaining color layers, making the judgment result more accurate.
[0077] For the stamp and icon layers, no separate content recognition is required. Once the text content is confirmed to be correct, the stamp and icon layers can be added. The icon layer can also be directly image-recognized to determine the bank category.
[0078] In step S23, based on the initial graphic coordinate system, the continuity of pixels in each color layer of the first filtering result is judged, and continuous pixels are recorded as a pixel set, including the following steps:
[0079] S231, Based on the initial graphic coordinate system, the pixels of each color layer in the first screening result are marked with coordinates to obtain pixel coordinates;
[0080] S232, perform continuity judgment on pixel coordinates, record two pixels with adjacent coordinates as continuous, count the pixel coordinates with continuity, and obtain the pixel set;
[0081] S233, construct a rectangular region that matches the pixel set based on the edge pixels of the pixel set, and use the length information of the rectangular region as the length information of the pixel set;
[0082] S234, record the coordinates corresponding to the center position of the rectangular region as the position coordinates of the pixel set.
[0083] In this embodiment, the position coordinates of pixels in each color layer of the first screening result are marked using an initial graphic coordinate system, thus ensuring the uniformity of the position coordinates. By constructing rectangular regions for the pixel sets, the spatial information occupied by the corresponding pixel sets is determined, making the judgment results more accurate when distinguishing between the text layer and the structure layer. This provides a precise data foundation for subsequent bank electronic receipt recognition, resulting in more accurate recognition results.
[0084] For example, on electronic bank receipts, both text and lines have a length and width greater than a single pixel, thus each type of text and line has its own independent pixel set. Therefore, by performing continuous pixel statistics, the chaotic pixels are divided into individual pixel sets. When recognizing color layers, only individual pixel sets need to be identified, reducing the difficulty of pixel recognition.
[0085] For example, if color layer A is a structural layer, then there will be certain lines on this structural layer to isolate and divide the content on the bank electronic receipt. At this time, there are at least several consecutive pixels on the line, such as: pixel 1: (0,0), pixel 2: (0,1), pixel 3: (0,2), etc., thus forming a straight line.
[0086] Since pixels are planar, the width and height information of a single pixel is 1px and 1px respectively. However, if the width is 10 pixels and the height is 100 pixels, then the width and height information of the rectangular space are 10px and 100px respectively. In this case, the position coordinates of the rectangular space are (5px, 50px).
[0087] In step S4, based on the text in the first text content, a corresponding text image is generated, and the generated text image is compared with the corresponding text in the receipt to be processed to determine the structural consistency of the text image, including the following steps:
[0088] S41, based on the text in the first text content, mark the corresponding generated text image as a standard image;
[0089] S42, Based on the generated text image, select the image of the corresponding text in the receipt to be processed to obtain the image to be matched;
[0090] S43, perform proportional matching between the standard image and the image to be matched to determine whether the image to be matched can overlap with the standard image;
[0091] S44. If the image to be matched coincides with the standard image, the structure of the text image is determined to be consistent; otherwise, the structure of the text image is determined to be inconsistent.
[0092] In this embodiment, by using the first text content obtained from recognition, a standard image is generated for the recognized text, thereby determining the standard image of the corresponding text. The standard image is then used as a benchmark to compare with the real image of the text, thereby determining whether there is any distortion in the corresponding area. By performing distortion recognition on the image of the receipt to be recognized, the accuracy of the first text content obtained from the image of the receipt to be recognized is determined, thus realizing the text correctness verification of the first text content and improving the accuracy of text recognition of the receipt to be recognized.
[0093] Exemplarily, after determining the text layer and performing optical character recognition on the text layer, to ensure the correctness and comprehensiveness of the recognized text, by using the already recognized text, a standard image of the text is generated, and the standard image is compared with the original image in the receipt to be processed, so as to determine whether there are deformations (such as perspective, distortion, wrinkles, etc.) in the receipt to be recognized. If the generated standard image is the same as the image to be matched (only cases of equal-proportion magnification or reduction exist), it can be directly determined that the image to be matched has no deformation. Coupled with the fact that the matching results of all recognized texts are the same, it can be inferred that the receipt to be processed has no deformation, and further it can be indicated that the text recognition of the text layer of the receipt to be processed is accurate and there is no recognition error. This ensures the accuracy of the judgment result when subsequently reorganizing and semantically judging the recognized text.
[0094] For example, the first text content includes the character "口". Suppose the standard image of this character is a square image with a length and width of 10 px each. If the length and width of "口" in the receipt to be processed are both 15 px square images, it is determined that the "口" image in the receipt to be processed is an equal-proportion magnification of the standard "口" image. If the base length of "口" in the receipt to be processed is 15 px and 5 px, and the height is 5 px of an isosceles trapezoid, it is determined that the "口" image in the receipt to be processed is inconsistent with the standard "口" image, that is, it indicates that the receipt to be processed has a deformation. Therefore, the text layer of the receipt to be processed needs to be reprocessed to ensure the accuracy of recognition.
[0095] In step S5, if the structure between the generated text image and the receipt to be processed is inconsistent, the receipt to be processed is corrected according to the text image, and the corrected receipt to be processed is subjected to text recognition to obtain the second text content, including the following steps:
[0096] S51, if the standard image and the image to be matched do not overlap, based on the image to be matched, the standard image is deformed until the standard image overlaps with the image to be matched, and the image deformation process of the standard image is recorded to obtain the first deformation operation data;
[0097] S52, based on the purpose of each operation step in the first deformation operation data, the first deformation operation data is simplified to obtain the second deformation operation data;
[0098] S53, based on the second deformation operation data, the receipt to be processed is adjusted reversely to obtain the corrected receipt to be processed.
[0099] In this embodiment, when it is determined that the standard image and the image to be matched do not overlap, the standard image is deformed based on the image to be matched, which reduces the influence of interference factors and makes the image deformation more standardized and controllable. At the same time, the first deformation operation data is sorted out to remove redundant operations, simplify the deformation steps, reduce subsequent correction steps and improve correction efficiency.
[0100] For example, since the standard image and the image to be matched do not overlap, it indicates that the image to be matched is deformed. At the same time, since the image to be matched is a deformed image, it is not regular. Therefore, by performing regular deformation on the standard image, the amount of interference is reduced, making the change more accurate.
[0101] The standard image is deformed using built-in deformation categories, such as perspective and folding, until the deformed image is identical to the image to be matched. All deformation operations performed on the standard image then constitute the deformation operations for the image to be matched. Since the deformation of the standard image is based on the difference between the result of each deformation step and the image to be matched, redundant or invalid steps may exist between deformation steps. Therefore, it is necessary to organize the first deformation operation data to ensure the simplification of the second deformation operation data, reduce invalid processing data during image correction, and lower the system's resource requirements.
[0102] In step S6, if the generated text image has the same structure as the receipt to be processed, the built-in structure table is corrected according to the structure layer to obtain structure data. Then, the first text content or the second text content is classified and statistically analyzed according to the structure data to obtain the text to be verified, including the following steps:
[0103] S61, Match the built-in structure table based on the initial graphic data to construct a first structure table corresponding to the initial graphic data;
[0104] S62, based on the structure layer, determine the position coordinates of the lines in the corresponding structure layer, and match the first structure table according to the position coordinates of the lines to determine whether the position coordinates of the lines coincide with the cell border lines of the first structure table;
[0105] S63, if the position coordinates of the structural layer lines coincide with the cell border lines of the first structural table, then the first structural table is divided according to the structural layer lines to obtain the divided cell blocks, and the divided cell blocks are merged to obtain the second structural table.
[0106] S64. If the position coordinates of the structural layer lines do not coincide with the cell border lines of the first structural table, then the cells in the table are proportionally reduced according to the position coordinates of the structural layer lines until the position coordinates of the structural layer lines coincide with the cell border lines of the proportionally reduced first structural table. Then, the adjusted first structural table is divided and cells are merged according to the structural layer lines to obtain the second structural table.
[0107] S65, extract table features from the second structure table to obtain structure data.
[0108] In this embodiment, disordered text content is transformed into ordered text content by constructing a structured table. By matching the position coordinates of the lines in the structure layer with the first structured table, the matching degree between the first structured table and the bank electronic receipt to be processed in the real state is determined. By adjusting the size of the cells in the first structured table, the adjusted second structured table can overlap with the lines in the structure layer, thereby ensuring that the second structured content can accurately divide the recognized text content. This makes the combined result more accurate when combining the recognized text content into textual language, thus improving the accuracy of bank electronic receipt recognition.
[0109] For example, since bank electronic receipts are formatted according to an ordered table structure, the overall format of the bank electronic receipt matches that of a table. Therefore, a first structural table of the same size is constructed based on the bank electronic receipt to be processed. This first structural table is the same size as the bank electronic receipt, but it contains multiple proportionally sized cells.
[0110] The structure layer is constructed from multiple lines. Therefore, by matching the coordinates of the line with the first structure table, it can be determined whether the cell size of the first structure table matches the line of the structure layer. If the cell size matches the structure layer, it means that no cell adjustment is needed. At this time, it is only necessary to merge the multiple cells that are divided into a block to construct the overall cell distribution.
[0111] If the cell size does not match the structure layer, it indicates that the cell size needs to be adjusted. At this time, the size of the cells in the first structure table is adjusted as a whole to ensure that the adjusted cell size matches the lines in the structure layer, thereby constructing the second structure table. The second structure table then overlaps with the lines on the structure layer, and the text is distributed in the cells of the second structure table.
[0112] If the second structure table accurately divides the text content of each region, the semantics of the resulting text content will also be coherent.
[0113] In step S7, if the semantics of the text to be verified are not coherent, then based on the position information of the structural data and the text to be verified on the corresponding structure, a combination judgment is made on the text to be verified to determine whether the combination result has semantic coherence, and the combination results with semantic coherence are distributed and stored, including the following steps:
[0114] S71, If the semantics of the text to be verified are not coherent, then the text to be verified that is not semantically coherent is marked as the text to be analyzed, and the position coordinates of the text to be analyzed are marked to obtain the coordinate data to be analyzed.
[0115] S72, perform correlation matching between the text to be analyzed and the built-in return receipt keywords to determine the relevant keywords that can be matched for each character in the text to be analyzed, and record them as the keywords to be analyzed; where the keywords to be analyzed include related words and related words; for example, if the keyword is "recipient", and "receive" is a related word, then "payment" and "person" are related words. Here, related words refer to the words used for matching in the text to be scored.
[0116] S73, Based on structural data, obtain the coordinate data of the related words in the keyword to be analyzed, and select the surrounding text of the related words according to the coordinate data to be analyzed to obtain the matching words;
[0117] S74, match the word to be matched with the related words adjacent to the related words in the keyword to be analyzed. If the match is successful, select the adjacent text at the corresponding position of the word to be matched according to the positional relationship between the word to be matched and the related words, and match the selected text with the next related word in the keyword to be analyzed until the keyword to be analyzed is successfully matched.
[0118] S75, identify the successfully matched keywords to be analyzed, and segment the text to be analyzed based on the keywords to be analyzed and the corresponding coordinate data of the keywords to be analyzed, to obtain the segmented text;
[0119] S76 performs text combination and semantic recognition on the segmented text, and distributes the combination results with semantic coherence.
[0120] In this embodiment, the text to be analyzed is matched with built-in receipt keywords, thereby leveraging the unique characteristics of bank electronic receipts to further segment the text. Each character in the text is associated with a receipt keyword to determine the possible keywords corresponding to each character. Then, based on the positional coordinates between characters, adjacent characters are selected, and the selected results are matched with the corresponding associated characters of the relevant keywords to determine the successfully matched keywords for each character. The text is then spatially segmented using the positions of the successfully matched keywords, ensuring that each text statement is in a separate cell. This guarantees the accuracy of the segmented text statements and improves the accuracy of bank electronic receipt recognition.
[0121] For example, when there is a hidden text structure in a bank electronic receipt, the text in the divided cells will be messy, resulting in incoherent semantics in the combined text statements. Therefore, it is necessary to further divide the cells in this case to ensure that the text in the divided cells is ordered and correct.
[0122] For example, if a large cell contains "payee", "name", "contact information", "address", "business", etc., and the text is formatted directly without further division, the formatting result might be "payee name XXX, contact information XXXXX, business XXXX, address XXXXXX". In this case, when performing semantic judgment on the text to be validated, the result is that the semantics are not coherent.
[0123] In the actual division, the "payee" is in one cell, the "name" is in one cell, "XXX" is in one cell, the "address" is in one cell, and so on. Among them, "payee", "name", "contact information", etc. are all common keywords in bank electronic receipts. Therefore, keyword matching can be performed on each character in the text to be verified. For example, for the character "收", the matching keyword can be "payee" or "receiving account number". After determining the keywords that can be matched, select the adjacent text in the position area where the character "收" is located. For example, the adjacent characters are "款", "姓", "地", etc. Match "款", "姓", "地" with the "款" in the keyword "payee" to determine if the match is successful, and match "款", "姓", "地" with the "款" in the keyword "receiving account number" to determine if the match is successful. Since the "款" in the adjacent text matches the "款" in "payee" and "receiving account number", select the next character adjacent to "款" according to the position relationship between "收" and "款". For example, if the position relationship between "收" and "款" is from top to bottom, select the character below "款". If the character below "款" is "人", then match "人" with the "人" in the keyword "payee" and the "账" in the keyword "receiving account number". Since "人" matches the "人" in the keyword "payee", it is determined that the keyword corresponding to "收" is "payee", and the arrangement direction is from top to bottom.
[0124] At this time, if the arrangement directions of all the successfully matched keywords obtained by judgment are from top to bottom, then rearrange the text to be verified from top to bottom, and perform semantic recognition on the rearranged text.
[0125] If the arrangement order of the successfully matched keywords obtained by judgment has both left-to-right and top-to-bottom situations, then re-divide the text to be verified according to the position of the keyword, and obtain the text adjacent to the top-to-bottom arranged keyword and the text adjacent to the left-to-right arranged keyword. And the segmentation of the text to be verified conforms to the cell segmentation principle, that is, by splitting and merging cells, the segmentation of the text to be verified is realized, making the segmentation result more accurate.
[0126] Once all the text on the bank's electronic receipt can be accurately identified, the only remaining issue is how to format it, i.e., the combination of text. There are only two ways to combine text: one is to type it from left to right, and the other is to type it from top to bottom. When semantic inconsistencies occur when arranging all the text from left to right, it indicates that there may be two lines of text in one cell combined with text in an adjacent cell. For example, a cell might contain the message "There are only two combinations of text: one is typing from left to right, and the other is typing from top to bottom." The first line reads "There are only two combinations of text: one is typing from left to right," and the second line reads "typing, and the other is typing from top to bottom." An adjacent cell might contain the message "For example, a cell contains: ." The combined result of these two cells would be "There are only two combinations of text: one is typing from left to right, and the other is typing from top to bottom." In this case, the semantics are disjointed. However, due to the unique nature of bank electronic receipts, each form contains keywords such as "payee," "account number," and "amount." This allows for text segmentation based on these keywords. For example, "payee" would occupy a separate cell, separating the information preceding and following the payee's name, preventing consecutive content in different cells. This ensures more accurate recognition of the organized bank electronic receipt.
[0127] Compared with existing semantic analysis-based methods for recognizing electronic bank receipts, this invention improves the system's operating efficiency and ensures the accuracy of recognizing electronic bank receipts.
[0128] The above are all preferred embodiments of this application, and are not intended to limit the scope of protection of this application. Therefore, all equivalent changes made in accordance with the structure, shape and principle of this application should be covered within the scope of protection of this application.
Claims
1. A method for recognizing a bank electronic receipt based on semantic analysis, characterized in that, The method comprises the following steps: acquire a bank electronic receipt to be processed, and mark it as a to-be-processed receipt, perform color extraction and statistics on the to-be-processed receipt, and separate color layers corresponding to different colors; screen and analyze the color layers to obtain a text layer for text recognition and a structure layer for structure recognition; perform text recognition on the text layer based on an optical character recognition technology, determine the text content of the text layer, and mark it as first text content; generate a corresponding text image based on the text in the first text content, compare the generated text image with the corresponding text in the to-be-processed receipt, and determine the structural consistency of the text image; if the structure between the generated text image and the to-be-processed receipt is inconsistent, correct the to-be-processed receipt according to the text image, perform text recognition on the corrected to-be-processed receipt, and obtain second text content; if the structure between the generated text image and the to-be-processed receipt is consistent, modify an internal structure table according to the structure layer to obtain structure data, and classify and count the first text content or the second text content according to the structure data to obtain to-be-checked text; perform semantic recognition on the to-be-checked text to determine the coherence of the semantic of the to-be-checked text, if the semantic of the to-be-checked text does not have coherence, perform combination judgment on the to-be-checked text according to the position information of the structure data and the to-be-checked text on the corresponding structure, determine whether the combination result has semantic coherence, and perform distributed storage on the combination result having semantic coherence.
2. The method of claim 1, wherein the method further comprises: The method of screening and analyzing the color layers to obtain a text layer for text recognition and a structure layer for structure recognition comprises the following steps: perform layer size judgment on the to-be-processed receipt to obtain initial image data, and construct an initial image coordinate system based on the initial image data; compare the color layers with the initial image data, determine the pixel proportion of each color layer, compare the pixel proportion of each color layer with an internal proportion threshold, retain the color layer whose pixel proportion is less than the proportion threshold, and obtain a first screening result; based on the initial image coordinate system, perform continuity judgment on the pixel points in each color layer in the first screening result, and mark the continuous pixel points as a pixel set; perform length judgment on the pixel set of each color layer in the first screening result, if the length of the pixel set is greater than an internal length threshold, determine that the color layer corresponding to the position region is a structure layer; perform position coordinate selection on the color layer whose pixel set length is less than the internal length threshold, and compare the position coordinates of each pixel set with each other to determine whether the intervals between the pixel sets are the same; if the intervals between the position coordinates are the same, determine that the color layer corresponding to the position coordinates is a text layer.
3. The method of claim 2, wherein the method further comprises: The method of performing continuity judgment on the pixel points in each color layer in the first screening result based on the initial image coordinate system, and marking the continuous pixel points as a pixel set comprises the following steps: based on the initial image coordinate system, mark the pixels of each color layer in the first screening result with coordinates to obtain pixel coordinates; perform continuity judgment on the pixel coordinates, mark two pixels of adjacent coordinates as continuous, and count the pixel coordinates having continuity to obtain a pixel set; The edge pixels are constructed based on the pixel set, a rectangular region is constructed according to the pixel set, and length information of the rectangular region is taken as length information of the pixel set; A coordinate value corresponding to a center position of the rectangular region is taken as a position coordinate of the pixel set.
4. The method of claim 1, wherein the method further comprises: The structure consistency of the character image is determined by comparing the generated character image with the corresponding character in the to-be-processed return single based on the character in the first text content, including: The generated character image is marked as a standard image based on the character in the first text content; The image of the corresponding character in the to-be-processed return single is selected based on the generated character image, to obtain a to-be-matched image; The standard image is matched with the to-be-matched image in a same proportion, to determine whether the to-be-matched image can coincide with the standard image; If the to-be-matched image coincides with the standard image, it is determined that the structure of the character image has consistency, otherwise it is determined that the structure of the character image does not have consistency.
5. The method of claim 4, wherein the method further comprises: If the structure of the generated character image and the to-be-processed return single is inconsistent, the to-be-processed return single is corrected according to the character image, and the corrected to-be-processed return single is subjected to character recognition to obtain second text content, including: If the standard image does not coincide with the to-be-matched image, the standard image is deformed based on the to-be-matched image until the standard image coincides with the to-be-matched image, and a record is made of the image deformation process of the standard image to obtain first deformation operation data; The first deformation operation data is simplified based on a purpose of each operation step in the first deformation operation data to obtain second deformation operation data; The to-be-processed return single is adjusted in a reverse direction based on the second deformation operation data to obtain the corrected to-be-processed return single.
6. The method of claim 1, wherein the method further comprises: identifying the bank electronic statement based on the semantic analysis. 5 If the structure of the generated character image and the to-be-processed return single is consistent, the structure table built-in in the structure layer is modified to obtain structure data, and the first text content or the second text content is classified and counted according to the structure data to obtain to-be-verified text, including: The built-in structure table is matched based on the initial graphic data to construct a first structure table corresponding to the initial graphic data; The position coordinates of the lines in the structure layer are determined based on the structure layer, and the position coordinates of the lines are matched with the cell border lines of the first structure table to determine whether the position coordinates of the lines coincide with the cell border lines of the first structure table; If the position coordinates of the lines in the structure layer coincide with the cell border lines of the first structure table, the first structure table is divided according to the lines of the structure layer to obtain divided cell blocks, and the divided cell blocks are merged to obtain a second structure table; If the position coordinates of the lines in the structure layer do not coincide with the cell border lines of the first structure table, the cells in the table are reduced in a same proportion according to the position coordinates of the lines in the structure layer until the position coordinates of the lines in the structure layer coincide with the border lines of the cells in the first structure table after the reduction, and the adjusted first structure table is divided and the cells are merged according to the lines of the structure layer to obtain the second structure table; The structure data is obtained by extracting table features from the second structure table.
7. The method of claim 1, wherein the method further comprises: identifying the bank electronic statement based on the semantic analysis. If the semantics of the to-be-verified text do not have coherence, then the to-be-verified text is combined and judged according to the position information of the structure data and the to-be-verified text on the corresponding structure, whether the combination result has semantic coherence is determined, and the combination result with semantic coherence is stored in a distributed manner, including: If the semantics of the to-be-verified text do not have coherence, then the to-be-verified text without semantic coherence is marked as to-be-analyzed text, the position coordinates of the characters in the to-be-analyzed text are marked to obtain to-be-analyzed coordinate data; The to-be-analyzed text is associated with the built-in reply single keyword for matching to determine the keyword that can be matched by each character in the to-be-analyzed text, and the keyword is recorded as to-be-analyzed keyword; wherein the to-be-analyzed keyword includes associated words and associated words; Based on the structure data, to-be-analyzed coordinate data of the associated words in the to-be-analyzed keyword is obtained, and the surrounding characters of the associated words are selected according to the to-be-analyzed coordinate data to obtain to-be-matched characters; The to-be-matched characters are matched with the associated words adjacent to the associated words in the to-be-analyzed keyword, if the matching is successful, then the to-be-matched characters are selected according to the position relationship between the to-be-matched characters and the associated words, and the selected characters are matched with the next associated words in the to-be-analyzed keyword until the to-be-analyzed keyword is matched successfully; The to-be-analyzed keyword that matches successfully is determined, and the to-be-analyzed text is segmented according to the to-be-analyzed keyword and the to-be-analyzed coordinate data of the corresponding keyword to obtain segmented text; The segmented text is combined and the semantics is recognized, and the combination result with semantic coherence is stored in a distributed manner.