Bill dislocation character recognition method and device, equipment and storage medium
By cutting the bill images and judging the easily misaligned area, combined with end-to-end image understanding technology and OCR technology, the problem of low accuracy of recognition of financial bills is solved, and higher recognition accuracy and lower error rate are achieved.
Patent Information
- Application Number
- CN202510094710.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-16
AI Technical Summary
In the prior art, the accuracy of identifying misaligned content of financial bills is low, resulting in errors in business processes.
By acquiring the ticket image and performing area cutting according to the field relative area parameter configuration file, it is determined whether the field slice image belongs to the easily misaligned area. If it belongs to a dislocation-prone area, the processing is performed using end-to-end image understanding technology and OCR technology respectively, the first prediction result and the second prediction result are obtained, and the similarity comparison is performed to determine the recognition result.
It effectively improves the accuracy of identifying misaligned content in financial notes and reduces errors in business processes.
Smart Images

Figure CN120014651A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing technology, and in particular to a method, device, equipment and storage medium for recognizing misplaced characters on bills. Background Art
[0002] Traditional optical character recognition (OCR) technology belongs to single visual modality processing technology (single modality technology), which includes a text detection module and a text recognition module. The text detection module is used to locate the text area, which is output in units of lines of text. The output coordinates are quadrilateral boxes, and the coordinates are expressed in horizontal and vertical coordinate systems, with the upper left corner of the entire image as the origin and the lower right corner as (image width, image height). Common quadrilateral boxes use four vertices for output. The text recognition module is used to identify text in the text area, which is output in units of characters.
[0003] For standard bill recognition, OCR technology can output the content of each line of text in the bill image, and extract element information according to natural language processing technology (such as common regular expression methods), and output it in a standard key-value pair structured manner. However, in financial bills (such as checks, bills of exchange, settlement business authorization letters, and deposit slips), there will be a situation where the basic template and content are printed twice, and it is also easy to cause the account text content to be misaligned or overlapped with the grid. Although using OCR technology for recognition can improve the overall input efficiency, it will cause errors in the business process due to the omission of account character content due to over-grid.
[0004] Therefore, there is an urgent need for a method for recognizing misplaced characters on bills that can effectively improve the accuracy of recognizing misplaced content on financial bills. Summary of the invention
[0005] The main purpose of the present invention is to provide a method, device, equipment and storage medium for identifying misplaced characters on bills, aiming to solve the technical problem of low accuracy in identifying misplaced content on financial bills in the prior art.
[0006] To achieve the above object, the present invention provides a method for identifying misplaced characters on bills, the method comprising the following steps:
[0007] Acquire a bill image, and perform region cutting on the bill image according to a field relative region parameter configuration file to obtain a plurality of field slice images;
[0008] Based on the field error-prone status configuration file, determining whether the field slice graph belongs to an error-prone area;
[0009] If the field slice image belongs to an easily misplaced area, the field slice image is processed by end-to-end image understanding technology and OCR technology respectively to obtain a first prediction result and a second prediction result;
[0010] A similarity comparison is performed on the first prediction result and the second prediction result, and a recognition result is determined based on the similarity comparison result.
[0011] Optionally, after the step of judging whether the field slice diagram belongs to a misalignment-prone area based on the field error-prone status configuration file, the step further includes:
[0012] If the field slice image does not belong to the easily misplaced area, extracting the text position information and text content information of the field slice image by using OCR technology;
[0013] Regular expressions are used to extract the text position information and the text content information to obtain a recognition result.
[0014] Optionally, the step of acquiring a bill image and performing region cutting on the bill image according to a field relative region parameter configuration file to obtain a plurality of field slice images includes:
[0015] Use high-speed scanner equipment to image the bill and obtain the bill image;
[0016] Determine the relative area coordinate information corresponding to each field in the bill according to the field relative area parameter configuration file;
[0017] The bill image is region-cut based on relative region coordinate information corresponding to each field in the bill to obtain a plurality of field slice images.
[0018] Optionally, if the field slice image belongs to an easily misplaced area, the step of processing the field slice image by end-to-end image understanding technology and OCR technology respectively to obtain a first prediction result and a second prediction result includes:
[0019] If the field slice diagram belongs to an easily misplaced area, determining the field corresponding to the field slice diagram, and using the field name corresponding to the field slice diagram as a prompt word;
[0020] Based on the field slice graph and the prompt word, using end-to-end image understanding technology, obtaining a first prediction result corresponding to the field slice graph;
[0021] Based on the field slice diagram, using OCR technology, obtaining text information of the field slice diagram;
[0022] The text information is extracted using a regular expression to obtain a second prediction result corresponding to the field slice diagram.
[0023] Optionally, the step of performing a similarity comparison on the first prediction result and the second prediction result, and determining the recognition result based on the similarity comparison result, includes:
[0024] Performing a similarity comparison between the first prediction result and the second prediction result according to a character dimension to obtain a similarity comparison result;
[0025] If the similarity comparison result indicates that the similarity between the first prediction result and the second prediction result is not 1, matching the first prediction result and the second prediction result with a corpus to obtain a matching result;
[0026] When the matching result indicates that the matching degree between the first prediction result and the second prediction result and the corpus in the corpus is 1, determining the similarity between the first prediction result and the second prediction result;
[0027] The similarity and the first prediction result or the second prediction result are used as recognition results.
[0028] Optionally, after the step of matching the first prediction result and the second prediction result with a corpus to obtain a matching result, the method further includes:
[0029] When the matching result indicates that the matching degree between the first prediction result and the corpus in the corpus is 1 and the matching degree between the second prediction result and the corpus in the corpus is not 1, determining the similarity between the first prediction result and the second prediction result, and using the similarity and the first prediction result as the recognition result;
[0030] When the matching result indicates that the matching degree between the second prediction result and the corpus in the corpus is 1 and the matching degree between the first prediction result and the corpus in the corpus is not 1, determining the similarity between the first prediction result and the second prediction result, and using the similarity and the second prediction result as the recognition result;
[0031] When the matching result indicates that the matching degree between the first prediction result and the second prediction result and the corpus in the corpus is not 1, the similarity between the first prediction result and the second prediction result is determined, and the similarity and a null value are used as recognition results.
[0032] Optionally, after the step of performing a similarity comparison on the first prediction result and the second prediction result and determining the recognition result based on the similarity comparison result, the method further includes:
[0033] Review and correct the recognition result based on the recognition result to obtain a processed recognition result;
[0034] The processed recognition results are fed back to the corpus to obtain an updated corpus.
[0035] In addition, to achieve the above-mentioned purpose, the present invention also proposes a bill misplaced character recognition device, the device comprising:
[0036] A region cutting module is used to obtain a bill image and perform region cutting on the bill image according to a field relative region parameter configuration file to obtain a plurality of field slice images;
[0037] A region judgment module, used to judge whether the field slice graph belongs to a misalignment region based on a field error-prone status configuration file;
[0038] A prediction output module, configured to process the field slice image by end-to-end image understanding technology and OCR technology respectively to obtain a first prediction result and a second prediction result if the field slice image belongs to an easily misplaced area;
[0039] The result output module is used to perform a similarity comparison between the first prediction result and the second prediction result, and determine a recognition result based on the similarity comparison result.
[0040] In addition, to achieve the above-mentioned purpose, the present invention also proposes a bill misplaced character recognition device, which includes: a memory, a processor, and a bill misplaced character recognition program stored in the memory and executable on the processor, wherein the bill misplaced character recognition program is configured to implement the steps of the bill misplaced character recognition method as described above.
[0041] In addition, to achieve the above-mentioned purpose, the present invention also proposes a storage medium, on which a bill misplaced character recognition program is stored, and when the bill misplaced character recognition program is executed by a processor, the steps of the bill misplaced character recognition method as described above are implemented.
[0042] The present invention discloses the method of obtaining a bill image, and performing regional cutting on the bill image according to a field relative region parameter configuration file to obtain a plurality of field slice graphs; judging whether the field slice graph belongs to a misplaced region based on a field error-prone state configuration file; if the field slice graph belongs to a misplaced region, processing the field slice graph by end-to-end image understanding technology and OCR technology respectively to obtain a first prediction result and a second prediction result; performing a similarity comparison on the first prediction result and the second prediction result, and determining a recognition result based on the similarity comparison result. Since the present invention performs regional cutting on the bill image according to the field relative region parameter configuration file, and obtains a first prediction result and a second prediction result based on the end-to-end image understanding technology and OCR technology respectively when the field slice graph belongs to a misplaced region, and then determines a recognition result based on the similarity comparison result between the first prediction result and the second prediction result, compared with the prior art, the present invention effectively improves the accuracy of identifying misplaced content of financial bills. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 It is a flow chart of the first embodiment of the method for recognizing misplaced characters in bills of the present invention;
[0044] Figure 2 It is a flow chart of the second embodiment of the method for recognizing misplaced characters in bills of the present invention;
[0045] Figure 3 It is a schematic diagram of the overall process of the method for recognizing misplaced characters in bills of the present invention;
[0046] Figure 4 A schematic flow chart of a third embodiment of a method for recognizing misplaced characters on bills according to the present invention;
[0047] Figure 5 It is a similarity comparison and result reflux diagram in the bill misplaced character recognition method of the present invention;
[0048] Figure 6 It is a structural block diagram of the first embodiment of the bill misaligned character recognition device of the present invention;
[0049] Figure 7 It is a structural diagram of a bill misaligned character recognition device in a hardware operating environment involved in an embodiment of the present invention.
[0050] The realization of the purpose, functional features and advantages of the present invention will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0051] It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention.
[0052] The embodiment of the present invention provides a method for recognizing misplaced characters in bills, referring to Figure 1 , Figure 1 It is a flow chart of the first embodiment of the method for recognizing misplaced characters in bills of the present invention.
[0053] In this embodiment, the bill misplaced character recognition method includes steps S10 to S40:
[0054] Step S10: Acquire a bill image, and perform region cutting on the bill image according to a field relative region parameter configuration file to obtain a plurality of field slice images.
[0055] It should be noted that the execution subject of this embodiment can be a computer server device with data processing, network communication and program running functions used in the recognition scenario of misplaced content of bills, such as a server, a tablet computer, a smart phone, a smart watch, etc., or an electronic device capable of realizing the above functions, a bill misplaced character recognition device, etc. The following takes the bill misplaced character recognition device as an example to illustrate this embodiment and the following embodiments.
[0056] It should be understood that a bill image generally refers to an image or photo of a bill such as an invoice or receipt, and is a type of document image.
[0057] It should be explained that the field relative area parameter configuration file may be a file used to record the relative area coordinate information corresponding to each field in the bill image.
[0058] In a specific implementation, a high-speed scanner can be used to image the bill to obtain a bill image; the relative area coordinate information corresponding to each field in the bill is determined according to a field relative area parameter configuration file; the bill image is region-cut based on the relative area coordinate information corresponding to each field in the bill to obtain multiple field slices.
[0059] For example, a high-speed scanner is used to image the bill and collect the bill image; then the field relative area parameter configuration file is read, which records the relative area coordinate information of each field. The relative area coordinates can be appropriately expanded by 2 times the area of the original field area. The coordinate description method can use the horizontal and vertical coordinates of the upper left corner of the whole image as (0, 0), and the horizontal and vertical coordinates of the lower right corner as (the maximum width pixel value of the image, the maximum length pixel value of the image) as the reference system; finally, the bill image is cut into different field areas according to the field relative area parameter configuration file to obtain a fixed number of field slices.
[0060] It should be understood that the relative area slicing technology is used to perform regional cutting of the bill image according to the field relative area parameter configuration file. This technology does not require precise cutting because the character misalignment itself means floating in a certain area. The overall cutting setting uses relative areas, which has the advantages of being simple, fast in cutting, and does not require model training. If image edge detection or a special cutting model is used, it will take a long time and the content across grids will be cut off.
[0061] In addition, compared with whole-image recognition, regional cutting of bill images before recognition requires less training data for OCR technology and end-to-end image understanding technology, and the time spent on training and reasoning will also be reduced. The high-speed scanner scene has the advantage of fixed image angles, and the processing speed of the relative regional cutting solution set in advance is faster. If the entire image is subjected to text detection or cell (wired table) detection algorithms and then cut, the computing resources occupied and the response time will increase.
[0062] Step S20: Based on the field error-prone status configuration file, determine whether the field slice diagram belongs to an error-prone area.
[0063] It should be noted that the above-mentioned field error-prone status configuration file may be a file used to record whether each field in the bill image belongs to an error-prone area or a normal area.
[0064] It should be understood that the field error-prone status configuration file is a dynamic configuration library. It is not only for account fields, but can also be flexibly configured in real time according to the user's misalignment of the scene. It has the advantages of personalized change expansion and real-time execution.
[0065] Step S30: If the field slice image belongs to an easily misplaced area, the field slice image is processed by end-to-end image understanding technology and OCR technology respectively to obtain a first prediction result and a second prediction result.
[0066] It should be noted that end-to-end image understanding technology is a type of multimodal model, involving the joint processing of visual modality and text modality. At present, end-to-end image understanding technology (also called multimodal technology) integrates text detection and text recognition. Only one model can be used to complete the two modalities of image (visual modality data) and instruction text (text modality data) as input parameters and output recognition results. The advantage of this method is that it can integrate visual features and text features for comprehensive output, but the disadvantage is that the output text does not appear in the image, but the output text is semantically consistent. The model will predict based on the continuity of the text and understand that the account number is a string of numbers with a fixed number of digits. It will automatically complete the last few digits, that is, the image and text are inconsistent. This phenomenon is called hallucination.
[0067] In addition, OCR technology belongs to a single visual modality processing technology (hereinafter referred to as single modality technology), which can output the content of each line of text in the bill image, and extract element information according to natural language processing technology (such as the common regular expression method), and output it in a standard key-value pair structured manner. However, in financial bills (such as checks, bills of exchange, settlement business authorization letters, and deposit slips), there will be a situation where the basic template and content are printed twice, and it is also easy to cause the account text content to be misaligned or overlapped with the grid. Although the use of OCR technology for recognition can improve the overall input efficiency, it will cause errors in the business process due to the omission of account character content due to over-grid.
[0068] In a specific implementation, if the field slice diagram belongs to an easily misplaced area, the field corresponding to the field slice diagram is determined, and the field name corresponding to the field slice diagram is used as a prompt word; based on the field slice diagram and the prompt word, end-to-end image understanding technology is used to obtain a first prediction result corresponding to the field slice diagram; based on the field slice diagram, OCR technology is used to obtain text information of the field slice diagram; regular expressions are used to extract information from the text information to obtain a second prediction result corresponding to the field slice diagram.
[0069] Step S40: performing a similarity comparison between the first prediction result and the second prediction result, and determining a recognition result based on the similarity comparison result.
[0070] It should be noted that if the field slice diagram belongs to an area that is prone to misalignment, the first prediction result and the second prediction result are obtained based on the end-to-end image understanding technology and the OCR technology respectively, and then the first prediction result and the second prediction result are compared for similarity, and the recognition result is determined based on the similarity comparison result. Compared with the prior art, the present embodiment adopts a combination of OCR technology and end-to-end image understanding technology, which can retain the accuracy of the original OCR technology while performing separate branch processing on error-prone field samples. This parallel method can not only improve the overall accuracy, but also does not increase the processing time.
[0071] It should be explained that the recognition result output by this embodiment uses similarity to replace the original confidence of a single model (the model corresponding to OCR technology or end-to-end image understanding technology). Because the confidence is the result probability value output by a single model, this probability value is strongly correlated with the training corpus. The prediction results of multiple models are used to calculate the similarity, which is more of a reference as the overall confidence.
[0072] The present embodiment discloses obtaining a bill image, and performing regional segmentation on the bill image according to a field relative region parameter configuration file to obtain a plurality of field slice graphs; judging whether the field slice graph belongs to a misplaced region based on a field error-prone state configuration file; if the field slice graph belongs to a misplaced region, processing the field slice graph by end-to-end image understanding technology and OCR technology respectively to obtain a first prediction result and a second prediction result; performing a similarity comparison on the first prediction result and the second prediction result, and determining a recognition result based on the similarity comparison result. Since the present embodiment performs regional segmentation on the bill image according to the field relative region parameter configuration file, and obtains a first prediction result and a second prediction result based on the end-to-end image understanding technology and OCR technology respectively when the field slice graph belongs to a misplaced region, and then determines a recognition result based on the similarity comparison result of the first prediction result and the second prediction result, compared with the prior art, the present embodiment effectively improves the accuracy of identifying misplaced content of financial bills.
[0073] refer to Figure 2 , Figure 2 It is a flow chart of the second embodiment of the method for recognizing misplaced characters in bills of the present invention.
[0074] Based on the first embodiment, in this embodiment, after step S20, steps S30' to S40' are further included:
[0075] Step S30 ′: if the field slice image does not belong to the easily misaligned area, extract the text position information and text content information of the field slice image by using OCR technology.
[0076] Step S40 ′: extracting the text position information and the text content information using regular expressions to obtain a recognition result.
[0077] It should be noted that if the field slice diagram does not belong to an easily misplaced area, it means that the field slice diagram belongs to a normal area. Therefore, the text position information and text content information of the field slice diagram are directly extracted through OCR technology, and then the text position information and the text content information are extracted using regular expressions to obtain recognition results, which not only ensures the recognition accuracy, but also improves the recognition efficiency.
[0078] For example, refer to Figure 3 , Figure 3The figure is a schematic diagram of the overall process of the method for recognizing misplaced characters in bills of the present invention. As can be seen from the figure, a paper bill is photographed by a high-speed scanner to obtain a bill image; then, a relative area cut is performed according to the field relative area parameter configuration file to obtain slices of each field (i.e., field slice map); then, based on the field error-prone state library (i.e., field error-prone state configuration file), it is determined whether the slice is an easy-to-misplace area; if it is not an easy-to-misplace area, element information is extracted through OCR technology and regular expressions to obtain a recognition result (field output); if it is an easy-to-misplace area, a combination of end-to-end image understanding technology and OCR technology is used (at this time, the input parameter of the end-to-end image understanding technology is the field slice map and prompt words such as "account number", and the output parameter is the prediction result 1 of the model, i.e., the first prediction result, such as "571332209"; the input parameter of the OCR technology is the field slice map, and the output parameter is all the text information in the field slice map, such as "account number or address\t5171332209\tA / C No.orAddress", this string of characters will be used as the input parameter of the regular expression, and the output parameter is prediction result 2, that is, the second prediction result, such as: "5171332209".); then the prediction result 1 and the prediction result 2 are compared for similarity, and the recognition result is determined based on the similarity comparison result.
[0079] This embodiment discloses obtaining a bill image, and performing regional cutting on the bill image according to a field relative area parameter configuration file to obtain multiple field slice diagrams; judging whether the field slice diagram belongs to a misplaced area based on a field error-prone status configuration file; if the field slice diagram does not belong to a misplaced area, extracting the text position information and text content information of the field slice diagram through OCR technology; and extracting information from the text position information and the text content information using regular expressions to obtain recognition results. Since this embodiment performs regional cutting on the bill image according to the field relative area parameter configuration file, and directly extracts element information through OCR technology and regular expressions when the field slice diagram does not belong to a misplaced area to obtain recognition results, compared to the prior art, this embodiment not only ensures the accuracy of recognition, but also improves the efficiency of recognition.
[0080] refer to Figure 4 , Figure 4 It is a flow chart of the third embodiment of the method for recognizing misplaced characters in bills of the present invention.
[0081] Based on the above embodiments, in this embodiment, step S40 includes steps S401 to S404:
[0082] Step S401: performing a similarity comparison between the first prediction result and the second prediction result according to the character dimension to obtain a similarity comparison result.
[0083] Step S402: If the similarity comparison result indicates that the similarity between the first prediction result and the second prediction result is not 1, the first prediction result and the second prediction result are matched with a corpus to obtain a matching result.
[0084] Step S403: when the matching result indicates that the matching degree between the first prediction result and the second prediction result and the corpus in the corpus is 1, determine the similarity between the first prediction result and the second prediction result.
[0085] Step S404: taking the similarity and the first prediction result or the second prediction result as a recognition result.
[0086] It should be noted that if the similarity comparison result indicates that the similarity between the first prediction result and the second prediction result is 1, the similarity and the first prediction result can be used as the recognition result.
[0087] It should be explained that when the matching result indicates that the matching degree between the first prediction result and the corpus in the corpus is 1 and the matching degree between the second prediction result and the corpus in the corpus is not 1, the similarity between the first prediction result and the second prediction result is determined, and the similarity and the first prediction result are used as recognition results.
[0088] When the matching result indicates that the matching degree between the second prediction result and the corpus in the corpus is 1 and the matching degree between the first prediction result and the corpus in the corpus is not 1, the similarity between the first prediction result and the second prediction result is determined, and the similarity and the second prediction result are used as recognition results.
[0089] When the matching result indicates that the matching degree between the first prediction result and the second prediction result and the corpus in the corpus is not 1, the similarity between the first prediction result and the second prediction result is determined, and the similarity and a null value are used as recognition results.
[0090] In a specific implementation, in order to ensure the reliability of the recognition result, after the step of performing a similarity comparison between the first prediction result and the second prediction result and determining the recognition result based on the similarity comparison result, it also includes: reviewing and correcting based on the recognition result to obtain a processed recognition result; and returning the processed recognition result to the corpus to obtain an updated corpus.
[0091] It needs to be explained that the corpus is a dynamic corpus, to which new corpus texts can be added at any time. Through two methods, manual injection and process closed-loop reflux, it can continuously accumulate to form an expert knowledge base, making the output results of the entire process more reliable.
[0092] For example, refer to Figure 5 , Figure 5 This is a similarity comparison and result reflux diagram in the bill misaligned character recognition method of the present invention. In the figure, the prediction result 1 (i.e., the first prediction result) and the prediction result 2 (i.e., the second prediction result) are compared for similarity. If the prediction result 1 and the prediction result 2 are absolutely matched according to the character dimension and the similarity reaches 100% (i.e., the similarity is 1), the prediction result 1 and the similarity 100% are output, and the process ends; if it is not 100% (i.e., the similarity is not 1), the corpus is read. The data sources of the corpus are divided into two, one is manual import, and the other is based on the existing process, and the model output results are manually reviewed and automatically refluxed to the corpus. Then, prediction result 1 and prediction result 2 are matched with the corpus. If each of them matches one of the corpuses 100% (i.e., the matching degree is 1), the similarity between prediction result 1 and prediction result 2 is calculated, and prediction result 1 or prediction result 2 is randomly output, as well as the similarity, and this process ends. If prediction result 1 matches 100% (i.e., the matching degree between the first prediction result and the corpus is 1 and the matching degree between the second prediction result and the corpus is not 1), the similarity between prediction result 1 and prediction result 2 is calculated, and prediction result 1 and the similarity are output, and this process ends. If prediction result 2 matches 1 00% (i.e., the matching degree between the second prediction result and the corpus is 1 and the matching degree between the first prediction result and the corpus is not 1), the similarity between prediction result 1 and prediction result 2 is calculated, prediction result 2 and the similarity are output, and this process ends; if both prediction results 1 and 2 are not in the corpus (i.e., the matching degree between the first prediction result and the second prediction result and the corpus is not 1), the similarity between prediction result 1 and prediction result 2 is calculated, and a null value and the similarity are output. In addition, in the figure, the corpus can be continuously and automatically expanded through the data reflux of the output results, so that the reliability of the output results will become higher and higher.
[0093] The present embodiment discloses obtaining a bill image, and performing region cutting on the bill image according to a field relative region parameter configuration file to obtain a plurality of field slice graphs; judging whether the field slice graph belongs to a misplacement-prone region based on a field error-prone state configuration file; if the field slice graph belongs to a misplacement-prone region, processing the field slice graph by end-to-end image understanding technology and OCR technology respectively to obtain a first prediction result and a second prediction result; performing a similarity comparison on the first prediction result and the second prediction result according to a character dimension to obtain a similarity comparison result; if the similarity comparison result indicates that the similarity between the first prediction result and the second prediction result is not 1, The first prediction result and the second prediction result are matched with the corpus to obtain a matching result; when the matching result indicates that the matching degree of the first prediction result and the second prediction result with the corpus is 1, the similarity between the first prediction result and the second prediction result is determined; the similarity and the first prediction result or the second prediction result are used as the recognition result. Compared with the prior art, this embodiment uses the prediction results of the end image understanding technology and the OCR technology to calculate the similarity. When the similarity between the first prediction result and the second prediction result is not 1, the first prediction result and the second prediction result are matched with the corpus, which effectively ensures the accuracy and reliability of the recognition result.
[0094] In addition, an embodiment of the present invention further proposes a storage medium, on which a bill misplaced character recognition program is stored. When the bill misplaced character recognition program is executed by a processor, the steps of the bill misplaced character recognition method described above are implemented.
[0095] Reference Figure 6 , Figure 6 This is a structural block diagram of the first embodiment of the bill misaligned character recognition device of the present invention.
[0096] like Figure 6 As shown, the bill misaligned character recognition device proposed in the embodiment of the present invention includes: a region cutting module 601, a region judging module 602, a prediction output module 603 and a result output module 604.
[0097] The region cutting module 601 is used to obtain a bill image, and perform region cutting on the bill image according to a field relative region parameter configuration file to obtain a plurality of field slice images.
[0098] The region determination module 602 is used to determine whether the field slice diagram belongs to a prone-to-misalignment region based on a field prone-to-misalignment status configuration file.
[0099] The prediction output module 603 is used to process the field slice image by end-to-end image understanding technology and OCR technology respectively to obtain a first prediction result and a second prediction result if the field slice image belongs to an easily misaligned area.
[0100] The result output module 604 is used to perform a similarity comparison between the first prediction result and the second prediction result, and determine a recognition result based on the similarity comparison result.
[0101] The area cutting module 601 is also used to image the bill using a high-definition scanner to obtain a bill image; determine the relative area coordinate information corresponding to each field in the bill according to the field relative area parameter configuration file; and perform area cutting on the bill image based on the relative area coordinate information corresponding to each field in the bill to obtain multiple field slice diagrams.
[0102] The prediction output module 603 is also used to determine the field corresponding to the field slice diagram if the field slice diagram belongs to an easily misplaced area, and use the field name corresponding to the field slice diagram as a prompt word; based on the field slice diagram and the prompt word, using end-to-end image understanding technology, obtain a first prediction result corresponding to the field slice diagram; based on the field slice diagram, using OCR technology, obtain text information of the field slice diagram; use regular expressions to extract information from the text information to obtain a second prediction result corresponding to the field slice diagram.
[0103] The present device embodiment discloses obtaining a bill image, and performing regional segmentation on the bill image according to a field relative region parameter configuration file to obtain a plurality of field slice graphs; judging whether the field slice graph belongs to a misplaced region based on a field error-prone state configuration file; if the field slice graph belongs to a misplaced region, processing the field slice graph through end-to-end image understanding technology and OCR technology respectively to obtain a first prediction result and a second prediction result; performing a similarity comparison on the first prediction result and the second prediction result, and determining a recognition result based on the similarity comparison result. Since the present device embodiment performs regional segmentation on the bill image according to the field relative region parameter configuration file, and obtains a first prediction result and a second prediction result based on the end-to-end image understanding technology and OCR technology respectively when the field slice graph belongs to a misplaced region, and then determines a recognition result based on the similarity comparison result of the first prediction result and the second prediction result, compared with the prior art, the present device embodiment effectively improves the accuracy of identifying misplaced content of financial bills.
[0104] Based on the first embodiment of the apparatus for recognizing misplaced characters on bills of the present invention, a second embodiment of the apparatus for recognizing misplaced characters on bills of the present invention is proposed.
[0105] In this embodiment, the result output module 604 is also used to extract the text position information and text content information of the field slice image through OCR technology if the field slice image does not belong to the easy-to-misalignment area; use regular expressions to extract the text position information and the text content information to obtain the recognition result.
[0106] Other embodiments or specific implementations of the bill misaligned character recognition device of the present invention can refer to the above-mentioned method embodiments, which will not be described in detail here.
[0107] The present application provides a device for recognizing misplaced characters in bills, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method for recognizing misplaced characters in bills in the above-mentioned embodiment 1.
[0108] Reference below Figure 7 , which shows a schematic diagram of the structure of a bill misalignment character recognition device suitable for implementing the embodiment of the present application. The bill misalignment character recognition device in the embodiment of the present application may include but is not limited to mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 7 The bill misaligned character recognition device shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.
[0109] like Figure 7As shown, the bill misalignment character recognition device may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM: Read Only Memory) 1002 or a program loaded from a storage device 1003 to a random access memory (RAM: Random Access Memory) 1004. Various programs and data required for the operation of the bill misalignment character recognition device are also stored in the RAM 1004. The processing device 1001, the ROM 1002, and the RAM 1004 are connected to each other via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to the I / O interface 1006: input devices 1007 including, for example, a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; storage devices 1003 including, for example, a magnetic tape, a hard disk, etc.; and communication devices 1009. The communication device 1009 can allow the bill misalignment character recognition device to communicate wirelessly or wired with other devices to exchange data. Although the bill misalignment character recognition device with various systems is shown in the figure, it should be understood that it is not required to implement or have all the systems shown. More or fewer systems may be implemented or have alternatively.
[0110] In particular, according to the embodiments disclosed in the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device, or installed from a storage device 1003, or installed from a ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are executed.
[0111] The bill misplaced character recognition device provided by the present application adopts the bill misplaced character recognition method in the above embodiment, which can solve the technical problem of low accuracy of misplaced content recognition of financial bills in the prior art. Compared with the prior art, the beneficial effects of the bill misplaced character recognition device provided by the present application are the same as the beneficial effects of the bill misplaced character recognition method provided by the above embodiment, and other technical features of the bill misplaced character recognition device are the same as the features disclosed in the method of the previous embodiment, which will not be repeated here.
[0112] It should be understood that the various parts disclosed in this application can be implemented by hardware, software, firmware or a combination thereof. In the description of the above embodiments, specific features, structures, materials or characteristics can be combined in any one or more embodiments or examples in a suitable manner.
[0113] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
[0114] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or system including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or system. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or system including the element.
[0115] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.
[0116] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as a read-only memory / random access memory, a magnetic disk, or an optical disk), and includes a number of instructions for a terminal device (which can be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods described in each embodiment of the present invention.
[0117] The above are only preferred embodiments of the present invention, and are not intended to limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.
Claims
1. A method for recognizing misplaced characters on bills, characterized in that: The method comprises: Acquire a bill image, and perform region cutting on the bill image according to a field relative region parameter configuration file to obtain a plurality of field slice images; Based on the field error-prone status configuration file, determining whether the field slice graph belongs to an error-prone area; If the field slice image belongs to an easily misplaced area, the field slice image is processed by end-to-end image understanding technology and OCR technology respectively to obtain a first prediction result and a second prediction result; A similarity comparison is performed on the first prediction result and the second prediction result, and a recognition result is determined based on the similarity comparison result.
2. The method for recognizing misplaced characters on bills as claimed in claim 1, characterized in that: After the step of judging whether the field slice diagram belongs to the error-prone area based on the field error-prone status configuration file, the method further includes: If the field slice image does not belong to the easily misplaced area, extracting the text position information and text content information of the field slice image by using OCR technology; Regular expressions are used to extract the text position information and the text content information to obtain a recognition result.
3. The method for recognizing misplaced characters on bills as claimed in claim 1, characterized in that: The step of acquiring a bill image and performing region cutting on the bill image according to a field relative region parameter configuration file to obtain a plurality of field slice images includes: Use high-speed scanner equipment to image the bill and obtain the bill image; Determine the relative area coordinate information corresponding to each field in the bill according to the field relative area parameter configuration file; The bill image is region-cut based on relative region coordinate information corresponding to each field in the bill to obtain a plurality of field slice images.
4. The method for recognizing misplaced characters on bills as claimed in claim 1, characterized in that: If the field slice image belongs to an easily misplaced area, the step of processing the field slice image by end-to-end image understanding technology and OCR technology to obtain a first prediction result and a second prediction result includes: If the field slice diagram belongs to an easily misplaced area, determining the field corresponding to the field slice diagram, and using the field name corresponding to the field slice diagram as a prompt word; Based on the field slice graph and the prompt word, using end-to-end image understanding technology, obtaining a first prediction result corresponding to the field slice graph; Based on the field slice diagram, using OCR technology, obtaining text information of the field slice diagram; The text information is extracted using a regular expression to obtain a second prediction result corresponding to the field slice diagram.
5. The method for recognizing misplaced characters on bills as claimed in claim 1, characterized in that: The step of performing a similarity comparison between the first prediction result and the second prediction result and determining a recognition result based on the similarity comparison result includes: Performing a similarity comparison between the first prediction result and the second prediction result according to a character dimension to obtain a similarity comparison result; If the similarity comparison result indicates that the similarity between the first prediction result and the second prediction result is not 1, matching the first prediction result and the second prediction result with a corpus to obtain a matching result; When the matching result indicates that the matching degree between the first prediction result and the second prediction result and the corpus in the corpus is 1, determining the similarity between the first prediction result and the second prediction result; The similarity and the first prediction result or the second prediction result are used as recognition results.
6. The method for recognizing misplaced characters on bills as claimed in claim 5, characterized in that: After the step of matching the first prediction result and the second prediction result with the corpus to obtain the matching result, the method further includes: When the matching result indicates that the matching degree between the first prediction result and the corpus in the corpus is 1 and the matching degree between the second prediction result and the corpus in the corpus is not 1, determining the similarity between the first prediction result and the second prediction result, and using the similarity and the first prediction result as the recognition result; When the matching result indicates that the matching degree between the second prediction result and the corpus in the corpus is 1 and the matching degree between the first prediction result and the corpus in the corpus is not 1, determining the similarity between the first prediction result and the second prediction result, and using the similarity and the second prediction result as the recognition result; When the matching result indicates that the matching degree between the first prediction result and the second prediction result and the corpus in the corpus is not 1, the similarity between the first prediction result and the second prediction result is determined, and the similarity and a null value are used as recognition results.
7. The method for recognizing misplaced characters on bills according to any one of claims 1 to 6, characterized in that: After the step of comparing the first prediction result and the second prediction result for similarity and determining the recognition result based on the similarity comparison result, the method further includes: Review and correct the recognition result based on the recognition result to obtain a processed recognition result; The processed recognition results are fed back to the corpus to obtain an updated corpus.
8. A device for recognizing misplaced characters on bills, characterized in that: The device comprises: A region cutting module is used to obtain a bill image and perform region cutting on the bill image according to a field relative region parameter configuration file to obtain a plurality of field slice images; A region judgment module, used to judge whether the field slice graph belongs to a misalignment region based on a field error-prone status configuration file; A prediction output module, configured to process the field slice image by end-to-end image understanding technology and OCR technology respectively to obtain a first prediction result and a second prediction result if the field slice image belongs to an easily misplaced area; The result output module is used to perform a similarity comparison between the first prediction result and the second prediction result, and determine a recognition result based on the similarity comparison result.
9. A bill misaligned character recognition device, characterized in that: The device comprises: a memory, a processor and a bill misplaced character recognition program stored in the memory and executable on the processor, wherein the bill misplaced character recognition program is configured to implement the steps of the bill misplaced character recognition method as described in any one of claims 1 to 7.
10. A storage medium, characterized in that: The storage medium stores a bill misplaced character recognition program, and when the bill misplaced character recognition program is executed by the processor, the steps of the bill misplaced character recognition method as described in any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Image information identification method and device
CN120564204A