A multi-category invoice recognition method, system, electronic device and storage medium
By performing direction correction, edge corner extraction and OCR identification optimization on multi-category invoices, combined with invoice layout analysis and filter correction, the problem that traditional invoice identification solutions cannot adapt to multi-category invoices is solved, and high accuracy and robust invoice identification are achieved.
Patent Information
- Application Number
- CN202410606931.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-16
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2044-05-16
AI Technical Summary
Traditional invoice identification solutions are rigid and not universal, and cannot adapt to the problem of multi-category invoices.
By correcting the direction and extracting edge corner points of the invoice image to be identified, the invoice image is optimized, and the PaddleOCR model is used to optimize OCR recognition and optimization to obtain the initial text recognition information. Then, the invoice layout analysis is performed based on the initial text identification information, the invoice text classification information is obtained, and filtered and corrected to obtain the target invoice identification results.
It realizes flexible and adaptable multi-category invoice recognition, which improves the accuracy and robustness of invoice recognition.
Smart Images

Figure CN118675181B_ABST
Abstract
Description
Background Art
[0002] The invoice recognition task is a KIE (Key Information Extraction) task, and its main purpose is to accurately extract the key-value information of the keyword fields in the invoice image. The traditional key information extraction scheme for invoice recognition is to use relevant image preprocessing means to improve the clarity of the image, use the position information of the fixed invoice layout to achieve regional segmentation of the image, identify the segmented image through the optical character recognition OCR (Optical Character Recognition) algorithm, and transmit the recognized text information into the rule engine for key information extraction. Through the iteration of the OCR algorithm, the invoice recognition task is finally achieved. However, the traditional invoice recognition scheme is rigid and not universal, and cannot adapt to the problem of multi-category invoices.
[0003] Therefore, it is urgent to provide a technical solution to solve the above technical problems. Summary of the Invention
[0004] To solve the above technical problems, the present invention provides a multi-category invoice recognition method, system, electronic device, and storage medium.
[0005] In a first aspect, the present invention provides a multi-category invoice recognition method, and the technical solution of this method is as follows:
[0006] Perform direction correction and edge corner point extraction on the invoice image to be recognized to obtain an optimized invoice image, and perform OCR recognition optimization on the optimized invoice image to obtain initial text recognition information;
[0007] Based on the initial text recognition information, perform invoice layout analysis to obtain invoice text classification information, and perform filtering and correction on the invoice text classification information to obtain the target invoice recognition result.
[0008] The beneficial effects of a multi-category invoice recognition method of the present invention are as follows:
[0009] The method of the present invention can flexibly and adaptively achieve multi-category invoice recognition, and improve the accuracy and robustness of invoice recognition.
[0010] On the basis of the above solution, a multi-category invoice recognition method of the present invention can also be improved as follows.
[0011] In an optional manner, the step of performing direction correction and edge corner point extraction on the invoice image to be recognized to obtain an optimized invoice image includes:
[0012] Using PaddleClas, the direction of the invoice image to be recognized is corrected to obtain a corrected invoice image, and an edge corner point extraction algorithm is used to extract edge corner points from the corrected invoice image to obtain the optimized invoice image.
[0013] In an alternative embodiment, the step of performing OCR recognition optimization on the optimized invoice image to obtain initial text recognition information includes:
[0014] Using the trained PaddleOCR model, perform OCR recognition optimization on the optimized invoice image to obtain the initial text recognition information.
[0015] In an alternative embodiment, the PaddleOCR model includes: a text detection module, a text direction classifier module, and a text recognition module;
[0016] Among them, the text detection module is used to identify the text area in the image and generate a text bounding box of the image; the text direction classifier module is used to identify the angle of the text area relative to the horizontal direction to enable text recognition in the horizontal direction; the text recognition module is used to use the CRNN network to recognize each character in the text area to obtain text recognition information.
[0017] In an alternative embodiment, the step of performing invoice layout analysis based on the initial text recognition information to obtain invoice text classification information includes:
[0018] Based on the attention mechanism, obtain the information related to the decoder in the initial text recognition information, and use the encoder to encode the initial text recognition information to obtain an encoded vector;
[0019] Use the RNN network, and based on the encoded vector and the information related to the decoder in the initial text recognition information, generate the invoice text classification information.
[0020] In an alternative embodiment, the step of filtering and correcting the invoice text classification information to obtain the target invoice recognition result includes:
[0021] Filter the invalid characters in the invoice text classification information, and convert and correct the Chinese capital amount information in the invoice text classification information to obtain the target invoice recognition result.
[0022] In a second aspect, the present invention provides a multi-category invoice recognition system, and the technical solution of the system is as follows:
[0023] A multi-category invoice recognition system includes: a first recognition unit and a second recognition unit;
[0024] The first recognition unit is used for: performing orientation correction and edge corner point extraction on the invoice image to be recognized to obtain an optimized invoice image, and performing OCR recognition optimization on the optimized invoice image to obtain initial text recognition information;
[0025] The second recognition unit is used for: performing invoice layout analysis based on the initial text recognition information to obtain invoice text classification information, and performing filtering and correction on the invoice text classification information to obtain a target invoice recognition result.
[0026] The beneficial effects of a multi-category invoice recognition system of the present invention are as follows:
[0027] The system of the present invention can flexibly and adaptively achieve multi-category invoice recognition, and improves the accuracy and robustness of invoice recognition.
[0028] On the basis of the above solution, a multi-category invoice recognition system of the present invention can also be improved as follows.
[0029] In an optional manner, the step of performing orientation correction and edge corner point extraction on the invoice image to be recognized in the first recognition unit to obtain an optimized invoice image includes:
[0030] Using PaddleClas to perform orientation correction on the invoice image to be recognized to obtain a corrected invoice image, and using a key point extraction algorithm to perform edge corner point extraction on the corrected invoice image to obtain the optimized invoice image.
[0031] In a third aspect, the technical solution of an electronic device of the present invention is as follows:
[0032] It includes a memory, a processor, and a program stored on the memory and running on the processor. When the processor executes the program, it implements the steps of the multi-category invoice recognition method of the present invention.
[0033] In a fourth aspect, the technical solution of a computer-readable storage medium provided by the present invention is as follows:
[0034] Instructions are stored in the computer-readable storage medium. When the computer-readable storage medium reads the instructions, it causes the computer-readable storage medium to execute the steps of the multi-category invoice recognition method of the present invention.
[0035] The above description is only an overview of the technical solution of the present invention. In order to be able to understand the technical means of the present invention more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present invention more obvious and understandable, the specific embodiments of the present invention are given below. Description of the Drawings
[0036] The accompanying drawings are only used to illustrate the embodiments and are not considered to limit the present invention. Moreover, throughout the accompanying drawings, the same reference signs are used to denote the same components. In the drawings:
[0037] Figure 1 It is a schematic flowchart of an embodiment of a multi-category invoice recognition method provided by the present invention;
[0038] Figure 2 It is a schematic diagram of invoice images of three different invoice types;
[0039] Figure 3 It is a schematic diagram of the principle of invoice image orientation correction;
[0040] Figure 4 It is a schematic diagram of the principle of key point extraction;
[0041] Figure 5 It is a schematic diagram of the principle of OCR recognition optimization;
[0042] Figure 6 It is a schematic flowchart of the training process of the PaddleOCR model;
[0043] Figure 7 It is a schematic diagram of the principle of layout analysis;
[0044] Figure 8 It is a schematic diagram of the principle of filtering and correction;
[0045] Figure 9 It is an overall optimization flowchart;
[0046] Figure 10 It is one of the schematic diagrams of the results of accuracy rate testing;
[0047] Figure 11 It is another schematic diagram of the results of accuracy rate testing;
[0048] Figure 12 It is a schematic structural diagram of an embodiment of a multi-category invoice recognition system provided by the present invention;
[0049] Figure 13 It is a schematic structural diagram of an embodiment of an electronic device provided by the present invention. Detailed Embodiments
[0050] The exemplary embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments set forth herein.
[0051] Figure 1 It shows a schematic flowchart of an embodiment of a multi-category invoice recognition method provided by the present invention. As Figure 1As shown in the figure, it includes the following steps:
[0052] S1. Perform direction correction and edge corner point extraction on the invoice image to be recognized to obtain an optimized invoice image, and perform OCR recognition optimization on the optimized invoice image to obtain initial text recognition information.
[0053] Among them, the invoice image to be recognized is any type of invoice image that needs to be recognized in this embodiment. Figure 2 Figure 8 shows invoice images of three different invoice types. The optimized invoice image is the invoice image after direction correction and edge corner point extraction. The initial text recognition information is the text recognition information obtained after performing OCR recognition optimization on the invoice image. The text recognition information includes, but is not limited to: invoice code, invoice number, invoicing date, purchaser name, seller name, purchaser taxpayer identification number, seller taxpayer identification number, and total amount of tax and price (in lowercase), etc.
[0054] S2. Perform invoice layout analysis based on the initial text recognition information to obtain invoice text classification information, and perform filtering and correction on the invoice text classification information to obtain the target invoice recognition result.
[0055] Among them, invoice layout analysis is used to classify key fields for invoices of multiple layout types, realize the generation of corresponding layout information sequences for the input invoice image and the OCR recognition result output, so as to analyze the complex mapping relationship between the image and the layout information, and at the same time support the recognition of multiple types of invoices. The target invoice recognition result contains the recognition information including the invoice key fields finally recognized.
[0056] In an optional manner, the step of performing direction correction and edge corner point extraction on the invoice image to be recognized to obtain an optimized invoice image includes:
[0057] Use PaddleClas to perform direction correction on the invoice image to be recognized to obtain a corrected invoice image, and use a key point extraction algorithm to perform edge corner point extraction on the corrected invoice image to obtain the optimized invoice image.
[0058] Among them, since traditional image preprocessing only uses relatively basic graphics algorithms, although multiple preprocessing means are used, in the case of dealing with some complex scenarios, such as invoice distortion or tilt, the preprocessing may still be insufficient to handle various image transformations. Therefore, as Figure 3 shown in the figure, in this embodiment, fine-tuning is performed on the pre-trained model provided by the open-source PaddleClas to perform direction correction on the invoice image to be recognized, realizing the function of multi-direction correction of the invoice image, and ensuring that the text content in the corrected invoice image output after the model processes the invoice image to be recognized in any direction conforms to the human reading habit.
[0059] Among them, as Figure 4 shown, the key point extraction algorithm realizes the function of accurately extracting the edge corner points of the invoice in the corrected invoice image, so that in the case of invoice distortion or tilt, the invoice data can be adaptively extracted and restored to obtain an optimized invoice image.
[0060] In an optional manner, the step of performing OCR recognition optimization on the optimized invoice image to obtain initial text recognition information includes:
[0061] Using the trained PaddleOCR model, perform OCR recognition optimization on the optimized invoice image to obtain the initial text recognition information.
[0062] Among them, the PaddleOCR model includes: a text detection module, a text direction classifier module, and a text recognition module; the text detection module is used to identify the text area in the image and generate a text bounding box of the image; the text direction classifier module is used to identify the angle of the text area relative to the horizontal direction so that text recognition can be performed in the horizontal direction; the text recognition module is used to use the CRNN network to recognize each character in the text area to obtain text recognition information.
[0063] Specifically, the text detection module adopts the DB (Differentiable Binarization) model, aiming to accurately locate and identify the text area in the image. This module can identify the text bounding box in the image and separate the text in the image from the background. The function of the text direction classifier module is to classify the direction of the text in the text area so that subsequent text recognition can be performed in the horizontal text direction. The text recognition module is the core component of PaddleOCR and is responsible for converting the detected text area into text that can be understood by a computer. The text recognition module adopts a convolutional recurrent neural network (CRNN), divides the text area in the image into individual characters, and recognizes each character. By learning the relationships and orders between characters, the text recognition module can output text recognition information (as Figure 5 shown).
[0064] It should be noted that the training process of the PaddleOCR model is as Figure 6 shown.
[0065] In an optional manner, the step of performing invoice layout analysis based on the initial text recognition information to obtain invoice text classification information includes:
[0066] Based on the attention mechanism, obtain the information related to the decoder in the initial text recognition information, and use the encoder to encode the initial text recognition information to obtain an encoded vector. Use an RNN network, and based on the encoded vector and the information related to the decoder in the initial text recognition information, generate the invoice text classification information.
[0067] Among them, the attention mechanism is used to capture the information in the input sequence (initial text recognition information) related to the current moment of the decoder to help the decoder generate accurate output invoice text classification information. The encoder uses a recurrent neural network (RNN) to capture sequence information. Especially for text data, it encodes information such as the text results recognized by OCR and the text region position coordinates returned by text detection in the input sequence into a fixed-length encoded vector. The decoder also uses a recurrent neural network (RNN) to generate the final output sequence, such as invoice text classification information (as Figure 7 shown).
[0068] In an alternative way, the steps of filtering and correcting the invoice text classification information to obtain the target invoice recognition result include:
[0069] Filter the invalid characters in the invoice text classification information, and convert and correct the Chinese capital amount information in the invoice text classification information to obtain the target invoice recognition result.
[0070] Among them, by processing redundant fields, such as "No" in the invoice code, and redundant special punctuation marks, etc., extract valid information from the text, standardize the text format, and make it more suitable for subsequent invoice information recognition and processing. Extract valid characters from the Chinese capital amount through regular matching, including numeric characters and amount unit characters. According to the mapping relationship between Chinese numbers and amount units, convert the Chinese capital amount to the corresponding Arabic numeral amount, making it able to flexibly adapt to different amount formats, greatly improving the accuracy of the amount field, and finally obtaining the target invoice recognition result (as Figure 8 shown).
[0071] Figure 9The overall optimization flowchart of this embodiment is shown. Based on the open-source PaddleClas deep learning algorithm and the self-developed invoice key point extraction algorithm, the image data preprocessing module of this embodiment is constructed to realize multi-angle transformation of complex images and adaptive extraction and restoration of images in the scenarios of invoice distortion or inclination, enhancing the flexibility of the system. Based on the open-source OCR engine PaddleOCR, this embodiment designs and implements a high-precision Chinese-English digital OCR module. Through fine-tuning on a large amount of real data, iterative optimization is carried out for problems such as complex fonts, low contrast, or overlapping texts that often appear in the invoice recognition scenario, greatly improving the recognition accuracy in this scenario. The seq2seq layout analysis module designed in this embodiment supports keyword field classification for invoices of multiple layout types, realizes the generation of corresponding layout information sequences for the input invoice image and the OCR recognition result output, enables the system to analyze the complex mapping relationship between the image and the layout information, and at the same time supports multi-category invoice recognition, improving the generalization ability of the system. For special keyword fields with a high misrecognition rate, this embodiment designs and implements an invalid character filtering module and an amount auditing and correction module. By filtering invalid characters from the characters classified by the seq2seq layout analysis module and auditing and correcting relevant amount fields, the recognition accuracy of special keyword fields is greatly improved.
[0072] To prove that the accuracy of this solution is higher than that of the traditional invoice recognition solution, on the same data set, the traditional invoice recognition solution (original solution) and the general multi-category invoice recognition implemented in the present invention are respectively used to count the accuracy of the two in 8 keyword fields and conduct a comparative analysis to verify the effectiveness of the present invention:
[0073] 1) Since the traditional invoice recognition solution only supports the recognition of electronic invoices of type 2, an invoice data set INVOICE_TYPE2 is constructed, which contains a total of 2000 invoice data of type 2. The accuracy of the two solutions in 8 fields, namely invoice code, invoice number, issue date, purchaser name, seller name, purchaser taxpayer identification number, seller taxpayer identification number, and total amount of tax and price (in lowercase), is tested on the data set INVOICE_TYPE2, as Figure 10 shown. It can be Figure 10 seen that the accuracy of each field of this solution on the INVOICE_TYPE2 data set is completely better than that of the traditional solution.
[0074] 2) An invoice data set INVOICE_TYPE123 is constructed, which contains 2000 invoice data of each of types 1, 2, and 3. The accuracy of the two solutions in 8 fields, namely invoice code, invoice number, issue date, purchaser name, seller name, purchaser taxpayer identification number, seller taxpayer identification number, and total amount of tax and price (in lowercase), is tested on the data set INVOICE_TYPE123, asFigure 11 As shown. From Figure 11 it can be seen that the accuracy of each field in the INVOICE_TYPE123 dataset of this solution is completely better than that of the traditional solution.
[0075] In summary, the general multi-category invoice recognition system implemented in the solution based on this embodiment, and its accuracy is completely better than that of the traditional solution.
[0076] Figure 12 shows a schematic structural diagram of an embodiment of a multi-category invoice recognition system 200 provided by the present invention. As Figure 12 shown, the system 200 includes: a first recognition unit 210 and a second recognition unit 220;
[0077] The first recognition unit 210 is used to: perform direction correction and edge corner point extraction on the invoice image to be recognized to obtain an optimized invoice image, and perform OCR recognition optimization on the optimized invoice image to obtain initial text recognition information;
[0078] The second recognition unit 220 is used to: perform invoice layout analysis based on the initial text recognition information to obtain invoice text classification information, and perform filtering and correction on the invoice text classification information to obtain a target invoice recognition result.
[0079] In an optional manner, the step of performing direction correction and edge corner point extraction on the invoice image to be recognized in the first recognition unit 210 to obtain an optimized invoice image includes:
[0080] Using PaddleClas to perform direction correction on the invoice image to be recognized to obtain a corrected invoice image, and using a key point extraction algorithm to perform edge corner point extraction on the corrected invoice image to obtain the optimized invoice image.
[0081] In an optional manner, the step of performing OCR recognition optimization on the optimized invoice image in the first recognition unit 210 to obtain initial text recognition information includes:
[0082] Using the trained PaddleOCR model to perform OCR recognition optimization on the optimized invoice image to obtain the initial text recognition information.
[0083] In an optional manner, the PaddleOCR model includes: a text detection module, a text direction classifier module, and a text recognition module;
[0084] Among them, the text detection module is used to identify the text area in the image and generate the text bounding box of the image; the text direction classifier module is used to identify the angle of the text area relative to the horizontal direction so that text recognition is performed in the horizontal direction; the text recognition module is used to recognize each character in the text area by using the CRNN network to obtain text recognition information.
[0085] In an optional manner, the step of performing invoice layout analysis based on the initial text recognition information in the second recognition unit 220 to obtain invoice text classification information includes:
[0086] Based on the attention mechanism, obtain the information related to the decoder in the initial text recognition information, and use the encoder to encode the initial text recognition information to obtain an encoded vector;
[0087] Use the RNN network, and based on the encoded vector and the information related to the decoder in the initial text recognition information, generate the invoice text classification information.
[0088] In an optional manner, the step of filtering and correcting the invoice text classification information in the second recognition unit 220 to obtain the target invoice recognition result includes:
[0089] Filter the invalid characters in the invoice text classification information, and convert and correct the Chinese capital amount information in the invoice text classification information to obtain the target invoice recognition result.
[0090] The technical solution of this embodiment can flexibly and adaptively implement multi-category invoice recognition, and improve the accuracy and robustness of invoice recognition.
[0091] For the parameters and the steps for each module in the multi-category invoice recognition system 200 of this embodiment to implement the corresponding functions, reference can be made to the parameters and steps in the embodiment of the multi-category invoice recognition method in the above text, which will not be elaborated here.
[0092] As Figure 13 shown, an electronic device 300 according to an embodiment of the present invention, the electronic device 300 includes a processor 320, the processor 320 is coupled to a memory 310, and at least one computer program 330 is stored in the memory 310. The at least one computer program 330 is loaded and executed by the processor 320 so that the electronic device 300 implements any one of the above multi-category invoice recognition methods. Specifically:
[0093] The electronic device 300 can vary significantly due to different configurations or performances. It may include one or more processors 320 (Central Processing Units, CPUs) and one or more memories 310. Among them, at least one computer program 330 is stored in the one or more memories 310. The at least one computer program 330 is loaded and executed by the one or more processors 320, so that the electronic device 300 can implement any one of the multi-category invoice recognition methods provided in the above embodiments. Of course, the electronic device 300 may also have components such as a wired or wireless network interface, a keyboard, and an input / output interface for input and output. The electronic device 300 may also include other components for implementing device functions, which will not be elaborated here.
[0094] In an embodiment of the present invention, a computer-readable storage medium stores at least one computer program, and the at least one computer program is loaded and executed by a processor to enable a computer to implement any one of the above multi-category invoice recognition methods.
[0095] Optionally, the computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), a magnetic tape, a floppy disk, an optical data storage device, etc.
[0096] In an exemplary embodiment, a computer program product or a computer program is also provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the electronic device executes any one of the above multi-category invoice recognition methods.
[0097] It should be noted that the terms "first", "second", etc. in the specification and claims of this application are used to distinguish similar objects, and do not represent a limitation on a specific order or sequence. In appropriate cases, the order of use of similar objects can be interchanged, so that the embodiments of the present application described here can be implemented in an order other than the illustrated or described order.
[0098] Those skilled in the art of the present technology know that the present invention can be implemented as a system, a method, or a computer program product. Therefore, the present disclosure can be specifically implemented in the following forms, namely: it can be entirely hardware, entirely software (including firmware, resident software, microcode, etc.), or a combination of hardware and software, which is generally referred to as "circuit", "module", or "system" in this article. In addition, in some embodiments, the present invention can also be implemented in the form of a computer program product in one or more computer-readable media, which contain computer-readable program codes.
[0099] Any combination of one or more computer-readable media can be adopted. The computer-readable media can be computer-readable signal media or computer-readable storage media. The computer-readable storage media can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above. More specific examples (non-exhaustive list) of the computer-readable storage media include: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this application, the computer-readable storage media can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, device, or component.
[0100] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limitations on the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.
Claims
1. A multi-category invoice recognition method, characterized in that: include: Direction correction and edge corner point extraction are performed on the invoice image to be recognized to obtain an optimized invoice image, and OCR recognition optimization is performed on the optimized invoice image to obtain initial text recognition information; Performing invoice layout analysis based on the initial text recognition information to obtain invoice text classification information, and filtering and correcting the invoice text classification information to obtain a target invoice recognition result; The step of performing invoice layout analysis based on the initial text recognition information to obtain invoice text classification information includes: Based on the attention mechanism, information related to the decoder in the initial text recognition information is obtained, and the initial text recognition information is obtained by using the encoder for encoding to obtain an encoding vector; Generate the invoice text classification information by using an RNN network and based on the encoding vector and information related to the decoder in the initial text recognition information; The attention mechanism is used to capture the information in the initial text recognition information that is relevant to the decoder at the current moment, so as to help the decoder generate the invoice text classification information; the encoder uses a recurrent neural network to capture the sequence information of the text type, encode the text results recognized by OCR and the text area position coordinate information returned by text detection, and encode them into a fixed-length encoding vector; the decoder uses a recurrent neural network to generate the invoice text classification information based on the encoder's encoding vector; The steps of performing direction correction and edge corner point extraction on the invoice image to be recognized to obtain an optimized invoice image include: Using PaddleClas, the direction of the invoice image to be recognized is corrected to obtain a corrected invoice image, and the edge corner points of the corrected invoice image are extracted using a key point extraction algorithm to obtain the optimized invoice image; The step of performing OCR recognition optimization on the optimized invoice image to obtain initial text recognition information includes: Using the trained PaddleOCR model, perform OCR recognition optimization on the optimized invoice image to obtain the initial text recognition information; The PaddleOCR model includes: a text detection module, a text direction classifier module and a text recognition module; Among them, the text detection module is used to identify the text area in the image and generate the text boundary box of the image; the text direction classifier module is used to identify the angle of the text area relative to the horizontal direction so that text recognition can be performed in the horizontal direction; the text recognition module is used to use the CRNN network to recognize each character in the text area to obtain text recognition information.
2. The multi-category invoice recognition method according to claim 1, characterized in that: The step of filtering and correcting the invoice text classification information to obtain a target invoice recognition result includes: Invalid characters in the invoice text classification information are filtered, and the Chinese capital amount information in the invoice text classification information is converted and corrected to obtain the target invoice recognition result.
3. A multi-category invoice recognition system, characterized in that: include: a first identification unit and a second identification unit; The first recognition unit is used to: perform direction correction and edge corner point extraction on the invoice image to be recognized to obtain an optimized invoice image, and perform OCR recognition optimization on the optimized invoice image to obtain initial text recognition information; The second recognition unit is used to: perform invoice layout analysis based on the initial text recognition information to obtain invoice text classification information, and filter and correct the invoice text classification information to obtain a target invoice recognition result; The step of performing invoice layout analysis based on the initial text recognition information to obtain invoice text classification information in the second recognition unit includes: Based on the attention mechanism, information related to the decoder in the initial text recognition information is obtained, and the initial text recognition information is obtained by using the encoder for encoding to obtain an encoding vector; Generate the invoice text classification information by using an RNN network and based on the encoding vector and information related to the decoder in the initial text recognition information; The attention mechanism is used to capture the information in the initial text recognition information that is relevant to the decoder at the current moment, so as to help the decoder generate the invoice text classification information; the encoder uses a recurrent neural network to capture the sequence information of the text type, encode the text results recognized by OCR and the text area position coordinate information returned by text detection, and encode them into a fixed-length encoding vector; the decoder uses a recurrent neural network to generate the invoice text classification information based on the encoder's encoding vector; The step of performing direction correction and edge corner point extraction on the invoice image to be recognized in the first recognition unit to obtain an optimized invoice image includes: Using PaddleClas, the direction of the invoice image to be recognized is corrected to obtain a corrected invoice image, and the edge corner points of the corrected invoice image are extracted using a key point extraction algorithm to obtain the optimized invoice image; The step of performing OCR recognition optimization on the optimized invoice image to obtain initial text recognition information in the first recognition unit includes: Using the trained PaddleOCR model, perform OCR recognition optimization on the optimized invoice image to obtain the initial text recognition information; The PaddleOCR model includes: a text detection module, a text direction classifier module and a text recognition module; Among them, the text detection module is used to identify the text area in the image and generate the text boundary box of the image; the text direction classifier module is used to identify the angle of the text area relative to the horizontal direction so that text recognition can be performed in the horizontal direction; the text recognition module is used to use the CRNN network to recognize each character in the text area to obtain text recognition information.
4. An electronic device, characterized in that: The electronic device includes a processor, the processor is coupled to a memory, the memory stores at least one computer program, and the at least one computer program is loaded and executed by the processor so that the electronic device implements the multi-category invoice recognition method as described in claim 1 or 2.
5. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores at least one computer program, and the at least one computer program is loaded and executed by the processor so that the computer-readable storage medium implements the multi-category invoice identification method as claimed in claim 1 or 2.
Citation Information
Patent Citations
An image information extraction method and device
CN109034159A
Invoice information identification method and device, computer equipment and storage medium
CN114332883A
Medical text recognition method and device, computer equipment and storage medium
CN116912847A