Bill image identification method, apparatus and device, medium and program product
By combining field identifiers and text block detection models, the field identifiers and content location information of the invoice image are obtained, which solves the problem of low accuracy in invoice image recognition and achieves higher recognition accuracy.
Patent Information
- Application Number
- CN202511098878.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-06
- Publication Date
- 2025-11-11
AI Technical Summary
The problem of low accuracy in document image recognition in existing technologies.
By acquiring multiple field identifiers from the ticket image, and using a field identifier detection model and a text block detection model, the positional information of the field identifiers and text blocks is determined, thereby obtaining the field content.
It improves the recognition accuracy of invoice images, especially when the invoice type changes, it can still accurately locate field content.
Smart Images

Figure CN120932253A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of financial technology or other related fields, and in particular to a method, apparatus, device, medium and program product for recognizing bill images. Background Technology
[0002] Optical Character Recognition (OCR) is one of the more mature application areas of deep learning artificial intelligence. In OCR scenarios, computing devices can acquire images of documents to be processed and recognize the text within those images.
[0003] In related technologies, computing devices typically perform recognition processing on invoice images directly based on invoice detection models to obtain field identifiers and corresponding field contents in the invoice images.
[0004] However, this method suffers from low accuracy in recognizing invoice images. Summary of the Invention
[0005] This application provides a method, apparatus, device, medium, and program product for recognizing invoice images, in order to solve the technical problem of low accuracy in recognizing invoice images in related technologies.
[0006] In a first aspect, embodiments of this application provide a method for recognizing invoice images, comprising:
[0007] Obtain the image of the invoice to be identified, and obtain the identifiers of multiple fields corresponding to the invoice image;
[0008] Based on the field identifier detection model, the ticket image is processed to identify multiple field identifiers and obtain the location information of the multiple field identifiers.
[0009] Based on the text block detection model, the ticket image is processed to obtain the location information of multiple text blocks; among them, the text blocks are field identifiers or field content;
[0010] Based on the position information of multiple text blocks and the position information of multiple field identifiers, determine the field content corresponding to the multiple field identifiers;
[0011] The document recognition result corresponding to the document image includes multiple field identifiers and the field content corresponding to these field identifiers.
[0012] Secondly, embodiments of this application provide a device for recognizing invoice images, comprising:
[0013] The acquisition module is used to acquire the image of the ticket to be identified and to acquire multiple field identifiers corresponding to the ticket image;
[0014] The processing module is used to identify and process the ticket image based on the field identifier detection model, according to multiple field identifiers, to obtain the location information of multiple field identifiers;
[0015] The processing module is also used to recognize and process the ticket image based on the text block detection model to obtain the position information of multiple text blocks; wherein, the text blocks are field identifiers or field content;
[0016] The processing module is also used to determine the field content corresponding to multiple field identifiers based on the position information of multiple text blocks and the position information of multiple field identifiers;
[0017] The processing module is also used to determine the ticket recognition result corresponding to the ticket image, including multiple field identifiers and the field content corresponding to the multiple field identifiers.
[0018] Thirdly, embodiments of this application provide a computing device, including:
[0019] The processor, and the memory that is in communication with the processor;
[0020] Memory is used to store instructions that the computer executes;
[0021] The processor is used to execute computer execution instructions stored in memory to implement the method of the first aspect.
[0022] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the method of the first aspect.
[0023] Fifthly, embodiments of this application provide a computer program product, including a computer program, which, when executed by a processor, is used to implement the method of the first aspect.
[0024] This application provides a method, apparatus, device, medium, and program product for recognizing invoice images. In this method, a computing device can acquire an image of an invoice to be recognized. The computing device can acquire multiple field identifiers corresponding to the invoice image. The computing device can perform recognition processing on the invoice image based on a field identifier detection model, obtaining the position information of the multiple field identifiers. The computing device can also perform recognition processing on the invoice image based on a text block detection model, obtaining the position information of multiple text blocks; wherein, the text blocks are field identifiers or field content. The computing device can determine the field content corresponding to the multiple field identifiers based on the position information of the multiple text blocks and the position information of the multiple field identifiers. The computing device can determine that the invoice recognition result corresponding to the invoice image includes multiple field identifiers and the field content corresponding to the multiple field identifiers. Through the above methods, the recognition accuracy of invoice images is improved. Attached Figure Description
[0025] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0026] Figure 1 This is a schematic diagram illustrating an application scenario applicable to an embodiment of this application;
[0027] Figure 2a A flowchart illustrating an embodiment of a method for recognizing invoice images provided in this application;
[0028] Figure 2b This is a schematic diagram of an image segmentation scenario provided in an embodiment of this application;
[0029] Figure 2c This application provides a schematic diagram of a scenario for recognizing text blocks.
[0030] Figure 3 A schematic flowchart of a second embodiment of a method for recognizing invoice images provided in this application;
[0031] Figure 4 A flowchart illustrating a third embodiment of a method for recognizing invoice images provided in this application;
[0032] Figure 5 A schematic diagram of the structure of a ticket image recognition device provided in this application;
[0033] Figure 6 A schematic diagram of the structure of a computing device provided in this application.
[0034] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation
[0035] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0036] It should be noted that the user information (including but not limited to user identity information, user device information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, such as invoice images) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, storage, use, processing, transmission, provision, disclosure, and application of the relevant data all comply with the relevant laws, regulations, and standards of the relevant countries and regions, have taken necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation access points for users to choose to authorize or refuse.
[0037] Furthermore, the technical solution involved in this application, which conducts big data analysis on user information (including but not limited to biometric information, ID card information, consumption data, asset data, electronic terminal operation data, etc.) and uses artificial intelligence technology to make automated decisions, and makes decisions that have a significant impact on personal rights based on the results of automated decisions, provides users with corresponding operation entry points for users to choose to agree to or reject the results of automated decisions; if the user chooses to reject, the process will proceed to the expert decision-making process.
[0038] It should be noted that the method, apparatus, equipment, medium and program product for recognizing bill images provided in this application can be used in the field of financial technology or other related fields, and the application field of the method, apparatus, equipment, medium and program product for recognizing bill images in this application is not limited.
[0039] Optical Character Recognition (OCR) is one of the more mature application areas of deep learning artificial intelligence. In OCR scenarios, computing devices can acquire images of documents to be processed and recognize the text within those images.
[0040] In related technologies, computing devices typically perform recognition processing on invoice images directly based on invoice detection models to obtain field identifiers and corresponding field contents in the invoice images.
[0041] However, this method suffers from low accuracy in recognizing invoice images.
[0042] Based on the aforementioned technical problems, the technical concept of this application is as follows: A computing device can acquire an image of a document to be recognized. The computing device can acquire multiple field identifiers corresponding to the document image. The computing device can perform recognition processing on the document image based on a field identifier detection model, obtaining the position information of the multiple field identifiers. The computing device can perform recognition processing on the document image based on a text block detection model, obtaining the position information of multiple text blocks; wherein, the text blocks are field identifiers or field content. The computing device can determine the field content corresponding to the multiple field identifiers based on the position information of the multiple text blocks and the position information of the multiple field identifiers. The computing device can determine that the document recognition result corresponding to the document image includes multiple field identifiers and the field content corresponding to the multiple field identifiers.
[0043] The above methods improve the accuracy of document image recognition.
[0044] Figure 1 This is a schematic diagram illustrating an application scenario to which this application's embodiments apply. For example... Figure 1 As shown, this application scenario includes terminal device 10 and computing device 20.
[0045] In this application scenario:
[0046] The computing device 20 can acquire an image of the ticket to be identified. In one implementation, the ticket image can be sent by the terminal device 10.
[0047] The computing device 20 can obtain multiple field identifiers corresponding to the ticket image.
[0048] The computing device 20 can perform recognition processing on the ticket image based on the field identifier detection model and according to multiple field identifiers to obtain the location information of multiple field identifiers.
[0049] The computing device 20 can perform recognition processing on the ticket image based on the text block detection model to obtain the position information of multiple text blocks; wherein, the text blocks are field identifiers or field content.
[0050] The computing device 20 can determine the field content corresponding to multiple field identifiers based on the position information of multiple text blocks and the position information of multiple field identifiers.
[0051] The computing device 20 can determine that the ticket recognition result corresponding to the ticket image includes multiple field identifiers and the field contents corresponding to the multiple field identifiers. In one implementation, the computing device 20 can send the ticket recognition result corresponding to the ticket image to the terminal device 10.
[0052] It should be noted that, Figure 1 This is merely a schematic diagram illustrating one application scenario applicable to the embodiments of this application; this application does not necessarily imply any wrongdoing. Figure 1The actual form and interaction method of the various devices included are limited.
[0053] The method for recognizing invoice images according to this application will be described in detail below. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.
[0054] Figure 2a This is a schematic flowchart illustrating an embodiment of a method for recognizing invoice images provided in this application. See also... Figure 2a The method specifically includes the following steps:
[0055] S201: Obtain the image of the ticket to be identified, and obtain the identifiers of multiple fields corresponding to the ticket image.
[0056] In this embodiment, the computing device can acquire an image of the ticket to be identified.
[0057] In one implementation:
[0058] The computing device can acquire the ticket image sent by the terminal device.
[0059] In one implementation:
[0060] The computing device can acquire the image to be processed. This image can be sent by the terminal device.
[0061] A computing device can segment an image to be processed to obtain multiple initial ticket images. In one implementation, the computing device can segment the image to be processed based on a segmentation model to obtain multiple initial ticket images. For example, Figure 2b This is a schematic diagram illustrating an image segmentation scenario provided in an embodiment of this application. For example... Figure 2b As shown, the computing device can perform segmentation processing on the image to be processed, resulting in 8 initial ticket images.
[0062] For any initial ticket image, the computing device can perform skew correction processing on the initial ticket image to obtain the ticket image. In one implementation, the computing device can perform skew correction processing on the initial ticket image based on a rotation model to obtain the ticket image.
[0063] In addition, the computing device can obtain multiple field identifiers (also known as keywords) corresponding to the ticket image.
[0064] In one implementation:
[0065] The computing device can store preset mapping relationships. These preset mapping relationships indicate the mapping relationship between document types and field identifiers. Understandably, one document type can correspond to multiple field identifiers.
[0066] The computing device can obtain the ticket type from the ticket image. In one implementation, the computing device can obtain the ticket type sent by the terminal device. The ticket type is input by the user into the terminal device.
[0067] The computing device can query the preset mapping relationship based on the ticket type to determine multiple field identifiers.
[0068] For example, a computing device can query the preset mapping relationship based on the invoice type being a value-added tax invoice, and determine multiple field identifiers including invoice number, invoice code, and amount.
[0069] S202: Based on the field identifier detection model, the ticket image is processed to identify multiple field identifiers and obtain the location information of multiple field identifiers.
[0070] In this embodiment, the computing device can perform recognition processing on the ticket image based on the field identifier detection model (also known as the keyword detection model) and multiple field identifiers (also known as keywords) to obtain the location information of multiple field identifiers.
[0071] It should be noted that the field identifier detection model can be obtained by training an untrained field identifier detection model based on multiple ticket image samples, multiple field identifier samples corresponding to each ticket image sample, and the position information of each field identifier sample.
[0072] It should also be noted that multiple ticket image samples can be samples with background textures or other interference conditions (such as stains).
[0073] In one implementation:
[0074] The computing device can perform recognition processing on the ticket image based on the field identifier detection model to obtain multiple text information and the positional information of each text information. It should be noted that the positional information of each text information refers to the range of its horizontal and vertical coordinates. It should also be noted that in one implementation, the field identifier detection model can be an Optical Character Recognition (OCR) model.
[0075] For any given field identifier, the computing device can calculate the similarity between each piece of text and that field identifier. Understandably, the computing device calculates the similarity between each piece of text and that field identifier based on a field identifier detection model.
[0076] The computing device can identify target text information from multiple text messages that has a similarity greater than a preset similarity to the field identifier. For example, the preset similarity can be 90%. In essence, the computing device uses a field identifier detection model to identify target text information from multiple text messages that has a similarity greater than a preset similarity to the field identifier.
[0077] After identifying the target text information, the computing device can determine the location information of the target text information as the location information of the field identifier. In other words, based on a field identifier detection model, the computing device determines the location information of the target text information as the location information of the field identifier after identifying the target text information.
[0078] S203: Based on the text block detection model, the ticket image is processed to obtain the position information of multiple text blocks.
[0079] In this embodiment, the computing device can perform recognition processing on the ticket image based on a text block detection model to obtain the position information of multiple text blocks. In one implementation, the position information of multiple text blocks can be represented as a binary image.
[0080] It should be noted that the text block can be a field identifier or field content.
[0081] Figure 2c This is a schematic diagram illustrating a scenario for recognizing text blocks, provided as an embodiment of this application. For example... Figure 2c As shown, the computing device can perform recognition processing on the ticket image based on the text block detection model, obtaining four text blocks. After recognizing multiple text blocks, the computing device can determine the position information of each text block based on the text block detection model.
[0082] S204: Determine the field content corresponding to the multiple field identifiers based on the position information of multiple text blocks and the position information of multiple field identifiers.
[0083] In this embodiment, the computing device can determine the field content corresponding to multiple field identifiers based on the position information of multiple text blocks and the position information of multiple field identifiers.
[0084] In one implementation:
[0085] The computing device can determine the position information of the field content corresponding to multiple field identifiers based on the relative position relationship model, according to the position information of multiple field identifiers and multiple text blocks.
[0086] The computing device can obtain the field content corresponding to multiple field identifiers based on the location information of the field content corresponding to multiple field identifiers.
[0087] S205: Determine the ticket recognition result corresponding to the ticket image, including multiple field identifiers and the field content corresponding to the multiple field identifiers.
[0088] In this embodiment, the computing device can determine that the ticket recognition result corresponding to the ticket image includes: multiple field identifiers and the field content corresponding to the multiple field identifiers.
[0089] In one implementation, the computing device can store and process the ticket recognition results corresponding to the ticket image.
[0090] In one implementation, the computing device can send the ticket recognition result corresponding to the ticket image to the terminal device.
[0091] The beneficial effects of this embodiment are as follows: In this embodiment, the computing device can acquire an image of a document to be identified. The computing device can acquire multiple field identifiers corresponding to the document image. The computing device can perform recognition processing on the document image based on a field identifier detection model, according to the multiple field identifiers, to obtain the position information of the multiple field identifiers. The computing device can perform recognition processing on the document image based on a text block detection model, to obtain the position information of multiple text blocks; wherein, the text blocks are field identifiers or field content. The computing device can determine the field content corresponding to the multiple field identifiers based on the position information of the multiple text blocks and the position information of the multiple field identifiers. The computing device can determine that the document recognition result corresponding to the document image includes multiple field identifiers and the field content corresponding to the multiple field identifiers. Through the above methods, the recognition accuracy of the document image is improved.
[0092] Figure 3 This is a flowchart illustrating a second embodiment of a method for recognizing invoice images provided in this application. See also... Figure 3 The method specifically includes the following steps:
[0093] S301: Obtain the image of the ticket to be identified, and obtain the identifiers of multiple fields corresponding to the ticket image.
[0094] In this embodiment, the computing device can acquire the image of the ticket to be identified and acquire multiple field identifiers corresponding to the ticket image.
[0095] The specific implementation process is the same as that of S201, and will not be described in detail here.
[0096] S302: Based on the field identifier detection model, the ticket image is processed to identify multiple field identifiers and obtain the location information of multiple field identifiers.
[0097] In this embodiment, the computing device can perform recognition processing on the ticket image based on the field identifier detection model and according to multiple field identifiers to obtain the location information of multiple field identifiers.
[0098] The specific implementation process is the same as that of S202, and will not be described in detail here.
[0099] S303: Based on the text block detection model, the ticket image is processed to obtain the position information of multiple text blocks.
[0100] In this embodiment, the computing device can perform recognition processing on the ticket image based on the text block detection model to obtain the position information of multiple text blocks.
[0101] The text block represents either the field identifier or the field content.
[0102] The specific implementation process is the same as that of S203, and will not be described in detail here.
[0103] S304: Based on the relative positional relationship model, determine the positional information of the field content corresponding to multiple field identifiers according to the positional information of multiple field identifiers and multiple text blocks.
[0104] In this embodiment, the computing device can determine the position information of the field content corresponding to multiple field identifiers based on the position information of multiple field identifiers and multiple text blocks, according to the relative position relationship model. In one implementation, the relative position relationship model is obtained by the computing device through training an untrained relative position relationship model based on the position information of multiple field identifier samples, the position information of multiple text block samples, and the position information of the field content samples corresponding to the multiple field identifier samples.
[0105] S305: Obtain the field content corresponding to multiple field identifiers based on the location information of the field content corresponding to multiple field identifiers.
[0106] In this embodiment, the computing device can obtain the field content corresponding to multiple field identifiers based on the location information of the field content corresponding to multiple field identifiers.
[0107] In one implementation:
[0108] The computing device can split the ticket image based on the location information of the field content corresponding to multiple field identifiers to obtain sub-ticket images corresponding to each field identifier.
[0109] For any field identifier, the computing device can perform recognition processing on the sub-document image corresponding to the field identifier to obtain the field content corresponding to the field identifier. In one implementation, for any field identifier, the computing device can perform recognition processing on the sub-document image corresponding to the field identifier based on a recognition model to obtain the field content corresponding to the field identifier.
[0110] For example, regarding the field identifier—issue status—the computing device can segment the invoice image based on the location information of the field content corresponding to the field identifier (horizontal coordinate range: (90-92), vertical coordinate range: (5-10)) to obtain the sub-invoice image corresponding to the field identifier. The computing device can then perform recognition processing on the sub-invoice image corresponding to the field identifier based on a recognition model (also known as an OCR model) to obtain the field content corresponding to the field identifier.
[0111] In one implementation:
[0112] The computing device can perform recognition processing on a document image to obtain multiple field contents and determine the location information of each field. In one implementation, the computing device can perform recognition processing on the document image based on a recognition model (also known as an OCR model) to obtain multiple field contents. For any given field, the computing device can determine the location information of that field after identifying it.
[0113] For any given field identifier, the computing device can determine the field content corresponding to the field identifier from multiple field contents based on the location information of the field content corresponding to the field identifier, as well as the location information of multiple field contents. In one implementation, for any given field identifier, the computing device can determine the location information of the field content corresponding to the field identifier from the location information of multiple field contents as the target location information. The computing device can then determine the field content corresponding to the target location information as the field content corresponding to the field identifier.
[0114] S306: The ticket recognition result corresponding to the ticket image is determined to include multiple field identifiers and the field content corresponding to the multiple field identifiers.
[0115] In this embodiment, the computing device can determine that the ticket recognition result corresponding to the ticket image includes multiple field identifiers and the field content corresponding to the multiple field identifiers.
[0116] The beneficial effects of this embodiment are as follows: In this embodiment, the computing device can acquire the image of the ticket to be identified. The computing device can acquire multiple field identifiers corresponding to the ticket image. The computing device can perform recognition processing on the ticket image based on the field identifier detection model, according to the multiple field identifiers, to obtain the position information of the multiple field identifiers. The computing device can perform recognition processing on the ticket image based on the text block detection model, to obtain the position information of multiple text blocks; wherein, the text blocks are field identifiers or field content. The computing device can determine the position information of the field content corresponding to the multiple field identifiers based on the relative positional relationship model, according to the position information of the multiple field identifiers and the position information of the multiple text blocks. The computing device can obtain the field content corresponding to the multiple field identifiers based on the position information of the field content corresponding to the multiple field identifiers. The computing device can determine that the ticket recognition result corresponding to the ticket image includes multiple field identifiers and the field content corresponding to the multiple field identifiers. Through the above method, even when the ticket type (ticket format) of the ticket image changes, the computing device can still accurately locate the position information of the field content corresponding to the multiple field identifiers based on the relative positional relationship model, and thus accurately determine the field content corresponding to each field identifier, improving the recognition accuracy of the ticket image.
[0117] Figure 4 This is a flowchart illustrating a third embodiment of a method for recognizing invoice images provided in this application. See also... Figure 4 The method specifically includes the following steps:
[0118] S401: Obtain the image of the ticket to be identified, and obtain the identifiers of multiple fields corresponding to the ticket image.
[0119] In this embodiment, the computing device can acquire the image of the ticket to be identified and acquire multiple field identifiers corresponding to the ticket image.
[0120] The specific implementation process is the same as that of S201, and will not be described in detail here.
[0121] S402: Based on the field identifier detection model, the ticket image is processed to identify multiple field identifiers and obtain the location information of multiple field identifiers.
[0122] In this embodiment, the computing device can perform recognition processing on the ticket image based on the field identifier detection model and according to multiple field identifiers to obtain the location information of multiple field identifiers.
[0123] The specific implementation process is the same as that of S202, and will not be described in detail here.
[0124] S403: Based on the text block detection model, the ticket image is processed to obtain the position information of multiple text blocks.
[0125] In this embodiment, the computing device can perform recognition processing on the ticket image based on the text block detection model to obtain the position information of multiple text blocks.
[0126] The text block represents either the field identifier or the field content.
[0127] The specific implementation process is the same as that of S203, and will not be described in detail here.
[0128] S404: Based on the relative positional relationship model, determine the positional information of the field content corresponding to multiple field identifiers according to the positional information of multiple field identifiers and multiple text blocks.
[0129] In this embodiment, the computing device can determine the position information of the field content corresponding to multiple field identifiers based on the relative position relationship model, according to the position information of multiple field identifiers and multiple text blocks.
[0130] Its specific implementation process is the same as S304. It will not be described in detail here.
[0131] S405: Based on the location information of the field content corresponding to multiple field identifiers, obtain the field content corresponding to multiple field identifiers.
[0132] In this embodiment, the computing device can obtain the field content corresponding to multiple field identifiers based on the location information of the field content corresponding to multiple field identifiers.
[0133] The specific implementation process is the same as that of S305. It will not be described in detail here.
[0134] S406: For any field identifier, retrieve the field content format corresponding to the field identifier.
[0135] In this embodiment, the computing device can store the field content format corresponding to each field identifier.
[0136] For any field identifier, the computing device can obtain the field content format corresponding to the field identifier. For example, for the field identifier "ID number", the computing device can obtain the field content format corresponding to the field identifier as "18 digits or 17 digits plus 1 uppercase letter X".
[0137] S407: If the field content corresponding to the field identifier is determined to be in accordance with the format of the field content corresponding to the field identifier, the bill result shall include the field identifier and the field content corresponding to the field identifier.
[0138] In this embodiment, the computing device can determine whether the field content corresponding to the field identifier conforms to the field content format corresponding to the field identifier.
[0139] The computing device can determine the bill result, including the field identifier and the field content corresponding to the field identifier, if the field content conforms to the format of the field content corresponding to the field identifier.
[0140] In one implementation, if the computing device determines that the field content corresponding to the field identifier does not conform to the format of the field content corresponding to the field identifier, it can send the field identifier and the corresponding field content to the terminal device. The terminal device can then display the field identifier and the corresponding field content. The terminal device can also obtain the correction field content corresponding to the field identifier input by the user and send it to the computing device. The computing device can determine that the ticket result includes the field identifier and the corresponding correction field content.
[0141] In one implementation, the computing device can determine that if the field content corresponding to a field identifier does not conform to the format of the field content corresponding to the field identifier, the field identifier will be identified as the target field identifier. The computing device can determine whether there exists any field content among the field content corresponding to multiple field identifiers that matches the format of the field content corresponding to the target field identifier. If so, the field content matching the format of the field content corresponding to the target field identifier will be identified as the target field content corresponding to the target field identifier. The computing device can determine that the document result includes the target field identifier and the target field content corresponding to the target field identifier.
[0142] The beneficial effects of this embodiment are as follows: In this embodiment, the computing device can acquire an image of a ticket to be identified and obtain multiple field identifiers corresponding to the ticket image. The computing device can perform recognition processing on the ticket image based on a field identifier detection model, according to the multiple field identifiers, to obtain the position information of the multiple field identifiers. The computing device can perform recognition processing on the ticket image based on a text block detection model, to obtain the position information of multiple text blocks; wherein, the text blocks are field identifiers or field content. The computing device can determine the position information of the field content corresponding to the multiple field identifiers based on a relative positional relationship model, according to the position information of the multiple field identifiers and the position information of the multiple text blocks. The computing device can obtain the field content corresponding to the multiple field identifiers based on the position information of the field content corresponding to the multiple field identifiers. For any field identifier, the computing device can obtain the format of the field content corresponding to the field identifier. If the field content corresponding to the field identifier conforms to the format of the field content corresponding to the field identifier, the computing device can determine that the ticket result includes the field identifier and the field content corresponding to the field identifier. By verifying the field content corresponding to the field identifier based on the above-mentioned format of the field content corresponding to the field identifier, the accuracy of determining the ticket recognition result is improved.
[0143] The following are embodiments of the apparatus described in this application, which can be used to execute the embodiments of the method described in this application. For details not disclosed in the apparatus embodiments of this application, please refer to the embodiments of the method described in this application.
[0144] Figure 5 This is a schematic diagram of the structure of a document image recognition device provided in this application. Figure 5 As shown, the document image recognition device 50 includes an acquisition module 51 and a processing module 52.
[0145] The acquisition module 51 is used to acquire the image of the ticket to be identified and to acquire multiple field identifiers corresponding to the ticket image;
[0146] Processing module 52 is used to perform recognition processing on the ticket image based on the field identifier detection model and according to multiple field identifiers to obtain the position information of multiple field identifiers;
[0147] The processing module 52 is also used to perform recognition processing on the ticket image based on the text block detection model to obtain the position information of multiple text blocks; wherein, the text block is a field identifier or field content;
[0148] The processing module 52 is also used to determine the field content corresponding to the multiple field identifiers based on the position information of the multiple text blocks and the position information of the multiple field identifiers;
[0149] The processing module 52 is also used to determine the ticket recognition result corresponding to the ticket image, which includes multiple field identifiers and the field content corresponding to the multiple field identifiers.
[0150] The document image recognition device provided in this application embodiment can execute the technical solution shown in the above method embodiment, where the execution subject is a computing device. Its implementation principle and beneficial effects are similar, and will not be repeated here.
[0151] In one implementation, the processing module 52 is specifically used for:
[0152] Based on the relative positional relationship model, the positional information of the field content corresponding to the multiple field identifiers is determined according to the positional information of multiple field identifiers and multiple text blocks;
[0153] Based on the location information of the field content corresponding to multiple field identifiers, obtain the field content corresponding to multiple field identifiers.
[0154] The document image recognition device provided in this application embodiment can execute the technical solution shown in the above method embodiment, where the execution subject is a computing device. Its implementation principle and beneficial effects are similar, and will not be repeated here.
[0155] In one implementation, the processing module 52 is specifically used for:
[0156] Based on the location information of the field content corresponding to multiple field identifiers, the ticket image is split to obtain sub-ticket images corresponding to each field identifier;
[0157] For any field identifier, the corresponding sub-document image is recognized and processed to obtain the field content corresponding to the field identifier.
[0158] The document image recognition device provided in this application embodiment can execute the technical solution shown in the above method embodiment, where the execution subject is a computing device. Its implementation principle and beneficial effects are similar, and will not be repeated here.
[0159] In one implementation, the processing module 52 is specifically used for:
[0160] The document image is processed to obtain multiple fields and the location information of each field is determined.
[0161] For any field identifier, determine the field content corresponding to the field identifier from multiple field contents based on the location information of the field content corresponding to the field identifier and the location information of multiple field contents.
[0162] The document image recognition device provided in this application embodiment can execute the technical solution shown in the above method embodiment, where the execution subject is a computing device. Its implementation principle and beneficial effects are similar, and will not be repeated here.
[0163] In one implementation, the processing module 52 is specifically used for:
[0164] Based on the field identifier detection model, the ticket image is processed to obtain multiple text information and the position information of each text information;
[0165] For any given field identifier, the text information with a similarity greater than a preset similarity among multiple text information is identified as the target text information;
[0166] The location information of the target text information is determined as the location information of the field identifier.
[0167] The document image recognition device provided in this application embodiment can execute the technical solution shown in the above method embodiment, where the execution subject is a computing device. Its implementation principle and beneficial effects are similar, and will not be repeated here.
[0168] In one implementation, the processing module 52 is specifically used for:
[0169] Obtain the ticket type from the ticket image;
[0170] Based on the invoice type, query the preset mapping relationship to determine multiple field identifiers; whereby the preset mapping relationship indicates the mapping relationship between the invoice type and the field identifiers.
[0171] The document image recognition device provided in this application embodiment can execute the technical solution shown in the above method embodiment, where the execution subject is a computing device. Its implementation principle and beneficial effects are similar, and will not be repeated here.
[0172] In one implementation, module 51 is specifically used for:
[0173] Obtain the image to be processed;
[0174] The image to be processed is segmented to obtain multiple initial ticket images;
[0175] The initial ticket image is corrected to obtain the ticket image.
[0176] The document image recognition device provided in this application embodiment can execute the technical solution shown in the above method embodiment, where the execution subject is a computing device. Its implementation principle and beneficial effects are similar, and will not be repeated here.
[0177] In one implementation, the processing module 52 is specifically used for:
[0178] For any field identifier, retrieve the field content format corresponding to the field identifier;
[0179] If the field content corresponding to the field identifier is determined to be in accordance with the format of the field content corresponding to the field identifier, the bill result is determined to include the field identifier and the field content corresponding to the field identifier.
[0180] The document image recognition device provided in this application embodiment can execute the technical solution shown in the above method embodiment, where the execution subject is a computing device. Its implementation principle and beneficial effects are similar, and will not be repeated here.
[0181] Figure 6 This is a schematic diagram of the structure of a computing device provided in this application. Figure 6 As shown, the computing device 60 includes a processor 61 and a memory 62; wherein the processor 61 is communicatively connected to the memory 62, and the memory 62 is used to store computer execution instructions; the processor 61 is configured to execute the technical solutions in any of the foregoing method embodiments by executing the computer execution instructions stored in the memory 62.
[0182] Optionally, the memory 62 can be either standalone or integrated with the processor 61. Optionally, when the memory 62 is a device independent of the processor 61, the computing device 60 may further include a bus for connecting the aforementioned devices.
[0183] The computing device is used to execute the technical solutions in any of the foregoing method embodiments. Its implementation principle and technical effect are similar, and will not be described again here.
[0184] This application also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the technical solutions of any of the foregoing method embodiments.
[0185] This application also provides a computer program product, including a computer program, which, when executed by a processor, is used to implement the technical solutions of any of the foregoing method embodiments.
[0186] Those skilled in the art will understand that all or part of the steps of the above-described method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments; and the aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.
[0187] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.
Claims
1. A method for recognizing invoice images, characterized in that, include: Obtain the image of the ticket to be identified, and obtain the multiple field identifiers corresponding to the ticket image; Based on the field identifier detection model, the ticket image is processed for recognition according to the multiple field identifiers to obtain the location information of the multiple field identifiers; Based on the text block detection model, the ticket image is processed to obtain the position information of multiple text blocks; wherein, the text blocks are field identifiers or field content; Based on the position information of the multiple text blocks and the position information of the multiple field identifiers, determine the field content corresponding to the multiple field identifiers; The determination of the ticket recognition result corresponding to the ticket image includes the multiple field identifiers and the field content corresponding to the multiple field identifiers.
2. The method according to claim 1, characterized in that, The step of determining the field content corresponding to the multiple field identifiers based on the position information of the multiple text blocks and the position information of the multiple field identifiers includes: Based on the relative positional relationship model, the positional information of the field content corresponding to the multiple field identifiers is determined according to the positional information of the multiple field identifiers and the positional information of the multiple text blocks; Based on the location information of the field content corresponding to the multiple field identifiers, obtain the field content corresponding to the multiple field identifiers.
3. The method according to claim 2, characterized in that, The step of obtaining the field content corresponding to the multiple field identifiers based on the position information of the field content corresponding to the multiple field identifiers includes: Based on the location information of the field content corresponding to the multiple field identifiers, the ticket image is split to obtain sub-ticket images corresponding to each field identifier; For any field identifier, the sub-document image corresponding to the field identifier is identified and processed to obtain the field content corresponding to the field identifier.
4. The method according to claim 2, characterized in that, The step of obtaining the field content corresponding to the multiple field identifiers based on the position information of the field content corresponding to the multiple field identifiers includes: The ticket image is processed to obtain multiple field contents, and the location information of each field content is determined; For any field identifier, based on the location information of the field content corresponding to the field identifier and the location information of multiple field contents, the field content corresponding to the field identifier is determined from multiple field contents.
5. The method according to claim 1, characterized in that, The field identifier-based detection model performs recognition processing on the ticket image based on the multiple field identifiers to obtain the location information of the multiple field identifiers, including: Based on the field identifier detection model, the ticket image is processed to obtain multiple text information and the position information of each text information; For any field identifier, among multiple text messages, the text messages with a similarity greater than a preset similarity to the field identifier are identified as target text messages; The location information of the target text information is determined as the location information of the field identifier.
6. The method according to claim 1, characterized in that, The step of obtaining multiple field identifiers corresponding to the ticket image includes: Obtain the ticket type of the ticket image; Based on the bill type, a preset mapping relationship is queried to determine the multiple field identifiers; wherein, the preset mapping relationship indicates the mapping relationship between the bill type and the field identifiers.
7. The method according to claim 1, characterized in that, The process of obtaining the image of the ticket to be identified includes: Obtain the image to be processed; The image to be processed is segmented to obtain multiple initial ticket images; The initial ticket image is subjected to skew correction processing to obtain the ticket image.
8. The method according to claim 1, characterized in that, The determination of the ticket recognition result corresponding to the ticket image includes multiple field identifiers and the field content corresponding to the multiple field identifiers, including: For any field identifier, obtain the field content format corresponding to the field identifier; If the field content corresponding to the field identifier is determined to conform to the format of the field content corresponding to the field identifier, the bill result is determined to include the field identifier and the field content corresponding to the field identifier.
9. A device for recognizing ticket images, characterized in that, include: The acquisition module is used to acquire the image of the ticket to be identified and to acquire multiple field identifiers corresponding to the ticket image; The processing module is used to perform recognition processing on the ticket image based on the field identifier detection model, according to the multiple field identifiers, to obtain the location information of the multiple field identifiers; The processing module is further configured to perform recognition processing on the ticket image based on a text block detection model to obtain the position information of multiple text blocks; wherein, the text blocks are field identifiers or field content; The processing module is further configured to determine the field content corresponding to the multiple field identifiers based on the position information of the multiple text blocks and the position information of the multiple field identifiers; The processing module is further configured to determine that the ticket recognition result corresponding to the ticket image includes the plurality of field identifiers and the field content corresponding to the plurality of field identifiers.
10. A computing device, characterized in that, include: A processor, and a memory communicatively connected to the processor; The memory is used to store computer-executed instructions; The processor is configured to execute computer execution instructions stored in the memory to implement the method of any one of claims 1-8.
11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the method of any one of claims 1-8.
12. A computer program product, characterized in that, Includes a computer program, which, when executed by a processor, is used to implement the method of any one of claims 1-8.