Bill image processing method and device, storage medium and electronic equipment
By combining annotation and evaluation models, the system automatically annotates and verifies invoice images, solving the problem of low accuracy in manual annotation and improving the efficiency and reliability of invoice data management.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-02-13
- Publication Date
- 2026-04-03
AI Technical Summary
Existing technologies rely on manually set rules to annotate ticket images, which results in low annotation accuracy.
The annotation model automatically annotates ticket images based on prior knowledge and customized prompts, and uses an independent evaluation model for evaluation and verification, thus optimizing the annotation process.
It improves the accuracy and efficiency of document image annotation, reduces the need for manual intervention, and enhances the reliability of document data management.
Smart Images

Figure CN121786592A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of financial technology, and more specifically, to a method and apparatus for processing bill images, a storage medium, and an electronic device. Background Technology
[0002] In the daily operations of financial institutions, bill processing is a tedious yet crucial task. Currently, relevant bill processing methods mainly rely on manually setting rules to annotate bill images, which is not only time-consuming and error-prone but also inefficient and inaccurate.
[0003] Currently, there is no effective solution to the problem of low accuracy in labeling document images due to the reliance on manually set rules in related technologies. Summary of the Invention
[0004] The main objective of this application is to provide a method, apparatus, storage medium, and electronic device for processing invoice images, in order to solve the problem of low accuracy in labeling in related technologies that rely on manually set rules for labeling invoice images.
[0005] To achieve the above objectives, according to one aspect of this application, a method for processing invoice images is provided. The method includes: acquiring an invoice image to be labeled, prior knowledge corresponding to the invoice image to be labeled, and a first prompt word; labeling the invoice image to be labeled using a labeling model based on the invoice image to be labeled, the prior knowledge, and the first prompt word to obtain a first labeling result, wherein the labeling model is trained on a multimodal large model using a first sample dataset; acquiring a second prompt word, and evaluating the first labeling result using an evaluation model based on the invoice image to be labeled, the first labeling result, and the second prompt word to obtain a first evaluation result, wherein the evaluation model is trained on a multimodal large model using a second sample dataset; and determining a target labeling result corresponding to the invoice image to be labeled based on the first evaluation result.
[0006] Furthermore, the annotation model performs annotation processing on the invoice image to be annotated, prior knowledge, and the first prompt word to obtain the first annotation result, which includes: processing the invoice image to be annotated, prior knowledge, and the first prompt word through the modality fusion layer of the annotation model to obtain a first fused feature vector representation; processing the first fused feature vector representation through the feature extraction layer of the annotation model to obtain a first feature vector representation; processing the first feature vector representation through the information extraction layer of the annotation model to obtain the annotation information of the target field; and processing the annotation information of the target field through the output layer of the annotation model to obtain the first annotation result.
[0007] Furthermore, the evaluation model evaluates the first annotation result based on the image of the ticket to be annotated, the first annotation result, and the second prompt word to obtain the first evaluation result, which includes: processing the image of the ticket to be annotated, the first annotation result, and the second prompt word through the modal fusion layer of the evaluation model to obtain the second fusion feature vector representation; processing the second fusion feature vector representation through the evaluation logic layer of the evaluation model to obtain the initial evaluation result; and processing the initial evaluation result through the decision output layer of the evaluation model to obtain the first evaluation result.
[0008] Further, based on the first evaluation result, determining the target annotation result corresponding to the invoice image to be annotated includes: if the first evaluation result indicates that the first annotation result is unqualified, generating a third prompt word based on the first annotation result, the first evaluation result, the first prompt word, and prior knowledge; updating the first prompt word and prior knowledge using the third prompt word through the target large model to obtain the updated first prompt word and updated prior knowledge, wherein the target large model is obtained by training the initial large model using the third sample dataset; repeatedly executing the steps of annotating the invoice image to be annotated using the annotation model based on the invoice image to be annotated, the updated prior knowledge, and the updated first prompt word to obtain the second annotation result, and evaluating the second annotation result using the evaluation model based on the invoice image to be annotated, the second annotation result, and the second prompt word to obtain the second evaluation result, until the preset conditions are met, and the target annotation result is obtained.
[0009] Furthermore, the method also includes: if the first evaluation result indicates that the first annotation result is qualified, then the first annotation result is used as the target annotation result.
[0010] Furthermore, before acquiring the image of the invoice to be labeled, the prior knowledge corresponding to the image of the invoice to be labeled, and the first prompt word, the method also includes: performing optical character recognition processing on the image of the invoice to be labeled to obtain the text information corresponding to the image of the invoice to be labeled; and performing format conversion on the text information to obtain the prior knowledge.
[0011] Furthermore, before obtaining the image of the invoice to be labeled, the prior knowledge corresponding to the image of the invoice to be labeled, and the first prompt word, the method further includes: obtaining a prompt word template corresponding to the invoice type of the image of the invoice to be labeled; and filling the prompt word template according to the prior knowledge to obtain the first prompt word.
[0012] To achieve the above objectives, according to another aspect of this application, a device for processing a ticket image is provided. The device includes: a first acquisition unit, configured to acquire a ticket image to be labeled, prior knowledge corresponding to the ticket image to be labeled, and a first prompt word; a first processing unit, configured to perform labeling processing on the ticket image to be labeled based on the ticket image to be labeled, the prior knowledge, and the first prompt word using a labeling model to obtain a first labeling result, wherein the labeling model is trained on a multimodal large model using a first sample dataset; a second processing unit, configured to acquire a second prompt word and evaluate the first labeling result based on the ticket image to be labeled, the first labeling result, and the second prompt word using an evaluation model to obtain a first evaluation result, wherein the evaluation model is trained on a multimodal large model using a second sample dataset; and a first determination unit, configured to determine the target labeling result corresponding to the ticket image to be labeled based on the first evaluation result.
[0013] Further, the first processing unit includes: a first processing subunit, used to process the image of the ticket to be labeled, prior knowledge, and the first prompt word through the modality fusion layer of the annotation model to obtain a first fused feature vector representation; a second processing subunit, used to process the first fused feature vector representation through the feature extraction layer of the annotation model to obtain a first feature vector representation; a third processing subunit, used to process the first feature vector representation through the information extraction layer of the annotation model to obtain the annotation information of the target field; and a fourth processing subunit, used to process the annotation information of the target field through the output layer of the annotation model to obtain a first annotation result.
[0014] Furthermore, the second processing unit includes: a fifth processing subunit, used to process the image of the ticket to be labeled, the first labeling result, and the second prompt word through the modal fusion layer of the evaluation model to obtain a second fused feature vector representation; a sixth processing subunit, used to process the second fused feature vector representation through the evaluation logic layer of the evaluation model to obtain an initial evaluation result; and a seventh processing subunit, used to process the initial evaluation result through the decision output layer of the evaluation model to obtain a first evaluation result.
[0015] Further, the first determining unit includes: an eighth processing subunit, used to generate a third prompt word based on the first annotation result, the first evaluation result, the first prompt word, and prior knowledge when the first evaluation result indicates that the first annotation result is unqualified; a ninth processing subunit, used to update the first prompt word and prior knowledge based on the third prompt word using the target large model, to obtain the updated first prompt word and the updated prior knowledge, wherein the target large model is obtained by training the initial large model using the third sample dataset; and a determining subunit, used to repeatedly perform the steps of annotating the invoice image to be annotated using the annotation model based on the invoice image to be annotated, the updated prior knowledge, and the updated first prompt word to obtain a second annotation result, and evaluating the second annotation result using the evaluation model based on the invoice image to be annotated, the second annotation result, and the second prompt word to obtain a second evaluation result, until the preset conditions are met and the target annotation result is obtained.
[0016] Furthermore, the device also includes a second determining unit, used to take the first annotation result as the target annotation result if the first evaluation result indicates that the first annotation result is qualified.
[0017] Furthermore, the device also includes: a third processing unit, used to perform optical character recognition processing on the image of the ticket to be annotated before acquiring the image of the ticket to be annotated, the prior knowledge corresponding to the image of the ticket to be annotated, and the first prompt word, to obtain the text information corresponding to the image of the ticket to be annotated; and a fourth processing unit, used to perform format conversion on the text information to obtain the prior knowledge.
[0018] Furthermore, the device also includes: a second acquisition unit, used to acquire a prompt word template corresponding to the type of the ticket image to be labeled before acquiring the image of the ticket to be labeled, the prior knowledge corresponding to the image of the ticket to be labeled, and the first prompt word; and a fifth processing unit, used to fill the prompt word template according to the prior knowledge to obtain the first prompt word.
[0019] According to another aspect of the present invention, an electronic device is also provided, comprising: a memory storing an executable program; and a processor for running the program, wherein the program executes the document image processing method described above during runtime.
[0020] According to another aspect of the present invention, a computer-readable storage medium is also provided, the storage medium storing a program, wherein, when the program is running, the device where the storage medium is located executes the document image processing method described above.
[0021] In this embodiment, the following steps are employed: obtaining the image of the invoice to be labeled, the prior knowledge corresponding to the image of the invoice to be labeled, and a first prompt word; using a labeling model to label the image of the invoice to be labeled based on the image of the invoice to be labeled, the prior knowledge, and the first prompt word to obtain a first labeling result, wherein the labeling model is trained on a multimodal large model using a first sample dataset; obtaining a second prompt word, and using an evaluation model to evaluate the first labeling result based on the image of the invoice to be labeled, the first labeling result, and the second prompt word to obtain a first evaluation result, wherein the evaluation model is trained on a multimodal large model using a second sample dataset; and determining the target labeling result corresponding to the image of the invoice to be labeled based on the first evaluation result. This solves the technical problem in related technologies where labeling invoice images relies on manually set rules, resulting in low labeling accuracy. In this solution, the invoice image is automatically labeled by a labeling model based on prior knowledge and customized prompt words, and then evaluated and verified by another independently trained evaluation model. This effectively reduces the need for manual intervention, improves labeling accuracy and efficiency, and thus enhances the efficiency and reliability of invoice data management. Attached Figure Description
[0022] The accompanying drawings, which form part of this application, are used to provide a further understanding of this application. The illustrative embodiments and descriptions of this application are used to explain this application and do not constitute an undue limitation of this application. In the drawings:
[0023] Figure 1 A hardware block diagram of a computer terminal for implementing a method for processing ticket images is shown.
[0024] Figure 2 This is a flowchart of a method for processing invoice images according to an embodiment of this application;
[0025] Figure 3 This is a schematic diagram of the document image annotation process provided in the embodiments of this application;
[0026] Figure 4 This is a schematic diagram of a bill image processing apparatus provided according to an embodiment of this application;
[0027] Figure 5 This is a structural block diagram of an electronic device according to an embodiment of this application. Detailed Implementation
[0028] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.
[0029] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0030] It should be noted that the information collected in this application (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for display, data used for analysis, etc.) are information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, storage, use, processing, transmission, provision, disclosure, and application of this data all comply with relevant laws, regulations, and standards, necessary confidentiality measures have been taken, and they do not violate public order and good morals. Corresponding access points are provided for users to choose to authorize or refuse. For example, interfaces are set up between this system and relevant users or organizations, providing users with corresponding access points to choose to agree to or refuse automated decision-making results; if the user chooses to refuse, the process proceeds to the expert decision-making stage.
[0031] Example 1
[0032] According to an embodiment of this application, a method embodiment for processing a ticket image is also provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0033] The method embodiment provided in Embodiment 1 of this application can be executed on a mobile terminal, computer terminal, or similar computing device. Figure 1A hardware block diagram of a computer terminal (or mobile device) for implementing a method for processing ticket images is shown. Figure 1 As shown, the computer terminal 10 (or mobile device) may include one or more processors 102 (shown as 102a, 102b, ..., 102n in the figure) 102 (processor 102 may include, but is not limited to, a microprocessor MCU or a programmable logic device FPGA, etc.), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of a BUS bus), a network interface, a power supply, and / or a camera. Those skilled in the art will understand that... Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the aforementioned electronic device. For example, computer terminal 10 may also include... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.
[0034] It should be noted that the aforementioned one or more processors 102 and / or other data processing circuits are generally referred to herein as "data processing circuits". These data processing circuits may be embodied, in whole or in part, in software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuits may be a single, independent processing module, or may be integrated, in whole or in part, into any other element within the computer terminal 10 (or mobile device). As involved in the embodiments of this application, the data processing circuits serve as a processor control mechanism (e.g., selection of a variable resistor termination path connected to an interface).
[0035] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the document image processing method in this embodiment. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, thereby implementing the aforementioned document image processing method. The memory 104 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to the computer terminal 10 via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0036] The transmission device 106 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the communication provider of the computer terminal 10. In one example, the transmission device 106 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 106 may be a Radio Frequency (RF) module, used for wireless communication with the Internet.
[0037] The display may be a touchscreen liquid crystal display (LCD) that allows the user to interact with the user interface of the computer terminal 10 (or mobile device).
[0038] Under the aforementioned operating environment, this application provides the following: Figure 2 The method for processing the ticket image shown. Figure 2 This is a flowchart of a method for processing a ticket image according to Embodiment 1 of this application. The method for processing the ticket image includes:
[0039] Step S201: Obtain the image of the ticket to be labeled, the prior knowledge corresponding to the image of the ticket to be labeled, and the first prompt word;
[0040] Step S202: The annotation model is used to annotate the ticket image to be annotated based on the prior knowledge and the first prompt word to obtain the first annotation result. The annotation model is trained on a multimodal large model using the first sample dataset.
[0041] Step S203: Obtain the second prompt word, and evaluate the first annotation result based on the image of the ticket to be annotated, the first annotation result, and the second prompt word using the evaluation model to obtain the first evaluation result. The evaluation model is obtained by training a multimodal large model using the second sample dataset.
[0042] Step S204: Based on the first evaluation result, determine the target annotation result corresponding to the invoice image to be annotated.
[0043] Optionally, the document image processing system first acquires the document image to be labeled, the prior knowledge corresponding to the document image to be labeled, and the first prompt word. For example, a high-resolution image of a business authorization letter is acquired using a high-resolution scanner or camera as the document image to be labeled; the text in the document image to be labeled obtained through optical character recognition is used as prior knowledge; based on the prior knowledge, a preset prompt word template is filled to obtain the first prompt word, which is used to guide the multimodal large model to perform specific information extraction tasks, including but not limited to fields such as business type, voucher number, and authorization date.
[0044] Optionally, the multimodal large model is pre-trained using labeled ticket images (i.e., the first sample dataset) to enable it to understand and label all key fields on the tickets. The ticket images to be labeled, prior knowledge, and the first prompt words are input into the trained labeling model for labeling processing. The model outputs the first labeling result, such as a lightweight data exchange format file containing information on all key fields.
[0045] Optionally, the generated annotation results are populated into another preset prompt word template to construct a second prompt word, which guides the multimodal large model to check the correctness of the first annotation result. Using another dataset of ticket images containing annotation results and their evaluation labels (i.e., the second sample dataset), a multimodal large model is trained to perform quality evaluation based on the prompt word, ticket image, and annotation results. The ticket image to be annotated, the first annotation result, and the second prompt word are input into the trained evaluation model for evaluation. The model outputs a first evaluation result, indicating the quality of the first annotation result, such as "True" indicating the result is correct and requires no modification; and "False" indicating an error exists and further correction is needed.
[0046] Optionally, based on the first evaluation result, the target annotation result corresponding to the ticket image to be annotated is determined. For example, if the first evaluation result is "True", the first annotation result is directly used as the target annotation result; if the first evaluation result is "False", the first prompt word and prior knowledge are updated, and the automatic annotation and evaluation process is performed again until the evaluation result is "True" or the preset loop count limit is reached, and the annotation result after the end of the previous loop is used as the final target annotation result.
[0047] In an alternative embodiment, the following can be employed: Figure 3 The diagram shown illustrates how to automatically annotate invoice images. Figure 3 This is a schematic diagram of the document image annotation process provided in the embodiments of this application, such as... Figure 3As shown, the process mainly includes the following: Prior knowledge of the invoice image to be labeled is obtained through optical character recognition. This prior knowledge, the invoice image, and prompt word 1 are then input into a multimodal large model, which outputs the labeling results, such as field labels in JSON format. Next, the labeling results, the invoice image, and prompt word 2 are input into another multimodal large model for review and evaluation. If the results are satisfactory, the model outputs the labeling results. If not, the evaluation results, suggestions, and prompt word 3 are input into a large language model to update prompt word 1 and the prior knowledge. Then, the updated prompt word 1, the updated prior knowledge, and the invoice image are again input into the multimodal large model for automatic labeling and evaluation. The loop ends when a satisfactory labeling result is output. Alternatively, if the set number of loops is reached or resources are exhausted, the labeling result from the previous loop is used as the final labeling result.
[0048] In an optional embodiment, taking the automatic annotation of a business authorization form as an example, an example of prompt word 1 is shown below:
[0049] The given optical character recognition (OCR) information is: [Prior information obtained from OCR]. You are a multimodal information extraction expert. The given prior information was extracted from the given image using an OCR tool. Utilizing your powerful multimodal capabilities, combine the given prior information and the image to extract specific field information from the image with fine-grained precision. Please note the following during extraction:
[0050] 1. The output should be in the standard JSON key-value pair format. Note that single quotes should not be used.
[0051] 2. Fields to be extracted: Business type, voucher number, entrustment date, full name of the entruster, full name of the payee, account number or address of the entruster, account number or address of the payee, name of the payee's bank, amount in words, amount in figures, company seal, purpose, and bank information.
[0052] 3. The business type field contains multiple checkboxes, but only one checkbox will be selected when checked. The selected checkbox contains the information for that field.
[0053] 4. The given prior information does not include signature information. You need to carefully identify and retain the company or organization name in the signature.
[0054] 5. Bank fields contain multiple checkboxes. Checkboxes may be selected or not. When a checkbox is not selected, the field information is empty. When a checkbox is selected, only one checkbox will be selected, and the selected checkbox contains the information for that field.
[0055] 6. For the voucher number, only the last few digits need to be retained.
[0056] 7. For amounts in lowercase, the actual amount should be rounded to two decimal places.
[0057] In an optional embodiment, the prompt word 2 used to evaluate whether the annotation result is qualified is exemplified as follows:
[0058] Annotation Results: [Generated Annotation Results] You are a multimodal information evaluation expert. Please evaluate and verify the accuracy and quality of the annotation results based on the given annotation results and the following precautions. Output True if qualified, and False if unqualified. Please note the following during extraction:
[0059] 1. The output should be in the standard JSON key-value pair format. Note that single quotes should not be used.
[0060] 2. Fields to be extracted: Business type, voucher number, entrustment date, full name of the entruster, full name of the payee, account number or address of the entruster, account number or address of the payee, name of the payee's bank, amount in words, amount in figures, company seal, purpose, and bank information.
[0061] 3. The business type field contains multiple checkboxes, but only one checkbox will be selected when checked. The selected checkbox contains the information for that field.
[0062] 4. The given prior information does not include signature information. You need to carefully identify and retain the company or organization name in the signature.
[0063] 5. Bank fields contain multiple checkboxes. Checkboxes may be selected or not. When a checkbox is not selected, the field information is empty. When a checkbox is selected, only one checkbox will be selected, and the selected checkbox contains the information for that field.
[0064] 6. For the voucher number, only the last few digits need to be retained.
[0065] 7. For amounts in lowercase, the actual amount should be rounded to two decimal places.
[0066] In an optional embodiment, the prompt word 1 and the prior knowledge prompt word 3 are used for updating, as shown in the following example:
[0067] Annotation results: [Generated annotation results];
[0068] Evaluation results: [Generated evaluation results];
[0069] Prompt 1: [Prompt 1];
[0070] Prior knowledge: [Prior knowledge];
[0071] You are an expert in prompt word optimization and prior knowledge update. Based on the given annotation results, evaluation results, prompt word 1, and prior knowledge, first determine whether the annotation results are qualified. If qualified, output the annotation results directly. If not qualified, optimize prompt word 1 and update prior knowledge based on the annotation results and evaluation results, and output them.
[0072] In summary, this solution addresses the technical problem of low accuracy in labeling invoice images due to reliance on manually set rules in related technologies. Our solution automatically labels invoice images using a labeling model based on prior knowledge and customized prompts, and then evaluates and verifies the labeling using an independently trained evaluation model. This effectively reduces the need for manual intervention, improves labeling accuracy and efficiency, and thus enhances the efficiency and reliability of invoice data management.
[0073] Optionally, in the document image processing method provided in this application embodiment, the annotation model performs annotation processing on the document image to be annotated based on the document image to be annotated, prior knowledge, and a first prompt word to obtain a first annotation result, including: processing the document image to be annotated, prior knowledge, and the first prompt word through the modality fusion layer of the annotation model to obtain a first fused feature vector representation; processing the first fused feature vector representation through the feature extraction layer of the annotation model to obtain a first feature vector representation; processing the first feature vector representation through the information extraction layer of the annotation model to obtain the annotation information of the target field; and processing the annotation information of the target field through the output layer of the annotation model to obtain the first annotation result.
[0074] In an optional embodiment, the original invoice image to be labeled is preprocessed by the preprocessing layer of the labeling model, which performs size standardization, grayscale conversion, contrast enhancement, and noise reduction to obtain a preprocessed invoice image to be labeled. This preprocessed image is then input into the modality fusion layer, which integrates the image information and text information into a unified representation that the model can process, resulting in a first fused feature vector representation. For example, the modality fusion layer first uses a convolutional neural network to extract features from the invoice image to be labeled, converting the image into a set of visual feature vectors that reflect its content and structure. Then, it uses natural language processing technology to convert prior knowledge and the first prompt word into semantic feature vectors, capturing the inherent meaning and contextual relationships of the text. Finally, it fuses the image visual feature vectors with the text semantic feature vectors (e.g., by splicing, weighted summation, attention mechanism, or multimodal encoder) to obtain a fused feature vector representation.
[0075] In an optional embodiment, the first fused feature vector representation is analyzed in depth by the feature extraction layer of the annotation model to extract deep feature vectors (i.e., the first feature vector representation); the features are decoded by the information extraction layer of the annotation model using a sequence model or a specific field detection network, and converted into annotation information of a specific field to obtain the annotation information of the target field, wherein the target field may be business type, amount, etc.; the annotation information of the target field is formatted, error checked and corrected by the output layer of the annotation model, such as being converted into JSON format, to obtain the first annotation result.
[0076] The annotation model can effectively parse and understand complex ticket images. By combining prior knowledge with specific prompts, it can output high-quality annotation results, providing an accurate data foundation for subsequent evaluation and optimization steps.
[0077] Optionally, in the document image processing method provided in this application embodiment, the evaluation model evaluates the first annotation result based on the document image to be annotated, the first annotation result, and the second prompt word to obtain the first evaluation result, including: processing the document image to be annotated, the first annotation result, and the second prompt word through the modal fusion layer of the evaluation model to obtain a second fused feature vector representation; processing the second fused feature vector representation through the evaluation logic layer of the evaluation model to obtain an initial evaluation result; and processing the initial evaluation result through the decision output layer of the evaluation model to obtain the first evaluation result.
[0078] In an optional embodiment, the modal fusion layer of the evaluation model performs multimodal fusion of image data, text annotation results, and guidance prompts, transforming them into a form that the model can process. For example, JSON text is converted into an embedded representation and fused with image features in parallel or sequentially to obtain a second fused feature vector representation. The evaluation logic layer of the evaluation model uses trained evaluation logic to evaluate the fused features, determining whether the annotation results meet the expected standards and business rules. For example, by combining the evaluation criteria contained in the second prompts, a field-by-field comparative analysis is performed to identify the consistency between the annotation results and the image content, and an initial evaluation result is output. Then, the decision output layer of the evaluation model performs a threshold determination on the initial evaluation result to decide whether to accept the first annotation result. For example, a score higher than a certain value is considered acceptable, and vice versa, resulting in the first evaluation result. For example, the evaluation result is represented by a Boolean value (True / False). If it is True, it indicates that the annotation result is acceptable; if it is False, it indicates that the annotation result is unacceptable, and error categories or correction suggestions may be provided as feedback information.
[0079] The evaluation model enables the understanding and processing of fused multimodal information, and allows for refined evaluation based on specific business rules and standards, thus providing accurate guidance for the iterative optimization of the annotation process.
[0080] Optionally, in the method for processing invoice images provided in this application embodiment, determining the target annotation result corresponding to the invoice image to be annotated based on the first evaluation result includes: if the first evaluation result indicates that the first annotation result is unqualified, generating a third prompt word based on the first annotation result, the first evaluation result, the first prompt word, and prior knowledge; updating the first prompt word and prior knowledge using the third prompt word through a target large model to obtain an updated first prompt word and updated prior knowledge, wherein the target large model is obtained by training an initial large model using a third sample dataset; repeatedly executing the steps of annotating the invoice image to be annotated using an annotation model based on the invoice image to be annotated, the updated prior knowledge, and the updated first prompt word to obtain a second annotation result, and evaluating the second annotation result using an evaluation model based on the invoice image to be annotated, the second annotation result, and the second prompt word to obtain a second evaluation result, until a preset condition is met to obtain the target annotation result.
[0081] Optionally, in the method for processing ticket images provided in the embodiments of this application, the method further includes: if the first evaluation result indicates that the first annotation result is qualified, then the first annotation result is used as the target annotation result.
[0082] In an optional embodiment, if the first evaluation result indicates that the first annotation result is qualified, for example, if the first evaluation result is "True", then the first annotation result is directly used as the target annotation result.
[0083] In an optional embodiment, if the first evaluation result indicates that the first annotation result is unqualified, for example, if the first evaluation result is "False", then a third prompt word is generated based on the first annotation result, the first evaluation result, the first prompt word, and prior knowledge to guide the target large model to update the first prompt word and prior knowledge. The target large model updates the first prompt word and prior knowledge based on the third prompt word, resulting in updated first prompt words and updated prior knowledge. This model is trained using a third sample dataset (including sample ticket images, corresponding prior knowledge, correct annotation results, incorrect annotation results, evaluation results, and optimized prompt words and optimized prior knowledge instances) to optimize prompt words and update prior knowledge. For example, the target large model adjusts the expression of the first prompt word based on the input error feedback, and can also correct or supplement prior knowledge. Then, the automatic annotation and evaluation process is performed again until preset conditions are met, such as an evaluation result of "True", or the preset maximum number of iterations is reached. The annotation result after the previous iteration is taken as the final target annotation result.
[0084] Through iterative optimization mechanisms, the annotation accuracy of multimodal large models has been improved, thereby enhancing overall work efficiency and data quality.
[0085] Optionally, in the method for processing invoice images provided in this application embodiment, before obtaining the invoice image to be annotated, the prior knowledge corresponding to the invoice image to be annotated, and the first prompt word, the method further includes: performing optical character recognition processing on the invoice image to be annotated to obtain text information corresponding to the invoice image to be annotated; and performing format conversion on the text information to obtain prior knowledge.
[0086] In an optional embodiment, optical character recognition (OCR) processing is performed on the image of the invoice to be annotated, that is, the text information in the image is identified and extracted to obtain a preliminary text description. Then, based on business rules or historical data, the preliminary text description is transformed into a structured data form to form prior knowledge.
[0087] By combining prior knowledge, a data foundation is provided for subsequent intelligent labeling, enhancing the accuracy and efficiency of the model in identifying key fields of invoices.
[0088] Optionally, in the method for processing invoice images provided in this application embodiment, before obtaining the invoice image to be annotated, the prior knowledge corresponding to the invoice image to be annotated, and the first prompt word, the method further includes: obtaining a prompt word template corresponding to the invoice type of the invoice image to be annotated; and filling the prompt word template according to the prior knowledge to obtain the first prompt word.
[0089] In an optional embodiment, the system first identifies the type of the document to be annotated, such as a check, a business authorization letter, or other types of documents. Then, it selects a template that matches the document type from a pre-prepared prompt word template library and fills the prompt word template with prior knowledge to obtain the first prompt word.
[0090] By using prompt word templates that match specific ticket types and personalizing them with prior knowledge, customized first prompt words can be generated, enhancing the relevance and accuracy of multimodal large models in automatic ticket image annotation tasks, thereby improving overall processing efficiency and annotation quality.
[0091] The method for processing invoice images provided in this application includes the following steps: obtaining the invoice image to be labeled, the prior knowledge corresponding to the invoice image to be labeled, and a first prompt word; labeling the invoice image to be labeled using a labeling model based on the invoice image to be labeled, the prior knowledge, and the first prompt word to obtain a first labeling result, wherein the labeling model is trained on a multimodal large model using a first sample dataset; obtaining a second prompt word, and evaluating the first labeling result using an evaluation model based on the invoice image to be labeled, the first labeling result, and the second prompt word to obtain a first evaluation result, wherein the evaluation model is trained on a multimodal large model using a second sample dataset; and determining the target labeling result corresponding to the invoice image to be labeled based on the first evaluation result. This solves the technical problem in related technologies where labeling invoice images relies on manually set rules, resulting in low labeling accuracy. In this solution, the invoice image is automatically labeled by a labeling model based on prior knowledge and customized prompt words, and then evaluated and verified by another independently trained evaluation model. This effectively reduces the need for manual intervention, improves labeling accuracy and efficiency, and thus enhances the efficiency and reliability of invoice data management.
[0092] It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.
[0093] Example 2
[0094] This application also provides a device for processing invoice images. It should be noted that this device can be used to execute the invoice image processing method provided in this application. The following describes the invoice image processing device provided in this application.
[0095] According to an embodiment of this application, a bill image processing apparatus for implementing the above-described bill image processing method is also provided, such as... Figure 4 As shown, the device includes: a first acquisition unit 401, a first processing unit 402, a second processing unit 403, and a first determination unit 404.
[0096] The first acquisition unit 401 is used to acquire the image of the ticket to be labeled, the prior knowledge corresponding to the image of the ticket to be labeled, and the first prompt word;
[0097] The first processing unit 402 is used to annotate the ticket image to be annotated based on the annotation model, prior knowledge and the first prompt word to obtain the first annotation result. The annotation model is obtained by training a multimodal large model using the first sample dataset.
[0098] The second processing unit 403 is used to obtain the second prompt word and evaluate the first annotation result based on the image of the ticket to be annotated, the first annotation result and the second prompt word through the evaluation model to obtain the first evaluation result. The evaluation model is obtained by training a multimodal large model using the second sample dataset.
[0099] The first determining unit 404 is used to determine the target annotation result corresponding to the invoice image to be annotated based on the first evaluation result.
[0100] The document image processing apparatus provided in this application embodiment acquires a document image to be labeled, prior knowledge corresponding to the document image to be labeled, and a first prompt word through a first acquisition unit 401; a first processing unit 402 performs labeling processing on the document image to be labeled based on the document image to be labeled, prior knowledge, and the first prompt word through a labeling model to obtain a first labeling result, wherein the labeling model is obtained by training a multimodal large model using a first sample dataset; a second processing unit 403 acquires a second prompt word and evaluates the first labeling result based on the document image to be labeled, the first labeling result, and the second prompt word through an evaluation model to obtain a first evaluation result, wherein the evaluation model is obtained by training a multimodal large model using a second sample dataset; and a first determination unit 404 determines the target labeling result corresponding to the document image to be labeled based on the first evaluation result.
[0101] Optionally, in the document image processing apparatus provided in this application embodiment, the first processing unit includes: a first processing subunit, used to process the document image to be annotated, prior knowledge, and a first prompt word through the modal fusion layer of the annotation model to obtain a first fused feature vector representation; a second processing subunit, used to process the first fused feature vector representation through the feature extraction layer of the annotation model to obtain a first feature vector representation; a third processing subunit, used to process the first feature vector representation through the information extraction layer of the annotation model to obtain annotation information of the target field; and a fourth processing subunit, used to process the annotation information of the target field through the output layer of the annotation model to obtain a first annotation result.
[0102] Optionally, in the document image processing apparatus provided in this application embodiment, the second processing unit includes: a fifth processing subunit, used to process the document image to be annotated, the first annotation result, and the second prompt word through the modal fusion layer of the evaluation model to obtain a second fused feature vector representation; a sixth processing subunit, used to process the second fused feature vector representation through the evaluation logic layer of the evaluation model to obtain an initial evaluation result; and a seventh processing subunit, used to process the initial evaluation result through the decision output layer of the evaluation model to obtain a first evaluation result.
[0103] Optionally, in the document image processing apparatus provided in this application embodiment, the first determining unit includes: an eighth processing subunit, used to generate a third prompt word based on the first annotation result, the first evaluation result, the first prompt word, and prior knowledge when the first evaluation result indicates that the first annotation result is unqualified; a ninth processing subunit, used to update the first prompt word and prior knowledge based on the third prompt word using a target large model to obtain an updated first prompt word and updated prior knowledge, wherein the target large model is obtained by training an initial large model using a third sample dataset; and a determining subunit, used to repeatedly execute the steps of annotating the document image to be annotated using an annotation model based on the document image to be annotated, the updated prior knowledge, and the updated first prompt word to obtain a second annotation result, and evaluating the second annotation result using an evaluation model based on the document image to be annotated, the second annotation result, and the second prompt word to obtain a second evaluation result, until a preset condition is met and a target annotation result is obtained.
[0104] Optionally, in the bill image processing apparatus provided in the embodiments of this application, the apparatus further includes: a second determining unit, used to take the first annotation result as the target annotation result when the first evaluation result indicates that the first annotation result is qualified.
[0105] Optionally, in the document image processing apparatus provided in this application embodiment, the apparatus further includes: a third processing unit, used to perform optical character recognition processing on the document image to be annotated before acquiring the document image to be annotated, the prior knowledge corresponding to the document image to be annotated, and the first prompt word, to obtain text information corresponding to the document image to be annotated; and a fourth processing unit, used to perform format conversion on the text information to obtain prior knowledge.
[0106] Optionally, in the document image processing apparatus provided in this application embodiment, the apparatus further includes: a second acquisition unit, configured to acquire a prompt word template corresponding to the document type of the document image to be labeled before acquiring the document image to be labeled, the prior knowledge corresponding to the document image to be labeled, and the first prompt word; and a fifth processing unit, configured to fill the prompt word template according to the prior knowledge to obtain the first prompt word.
[0107] It should be noted that the first acquisition unit 401, the first processing unit 402, the second processing unit 403, and the first determination unit 404 mentioned above correspond to steps S201 to S204 in Embodiment 1. The four units and the corresponding steps implement the same instances and application scenarios, but are not limited to the content disclosed in Embodiment 1. It should be noted that the above modules or units can be hardware or software components stored in memory (e.g., memory 104) and processed by one or more processors (e.g., processors 102a, 102b, ..., 102n). The above units can also be part of a device and run in the computer terminal 10 provided in Embodiment 1.
[0108] Example 3
[0109] Embodiments of this application may provide an electronic device. Figure 5 This is a structural block diagram of an electronic device according to an embodiment of this application. Figure 5 As shown, the electronic device may include: one or more ( Figure 5 (Only one is shown) processor 502, memory 504, memory controller, and peripheral interface, wherein the peripheral interface is connected to the radio frequency module, audio module and display.
[0110] The memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the methods and apparatus in the embodiments of this application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, thereby implementing the above-described methods. The memory may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include memory remotely located relative to the processor, and these remote memories can be connected to the terminal via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0111] The processor can access information and applications stored in the memory via a transmission device to perform the following steps: acquiring the image of the ticket to be labeled, the prior knowledge corresponding to the image of the ticket to be labeled, and the first prompt word; labeling the image of the ticket to be labeled using a labeling model based on the image of the ticket to be labeled, the prior knowledge, and the first prompt word to obtain a first labeling result, wherein the labeling model is trained on a multimodal large model using a first sample dataset; acquiring a second prompt word, and evaluating the first labeling result using an evaluation model based on the image of the ticket to be labeled, the first labeling result, and the second prompt word to obtain a first evaluation result, wherein the evaluation model is trained on a multimodal large model using a second sample dataset; and determining the target labeling result corresponding to the image of the ticket to be labeled based on the first evaluation result.
[0112] The processor can access the information and application program stored in the memory via the transmission device to perform the following steps: processing the image of the ticket to be labeled, prior knowledge, and the first prompt word through the modal fusion layer of the annotation model to obtain a first fused feature vector representation; processing the first fused feature vector representation through the feature extraction layer of the annotation model to obtain a first feature vector representation; processing the first feature vector representation through the information extraction layer of the annotation model to obtain the annotation information of the target field; and processing the annotation information of the target field through the output layer of the annotation model to obtain a first annotation result.
[0113] The processor can call the information and application program stored in the memory through the transmission device to perform the following steps: process the image of the ticket to be annotated, the first annotation result, and the second prompt word through the modal fusion layer of the evaluation model to obtain a second fused feature vector representation; process the second fused feature vector representation through the evaluation logic layer of the evaluation model to obtain an initial evaluation result; and process the initial evaluation result through the decision output layer of the evaluation model to obtain a first evaluation result.
[0114] The processor can access the information and application program stored in the memory via the transmission device to execute the following steps: If the first evaluation result indicates that the first annotation result is unqualified, generate a third prompt word based on the first annotation result, the first evaluation result, the first prompt word, and prior knowledge; update the first prompt word and prior knowledge using the third prompt word through the target large model to obtain the updated first prompt word and updated prior knowledge, wherein the target large model is trained on the initial large model using a third sample dataset; repeatedly execute the steps of annotating the invoice image to be annotated using the annotation model based on the image to be annotated, the updated prior knowledge, and the updated first prompt word to obtain a second annotation result, and evaluating the second annotation result using the evaluation model based on the image to be annotated, the second annotation result, and the second prompt word to obtain a second evaluation result, until the preset conditions are met, and the target annotation result is obtained.
[0115] The processor can call the information and application stored in the memory through the transmission device to perform the following steps: if the first evaluation result indicates that the first annotation result is qualified, the first annotation result is used as the target annotation result.
[0116] The processor can call the information and application stored in the memory through the transmission device to perform the following steps: before acquiring the image of the ticket to be labeled, the prior knowledge corresponding to the image of the ticket to be labeled, and the first prompt word, perform optical character recognition processing on the image of the ticket to be labeled to obtain the text information corresponding to the image of the ticket to be labeled; perform format conversion on the text information to obtain the prior knowledge.
[0117] The processor can call the information and application stored in the memory through the transmission device to perform the following steps: before acquiring the image of the ticket to be labeled, the prior knowledge corresponding to the image of the ticket to be labeled, and the first prompt word, acquire the prompt word template corresponding to the ticket type of the image of the ticket to be labeled; fill the prompt word template according to the prior knowledge to obtain the first prompt word.
[0118] Those skilled in the art will understand that Figure 5 The structure shown is for illustrative purposes only. Electronic devices can also be smartphones, tablets, handheld computers, mobile internet devices (MIDs), PADs, and other terminal devices. Figure 5 This does not limit the structure of the aforementioned electronic device. For example, electronic devices may also include components that are more... Figure 5 The more or fewer components shown (such as network interfaces, display devices, etc.), or having the same Figure 5 The different configurations shown.
[0119] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing the hardware related to the terminal device. The program can be stored in a computer-readable storage medium, which may include: flash drive, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0120] Example 4
[0121] Embodiments of this application also provide a computer-readable storage medium. Optionally, in this embodiment, the storage medium can be used to store the program code executed by the document image processing method provided in Embodiment 1.
[0122] Optionally, in this embodiment, the storage medium may be located in any computer terminal in a group of computer terminals in a computer network, or in any mobile terminal in a group of mobile terminals.
[0123] This application also provides a computer program product, which, when executed on a data processing device, is adapted to perform the steps of a method for processing a document image.
[0124] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0125] In the above embodiments of this application, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0126] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces; the indirect coupling or communication connection between units or modules may be electrical or other forms.
[0127] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0128] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0129] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard drive, magnetic disk, or optical disk.
[0130] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.
Claims
1. A method for processing a ticket image, characterized in that, include: Obtain the image of the invoice to be annotated, the prior knowledge corresponding to the image of the invoice to be annotated, and the first prompt word; The annotation model is used to annotate the invoice image to be annotated based on the invoice image to be annotated, the prior knowledge, and the first prompt word to obtain the first annotation result. The annotation model is obtained by training a multimodal large model using the first sample dataset. A second prompt word is obtained, and the first annotation result is evaluated by an evaluation model based on the image of the ticket to be annotated, the first annotation result, and the second prompt word to obtain a first evaluation result. The evaluation model is obtained by training a multimodal large model using a second sample dataset. Based on the first evaluation result, the target annotation result corresponding to the invoice image to be annotated is determined.
2. The method according to claim 1, characterized in that, The annotation model annotates the invoice image to be annotated based on the invoice image to be annotated, the prior knowledge, and the first prompt word, resulting in the following first annotation result: The modal fusion layer of the annotation model processes the image of the ticket to be annotated, the prior knowledge, and the first prompt word to obtain a first fused feature vector representation. The first fused feature vector representation is processed by the feature extraction layer of the labeled model to obtain the first feature vector representation; The first feature vector representation is processed by the information extraction layer of the annotation model to obtain the annotation information of the target field; The annotation information of the target field is processed by the output layer of the annotation model to obtain the first annotation result.
3. The method according to claim 1, characterized in that, The evaluation model evaluates the first annotation result based on the image of the ticket to be annotated, the first annotation result, and the second prompt word, resulting in a first evaluation result including: The evaluation model's modal fusion layer processes the image of the ticket to be labeled, the first labeling result, and the second prompt word to obtain a second fusion feature vector representation. The evaluation logic layer of the evaluation model processes the second fused feature vector representation to obtain the initial evaluation result; The initial evaluation result is processed by the decision output layer of the evaluation model to obtain the first evaluation result.
4. The method according to claim 1, characterized in that, Based on the first evaluation result, the target annotation result corresponding to the invoice image to be annotated is determined as follows: If the first evaluation result indicates that the first annotation result is unqualified, a third prompt word is generated based on the first annotation result, the first evaluation result, the first prompt word, and the prior knowledge. The target large model updates the first prompt word and the prior knowledge based on the third prompt word to obtain the updated first prompt word and the updated prior knowledge. The target large model is obtained by training the initial large model using the third sample dataset. The process of repeatedly annotating the invoice image to be annotated using the annotation model based on the invoice image to be annotated, the updated prior knowledge, and the updated first prompt word to obtain a second annotation result, and then evaluating the second annotation result using the evaluation model based on the invoice image to be annotated, the second annotation result, and the second prompt word to obtain a second evaluation result, is repeated until the preset conditions are met to obtain the target annotation result.
5. The method according to claim 4, characterized in that, The method further includes: If the first evaluation result indicates that the first annotation result is qualified, the first annotation result shall be used as the target annotation result.
6. The method according to claim 1, characterized in that, Before acquiring the image of the invoice to be labeled, the prior knowledge corresponding to the image of the invoice to be labeled, and the first prompt word, the method further includes: The image of the invoice to be annotated is subjected to optical character recognition processing to obtain the text information corresponding to the image of the invoice to be annotated. The text information is formatted to obtain the prior knowledge.
7. The method according to claim 1, characterized in that, Before acquiring the image of the invoice to be labeled, the prior knowledge corresponding to the image of the invoice to be labeled, and the first prompt word, the method further includes: Obtain the prompt word template corresponding to the type of the invoice in the image to be annotated; The prompt word template is filled in based on the prior knowledge to obtain the first prompt word.
8. A device for processing ticket images, characterized in that, include: The first acquisition unit is used to acquire the image of the invoice to be labeled, the prior knowledge corresponding to the image of the invoice to be labeled, and the first prompt word; The first processing unit is used to annotate the invoice image to be annotated based on the invoice image to be annotated, the prior knowledge, and the first prompt word using an annotation model to obtain a first annotation result, wherein the annotation model is obtained by training a multimodal large model using a first sample dataset; The second processing unit is used to obtain the second prompt word, and to evaluate the first annotation result based on the image of the ticket to be annotated, the first annotation result and the second prompt word through an evaluation model to obtain a first evaluation result. The evaluation model is obtained by training a multimodal large model using a second sample dataset. The first determining unit is used to determine the target annotation result corresponding to the invoice image to be annotated based on the first evaluation result.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored executable program, wherein, when the executable program is executed, it controls the device on which the computer-readable storage medium is located to perform the document image processing method according to any one of claims 1 to 7.
10. An electronic device, characterized in that, include: Memory, which stores executable programs; A processor for running the program, wherein the program, when running, performs the method for processing a ticket image according to any one of claims 1 to 7.