Bill voucher information extraction method, system and equipment based on multi-mode and OCR model fusion

By fusing multimodal and OCR models, the accuracy and robustness issues of bank voucher information extraction were resolved, enabling efficient and automated processing of complex bank vouchers.

CN120976947APending Publication Date: 2025-11-18ZHIWEI (SUZHOU) INFORMATION TECH CO LTD
View PDF 0 Cites 10 Cited by

Patent Information

Application Number
CN202510673056.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Traditional OCR systems struggle to adapt to non-standardized formats, mixed text and image layouts, and semantically ambiguous fields when processing bank vouchers, resulting in insufficient accuracy and robustness in information extraction.

Method used

By employing a multimodal and OCR model fusion approach, key fields can be accurately located and extracted through preprocessing, OCR recognition, multimodal model joint encoding, and error correction mechanisms.

Benefits of technology

It improves the accuracy and robustness of bank voucher information extraction, adapts to complex formats, reduces error rates, and enhances cross-type generalization performance and automation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120976947A_ABST
    Figure CN120976947A_ABST
Patent Text Reader

Abstract

The invention relates to a bill voucher information extraction method based on multi-mode and OCR model fusion. The method comprises the following steps: S1, obtaining an image of a bill voucher; s2, preprocessing the image; s3, identifying the preprocessed image by using an OCR engine to obtain the text content and the corresponding two-dimensional coordinates of each text block; s4, taking the recognized text segments and the original image as input, performing joint coding by using a pre-trained multi-modal model, evaluating and outputting the matching degree of each text segment and a predefined field category by the model, and determining candidate texts of each field and confidence of the candidate texts; s5, accurately positioning and extracting the key field, and verifying the consistency of the OCR output and the semantic result; s6, if the verification result conflicts or the identification reliability of a certain field is lower than a threshold value, error correction operation is carried out; and S7, outputting the structured bill voucher information. Through multi-modal fusion and iterative correction, the error rate of non-standard voucher information extraction is effectively reduced, and the method is suitable for various voucher formats and complex scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image recognition, and in particular to a method, system, and device for extracting invoice and voucher information based on the fusion of multimodal and OCR models. Background Technology

[0002] In banking processes, voucher images (such as remittance slips, receipts, bank acceptance drafts, receipts, and electronic bill printouts) are core paper or electronic documents. These vouchers typically need to be digitized and structured, with key fields (such as account name, amount, date, bank name, and transaction number) accurately extracted to facilitate automated approval, reconciliation, risk control, and archiving processes.

[0003] However, bank vouchers exhibit several prominent characteristics that make traditional OCR solutions ill-suited for information extraction. Highly non-standardized format: Bank vouchers originate from diverse sources, generated by different business systems, branches, or third-party systems, resulting in varying layout styles and field arrangements, making it difficult to establish a unified template. Mixed text and images with numerous interfering elements: Vouchers often contain seals, signatures, QR codes, barcodes, watermarks, horizontally and vertically intersecting text, background patterns, etc., all of which severely interfere with OCR. Ambiguous field semantics and lack of labels: Many vouchers do not explicitly label fields such as "amount" or "bank," instead directly displaying values ​​or using language (e.g., "Please make payment in XX year XX month"), making accurate extraction difficult solely through character matching. Therefore, while traditional OCR systems (such as Tesseract and PaddleOCR) perform well in ordinary text recognition, their adaptability, accuracy, and robustness are insufficient to meet the actual business needs in the unstructured scenario of bank vouchers.

[0004] In current banking operations, scanned vouchers often have inconsistent formats, complex layouts, and are mixed with text, seals, icons, etc. Traditional OCR technology is mainly effective for regular text and has significant limitations in processing non-standard voucher images: First, voucher formats are diverse, with inconsistent field positions, fonts, and layouts, making it difficult for OCR to directly extract target information; second, image quality varies (e.g., noise, blur, skew), often leading to recognition errors; third, vouchers often contain severe text-image mixing (e.g., barcodes, signatures, or seals), making it difficult for pure OCR models to understand contextual relationships and semantic logic, resulting in low field extraction accuracy. Existing solutions usually require manually defined rules or templates, which are difficult to handle the wide variety of non-standard format vouchers.

[0005] In summary, the existing technical solutions have the following drawbacks: 1. Rule / template-driven field extraction systems: These solutions involve manually configuring templates for different types of documents (i.e., fixed positions of fields in the image, format rules, keyword triggers, etc.). When the system identifies a document type, it extracts fields according to the preset template. For example, if a bank statement template is detected, the system will extract the value of the x-th row and y-th column region as the amount based on preset coordinates. Using a keyword matching strategy, such as recognizing the words "opening bank," the system will take the text to its right as the value. This method has high template maintenance costs, is difficult to generalize to a wide variety of new document types, and becomes ineffective due to field offsets or layout changes; it cannot handle documents without specified field names or fixed formats.

[0006] 2. Table Structure Detection + OCR Joint Recognition Scheme: This method combines table detection technology to first identify table lines and structure in the document, and then extracts fields based on row and column positions using the OCR output. Representative methods include first using table detection models (such as CascadeTabNet and TableNet) to detect table regions, then using text block grouping and row / column alignment techniques (such as projection histograms) to parse the structure, and finally mapping the meaning of fields based on the table header content and the OCR results. This approach is only suitable for well-structured table-based documents; it is ineffective for mixed text and image formats and non-standard tables, lacks language-level semantic understanding capabilities, and cannot handle semantically ambiguous fields.

[0007] 3. Some intelligent document parsing models based on visual document pre-training (such as LayoutLM, Donut, FormNet, etc.) are beginning to be used for structured document information extraction. These models input text, location, and visual features from images into a Transformer network to learn a joint image-text representation, and then perform field classification or question-answering extraction. Most of these models still rely on training with structured or semi-structured form documents; they lack adaptation to mixed image and text layouts, multimodal contextual understanding, and the semantic details of bank vouchers; and they remain quite sensitive to low-quality images and complex background interference. Summary of the Invention

[0008] The technical problem this invention aims to solve is to design a method, system, and device for extracting information from bank documents based on the fusion of multimodal and OCR models, thereby improving the accuracy and robustness of automatic extraction of bank document information. This addresses existing technical problems.

[0009] To address the aforementioned technical problems, this invention provides a method for extracting invoice and voucher information based on the fusion of multimodal and OCR models, specifically including the following steps: Step S1: Obtain an image of the invoice / voucher.

[0010] Step S2: Preprocess the acquired image, including image denoising, tilt correction and region detection processing on the original image.

[0011] Step S3: Use the OCR engine to perform text recognition on the preprocessed image to obtain the text content and corresponding two-dimensional coordinates of each text block.

[0012] Step S4: Take the text fragments obtained by OCR recognition and the original image as input, use a pre-trained multimodal model for joint encoding, evaluate and output the matching degree of each text fragment with the predefined field category, and determine the candidate texts and their confidence scores for each field.

[0013] Step S5: Combine the semantic matching results output by the multimodal model with the location information of the OCR text to accurately locate and extract key fields, and verify the consistency between the OCR output and the semantic results.

[0014] Step S6: Error correction judgment. If the verification results conflict or the recognition confidence of a certain field is lower than the threshold, an error correction operation is performed. After multiple rounds of iteration, the field value with high confidence is obtained.

[0015] Step S7: Output structured invoice and voucher information.

[0016] Furthermore, in step S2, during the preprocessing operation, edge detection and projection analysis are used to obtain text regions and segment candidate text blocks.

[0017] Furthermore, in step S2, during the preprocessing operation, the colorful tickets and vouchers are subjected to grayscale or binarization processing to filter out background textures and improve the OCR recognition rate.

[0018] Furthermore, in step S4, the multimodal model generates a semantic vector for each text region, and extracts visual features from the overall and local images. Through the interaction of visual and linguistic features, it determines the correlation between different text fragments and field labels (such as "bank name", "amount", etc.), and outputs the semantic representation of each candidate text fragment and its matching degree with the predefined field category.

[0019] Furthermore, in step S6, the error correction operation is as follows: re-extract the image of the region, adjust the contrast, and then perform OCR recognition again.

[0020] Furthermore, in step S6, the error correction operation can also be: using a language model to re-infer the correct text.

[0021] Furthermore, in step S6, the error correction operation can also be: re-analyze the area where the field is located, subdivide the area to re-identify characters, and at the same time perform logical verification with reference to the overall context of the invoice (such as other amount fields and contract information).

[0022] Furthermore, in step S6, during the aforementioned error correction operation, the key fields are validated using built-in rules, which include that the amount must be in numeric format and the date must conform to a specific format.

[0023] This invention also provides a system for extracting invoice and voucher information based on the fusion of multimodal and OCR models, comprising: Image acquisition module: Used to acquire images of invoices and vouchers.

[0024] Image preprocessing module: Used to preprocess the input scanned image of the voucher, automatically segmenting potential text blocks and graphic regions.

[0025] OCR Recognition Module: Employs a deep learning-based OCR engine to detect and recognize text in preprocessed images, thereby obtaining text content and corresponding location information.

[0026] Multimodal semantic matching module: It takes the text fragments obtained by OCR recognition and the original image as input, uses a pre-trained visual language model for joint encoding, judges the correlation between different text fragments and field labels (such as "bank of account", "amount" etc.), and outputs the semantic representation of each candidate text fragment and its matching degree with the predefined field category, and determines the candidate text and its confidence level for each field.

[0027] Field location module: Used to combine the semantic matching results output by the multimodal model with the location information of the OCR text to accurately locate and extract key fields.

[0028] Verification module: For each target field (such as "amount"), the system automatically locates the corresponding visual area based on language prompts and verifies whether the OCR output is consistent with the semantic result.

[0029] Error correction module: Used to correct errors when the verification results are conflicting or the recognition confidence of a certain field is lower than the threshold. After multiple rounds of iteration, the final field value with high confidence is obtained.

[0030] Output module: Used to output structured ticket and voucher information.

[0031] The present invention also provides an electronic device, comprising: At least one processor; and At least one memory communicatively connected to the processor; The memory stores instructions that can be executed by a processor, which are then executed by the processor to enable the electronic device to perform the aforementioned method for extracting invoice and voucher information based on the fusion of multimodal and OCR models.

[0032] The method for extracting information from negotiable instruments and vouchers in this invention is based on the fusion of a multimodal visual language model and an OCR model. This method utilizes a pre-trained visual language model to jointly analyze voucher images and OCR text information, combining contextual semantics and visual structure for field localization and content correction. It accurately extracts key fields (such as bank name, amount, date, payee, etc.) from complex images, achieving the following technical effects: (1) Strong adaptability to complex formats: The method of this invention does not rely on a fixed template and can automatically adapt to various voucher formats. Whether it is a bill, report or unstructured receipt, it can accurately locate the field position through semantic matching, showing good cross-format adaptability.

[0033] (2) High recognition accuracy: By fusing multimodal contextual information, the system significantly improves the recognition accuracy of key fields (such as bank name, amount, date, etc.). In testing, compared with traditional OCR methods, the error rate is significantly reduced, and the tolerance for low-quality images and mixed-format scenarios is higher.

[0034] (3) Excellent cross-type generalization performance: The multimodal pre-trained model has rich knowledge of text and graphics, and can generalize well to different types of text, fonts and layouts. The method of the present invention does not require manual design of rules for each voucher, and still maintains high extraction performance when faced with new voucher styles.

[0035] (4) High robustness and automation: The multi-round error correction mechanism in this invention effectively reduces missed and incorrect identifications, and can automatically correct even when initial identification errors occur, thus improving the stability of the system. The high degree of automation reduces the cost of manual intervention and is suitable for large-scale deployment and real-time processing requirements. Attached Figure Description

[0036] The specific embodiments of the present invention will be further explained below with reference to the accompanying drawings.

[0037] Figure 1 This is a flowchart of the method for extracting invoice and voucher information based on the fusion of multimodal and OCR models according to the present invention.

[0038] Figure 2 This is a block diagram of the document and voucher information extraction system based on the fusion of multimodal and OCR models of the present invention. Detailed Implementation Example 1

[0039] Combination Figure 1 The method for extracting invoice and voucher information based on the fusion of multimodal and OCR models in this embodiment specifically includes the following steps: Step S1: Obtain the image of the negotiable instrument. In this embodiment, specifically, a scanned image of a non-standard bank voucher is used as an example. This image contains a mixture of Chinese and English fonts and a background printed pattern.

[0040] Step S2: Preprocess the acquired image. This preprocessing includes image denoising, tilt correction, and region detection to improve subsequent recognition performance. Layout analysis technology is used to automatically segment potential text blocks and graphic regions, providing a foundation for OCR recognition and visual analysis.

[0041] In this embodiment, preferably, in step S2, during the preprocessing operation, edge detection and projection analysis are used to obtain the text region and segment out candidate text blocks.

[0042] In this embodiment, preferably, in step S2, during the preprocessing operation, the colorful tickets and vouchers are subjected to grayscale or binarization processing to filter out background textures and improve the OCR recognition rate.

[0043] Step S3: Use the OCR engine to perform text recognition on the preprocessed image to obtain the text content and corresponding two-dimensional coordinates of each text block.

[0044] Specifically, in this embodiment, text detection and recognition are performed on the image using OCR to obtain the text content and corresponding location information (boundary coordinates of text lines or words), yielding preliminary text results for subsequent semantic analysis. At this stage, preliminary text data such as "Bank: XXXX Bank", "Amount: ¥1,234.56", and "Date: May 8, 2025" may be obtained. The OCR results may contain some character recognition errors or missing tags.

[0045] Step S4: Take the text fragments obtained by OCR recognition and the original image as input, and use a pre-trained multimodal model (such as a Transformer network that fuses image and text information) for joint encoding. The model evaluates and outputs the matching degree of each text fragment with the predefined field category, and determines the candidate texts and their confidence scores for each field.

[0046] In this embodiment, preferably, in step S4, the multimodal model generates a semantic vector for each text region, and simultaneously extracts visual features from both the overall and local aspects of the image. Through the interaction of visual and linguistic features, it determines the correlation between different text fragments and field labels (such as "bank name", "amount", etc.), and outputs the semantic representation of each candidate text fragment and its matching degree with the predefined field category. For example, the model finds that the region containing the text "XXXX Bank" is highly semantically related to the "bank name" field; the text containing the "¥" symbol and numbers is highly related to the "amount" field. Based on the model output, the model determines the candidate texts and their confidence levels for each field.

[0047] Step S5: Combining the semantic matching results output by the multimodal model with the positional information of the OCR text, the key fields are precisely located and extracted. Based on the semantic matching results, the system returns to the image coordinates to accurately locate the visual region of each field. For example, the "Bank Name" field is located in the upper left printed area of ​​the image. For each target field (such as "Amount"), the system automatically locates the corresponding visual region based on language prompts and verifies the consistency between the OCR output and the semantic results.

[0048] Step S6: Error correction judgment. If the verification results conflict or the recognition confidence of a certain field is lower than the threshold, an error correction operation is performed. After multiple rounds of iteration, the field value with high confidence is obtained.

[0049] In this embodiment, preferably, in step S6, during the aforementioned error correction operation, key fields are validated using built-in rules. These built-in rules include that the amount must be in numerical format and the date must conform to a specific format. For the same field, if there is a discrepancy between the OCR and the semantic result (e.g., the OCR incorrectly identifies "agriculture" as "Nongxiao"), the system will utilize multiple rounds of error correction. The system will re-perform OCR recognition on the specific area, adjust preprocessing parameters, or perform adaptive correction by combining statistical rules and contextual logic (e.g., verifying the sum of the total amount and the itemized amounts).

[0050] In this embodiment, preferably, in step S6, the error correction operation is: re-extract the image of the region, adjust the contrast, and then perform OCR recognition again.

[0051] In this embodiment, preferably, in step S6, the error correction operation can also be: using a language model to re-infer the correct text.

[0052] In this embodiment, preferably, in step S6, the error correction operation can also be: re-analyze the area where the field is located, subdivide the area to re-identify characters, and at the same time perform logical verification with reference to the overall context of the bill / voucher (such as other amount fields, contract information).

[0053] In this embodiment, after entering the error correction process, the system will concentrate resources to re-analyze the area where the field is located. After one or more rounds of iteration, the system obtains the field value with high credibility.

[0054] Step S7: Output structured invoice and voucher information.

[0055] In this embodiment, after completing the above steps, the system outputs structured voucher information, such as "Bank: XXXX Bank", "Amount: ¥1,234.56", "Date: May 8, 2025", "Payee: Zhang San", etc. These results can be directly used by the bank's back-end system without manual verification.

[0056] The document information extraction method based on the fusion of multimodal and OCR models in this embodiment upgrades from single OCR to joint modeling of text and image semantics. By integrating visual language models (such as BLIP, VIT, etc.) and OCR results, it realizes contextual understanding and semantic-level reasoning of fields. It is used for key information identification and location in unstructured document images. The system is no longer limited to character recognition, but enhances contextual understanding through multimodal interaction. It has stronger reasoning and abstract extraction capabilities, especially for semantically ambiguous fields. The document and voucher information extraction method based on the fusion of multimodal and OCR models in this embodiment adopts a non-template-based field location mechanism. It does not rely on predefined templates, table borders or field position rules, but only on the field recognition mechanism built on text-image-semantic relevance. It supports the processing of non-standard document images with messy structure, mixed text and images, and no-boundary fields. Even when faced with voucher images with loose structure and messy fields, it can locate fields based on semantic relevance, adapt to unstructured and mixed text and image scenarios, and improve robustness.

[0057] The document information extraction method based on the fusion of multimodal and OCR models in this embodiment introduces a multi-round error correction and confidence mechanism. It designs a multi-round correction process based on recognition confidence, and combines logical rule constraints (such as the amount cannot be negative and the date cannot be future) and context verification to make the overall information extraction more accurate and reliable. Example 2

[0058] Combination Figure 2 The invoice and voucher information extraction system based on the fusion of multimodal and OCR models in this embodiment specifically includes the following modules: Image acquisition module: Used to acquire images of invoices and vouchers.

[0059] Image preprocessing module: This module preprocesses the input scanned voucher image, automatically segmenting potential text blocks and graphic regions to improve subsequent recognition results. It automatically segments potential text blocks and graphic regions using layout analysis technology, providing a foundation for OCR recognition and visual analysis.

[0060] The OCR recognition module uses a deep learning-based OCR engine to detect and recognize text in the preprocessed image, obtaining the text content and corresponding location information (boundary coordinates of text lines or words). This module outputs preliminary text results for subsequent semantic analysis.

[0061] Multimodal semantic matching module: Taking the text fragments obtained from OCR recognition and the original image as input, it uses a pre-trained visual language model (such as the Transformer network that integrates image and text information) for joint encoding. This model can understand the semantic context of the voucher and judge the correlation between different text fragments and field labels (such as "bank of account", "amount" etc.) through the interaction of visual and linguistic features. The model will output the semantic representation of each candidate text fragment and its matching degree with the predefined field category, and determine the candidate text and its confidence level for each field.

[0062] Field location module: Used to combine the semantic matching results output by the multimodal model with the location information of the OCR text to accurately locate and extract key fields.

[0063] Verification module: For each target field (such as "amount"), the system automatically locates the corresponding visual area based on language prompts and verifies whether the OCR output is consistent with the semantic result.

[0064] Error correction module: Used to correct errors when verification results conflict or the recognition confidence of a certain field is lower than the threshold. After multiple rounds of iteration, a field value with high confidence is obtained. If a conflict is found or the recognition confidence is low, a multi-round error correction mechanism is triggered: the system can re-perform OCR recognition on a specific area, adjust preprocessing parameters, or perform adaptive correction by combining statistical rules and contextual logic (such as checking the sum of the total amount and the itemized amounts).

[0065] Output module: Used to output structured ticket and voucher information. Example 3

[0066] The electronic device in this embodiment specifically includes: At least one processor; and At least one memory communicatively connected to the processor; The memory stores instructions that can be executed by a processor, which are then executed by the processor to enable the electronic device to perform the aforementioned method for extracting invoice and voucher information based on the fusion of multimodal and OCR models.

[0067] Many specific details have been set forth in the foregoing description to provide a thorough understanding of the present invention. However, the above description is merely a preferred embodiment of the present invention, and the present invention can be implemented in many other ways different from those described herein. Therefore, the present invention is not limited to the specific embodiments disclosed above. Furthermore, any person skilled in the art can make many possible variations and modifications to the technical solutions of the present invention, or modify them into equivalent embodiments, using the methods and techniques disclosed above, without departing from the scope of the present invention. Any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention, without departing from the content of the present invention, shall still fall within the protection scope of the present invention.

Claims

1. A method for extracting invoice and voucher information based on the fusion of multimodal and OCR models, characterized in that: Includes the following steps: Step S1: Obtain an image of the invoice / voucher; Step S2: Preprocess the acquired image, the preprocessing including image denoising, tilt correction and region detection of the original image; Step S3: Use the OCR engine to perform text recognition on the preprocessed image to obtain the text content and corresponding two-dimensional coordinates of each text block; Step S4: Take the text fragments obtained by OCR recognition and the original image as input, use a pre-trained multimodal model for joint encoding, evaluate and output the matching degree of each text fragment with the predefined field category, and determine the candidate texts and their confidence scores for each field. Step S5: Combine the semantic matching results output by the multimodal model with the location information of the OCR text to accurately locate and extract key fields, and verify the consistency between the OCR output and the semantic results; Step S6: Error correction judgment. If the verification results conflict or the recognition confidence of a certain field is lower than the threshold, an error correction operation is performed. After multiple rounds of iteration, the final field value with high confidence is obtained. Step S7: Output structured invoice and voucher information.

2. The method for extracting invoice and voucher information based on the fusion of multimodal and OCR models according to claim 1, characterized in that: In step S2, during the preprocessing operation, edge detection and projection analysis are used to obtain text regions and segment candidate text blocks.

3. The method for extracting invoice and voucher information based on the fusion of multimodal and OCR models according to claim 1, characterized in that: In step S2, during the preprocessing operation, the colorful tickets and vouchers are subjected to grayscale or binarization processing to filter out background textures.

4. The method for extracting invoice and voucher information based on the fusion of multimodal and OCR models according to claim 1, characterized in that: In step S4, the multimodal model generates a semantic vector for each text region, and extracts visual features from the overall and local parts of the image. Through the interaction of visual and linguistic features, it determines the correlation between different text fragments and field labels, and outputs the semantic representation of each candidate text fragment and its matching degree with the predefined field category.

5. The method for extracting invoice and voucher information based on the fusion of multimodal and OCR models according to claim 1, characterized in that: In step S6, the error correction operation is as follows: re-extract the image of the region, adjust the contrast, and then perform OCR recognition again.

6. The method for extracting invoice and voucher information based on the fusion of multimodal and OCR models according to claim 1, characterized in that: In step S6, the error correction operation is to use the language model to re-infer the correct text.

7. The method for extracting invoice and voucher information based on the fusion of multimodal and OCR models according to claim 1, characterized in that: In step S6, the error correction operation is as follows: re-analyze the area where the field is located, subdivide the area to re-identify characters, and at the same time perform logical verification with reference to the overall context of the invoice.

8. The method for extracting invoice and voucher information based on the fusion of multimodal and OCR models according to any one of claims 5-7, characterized in that: In step S6, during the error correction operation, key fields are validated using built-in rules, which include that the amount must be in numeric format and the date must conform to a specific format.

9. A system for extracting invoice and voucher information based on the fusion of multimodal and OCR models, characterized in that: include: Image acquisition module: used to acquire images of invoices and vouchers; Image preprocessing module: used to preprocess the input scanned voucher image, automatically segmenting potential text blocks and graphic regions; OCR Recognition Module: Employs a deep learning-based OCR engine to detect and recognize text in preprocessed images, thereby obtaining text content and corresponding location information; Multimodal semantic matching module: It takes the text fragments obtained by OCR recognition and the original image as input, uses a pre-trained visual language model for joint encoding, judges the correlation between different text fragments and field labels, and outputs the semantic representation of each candidate text fragment and its matching degree with the predefined field category, and determines the candidate text and its confidence level for each field. Field location module: Used to combine the semantic matching results output by the multimodal model with the location information of the OCR text to accurately locate and extract key fields; Verification module: For each target field, the system automatically locates the corresponding visual region based on language prompts and verifies whether the OCR output is consistent with the semantic result; Error correction module: Used to correct errors when the verification results are conflicting or the recognition confidence of a certain field is lower than the threshold. After multiple rounds of iteration, the final field value with high confidence is obtained. Output module: Used to output structured ticket and voucher information.

10. An electronic device, characterized in that: include: At least one processor; as well as At least one memory communicatively connected to the processor; The memory stores instructions that can be executed by a processor, which are executed by the processor to cause the electronic device to perform the invoice and voucher information extraction method based on multimodal and OCR model fusion as described in any one of claims 1-7.

Citation Information

Cited By

  • Multi-modal large model driven geometric image reading and analyzing method and system

    CN121438290A

  • Large model answer sheet OCR (Optical Character Recognition) method and system based on detection and character segmentation

    CN121545173A

  • Automatic recharging system based on AI message identification

    CN121563507A

  • Multi-modal fusion-based OCR (optical character recognition) information dynamic verification method and system and medium

    CN121708608A

  • Intelligent matching method of receipt file and voucher file and related device

    CN121980287A