Intelligent document recognition and structured data processing system based on visual language model
Patent Information
- Application Number
- TW115201703
- Authority / Receiving Office
- TW · TW
- Patent Type
- Utility models
- Current Assignee / Owner
- Filing Date
- 2026-02-25
- Publication Date
- 2026-07-11
- Estimated Expiration
- 2036-02-24
Smart Images

Figure IMG-2_DRAW_115201703-A0305-14-0001-1 
Figure IMG-2_DRAW_115201703-A0305-14-0002-2 
Figure IMG-2_DRAW_115201703-A0305-14-0003-3
Abstract
Description
Intelligent Document Recognition and Structured Data Processing System Based on Visual Language Model Intelligent Document Recognition and Structured Data Processing System Based on Visual Language Model Technical Field
[0001] The Intelligent Voucher Recognition and Structured Data Processing System Based on Visual Language Model refers to a voucher image processing and data extraction system, specifically an intelligent system that uses a visual language model combined with artificial intelligence (AI) computing devices to recognize and process voucher images of various formats. This system is applicable to fields such as document image processing, optical character recognition, financial technology and / or tax management. Prior Technology
[0002] When traditional Optical Character Recognition (OCR) technology is applied to the recognition of Taiwan's uniform invoices, the recognition rate often falls due to the following factors:
[0003] 1. Significant variation in image quality: In practical applications, voucher images are often blurred, distorted, and / or lack sufficient contrast due to factors such as shooting angle, light reflection, paper wrinkles or creases.
[0004] 2. Diverse Formats: Taiwan's tax system has various invoice formats, including traditional three-part invoices, three-part cash register invoices, electronic invoices, certificate copies, and handwritten receipts. The layout, column positions, and text arrangement of these various documents differ greatly.
[0005] 3. Fixed coordinate limitation: Traditional OCR systems mostly use "template comparison" technology, which predefines a fixed range of pixel coordinates to extract information from specific fields. This method is extremely dependent on the flatness and format consistency of the image, and cannot adapt to the various different formats in the actual environment.
[0006] 4. Lack of semantic understanding: Traditional OCR can only identify characters at the level of characters and cannot understand the logical relationships between fields (e.g., the total amount should equal the sales amount plus business tax), nor can it correct errors based on the overall semantics of the document.
[0007] 5. Synchronous blocking processing: Existing systems mostly use a synchronous processing architecture, requiring users to wait for each voucher image to be processed before taking the next voucher image, resulting in low batch processing efficiency. When users need to process a large number of vouchers (for example, when submitting expense reports at the end of each month), they must repeatedly execute the "take a picture → wait → view the result" loop, which is a poor user experience.
[0008] Therefore, there is an urgent need in this field for an intelligent voucher identification and structured data processing system that is applicable to image vouchers of different formats and can batch process large amounts of image voucher data. Summary of the Invention
[0009] This invention provides an intelligent credential recognition and structured data processing system based on a visual language model. The system includes a processor, an image input unit, a visual language model inference module, a text result repair module, a data verification module, a post-processing coordinator, and a display. The processor performs visual language model inference, image preprocessing, data verification, and post-processing operations. The image input unit, coupled to the processor, receives credential images and converts them into digital images. The visual language model inference module, coupled to the processor, includes an image encoder, a language decoder, and a parameter fine-tuning module. The image encoder uses a deep learning feature extractor to extract multiple image features, and the language decoder generates text results based on these image features. The text result repair module, coupled to the visual language model inference module, performs grammatical correction of the text results. The data verification module, coupled to the text result repair module, performs structural verification of the text results. A post-processing coordinator, coupled to the data verification module, uses the post-processed text results, after grammatical correction and / or structural verification, to generate company information, buyer information, date information, and / or amount information. A display, coupled to the post-processing coordinator, outputs the text result recognition results and compliance reports. Simple Explanation of the Diagram
[0010] Figure 1 is a block diagram of a smart credential recognition and structured data processing system based on a visual language model in at least one embodiment of this invention. Figure 2 is a block diagram of the task queue management module in at least one embodiment of this invention. Figure 3 is a flowchart of a method for identifying electronic invoices in at least one embodiment of this invention. Figure 4 is a flowchart of a method for identifying a three-part invoice in at least one embodiment of this invention. Figure 5 is a flowchart of a method for identifying a three-part cash register invoice in at least one embodiment of this invention. Figure 6 is a flowchart of a method for identifying a traditional uniform invoice in at least one embodiment of this invention. Figure 7 is a flowchart of the method for identifying the proof link in at least one embodiment of this invention. Implementation
[0011] The following describes the implementation of this invention through specific embodiments. Those skilled in the art can easily understand the spirit, advantages, and effects of this invention based on the content contained herein. However, the specific embodiments described herein are not intended to limit this invention. This invention can also be implemented or applied through other different implementation methods, and the details described herein can also be varied or modified according to different viewpoints and applications without departing from the spirit of this invention.
[0012] When the terms "comprising," "including," or "having" are used in this document to refer to a specific element, unless otherwise stated, they may include other materials, components, structures, areas, parts, devices, systems, steps, or connections, rather than excluding such other elements.
[0013] Unless otherwise expressly stated herein, the singular forms "a" and "the" used herein also include the plural forms, and the word "or" used herein is used interchangeably with "and / or". Furthermore, as used herein, the phrase "at least one" refers to one or more requirements and should be understood to mean at least one requirement selected from any one or more of the listed requirements, but does not necessarily include at least one of each of the listed requirements, and does not exclude any combination of the listed requirements.
[0014] In view of the deficiencies in previous technologies, this invention provides an intelligent credential recognition and structured data processing system based on a visual language model. The technical problems solved by this system include at least the following five points, but this invention is not limited to these.
[0015] Firstly, this invention overcomes the vulnerability of traditional coordinate positioning technology to image geometric deformation and format variations, achieving universal recognition capability for multiple document formats. In at least one embodiment of this invention, the system supports the recognition of five types of documents, including three-part invoices, three-part cash register invoices, electronic invoices, proof copies, and traditional invoices. Through the semantic understanding capability of the visual language model inference module, it does not require preset fixed coordinate ranges for each format. Using the visual language model for inference, it can be applied to five different document formats, greatly improving the practicality of this invention.
[0016] Secondly, this invention integrates the semantic understanding capabilities of a visual language model, enabling the system to intelligently identify fields based on the overall context of the document, rather than relying solely on fixed coordinates. For example, the system can identify the correct invoice date field based on contextual fields such as "Electronic Invoice Certificate," "November-December 2023," and "Invoice Number" in an electronic invoice, and digitize the invoice date field through Optical Character Recognition (OCR). With the assistance of the visual language model, the accuracy of OCR technology can be significantly improved. Compared to traditional OCR technology using fixed coordinates, the combination of visual language model and OCR can increase the accuracy of invoice field reading, making it applicable to different types of invoices and fields.
[0017] Thirdly, a multi-layered data verification and automatic correction mechanism is established, combining external information retrieval (e.g., unified number online query), internal logical rules (e.g., total amount verification), and fuzzy matching technology to improve recognition accuracy. For example, the total amount in an invoice should equal the sales amount plus business tax. Considering the relationship between these different fields, multi-layered data verification can be performed. When errors occur in the data logic of different fields (e.g., due to a blurry invoice image, the OCR interpretation of the total amount is incorrect), the system can combine external information retrieval, internal logical rules, and fuzzy matching technology to correct the incorrectly interpreted field (e.g., the total amount). Through this multi-layered data verification and automatic correction mechanism, interpretation errors caused by visual language models and OCR technology can be significantly reduced, improving the system's recognition accuracy.
[0018] Fourthly, the system achieves low-latency, real-time processing capabilities, enabling it to complete the identification, verification, and correction process for a single credential within P time. To achieve this, the system can utilize specialized processors, such as Graphics Processing Units (GPUs) or Neural Processing Units (NPUs). By using these specialized processors to execute the inference of the visual language model, system performance can be improved, allowing the system to complete the identification, verification, and correction process for a single credential within P time. In at least one embodiment, P can be 2 seconds. The specialized processor (GPU or NPU) can be located in the cloud or on the ground; this invention is not limited to these locations.
[0019] Fifthly, the system supports asynchronous processing of continuously captured receipts, allowing users to capture multiple receipts consecutively without waiting for each receipt to be processed. A task queue mechanism ensures sequential processing, improving user experience and overall throughput. The system can temporarily store multiple consecutively captured receipt images in a task queue, monitor the graphics card (including GPU) usage, and automatically schedule multiple processing tasks. When the graphics card is idle, the system automatically retrieves the next receipt image from the task queue and sends it to the processing queue. The processing queue executes the identification, verification, and correction processes for multiple receipt images in a first-in-first-out manner. Therefore, by using the system's task queue mechanism, users can continuously capture receipts without waiting for each receipt image to be processed, significantly improving user convenience.
[0020] Figure 1 is a block diagram of a visual language model-based intelligent credential recognition and structured data processing system 10 according to at least one embodiment of this invention. The visual language model-based intelligent credential recognition and structured data processing system 10 may include an image input unit 102, a QR (Quick Response) code decoding module 104, a processor 106, a memory storage unit 108, a visual language model inference module 110, a text result repair module 112, a data verification module 114, a post-processing coordinator 116, an external information retrieval module 118, and a display 120.
[0021] In at least one embodiment, the image input unit 102 is coupled to the processor 106 to receive a voucher image and convert it into a digital image. The image input unit 102 can be an electronic device with a camera function, such as a smartphone, camera, scanner, and / or tablet computer; this invention is not limited to these. The voucher image captured by the image input unit 102 can tolerate a certain degree of blurring, wrinkles, or reflections in the system 10 of this invention. Compared with the prior art, the system 10 of this invention can repair fields that may have been misread through data verification and post-processing calculations. In at least one embodiment, the image input unit 102 sequentially performs direct scanning of the original image, grayscale conversion, grayscale conversion after 2x magnification, and adaptive binarization after 3x magnification, wherein the adaptive binarization after 3x magnification uses a 21×21 pixel local window to calculate the threshold.
[0022] In at least one embodiment, the QR code decoding module 104 is coupled to the processor 106 and includes a preprocessing circuit 1002 and a QR code parsing circuit 1004. The preprocessing circuit 1002 performs multi-stage image preprocessing (including grayscale conversion, scaling, and adaptive binarization) to improve the recognition rate. The QR code parsing circuit 1004 parses a fixed 77-code format (including invoice text, date, amount, unified serial number, and other field information) and / or product details according to the Taiwan electronic invoice QR code encoding standard. When the captured voucher image contains a QR code (e.g., an electronic invoice), the QR code parsing circuit 1004 in the QR code decoding module 104 can interpret the QR code to generate electronic invoice-related field information such as invoice number, company name, company telephone number, and unified serial number. In at least one embodiment, the information interpreted by the QR code can be compared with the inference result of the visual language model inference module 110 to improve the accuracy of the QR code decoding module 104 in interpreting the QR code. In at least one embodiment, the QR code parsing circuit 1004 sequentially parses the invoice serial number (10 digits), the date of issuance (7 digits), the random code (4 digits), the sales amount (8 digits in hexadecimal), the total amount including tax (8 digits in hexadecimal), the buyer's unique number (8 digits), the seller's unique number (8 digits), the encrypted verification information (24 digits), and / or the item information (variable length). This invention is not limited to these.
[0023] In at least one embodiment, the processor 106 is used to perform visual language model inference, image preprocessing, data verification and / or post-processing operations. The processor 106 may include a central processing unit (CPU), a graphics processing unit (GPU) and / or a neural processing unit (NPU). The GPU and NPU can be used to accelerate visual language model inference, while the CPU can be used to perform operations other than neural networks, such as image preprocessing, data verification and / or post-processing operations.
[0024] In at least one embodiment, the memory storage unit 108 is coupled to the processor 106 and includes a first memory 1006 and a second memory 1008. The first memory 1006 stores a plurality of weight parameters of the visual language model (including basic model parameters and parameter adjustment parameters of the parameter fine-tuning module 1014). The second memory 1008 stores a company information cache database, which stores the corresponding company name, company phone number, and company address using a unified identification number as the key. The processor 106 can search for relevant information (company name, company phone number, and company address) of the company that issued the invoice through the unified identification number of the company in the image certificate. The plurality of weight parameters of the visual language model stored in the first memory 1006 can be provided to the visual language model inference module 110 to perform inference and generate inference results, while the company information cache database stored in the second memory 1008 can be provided to the processor 106 to search for company information, thereby improving the efficiency of image certificate recognition.
[0025] In at least one embodiment, the visual language model inference module 110 is coupled to the processor 106 and includes an image encoder 1010, a language decoder 1012, and a parameter-efficient fine-tuning module 1014. The image encoder 1010 uses a deep learning feature extractor (e.g., but not limited to, a visual transformer) architecture to extract multiple image features from an image document and encodes the image into a feature vector. The language decoder 1012 uses a Large Language Model (LLM) to generate text results based on the extracted multiple image features. The text results can be output in JSON format, but this invention is not limited to this. The text results can include, but are not limited to, invoice-related information such as invoice number, company name, company phone number, and company unified identification number. In some embodiments, the parameter-efficient fine-tuning module 1014 can be a low-rank adaptive layer. By using the parameter-efficient fine-tuning module 1014, the number of parameters required to train the visual language model is significantly reduced, making it easier and more convenient to fine-tune the visual language model without reducing efficiency.
[0026] In at least one embodiment, the text result repair module 112 is coupled to the visual language model inference module 110 to receive and perform syntax correction on the text results output by the visual language model inference module 110. The text result repair module 112 may be a JSON format repair module, but this invention is not limited thereto. The text result repair module 112 may include a syntax correction circuit, which is used to perform steps such as removing Markdown tags from the visual language model output, correcting Python literals to JSON standard format, converting curly quotes, and removing trailing commas. The text result repair module 112 may also include a structure verification circuit, which can be used to verify the integrity of the JSON structure. If JSON parsing fails, it is encapsulated into a backup format.
[0027] In at least one embodiment, the data verification module 114 is coupled to the text result repair module 112 to perform structural verification of the text results. The data verification module 114 may include a format verification unit 1016, a cross-column rule verification unit 1018, and a normalization processing unit 1020. The format verification unit 1016 verifies the format of each column according to regular expressions (e.g., the unified number must be 8 digits, and the month must be a number from 1 to 12). The cross-column rule verification unit 1018 checks the summation formula (e.g., the total amount should equal the sales amount plus business tax) and column mutual exclusion rules (e.g., the invoice number cannot be the same as the unified number). The normalization processing unit 1020 uniformly processes the thousands separator format, year conversion (e.g., the conversion between the Republic of China year and the Gregorian calendar year, 110th year of the Republic of China is 2021 AD), and removes leading zeros.
[0028] In at least one embodiment, the post-processing coordinator 116 is coupled to the data verification module 114, and uses the post-processed text results after grammatical correction and / or structural verification to generate company information, buyer information, date information, and / or amount information. The post-processing coordinator 116 may include a prefix processing unit, a company information processing unit, a buyer information processing unit, a date processing unit, and an amount processing unit. The prefix processing unit can force the invoice characters to uppercase; the company information processing unit can execute a fuzzy matching algorithm to combine cache queries and an external database interface; the buyer information processing unit can perform fuzzy comparison and verification of buyer information; the date processing unit can perform cross-validation of the invoice summary field to verify the year and month, and use an edit distance algorithm to select the most similar candidate value; and the amount processing unit can perform bidirectional calculation (e.g., centered on sales amount or total amount) and use string similarity to select the best correction scheme.
[0029] In at least one embodiment, the post-processing coordinator 116 can execute a fuzzy matching algorithm to calculate the similarity between the OCR result and the cached data. The fuzzy matching algorithm uses a weighted scoring mechanism. When the total score calculated by the weighted scoring mechanism is greater than a threshold, the system 10 will use the cached data. In at least another embodiment, the post-processing coordinator 116 can perform bidirectional computation, which calculates the total string similarity score between the two amount schemes and the OCR, and selects the amount scheme with the higher score to correct the amount information. In at least another embodiment, the post-processing coordinator 116 extracts the Republic of China year and month from the invoice summary field and performs cross-validation with the OCR recognition value. When the cross-validation does not match, the post-processing coordinator 116 calculates the edit distance to select the most similar candidate value to correct the date information.
[0030] In at least one embodiment, an external information retrieval module 118 is coupled to a post-processing coordinator 116. The external information retrieval module 118 may include an external database interface unit, a cache update unit, and a confidence value calculation unit. The external database interface unit sends HTTP requests to a third-party company information platform and retrieves the company name, company phone number, and company address corresponding to the unified identification number. The cache update unit writes the successfully retrieved data into the second memory 1008 of the memory storage unit 108 to avoid repeatedly executing network requests. The confidence value calculation unit calculates the similarity (confidence value) between the retrieval results from the external database interface and the OCR recognition results. The system 10 only uses the data from the external database interface when the confidence value is higher than a threshold.
[0031] In at least one embodiment, display 120 is coupled to post-processing coordinator 116 to output identification results and compliance reports of textual results (e.g., textual results in JSON format). Display 120 can output and display identification results and compliance reports in structured JSON format for use by accounting and / or tax systems.
[0032] Figure 2 is a block diagram of the task queue management module 20 in at least one embodiment of the present invention. In at least one embodiment, the intelligent credential recognition and structured data processing system 10 based on a visual language model may further include the task queue management module 20, which is coupled to the post-processing coordinator 116 and the display 120. The task queue management module 20 may include a credential receiving buffer 202, a task scheduler 204, a status tracking unit 206, a result temporary storage area 208, and an asynchronous notification mechanism 210.
[0033] In at least one embodiment, the voucher receiving buffer 202 is coupled to the image input unit 102 to receive a plurality of voucher images from the image input unit 102. A first-in-first-out (FIFO) queue structure is used to temporarily store the plurality of consecutively captured voucher images in the task queue. The FIFO queue structure allows the voucher images to be processed sequentially, making it convenient for users to view them.
[0034] In at least one embodiment, a task scheduler 204 is coupled to a credential receiving buffer 202, a processor 106, and a visual language model inference module 110. This scheduler monitors the utilization of the processor 106 (e.g., GPU, NPU, CPU) and automatically schedules multiple processing tasks. When a processor 106 is idle, the task scheduler automatically retrieves a credential image from the task queue and passes it to the processing queue. The task scheduler 204 may include a batch processing judgment unit. When the number of credential images accumulated in the task queue reaches a threshold (e.g., 5 images), the task scheduler initiates a batch inference mode. In batch inference mode, the task scheduler 204 retrieves multiple credential images at once and simultaneously passes them to the visual language model inference module 110 for parallel processing. This parallel processing increases the utilization of the processor 106 (e.g., GPU, graphics card) to over 85%, and doubles the throughput of the processor 106.
[0035] In at least one embodiment, the status tracking unit 206 is coupled to the task scheduler 204 to assign a universally unique identifier (UUID) to each voucher image and record the processing status of each voucher image. The processing status may include pending processing, processing in progress, completed, and / or error status. Users can query the processing progress and status of each voucher image through the UUID, facilitating user management and use. The status tracking unit 206 may include a status mapping table, which may include multiple fields: universally unique identifier (UUID), timestamp, processing status, and / or error message. The processing status may include four states: pending processing, processing in progress, completed, and error status. Users can query the processing status and result of any voucher image through the UUID.
[0036] In at least one embodiment, a result buffer 208 is coupled to the task scheduler 204 to store the processed credential image recognition results. The result buffer 208 adopts a ring buffer structure to retain the processing results of the most recent N tasks for user query. In at least one embodiment, an asynchronous notification mechanism 210 is coupled to the result buffer 208 to notify the front-end application when task processing is completed. The asynchronous notification mechanism 210 supports a postback function mode. In the postback function mode, when task processing is completed, a pre-registered postback function is executed, passing in a universally unique identifier (UUID) and the processing result. In at least one other embodiment, the asynchronous notification mechanism 210 can support an event subscription mode. In the event subscription mode, a publish and subscribe architecture is used. The front-end application needs to subscribe to the completion event of a specific UUID in advance. This allows the asynchronous notification mechanism 210 of the task queue management module 20 to publish a notification corresponding to the completion event of the specific UUID through network packets. All devices that have subscribed to the specific UUID in advance can receive the completion notification.
[0037] In at least one embodiment of this invention, the intelligent voucher recognition and structured data processing system 10 and the task queue management module 20 based on a visual language model achieve high accuracy, high format adaptability, automatic correction capability, low-latency real-time processing, controllable hardware costs, and asynchronous processing capability. In at least one embodiment, the high accuracy includes an overall field recognition accuracy of up to 95.3% and an accuracy of up to 98.5% for electronic invoices (assisted by QR codes), representing an improvement of approximately 12-18 percentage points compared to traditional coordinate positioning technology. In at least another embodiment, the high format adaptability includes the semantic understanding capability of the visual language model, eliminating the need to redefine fixed coordinate ranges for each new invoice format; fine-tuning can be completed with only a small number of samples (50-200 invoices per category) for model training. In at least one other embodiment, the automatic correction capability includes an improvement in company information accuracy of 7.2-7.8 percentage points (through cache matching and external database interface), an improvement in amount field accuracy of 3.1-3.7 percentage points (through bidirectional inference), and an improvement in date field accuracy of 4.4-5.7 percentage points (through cross-validation with summary field). In at least one other embodiment, low-latency real-time processing includes a processor 106 inference latency of less than 1.5 seconds / page, a total latency including post-processing of less than 2 seconds / page (including cache usage), a total latency including post-processing of less than 4 seconds / page (including query time from external database interface), and a single GPU throughput of 30-60 pages / minute. Controllable hardware costs include the ability to operate on consumer-grade GPUs (12GB VRAM), reducing the hardware threshold by approximately 75% compared to traditional full-precision models (which require a GPU with 48GB VRAM). The asynchronous processing capability allows users to continuously capture voucher images without waiting for processing to complete. The task scheduler 204 of the task queue management module 20 automatically adds voucher images to the task queue for sequential processing and supports batch processing mode. When a certain number of tasks accumulate in the task queue, batch processing mode can be activated, thereby increasing GPU utilization to over 85% and throughput by approximately twice that of single-image processing. The task queue management module 20 provides real-time status queries, allowing users to check the processing progress and results of any voucher image via UUID. The front-end interface is unblocked; the capture and display of voucher images are unaffected by back-end processing latency, improving user experience smoothness by approximately 60%.
[0038] Figure 3 is a flowchart of a method for identifying electronic invoices in at least one embodiment of this invention. When a user inputs an image of an electronic invoice through the image input unit 102, the intelligent voucher recognition and structured data processing system 10 based on a visual language model executes steps S302-S310.
[0039] Figure 4 is a flowchart of a method for identifying a three-part invoice in at least one embodiment of this invention. When a user inputs an image of a three-part invoice through the image input unit 102, the intelligent voucher recognition and structured data processing system 10 based on a visual language model executes steps S402-S412.
[0040] Figure 5 is a flowchart of a method for recognizing a three-part cash register invoice in at least one embodiment of this invention. When a user inputs an image of a three-part cash register invoice through the image input unit 102, the intelligent voucher recognition and structured data processing system 10 based on a visual language model executes steps S502-S512.
[0041] Figure 6 is a flowchart of a method for identifying a traditional uniform invoice in at least one embodiment of this invention. When a user inputs an image of a traditional uniform invoice (including handwritten information) through the image input unit 102, the intelligent voucher recognition and structured data processing system 10 based on a visual language model executes steps S602-S608.
[0042] Figure 7 is a flowchart of the method for identifying a certificate in at least one embodiment of this invention. When a user inputs a certificate image through the image input unit 102, the intelligent certificate identification and structured data processing system 10 based on the visual language model executes steps S702-S708.
[0043] The foregoing outlines the features of several embodiments, enabling those skilled in the art to fully understand the various aspects of this invention. Those skilled in the art should recognize that this invention provides a basis for designing or modifying other processes and structures to achieve substantially the same functionality and / or results as the embodiments described above. Furthermore, such equivalent configurations do not depart from the spirit and scope of this invention, and various changes, substitutions, and modifications can be made without departing from that spirit and scope.
[0044] 10: Intelligent Voucher Identification and Structured Data Processing System Based on Visual Language Model 102: Image Input Unit 104: QR code decoding module 106: Processor 108: Memory storage unit 110: Visual Language Model Inference Module 112: Text Result Repair Module 114: Data Verification Module 116: Post-processing coordinator 118: External Information Retrieval Module 120: Monitor 1002: Preprocessing circuit 1004: QR code parsing circuit 1006: First Memory 1008: Second Memory 1010: Image Encoder 1012: Language Decoder 1014: High-efficiency parameter fine-tuning module 1016: Format Validation Unit 1018: Cross-column rule validation unit 1020: Regularization Processing Unit 20: Task Queue Management Module 202: Credential Receive Buffer 204: Task Scheduler 206: State Tracking Unit 208: Result Storage Area 210: Asynchronous notification mechanism S302-S310, S402-S412, S502-S512, S602-S608, S702-S708: Steps
Claims
1. A smart credential recognition and structured data processing system based on a visual language model, comprising: A processor for performing a visual language model inference, an image preprocessing, a data verification, and a post-processing operation; an image input unit coupled to the processor for receiving a credential image and converting the credential image into a digital image; A visual language model inference module, coupled to the processor, includes: an image encoder for extracting a plurality of image features using a deep learning feature extractor; a language decoder for generating a text result based on the image features; and a parameter-efficient fine-tuning module; a text result repair module, coupled to the visual language model inference module, for performing grammatical correction on the text result; a data verification module, coupled to the text result repair module, for performing structural verification on the text result; a post-processing coordinator, coupled to the data verification module, for post-processing the text result after grammatical correction and / or structural verification to generate company information, buyer information, date information, and / or amount information; and a display, coupled to the post-processing coordinator, for outputting a recognition result and / or a compliance report of the text result.
2. The system as described in claim 1, further comprising a memory storage unit coupled to the processor, including: a first memory for storing a plurality of weight parameters of the visual language model; and a second memory for storing a company information cache database, wherein the company information cache database uses a unique identifier as a key to store a corresponding company name, a company telephone number, and / or a company address.
3. The system as described in claim 1, further comprising a QR (Quick Response) code decoding module coupled to the processor, including: A preprocessing circuit is used to perform a multi-stage image preprocessing to improve the recognition rate; And a QR code parsing circuit, used to parse a fixed 77 code format and / or a product detail according to a Taiwan electronic invoice QR code encoding standard.
4. The system as described in Request 1, wherein the parameter fine-tuning module is a low-rank adaptive layer.
5. The system as described in claim 1, further comprising: An external information retrieval module, coupled to the post-processing coordinator, sends an HTTP request to a third-party company information platform to retrieve a company name, a telephone number, and an address corresponding to a unique identifier, and calculates a confidence value to determine whether to use a query from an external database interface.
6. The system as described in claim 1, wherein the post-processing coordinator is further configured to: execute a fuzzy matching algorithm to calculate a similarity between an Optical Character Recognition (OCR) result and cached data, wherein, The fuzzy matching algorithm uses a weighted scoring mechanism. When the total score calculated by the weighted scoring mechanism is greater than a threshold, the cached data is used. A bidirectional operation is performed, which calculates the total similarity score between the two amount schemes and the character string of the optical character recognition, and selects the amount scheme with the higher score to correct the amount information. The algorithm also extracts the Republic of China year and January from an invoice summary field and performs a cross-validation with the recognition value of the optical character recognition. When the cross-validation does not match, the post-processing coordinator calculates an edit distance to select the most similar candidate value to correct the date information.
7. The system as described in claim 1, further comprising a task queue management module coupled to the post-processor coordinator and the display, the task queue management module comprising: A credential receiving buffer, coupled to the image input unit, is used to temporarily store multiple consecutively captured credential images in a task queue using a first-in-first-out queue structure; a task scheduler, coupled to the credential receiving buffer, the processor, and the visual language model inference module, is used to monitor the utilization of a display card and automatically schedule multiple processing tasks, wherein when the display card is idle, the credential image is automatically retrieved from the task queue and transferred to a processing queue; a status tracking unit, coupled to the task scheduler, is used to assign a unique identifier to each credential image and record a processing status; a result buffer, coupled to the task scheduler, is used to retain the processing results of one or more recent tasks using a circular buffer structure; and an asynchronous notification mechanism, coupled to the result buffer, is used to notify a front-end application when task processing is completed.
8. The system as described in claim 7, wherein the task scheduler is configured to initiate a batch inference mode when the number of credentials accumulated in a task queue reaches a threshold, wherein in the batch inference mode, the task scheduler retrieves multiple credential images at once and simultaneously inputs them into the visual language model inference module for parallel processing, the parallel processing increasing the utilization of the graphics card to over 85% and increasing the throughput of the graphics card by 2 times.
9. The system as described in claim 7, wherein the state tracking unit includes a state mapping table, the state mapping table including a plurality of fields: a universally unique identifier (UUID), a timestamp, the processing state and / or an error message, and wherein the processing state includes four states: a pending state, a processing state, a completed state and / or an error state.
10. The system as described in claim 7, wherein the asynchronous notification mechanism supports a postback function mode or an event subscription mode, wherein: When the synchronous notification mechanism supports the postback function mode, it is used to execute a pre-registered postback function when the task is completed, passing in a universally unique identifier (UUID) and the processing result; and when the synchronous notification mechanism supports the event subscription mode, it is used to use a publish and subscribe architecture, so that when the front-end application subscribes to a completion event of a specific UUID, the system publishes a notification corresponding to the completion event of the specific UUID through a network packet.