Bill information identification method, apparatus and device, and storage medium
By using a multimodal large model to identify and extract information on bill images, the cumbersome problem of extracting information separately for each bill in the prior art is solved, and efficient bill information recognition is achieved.
Patent Information
- Application Number
- CN202510081422.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-05-16
AI Technical Summary
The prior art requires individual structured extraction of each bill when extracting information from different types of bills, which is cumbersome and inefficient.
The image processing module in the multimodal large model is used to identify the bill image in text, and the language processing module extracts the recognition results to obtain the bill information.
There is no need to write a structured extraction method for each bill, which improves the efficiency of bill information recognition and can automatically identify and extract key information from different types of bills.
Smart Images

Figure CN120014649A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a bill information recognition method, device, equipment and storage medium. Background Art
[0002] With the development of the Internet and financial technology, the number of various types of bills has increased rapidly, such as train tickets, air ticket itineraries, VAT invoices, bank checks, etc. There are a large number of business-related information fields in various bills. In order to improve the management level of information in bills, it is necessary to extract the information of bills to digitize the bill information. Bill information extraction is mainly a process of automatically identifying and extracting key information on paper or electronic bills. For example, OCR optical character technology is used to identify bill number, date, amount, product name and other information. However, the text data obtained by OCR recognition is unstructured data and needs to be converted into structured data. Different types of bills have different formats and may involve different layouts and information organization methods. At this time, it is necessary to perform structured extraction for each bill separately, which is cumbersome and inefficient. Summary of the invention
[0003] Based on this, it is necessary to provide a bill information identification method, device, equipment and storage medium for the above technical problems to solve at least one of the above technical problems.
[0004] The present invention provides a bill information recognition method, comprising:
[0005] Get the bill image to be processed;
[0006] Using the image processing module in the multimodal large model to perform text recognition on the bill image to obtain a text recognition result;
[0007] The language processing module in the multimodal large model is used to extract information from the text recognition results to obtain bill information.
[0008] Optionally, according to a bill information recognition method provided by the present invention, after extracting information from the text recognition result using the language processing module in the multimodal large model to obtain the bill information, the method further comprises:
[0009] Determining a target amount according to the entry station information and the exit station information in the ticket information;
[0010] The target amount is compared with the amount data in the bill information to determine the final amount data according to the comparison result.
[0011] Optionally, according to a bill information recognition method provided by the present invention, the text recognition result includes text content information, text position information and relative positions between text contents;
[0012] The method of using the image processing module in the multimodal large model to perform text recognition on the bill image to obtain a text recognition result includes:
[0013] Determining a text area containing text content in the bill image;
[0014] Using the image processing module in the multimodal large model to perform text recognition on the text area to obtain text content information and text position information corresponding to the text content information;
[0015] The relative positions of the text contents are determined according to the text content information and the text position information.
[0016] Optionally, according to a bill information recognition method provided by the present invention, the extracting information from the text recognition result using the language processing module in the multimodal large model to obtain the bill information includes:
[0017] The language processing module is used to perform semantic entity recognition on the text content information, text position information and relative positions between text contents in the text recognition result to obtain the bill information.
[0018] Optionally, according to a bill information recognition method provided by the present invention, the method of extracting information from the text recognition result using the language processing module in the multimodal large model to obtain the bill information comprises:
[0019] Determine the mapping relationship between each data item in the bill information and each data item in the preset bill entry template;
[0020] Assembling the bill information using the mapping relationship to obtain bill entry data in a target format;
[0021] The bill in the target format is entered into a preset database.
[0022] Optionally, according to a bill information recognition method provided by the present invention, the method further comprises:
[0023] Acquire a bill sample, wherein the data items in the bill sample are marked with text labels;
[0024] Input the bill sample into the multimodal large model, use the image processing module to perform text recognition on the bill sample, and use the language processing module to extract information from the recognition result to obtain text extraction information;
[0025] The multimodal large model is fine-tuned based on the text extraction information and the text labels.
[0026] Optionally, according to a bill information recognition method provided by the present invention, after acquiring the bill image to be processed, the method further includes:
[0027] The bill image is preprocessed and data enhanced.
[0028] The present invention also provides a bill information recognition device, comprising:
[0029] An acquisition module, used for acquiring the bill image to be processed;
[0030] A recognition module, used to perform text recognition on the bill image using the image processing module in the multimodal large model to obtain a text recognition result;
[0031] The extraction module is used to extract information from the text recognition result using the language processing module in the multimodal large model to obtain the bill information.
[0032] The present invention also provides a computer device, comprising a memory, a processor, and computer-readable instructions stored in the memory and executable on the processor, wherein the processor implements the above-mentioned bill information identification method when executing the computer-readable instructions.
[0033] The present invention also provides one or more readable storage media storing computer-readable instructions, and the computer-readable instructions implement the above-mentioned bill information identification method when executed by a processor.
[0034] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the method for identifying bill information as described above is implemented.
[0035] The above-mentioned bill information recognition method, device, equipment and storage medium include: obtaining a bill image to be processed; using the image processing module in the multimodal large model to perform text recognition on the bill image to obtain a text recognition result; using the language processing module in the multimodal large model to extract information from the text recognition result to obtain bill information. The present invention uses a multimodal large model to perform image recognition on bill images, and combines the multimodal large model to perform natural language processing on text content information, text location information and other information to extract bill information, thereby eliminating the need to write a structured method for each bill to extract key information, thereby improving the efficiency of bill information recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative labor.
[0037] Figure 1 It is a flow chart of a bill information identification method in one embodiment of the present invention;
[0038] Figure 2 It is a structural schematic diagram of a bill information identification device in one embodiment of the present invention;
[0039] Figure 3 is a schematic diagram of a computer device in one embodiment of the present invention. DETAILED DESCRIPTION
[0040] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0041] The terms used in one or more embodiments of the present invention are only for the purpose of describing specific embodiments, and are not intended to limit one or more embodiments of the present invention. The singular forms of "a", "said" and "the" used in one or more embodiments of the present invention are also intended to include plural forms, unless the context clearly indicates other meanings. It should also be understood that the term "and / or" used in one or more embodiments of the present invention refers to and includes any or all possible combinations of one or more associated listed items.
[0042] It should be understood that, although the terms first, second, etc. may be used to describe various information in one or more embodiments of the present invention, these information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of the present invention, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "when..." or "when...".
[0043] In one embodiment, specifically, Figure 1 As shown, Figure 1 1 is a flow chart of a method for identifying bill information in an embodiment of the present invention. The present invention provides a method for identifying bill information, comprising the following steps:
[0044] Step S11, obtaining a bill image to be processed;
[0045] The embodiment of the present invention does not limit the implementation scenario. There may be different bills in different scenarios. For example, in a freight scenario, the bill image may be a high-speed toll bill, etc., and may also include other types of bills such as vehicle maintenance invoices.
[0046] In the embodiment of the present invention, the bill can be photographed by a camera device such as a mobile phone, or scanned by a scanning device such as a scanner to obtain the bill image. An image upload port can be provided in the bill entry interface of the bill entry system, and the customer can upload the bill image to the bill entry system through the APP end, or the business personnel can upload the bill image to the bill entry system through the upload port of the counter end.
[0047] In actual situations, the image files obtained by scanning may have blurred images, multiple binary image noise points, and overlapping seals and texts in binary images. In addition, the image files obtained by taking photos may be restricted by various factors such as integrity, flatness, clarity, and tilt angle. In addition, the image files are also affected to varying degrees by the noise contained in the original bills themselves (for example, seal coverage, printer skipping, light ink, invoice serialization, folding marks, etc.). These situations will affect the text detection and recognition of image files. Therefore, after obtaining the bill image, the bill image is preprocessed and data enhanced. Among them, preprocessing includes grayscale conversion, normalization, denoising and other processing methods, and data enhancement processing includes rotation, flipping, cropping, scaling and color conversion and other processing methods.
[0048] Step S12, using the image processing module in the multimodal large model to perform text recognition on the bill image to obtain a text recognition result;
[0049] It should be noted that a multimodal large model refers to a model that can process and understand multiple different types of data input (such as text, images, voice, etc.). These models usually include multiple processing modules, each of which is responsible for data processing tasks in different modalities. For example, a multimodal large model includes an image processing module and a language processing module, where the image processing module can perform image recognition on bill images, and the language processing module can extract information from text recognition results through natural language understanding technology.
[0050] Specifically, the image processing module in the multimodal large model is used to perform text recognition on the bill image. Optionally, in order to improve the recognition accuracy, the text area containing text content in the bill image is first detected; the image processing module in the multimodal large model is used to perform text recognition on the text area to obtain a text recognition result, wherein the text recognition result includes information such as text content information, text position information, and relative position between text contents.
[0051] Step S13, using the language processing module in the multimodal large model to extract information from the text recognition result to obtain bill information.
[0052] Specifically, the language processing module is used to perform semantic entity recognition on the text content information, text position information and relative positions between the text contents in the text recognition results, so as to extract useful information from the text recognition results and obtain the bill information. For example, in the bill scenario, named entity recognition (NER) technology extracts key information such as date, amount, invoice number, merchant name, etc.
[0053] The embodiment of the present invention, through the above scheme, includes: obtaining a bill image to be processed; using the image processing module in the multimodal large model to perform text recognition on the bill image to obtain a text recognition result; using the language processing module in the multimodal large model to extract information from the text recognition result to obtain bill information. The embodiment of the present invention uses a multimodal large model to perform image recognition on the bill image, and combines the multimodal large model to perform natural language processing on text content information, text location information and other information to extract bill information, thereby eliminating the need to write a structured method for each bill to extract key information, thereby improving the efficiency of bill information recognition.
[0054] In one embodiment of the present invention, after extracting information from the text recognition result using the language processing module in the multimodal large model to obtain the bill information, the method further includes:
[0055] The target amount is determined according to the entry station information and the exit station information in the ticket information; the target amount is compared with the amount data in the ticket information to determine the final amount data according to the comparison result.
[0056] It should be noted that, for the transportation scenario, the extracted ticket information includes information such as the entry station information, the exit station information, the amount data, etc. In order to ensure the accuracy of the amount data extracted from the ticket, in one embodiment, the target amount of the expressway is calculated according to the entry station information and the exit station information in the ticket information in accordance with the expressway charging rules. Then, the target amount is compared with the amount data in the ticket information to determine whether the recognized amount data is correct based on the comparison result, thereby determining the final amount data.
[0057] The embodiment of the present invention, through the above scheme, includes: determining the target amount according to the entry station information and the exit station information in the ticket information; comparing the target amount with the amount data in the ticket information to determine the final amount data according to the comparison result. The target amount of the highway is estimated according to the identified entry station information and the exit station information, and then judging whether the identified amount data is correct according to the calculated target amount, so as to determine the final amount data.
[0058] In one embodiment of the present invention, the image processing module in the multimodal large model is used to perform text recognition on the bill image to obtain a text recognition result, including:
[0059] Determine the text area containing text content in the bill image; use the image processing module in the multimodal large model to perform text recognition on the text area to obtain text content information and text position information corresponding to the text content information; determine the relative position between the text contents based on the text content information and the text position information.
[0060] Specifically, the text area in the bill image is extracted by using preset image segmentation techniques, such as edge detection (Canny algorithm), image threshold segmentation, etc., and deep learning methods such as convolutional neural network (CNN) can also be used to realize automatic recognition and positioning of text areas. Then, the image processing module in the multimodal large model is used to perform text recognition on the text area to obtain text content information and text position information corresponding to the text content information. The text content information includes all the text or words recognized in the image, and the text can be numbers, letters or Chinese characters in the bill. The text position information includes the position of each recognized text or word in the image (the coordinates of the bounding box). Furthermore, the relative positions of different text contents are calculated based on the text content information and the text position information. For example, the horizontal or vertical coordinate relationship of two text boxes is compared to calculate the relative positions of different text contents.
[0061] The embodiment of the present invention, through the above scheme, includes: determining the text area containing text content in the bill image; using the image processing module in the multimodal large model to perform text recognition on the text area to obtain text content information and text position information corresponding to the text content information; and determining the relative position between the text contents based on the text content information and the text position information. The multimodal large model is used to perform text recognition on the bill image to obtain information such as text content information, text position information, and the relative position between text contents. The multimodal large model can recognize different types of bills and improve the efficiency of bill information recognition.
[0062] In one embodiment of the present invention, before performing text recognition on the bill image using the image processing module in the multimodal large model and obtaining the text recognition result, the method further includes:
[0063] Obtain a bill sample, wherein the data items in the bill sample are marked with text labels; input the bill sample into the multimodal large model to perform text recognition on the bill sample using an image processing module, and extract information from the recognition result using a language processing module to obtain text extraction information; and fine-tune the multimodal large model based on the text extraction information and the text labels.
[0064] Specifically, a batch of bill samples are collected, wherein each bill sample contains information of various data items (such as amount, date, name, etc.), and the information is marked with text labels (such as "amount", "date", etc.).
[0065] Furthermore, the bill sample is input into the multimodal large model, and the image processing module of the multimodal large model (for example, a deep learning model of computer vision) is used to perform text recognition on the bill image to obtain preliminary text extraction results. Then, the language processing module (for example, natural language processing (NLP) technology) is used to perform text analysis and information extraction on the recognition results to obtain text extraction information. The model associates the recognized text extraction information with the text labels on the bill, and then further trains and fine-tunes the multimodal large model based on the text extraction information and the text labels in the bill sample. By fine-tuning the multimodal large model, it is achieved that it can accurately extract and understand various information on the bill.
[0066] In one embodiment of the present invention, after extracting information from the text recognition result using the language processing module in the multimodal large model to obtain the bill information, the following steps are included:
[0067] Determine the mapping relationship between each data item in the bill information and each data item in the preset bill entry template; assemble the bill information using the mapping relationship to obtain bill entry data in a target format; and enter the bill in the target format into a preset database.
[0068] Specifically, the mapping relationship between each data item in the bill information and each data item in the preset bill entry template is determined, wherein the preset bill entry template can be a template of the bill entry interface, and the preset bill entry template can be set according to actual needs, which is not limited in this embodiment. Further, according to the mapping relationship, each data item in the extracted bill information is mapped with each data item in the preset bill entry template to generate bill entry data in the target format; further, the bill in the target format is entered into the preset database.
[0069] The embodiment of the present invention, through the above scheme, includes: determining the mapping relationship between each data item in the bill information and each data item in the preset bill entry template; assembling the bill information using the mapping relationship to obtain the bill entry data in the target format; and entering the bill in the target format into the preset database. Automatic filling and entry of bill data is achieved, and the bill entry work is accelerated.
[0070] It should be understood that the order of execution of the steps in the above embodiment does not necessarily mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of the present invention.
[0071] In one embodiment, a bill information recognition device is provided, and the bill information recognition device corresponds one-to-one with the bill information recognition method in the above embodiment. Figure 2 As shown, Figure 2 1 is a schematic diagram of a structure of a bill information recognition device in one embodiment of the present invention, the bill information recognition device comprises:
[0072] An acquisition module 21 is used to acquire a bill image to be processed;
[0073] The recognition module 22 is used to perform text recognition on the bill image using the image processing module in the multimodal large model to obtain a text recognition result;
[0074] The extraction module 23 is used to extract information from the text recognition result using the language processing module in the multimodal large model to obtain the bill information.
[0075] The bill information recognition device also includes:
[0076] an amount determination module, for determining a target amount according to the entry station information and the exit station information in the ticket information;
[0077] The comparison module is used to compare the target amount with the amount data in the bill information to determine the final amount data according to the comparison result.
[0078] The bill information recognition device also includes:
[0079] The text recognition result includes text content information, text position information and relative positions between text contents.
[0080] The identification module 22 is also used for:
[0081] Determining a text area containing text content in the bill image;
[0082] Using the image processing module in the multimodal large model to perform text recognition on the text area to obtain text content information and text position information corresponding to the text content information;
[0083] The relative positions of the text contents are determined according to the text content information and the text position information.
[0084] The extraction module 23 is also used for:
[0085] The language processing module is used to perform semantic entity recognition on the text content information, text position information and relative positions between text contents in the text recognition result to obtain the bill information.
[0086] The bill information recognition device also includes:
[0087] A mapping relationship determination module, used to determine the mapping relationship between each data item in the bill information and each data item in the preset bill entry template;
[0088] An assembling module, used to assemble the bill information using the mapping relationship to obtain bill entry data in a target format;
[0089] The input module is used to input the bill in the target format into a preset database.
[0090] The bill information recognition device also includes:
[0091] A bill sample module, used to obtain a bill sample, wherein the data items in the bill sample are marked with text labels;
[0092] An image processing module, used for inputting the bill sample into the multimodal large model, so as to perform text recognition on the bill sample using the image processing module, and extract information from the recognition result using the language processing module to obtain text extraction information;
[0093] A model fine-tuning module is used to fine-tune the multimodal large model based on the text extraction information and the text label.
[0094] The bill information recognition device also includes:
[0095] The processing module is used to perform preprocessing and data enhancement processing on the bill image.
[0096] For the specific definition of the bill information identification device, please refer to the definition of the bill information identification method above, which will not be repeated here. Each module in the above-mentioned bill information identification device can be implemented in whole or in part by software, hardware and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.
[0097] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 3 As shown, Figure 3: is a schematic diagram of a computer device in an embodiment of the present invention. The computer device includes a processor, a memory, a network interface and a database connected by a device bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a readable storage medium and an internal memory. The readable storage medium stores an operating device, a computer-readable instruction and a database. The internal memory provides an environment for the operation of the operating device and the computer-readable instructions in the readable storage medium. The database of the computer device is used to store data involved in the bill information identification method. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer-readable instruction is executed by the processor, a bill information identification method is implemented. The readable storage medium provided in this embodiment includes a non-volatile readable storage medium and a volatile readable storage medium.
[0098] In one embodiment, a computer device is provided. The computer device may be a terminal device, and its internal structure diagram may be as follows: Figure 3 As shown. The computer device includes a processor, a memory, and a network interface connected through a device bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a readable storage medium. The readable storage medium stores computer-readable instructions. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer-readable instructions are executed by the processor, a method for identifying bill information is implemented. It includes: obtaining a bill image to be processed; using the image processing module in the multimodal large model to perform text recognition on the bill image to obtain a text recognition result; using the language processing module in the multimodal large model to extract information from the text recognition result to obtain bill information. The readable storage medium provided in this embodiment includes a non-volatile readable storage medium and a volatile readable storage medium.
[0099] In one embodiment, a computer device is provided, including a memory, a processor, and computer-readable instructions stored in the memory and executable on the processor, and when the processor executes the computer-readable instructions, the steps of the above-mentioned bill information recognition method are implemented. The method includes: obtaining a bill image to be processed; performing text recognition on the bill image using an image processing module in a multimodal large model to obtain a text recognition result; and extracting information from the text recognition result using a language processing module in the multimodal large model to obtain bill information.
[0100] In one embodiment, a readable storage medium is provided, the readable storage medium stores computer-readable instructions, and when the computer-readable instructions are executed by a processor, the steps of the above-mentioned bill information identification method are implemented. A person of ordinary skill in the art can understand that the implementation of all or part of the process in the above-mentioned embodiment method can be completed by instructing the relevant hardware through computer-readable instructions, and the computer-readable instructions can be stored in a non-volatile readable storage medium or a volatile readable storage medium. When the computer-readable instructions are executed, they can include the process of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0101] In one embodiment, a computer program product is provided, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the bill information identification method provided by the above-mentioned methods.
[0102] Those skilled in the art can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0103] The embodiments described above are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. Such modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention.
Claims
1. A bill information recognition method, characterized in that: include: Get the bill image to be processed; Using the image processing module in the multimodal large model to perform text recognition on the bill image to obtain a text recognition result; The language processing module in the multimodal large model is used to extract information from the text recognition result to obtain the bill information.
2. The bill information recognition method according to claim 1, characterized in that: After extracting information from the text recognition result using the language processing module in the multimodal large model to obtain the bill information, the method further includes: Determining a target amount according to the entry station information and the exit station information in the ticket information; The target amount is compared with the amount data in the bill information to determine the final amount data according to the comparison result.
3. The bill information recognition method according to claim 1, characterized in that: The text recognition result includes text content information, text position information and relative positions between text contents; The method of using the image processing module in the multimodal large model to perform text recognition on the bill image to obtain a text recognition result includes: Determining a text area containing text content in the bill image; Using the image processing module in the multimodal large model to perform text recognition on the text area to obtain text content information and text position information corresponding to the text content information; The relative positions of the text contents are determined according to the text content information and the text position information.
4. The bill information recognition method according to claim 1, characterized in that: The method of extracting information from the text recognition result using the language processing module in the multimodal large model to obtain bill information includes: The language processing module is used to perform semantic entity recognition on the text content information, text position information and relative positions between text contents in the text recognition result to obtain the bill information.
5. The bill information recognition method according to claim 1, characterized in that: The method of extracting information from the text recognition result by using the language processing module in the multimodal large model to obtain the bill information includes: Determine the mapping relationship between each data item in the bill information and each data item in the preset bill entry template; Assembling the bill information using the mapping relationship to obtain bill entry data in a target format; The bill in the target format is entered into a preset database.
6. The bill information recognition method according to claim 1, characterized in that: The method of using the image processing module in the multimodal large model to perform text recognition on the bill image and obtaining the text recognition result also includes: Acquire a bill sample, wherein the data items in the bill sample are marked with text labels; Input the bill sample into the multimodal large model, use the image processing module to perform text recognition on the bill sample, and use the language processing module to extract information from the recognition result to obtain text extraction information; The multimodal large model is fine-tuned based on the text extraction information and the text labels.
7. The bill information recognition method according to claim 1, characterized in that: After obtaining the bill image to be processed, the method further includes: The bill image is preprocessed and data enhanced.
8. A bill information recognition device, characterized in that: include: An acquisition module, used for acquiring the bill image to be processed; A recognition module, used to perform text recognition on the bill image using the image processing module in the multimodal large model to obtain a text recognition result; The extraction module is used to extract information from the text recognition result using the language processing module in the multimodal large model to obtain the bill information.
9. A computer device comprising a memory, a processor, and computer-readable instructions stored in the memory and executed on the processor, characterized in that: When the processor executes the computer-readable instructions, the bill information recognition method according to any one of claims 1 to 7 is implemented.
10. A readable storage medium having computer readable instructions stored thereon, characterized in that: When the computer-readable instructions are executed by a processor, the bill information recognition method according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
File structured information extraction method and device, equipment, medium and product
CN120849649A
File structured information extraction method, device, equipment, medium and product
CN120849649B
Foreign language bill image translation method and device based on artificial intelligence and medium
CN120913229A