Bill identification method, device and equipment and storage medium
By using multimodal feature extraction and decoder to determine the invoice recognition model, the problem of inaccurate invoice recognition is solved, especially when the character information is unclear or blurry, thus improving the accuracy of invoice recognition and the shopping experience.
Patent Information
- Application Number
- CN202511061763.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-31
- Publication Date
- 2026-01-02
- Estimated Expiration
- 2045-07-31
AI Technical Summary
Existing ticket recognition technologies are not accurate enough when dealing with unclear or blurry character information, especially irregular characters or font icons, leading to errors in ticket information extraction.
The trained invoice recognition model is used to extract multimodal features, combining visual and textual features. It utilizes a visual encoder, a pre-set large language model, and a cross-modal fusion device to generate multimodal features, and then uses an autoregressive language model decoder to determine the target merchant information.
It improves the accuracy of receipt recognition, especially when the character information is unclear or blurry, it can more accurately identify merchant information, reduce misidentification, and enhance the shopping experience.
Smart Images

Figure CN120564218B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of image recognition, and particularly relates to a bill identification method and device, equipment and a storage medium. BACKGROUND
[0002] With the continuous development of e-commerce, digital identification and analysis of bills have become increasingly important. Merchants or platforms usually provide bill identification and analysis services, and users can directly upload bill images to obtain corresponding bill information, thereby improving the shopping experience of users.
[0003] Currently, bill information in a bill image can be extracted through image recognition technology. However, in the image recognition process, the character information in the bill image is usually a font icon or a special-shaped character, and single image recognition can easily lead to incomplete character information recognition, which seriously affects the accuracy of bill identification. SUMMARY
[0004] The bill identification method, device, equipment and storage medium provided by the embodiments of the application can solve the problem of inaccurate bill identification.
[0005] In a first aspect, the embodiments of the application provide a bill identification method, comprising: obtaining an image of a bill to be identified; processing the image of the bill to be identified, a preset prompt word and preset standard merchant information by using a trained bill identification model to determine target merchant information corresponding to the bill to be identified; wherein the preset prompt word is used to instruct the trained bill identification model to perform multi-modal feature extraction on the image of the bill to be identified, and the multi-modal feature is generated based on visual features of the image of the bill to be identified and text features of the image of the bill to be identified; and the preset prompt word is also used to instruct the trained bill identification model to determine the target merchant information corresponding to the bill to be identified according to the multi-modal feature.
[0006] In a possible implementation manner of the first aspect, the multi-modal feature of the bill to be identified is a feature obtained by fusing different types of information of the bill to be identified. The different types of information can be complementary and aligned, so that the multi-modal feature of the bill to be identified can more accurately represent the image of the bill to be identified, and the obtained target merchant information is more accurate. The problem of inaccurate bill identification caused by using single modal for image recognition in the case that the image of the bill to be identified contains unclear or blurred character information, or the character information in the bill image is a font icon or a special-shaped character is solved.
[0007] In a second aspect, the embodiments of the application provide a bill identification device, comprising:
[0008] An obtaining module is configured to obtain an image of a bill to be identified;
[0009] The determining module is configured to determine target merchant information corresponding to the bill to be identified by processing the image of the bill to be identified, the preset prompt word, and the preset standard merchant information using the trained bill identification model; the preset prompt word is used to instruct the trained bill identification model to perform multi-modal feature extraction on the image of the bill to be identified, and the multi-modal feature is generated based on visual features of the image of the bill to be identified and text features of the image of the bill to be identified; and the preset prompt word is also used to instruct the trained bill identification model to determine the target merchant information corresponding to the bill to be identified according to the multi-modal feature.
[0010] In a third aspect, an electronic device is provided, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the method of the first aspect when executing the computer program.
[0011] In a fourth aspect, a computer readable storage medium is provided, which stores a computer program, and the computer program, when executed by a processor, implements the steps of the method of the first aspect or the second aspect.
[0012] In a fifth aspect, a computer program product is provided, which, when executed on an electronic device, causes the electronic device to perform the method of the first aspect.
[0013] It can be understood that the beneficial effects of the second aspect to the fifth aspect can be referred to the related description in the first aspect, and will not be repeated here. BRIEF DESCRIPTION OF DRAWINGS
[0014] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0015] Figure 1 is a flowchart of a bill identification method provided by the embodiments of the present application;
[0016] Figure 2 is a flowchart of a multi-modal feature determination process provided by the embodiments of the present application;
[0017] Figure 3 is a flowchart of a training process of a bill identification model provided by the embodiments of the present application;
[0018] Figure 4is a flowchart of a process of identifying an image of a bill to be identified by using a double identification mechanism, provided by an embodiment of the present application;
[0019] Figure 5 is a structural diagram of a bill identification device, provided by an embodiment of the present application;
[0020] Figure 6 is a structural diagram of another bill identification device, provided by an embodiment of the present application;
[0021] Figure 7 is a structural diagram of an electronic device, provided by an embodiment of the present application. DETAILED DESCRIPTION
[0022] In the following description, for the purposes of explanation and not limitation, specific details are set forth, such as particular system configurations, techniques, etc., in order to provide a thorough understanding of the embodiments of the application. However, it will be apparent to those skilled in the art that the application can be practiced in other embodiments that depart from these specific details. In other instances, detailed descriptions of well-known systems, devices, circuits, and methods are omitted so as not to obscure the description of the application with unnecessary detail.
[0023] It is to be understood that the terminology "includes", "has", "holds", "contains" and / or "comprising", when used in this specification and in the following claims, indicates the presence of the described features, integers, steps, operations, elements, and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0024] It is also to be understood that the terminology "and / or" when used in this specification and in the following claims, refers to at least one of the items, or any combination of one or more of the items, and includes any possible combination of the items.
[0025] As used in this specification and in the claims, the terms "if" and "when" can be interpreted to mean "upon" or "in response to a determination" or "in response to a detection" depending on the context. Similarly, the phrase "if it is determined" or "if [a described condition or event] is detected" can be interpreted to mean "upon determining" or "in response to determining" or "upon detecting [the described condition or event]" or "in response to detecting [the described condition or event]" depending on the context.
[0026] In addition, in the description of the specification and the appended claims, the terms "first", "second", "third", etc. are used only to distinguish descriptions, and cannot be understood as indicating or implying relative importance.
[0027] Reference within the specification of this application to“one embodiment” or“some embodiments” means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the application. The appearances of the phrase“in one embodiment” or“in some embodiments” in various places within specified descriptions in this specification are not necessarily all referring to the same embodiment, however, but can refer to one or more but not all embodiments. The terms“including,”“comprising,”“carrying,”“having,” and variations thereof are meant to cover a non-exclusive inclusion, subject to other explicit limitations only. The term“consisting essentially of’ means“including, but not limited to.”
[0028] In a bill identification and analysis service provided by a merchant or a platform, a user can directly obtain corresponding bill information according to an uploaded bill image, and the merchant or the platform can automatically implement bill analysis according to the identified bill information, thereby improving the shopping experience of the user. For example, in the scene of automatic accumulation of shopping bills, the merchant or the platform can identify the bill information such as the merchant name, transaction time, transaction amount, and transaction voucher according to the bill image uploaded by the user, and then implement automatic accumulation according to the bill information, thereby improving the user retention rate. Among them, the accurate identification of bill information in the bill image is crucial for bill analysis.
[0029] Currently, the bill information in the bill image is mainly extracted by the following methods: 1. Template matching technology: template matching is a traditional image recognition method based on rules. For shopping bills of different formats, a template is designed based on predefined rules and prior knowledge such as text position information. The input image is compared with the preset template by template matching to quickly locate and extract bill information. 2. Deep learning-based technical solution: automatically learn the text features in the bill image through a neural network model to realize text detection and extract bill information. 3. Technical solution based on OCR technology and large language model: extract the original text content from the image through OCR technology, and then combine the powerful semantic understanding ability of the large language model to correct the original text content, logical reasoning, information extraction, and structured output.
[0030] However, the above technical solution for extracting bill information in the bill image has the following disadvantages: 1. The traditional template matching technology cannot adapt to the diversity of shopping formats, has weak generalization ability, and is difficult to cope with frequently changing formats or complex layout document content, resulting in high error rate of bill information extraction. 2. The method based on deep learning needs to construct a large-scale labeled data set, including text box coordinate labeling, text content labeling, etc. The labeling process is complex and costly, especially for complex format document labeling efficiency is significantly reduced. Moreover, for low contrast, blurred handwriting and other degraded images, the model generalization ability is limited and the robustness is insufficient, and additional data enhancement or prior knowledge needs to be introduced for retraining. 3. The text errors (such as similar character confusion, picture blur or wrinkles causing partial font missing) in the original text content extracted from the image by the OCR technology will directly affect the semantic understanding of the large language model, forming an "error amplification effect", especially for identifying "merchant name" and other character information, it is easy to be misrecognized because of the traditional Chinese characters or artistic characters in the bill picture, thereby affecting the accuracy of bill recognition, which may eventually cause the shopping bill points to be accumulated to other merchants, affecting the significance. Therefore, the above method will lead to inaccurate bill recognition.
[0031] In order to improve the accuracy of bill recognition, the embodiment of the present application provides a bill recognition method, acquiring an image of a bill to be recognized; adopting a trained bill recognition model to process the image of the bill to be recognized, a preset prompt word and a preset standard merchant information, and determining target merchant information corresponding to the bill to be recognized. The multi-modal feature of the bill to be recognized is a feature obtained by fusing different types of information of the bill to be recognized. Different types of information can be complementary and aligned, capturing the relevance between visual features and text features in the image of the bill to be recognized. Specifically, in the case that the image of the bill to be recognized is blurred, occluded or has other image quality problems, resulting in incomplete extraction of text features, the corresponding clues can still be provided by visual features. Conversely, if the visual features are not clear enough, more accurate semantic support can be provided by the text features, so that the multi-modal feature of the bill to be recognized can more accurately represent the image of the bill to be recognized, and the obtained target merchant information is more accurate. The problem of inaccurate bill recognition caused by using a single mode for image recognition in the case that the image of the bill to be recognized contains unclear or blurred character information in the bill image, which is a font icon or a special-shaped character, is solved.
[0032] Figure 1 A flowchart of a bill recognition method provided by the embodiment of the present application is shown, as shown in Figure 1 The bill recognition method provided by the embodiment of the present application includes S101-S102, wherein:
[0033] S101, acquiring an image of a bill to be recognized.
[0034] The image of the to-be-identified ticket can be an image of a shopping receipt uploaded by a user through a terminal device (e.g., a mobile phone, a computer, etc.), can be a photo of a paper receipt printed by a POS machine, or can be a screenshot of a shopping order in an application.
[0035] In S102, the trained ticket identification model is used to process the image of the to-be-identified ticket, the preset prompt word, and the preset standard merchant information, to determine target merchant information corresponding to the to-be-identified ticket.
[0036] The preset prompt word is used to instruct the trained ticket identification model to perform multi-modal feature extraction on the image of the to-be-identified ticket. The multi-modal feature is generated based on visual features of the image of the to-be-identified ticket and text features of the image of the to-be-identified ticket. The preset prompt word is also used to instruct the trained ticket identification model to determine the target merchant information corresponding to the to-be-identified ticket according to the multi-modal feature. The preset prompt word can help the trained ticket identification model to quickly perform multi-modal feature extraction and determine the target merchant information, reduce the interference of irrelevant information in the ticket identification process, and thus improve the identification efficiency.
[0037] For example, the preset standard merchant information at least includes a complete merchant name.
[0038] In some embodiments, the trained ticket identification model is used to determine the target merchant information corresponding to the to-be-identified ticket mainly includes two processes: 1, a multi-modal feature extraction process of the to-be-identified ticket, and 2, a process of determining the target merchant information corresponding to the to-be-identified ticket by using the multi-modal feature.
[0039] The two processes will be described below.
[0040] The process of extracting the multi-modal feature of the to-be-recognized bill: mainly realized by three modules of the trained bill recognition model, including a vision encoder, a preset large language model (LLM), and a cross-modal fusioner. Specifically, based on the above three modules, the process of extracting the multi-modal feature of the to-be-recognized bill by the trained bill recognition model according to the indication included in the preset prompt word is as follows: the vision encoder converts the original visual signal (i.e., the image of the to-be-recognized bill) into structured visual features that can be understood by machines; the preset LLM processes the text information corresponding to the image of the to-be-recognized bill according to the preset prompt word to generate text features; the cross-modal fusioner, as a "bridge" connecting the two modalities of vision and text, fuses the visual features and the text features to realize the association and understanding of the multi-modal information of the to-be-recognized bill. In combination with the specific implementation process, the preset prompt word can also be used to specifically indicate the visual features of the image of the to-be-recognized bill in the vision encoder of the trained bill recognition model, indicate the text features of the image of the to-be-recognized bill in the preset LLM of the trained bill recognition model, and indicate the fusion of the visual features and the text features in the cross-modal fusioner of the trained bill recognition model to obtain the multi-modal feature of the to-be-recognized bill.
[0041] It can be understood that the multi-modal feature extraction model composed of the vision encoder, the preset LLM, and the cross-modal fusioner can also be included in the trained bill recognition model.
[0042] The process of determining the target merchant information corresponding to the to-be-recognized bill by using the multi-modal feature: according to the indication of the trained bill recognition model by the preset prompt word, the trained bill recognition model determines the target merchant information corresponding to the to-be-recognized bill according to the multi-modal feature. The multi-modal feature of the to-be-recognized bill can be used as a context to predict the highest probability word in the all possible merchant-related word table output by the trained multi-modal extraction model through the autoregressive language model decoder in the trained bill recognition model based on the multi-modal feature of the to-be-recognized bill, and determine the target merchant information.
[0043] The autoregressive language model decoder can be any of the following: any of the Generative Pre-trained Transformer (GPT) series, Text-To-Text Transfer Transformer (T5), or Bidirectional and Auto-Regressive Transformers (BART).
[0044] For example, the trained invoice recognition model is trained based on data in a preset format. Therefore, when the trained invoice recognition model receives the image of the invoice to be recognized, preset prompt words, and preset standard merchant information, it first converts these data into a preset format to adapt to the data format requirements that the trained invoice recognition model can handle. Then, it performs feature extraction or recognition on the preset format data. This embodiment uses JSON format as an example for illustration: [
[0046] {
[0047] "id": "REC_0001",
[0048] "conversations": [
[0049] {
[0050] "from": "user",
[0051] "value": " Please use the image of the shopping receipt and its OCR text information to extract the merchant name from the receipt according to the following steps: ... Receipt OCR text: [[('Starbucks Coffee China', 0.99421984), ... Candidate merchants: [StarbucksCoffeeChina, Chow Tai Fook, ...]"
[0052] },
[0053] {
[0054] "from": "assistant",
[0055] "value": "StarbucksCoffeeChina"
[0056] } ]
[0058] } ...... ]
[0060] The preset prompt word can be:
[0061] Please extract the merchant name in the shopping receipt according to the following steps combined with the shopping receipt picture and its OCR text information:
[0062] Recognize the merchant name: Recognize the visual features related to the LOGO and trademark in the picture, and complete the verification of the merchant name according to the provided
ticket OCR text
[0063] Match the merchant name: Match the recognized merchant name with the provided
candidate merchants
[0064] Strictly output the json format: ```json {"brand_match": "merchant name" or "-"}```.
[0065] The above prompt word is only an example and does not limit the embodiments of the present application. The preset prompt word is artificially designed to guide the trained bill recognition model in the embodiments of the present application, and indicates the trained bill recognition model in the present application to recognize the target merchant information in the image of the bill to be recognized.
[0066] The preset standard merchant information can be input by the user. For example, the complete merchant name of all merchants in a certain supermarket or shopping platform can be collected to obtain a standard merchant name list, and the standard merchant name list is used as the preset labeled merchant information.
[0067] To improve the accuracy of bill identification, the embodiment of the application provides a bill identification method. The trained bill identification model is used to process the image of the bill to be identified, the preset prompt word and the preset standard merchant information, and the target merchant information of the bill to be identified is obtained. The preset prompt word is used to instruct the trained bill identification model to perform multi-modal feature extraction on the image of the bill to be identified. The multi-modal feature is generated based on the visual feature of the bill to be identified image and the text feature of the bill to be identified image. The preset prompt word is also used to instruct the trained bill identification model to determine the target merchant information corresponding to the bill to be identified according to the multi-modal feature. Since the multi-modal feature of the bill to be identified is a feature obtained by fusing different types of information (for example: visual feature and text feature) of the bill to be identified, different types of information can be complementary and aligned, and the correlation between the visual feature and the text feature in the image of the bill to be identified can be captured. Specifically, in the case that the text feature extraction is incomplete due to blurring, occlusion or other image quality problems in the image of the bill to be identified, the corresponding clues can still be provided through the visual feature. Conversely, if the visual feature is not clear enough, more accurate semantic support can be provided through the text feature, so that the multi-modal feature of the bill to be identified can more accurately represent the image of the bill to be identified. The target merchant information obtained based on the multi-modal feature of the bill to be identified is also more accurate. The problem of inaccurate bill identification caused by using a single mode for image recognition in the case that the image of the bill to be identified contains unclear or blurred character information, which is a font icon or a special-shaped character, is solved.
[0068] In some embodiments, Figure 2 A flowchart of a multi-modal feature determination process is shown in Figure 2 As shown in the figure, the determination process of the multi-modal feature of the bill to be identified specifically includes S1021 to S1023, wherein:
[0069] S1021, using the visual encoder in the trained bill identification model, performing visual coding on the image of the bill to be identified to determine the visual feature sequence.
[0070] In some embodiments, the visual encoder in the trained bill identification model is used to segment the image of the bill to be identified to obtain a plurality of segmented image blocks. The visual feature sequence is determined according to the plurality of image blocks obtained by segmenting the image of the bill to be identified, the position code corresponding to each image block and the image modal identifier of each image block. The size of each image block can be the same or different.
[0071] Specifically, the visual encoder in the trained bill recognition model is used for image patching (Patch Embedding) on the image of the bill to be recognized, to obtain a plurality of image patches, map the plurality of image patches to a unified dimension size of the multi-modal feature extraction model, obtain an image patch sequence, add position encoding and image feature labels to each image patch in the image patch sequence, and form a visual feature sequence. The visual feature sequence is represented by a vector, and each vector represents the visual feature corresponding to the image patch in the image of the bill to be recognized.
[0072] It should be noted that the visual encoder can only process images of a specific size. In the case where the image size of the bill to be recognized is exactly the image size that the visual encoder can process, S1021 is directly executed. In the case where the image size of the bill to be recognized does not conform to the image size that the visual encoder can process, the image size of the bill to be recognized needs to be adjusted to the image size that the visual encoder can process before S1021 is executed.
[0073] For example, the embodiment of the present application is described by taking the implementation of the visual encoder using the improved ViT (Vision Transformer) architecture as an example. S1021 is described as follows: based on the improved ViT (Vision Transformer) architecture, the image size of the bill to be recognized is dynamically adjusted to the resolution adapted to the visual encoder, and the image of the bill to be recognized is normalized by using the mean and standard deviation to make the pixel distribution consistent. The image is cut into 14x14 non-overlapping small blocks, and each image block is flattened into one-dimensional data, mapped to d_model consistent with the preset large language model dimension through a linear layer to form patch embeddings, and then position encoding is added to each patch embedding to more accurately capture the spatial position information of the image. Finally, it is input into the multi-layer Transformer Encoder block. Each layer will perform self-association calculation and further feature conversion on these vectors.
[0074] S1022, using the preset large language model in the trained bill recognition model, text encoding is performed on the first text corresponding to the image of the bill to be recognized to determine a text feature sequence.
[0075] The first text includes the text corresponding to the image of the bill to be recognized, the preset prompt word, and the preset standard merchant information.
[0076] Exemplarily, the text corresponding to the image of the to-be-identified bill can include text information corresponding to a merchant name, a commodity name, a transaction amount, a quantity, an order number, a date, and the like in the to-be-identified bill, and can also include confidence of the text information and position information of the text information, and the like, wherein the confidence of the text information is used to reflect accuracy of a bill text recognition result.
[0077] In some embodiments, before step S1022, further comprising: identifying the image of the to-be-identified bill by a preset Optical Character Recognition (OCR) model in the trained bill identification model, to obtain text corresponding to the image of the to-be-identified bill.
[0078] The preset OCR model can be a PP-OCRv4 model, or can be another character recognition model, which is not limited in the embodiments of the present application.
[0079] In some embodiments, different OCR models can receive different input images, the image of the bill to be recognized can be preprocessed to obtain an image meeting the input requirements of the OCR model, and then the image meeting the input requirements of the OCR model is recognized by the OCR model to obtain the text corresponding to the image of the bill to be recognized. For example, taking the preset OCR model PP-OCRv4 as an example, the bill image can be cropped to obtain a cropped bill image, the cropped bill image can be converted into a Base64 encoding format, and the bill image in the Base64 encoding format can be encapsulated into a specified key-value pair structure (for example, {"key": ["image"], "value": [base64_str]}), so that the PP-OCRv4 model can recognize it. The PP-OCRv4 model can output a response result such as {'err_no': 0, 'err_msg':, 'key': ['result'], 'value': ["[[('Starbucks Coffee China', 0.99421984), [[159.0, 11.0], [375.0, 44.0],[371.0, 69.0], [155.0, 36.0]]],[('296.00', 0.9857364), [[233.0, 622.0],[271.0, 626.0], [269.0, 636.0], [231.0, 635.0]]]]"],...... 'tensors': []}. The text information corresponding to the bill image can be all or part of the information in the response result output by the PP-OCRv4 model. For example, the text corresponding to the image of the bill to be recognized can include the text content, confidence, position coordinates and other effective information contained in the 'value' field.
[0080] In some optional embodiments, the above-mentioned preset large language model can be one of ChatGLM4, Llama, Qwen2.5 and GPT.
[0081] In some embodiments, the first text is segmented by a segmenter in the preset large language model to obtain a text word sequence, and a text feature sequence is determined according to each text word in the text word sequence, a text position encoding of each text word and a text modal identifier of each text word.
[0082] For example, a specific word segmenter in a pre-defined large language model is used to convert the first text into a series of discrete tokens. This series of discrete tokens forms a token sequence. A token embedding table is used to map the token sequence into a sequence of word embedding vectors. This sequence of word embedding vectors has the same dimension as the pre-defined large language model, which is also d_model dimension, resulting in a text embedding sequence. Text position encoding and text modality pattern are then added to each text word in the text embedding sequence to obtain a text feature sequence. It is important to note that the text position encoding shares a unified position space with the visual position encoding to ensure consistency in positional information across different modalities, thereby avoiding positional conflicts between modalities and improving the stability and accuracy of multimodal fusion.
[0083] S1023, based on a cross-modal attention mechanism, fuses visual feature sequences and text feature sequences to generate multimodal features.
[0084] In some embodiments, a prefix concatenation strategy is employed, where visual features always precede text features, forming a visual-text joint sequence. This combined visual-text sequence is then input into one or more cross-modal Transformer layers. A "cross-association" mechanism allows image and text information to understand each other, establishing a visual-text association through cross-modal attention. For example, the token corresponding to "Starbucks" in the text will focus on the visual patch containing the logo in the image, achieving cross-modal information alignment.
[0085] In some embodiments, a cross-modal attention mechanism is used to calculate the attention weights of the image features and text features corresponding to the sample ticket image, thereby capturing the semantic alignment relationship between the image and the text. Based on the attention weights, the image feature sequence Vimage and the text feature sequence Vtext corresponding to the sample ticket image are weighted and fused to obtain the fused multimodal sequence, i.e., the multimodal features.
[0086] In some embodiments, before step S102, a trained invoice recognition model needs to be obtained. This application embodiment also provides a process for training the invoice recognition model and obtaining the trained invoice recognition model, such as... Figure 3 As shown, the process of training the initial invoice recognition model specifically includes:
[0087] S301, Obtain the training set of sample ticket images, sample prompts, and sample merchant information.
[0088] The sample bill image training set includes a plurality of sample bill images. The sample bill image training set can include bill images of different types. For example, sample shopping receipt images from different retailers, different sizes, and different print quality can be collected as sample bill images. The sample shopping receipt images can cover bill information in different languages, fonts, font sizes, and layouts. The sample shopping receipt images can include at least one of the following bill information: product name, price, quantity, order number, serial number, barcode, date, merchant information, etc. Meanwhile, the sample shopping receipt images can include photographed pictures of paper shopping receipts and screenshots of shopping using program applications, etc.
[0089] For example, shopping receipt images of more than 10 retail formats (e.g., shopping malls / convenience stores / catering / clothing, etc.) can be collected. The shopping receipt image collection channels can include: (a) photos taken by physical entity receipts: images of paper receipts taken by users using mobile phones are collected in the background. The photographed receipt pictures should cover multiple scenarios such as blur, tilt, folding, and partial occlusion; (b) electronic voucher screenshots: screenshots of order pages of mainstream e-commerce platforms (Meituan / Douyin / Jingdong) are intercepted. The order page screenshots must retain page LOGO, ad pop-up, and other interference elements.
[0090] For example, the sample shopping receipt images can also be classified by data. Specifically, the collected shopping receipt images are classified into the following three categories: 1. standard receipt (i.e., those clear, unfolded, and intact receipts selected from the photographed physical entity receipts); 2. screenshot receipt (i.e., electronic voucher screenshots); and 3. adversarial sample (i.e., photographed physical entity receipts with problems such as blur, folding, occlusion, and damage). In the embodiments of the present application, the above-mentioned three types of images can be collected in a fixed proportion, for example, 50% of standard receipts, 35% of screenshot receipts, and 15% of adversarial samples. The embodiments of the present application can also collect the above-mentioned three types of images in other proportions, which are not limited in the embodiments of the present application.
[0091] The sample bill image training set further includes target text corresponding to each sample bill image. Specifically, each sample bill image can be labeled according to a preset field type to obtain a key bill field, and then a required field is selected from the key bill field to obtain the target text corresponding to each sample bill image. The preset field type can include a name type, an amount type, a time type, and the like, and then the target text corresponding to each sample bill image obtained by labeling is recorded in a data table. Alternatively, a specified field in the key bill field can be labeled according to a need of a preset business analysis scenario, and then recorded in the data table. For example, assuming that the preset business analysis scenario is a bill automatic scoring scenario, the target text corresponding to each sample bill image can include a labeled store name, time, amount, and the like, which are required field types of the bill automatic scoring scenario.
[0092] It can be understood that the labeling can be implemented by using an OCR model. First, the OCR model is used to extract original OCR text corresponding to the sample bill image. The original OCR text includes character-level errors, sentence break errors, and format confusion errors. Second, according to the original OCR output text, text confidence, and text coordinate information, information of a key field to be extracted from each shopping bill, such as a merchant name, a transaction time, and an amount, is manually labeled. When the merchant name is manually labeled, the standard merchant data corresponding to the store is considered comprehensively. It needs to be noted whether the merchant name information in the original OCR text can be mapped to the collected standard merchant data. For example, if the overall merchant name recognition is affected by a wrong character or a few characters, the merchant name is labeled as “-”. If the overall recognition is not affected, the merchant name is labeled as the merchant name in the standard merchant data. When the order number or the transaction amount is manually labeled, the characters recognized in the original OCR text are used as a reference. The specified characters are labeled as the order number or the transaction amount. The data table is used to record the name of the bill picture file and the merchant name corresponding to each bill. When manually labeling, if the merchant name text in the picture can be recognized by the human eye, the actual standard merchant data is labeled. If the picture is blocked or incomplete and the merchant name cannot be recognized by the human eye, the merchant name is labeled as “-”.
[0093] S302, using an initial bill recognition model, processing each sample bill image, a sample prompt word, and sample merchant information to obtain target merchant information of each sample bill image.
[0094] The sample prompt word is used to instruct the initial bill recognition model to perform multi-modal feature extraction on each sample bill image. The multi-modal feature is generated based on a visual feature of the sample bill image and a text feature of the sample bill image. The sample prompt word is further used to instruct the initial bill recognition model to determine target merchant information corresponding to the sample bill image according to the multi-modal feature of the sample bill image.
[0095] The sample merchant information is preset merchant information corresponding to the sample bill image, for example, a complete set of merchant names of all merchant information in a shopping mall corresponding to a sample bill image.
[0096] It should be noted that the specific implementation of S302 can refer to the specific implementation of step S102 described above, and the embodiments of the present application will not be repeated here.
[0097] S303, according to the multi-modal feature of each sample bill image, the target text corresponding to each sample bill image, and the preset loss function, determining the loss value of each sample bill image.
[0098] The target text is the real merchant information corresponding to the sample bill image.
[0099] For example, the preset loss function is any one of the following loss functions: image-text contrast loss function, image-text matching loss function and language modeling loss function.
[0100] In some embodiments, during the process of training the initial bill recognition model, when the loss value meets the preset loss value requirement, the training of the initial bill recognition model can be stopped, and the trained bill recognition model is obtained.
[0101] For example, for each sample bill image, step S303 can be implemented by a self-recurrent decoder, which takes the multi-modal feature of the sample bill image as the initial context and generates the merchant information text, i.e. the target merchant information, by self-recurrently. The self-recurrent decoder is a multi-modal feature generated according to the previous visual + text feature, which predicts the next token by self-recurrence. The self-recurrent mechanism is: from the first token, after generating a token at each step, the token is added to the context, and the next token is predicted based on the updated context, until the end symbol (such as <end>). Specifically, the autoregressive decoder predicts the next token based on the initial context, choosing the token with the highest probability from the vocabulary (containing all possible merchant-related words) and adding it to the generated sequence. The updated sequence is then used as the new context for the next prediction, and this process continues until the complete merchant information text is predicted, serving as the sample merchant information to be confirmed. The loss value of the sample invoice image is determined using the sample merchant information to be confirmed, the target text corresponding to the sample invoice image, and a pre-set loss function.
[0102] Specifically, the language model (essentially an autoregressive decoder) uses the multimodal features as context and autoregressively predicts the next token (i.e., a word that may appear in the target merchant information).
[0103] The self-attention layer uses a causal mask to ensure that each position can only see itself and previous tokens, while the cross-attention layer is a key step for the language model to "understand" the joint image + text context. It allows each generated text token to query the entire fused context sequence and selectively focus on the most relevant part of the input sequence (source language, context) for the current generation step (e.g., when generating "Starbucks Coffee China", the source language's logo, Starbucks, needs to be focused on). The core of calculating the loss value of the sample invoice image lies in measuring the difference between the sample merchant information to be confirmed and the target text corresponding to the sample invoice image. Specifically, the loss function compares the difference between the model's prediction (Pj) at the reply start position and subsequent positions and the true merchant information (target text corresponding to the sample invoice image) position by position.
[0104] S304, updating the model parameters of the initial invoice recognition model using the loss value of each sample invoice image to obtain a trained multimodal feature extraction model.
[0105] The loss value is updated for the model parameters (including all learnable parameters of the encoder and decoder) in the initial invoice recognition model through the backpropagation mechanism. For example, the model parameters can include model weights (query_key_value), attention weights, and feedforward network (FFN) weights.
[0106] It should be noted that this process effectively guides the model to learn how to generate accurate and semantically relevant replies based on the input image information and context content. Through continuous optimization, the model gradually establishes deep semantic associations between images, context questions, and expected answers, thereby improving its understanding and generation capabilities in multimodal tasks.
[0107] To further improve the accuracy of bill identification, after the trained bill identification model provided in the embodiments of the present application obtains the target merchant information of the to-be-identified bill image, the target merchant information of the to-be-identified bill image can be further confirmed to obtain the verified target merchant information corresponding to the to-be-identified bill image. Specifically, the target merchant information of the to-be-identified bill image is matched with the preset standard merchant information. If the target merchant information of the to-be-identified bill image is included in the preset standard merchant information, the target merchant information of the to-be-identified bill image is determined as the verified target merchant information; if the target merchant information of the to-be-identified bill image is not included in the preset standard merchant information, the process of identifying the to-be-identified bill image using the trained bill identification model is returned to re-identify, or the user is prompted that the identification is failed.
[0108] It can be understood that the matching process of the target merchant information of the to-be-identified bill image and the preset standard merchant information can also be performed in the trained bill identification model. Correspondingly, the preset prompt word can also include: instructing the trained bill identification model to match the target merchant information of the to-be-identified bill image with the preset standard merchant information. Correspondingly, in the training process of the initial bill identification model, the sample prompt word can include instructing the initial bill identification model to match the target merchant information of each sample bill image with the preset standard merchant information.
[0109] To improve the accuracy of bill identification, the embodiments of the present application provide a bill identification method, obtaining an image of a to-be-identified bill; using a trained bill identification model, processing the image of the to-be-identified bill, a preset prompt word and a preset standard merchant information to obtain target merchant information of the to-be-identified bill, wherein the preset prompt word is used to instruct the trained bill identification model to perform multi-modal feature extraction on the image of the to-be-identified bill, the multi-modal feature is generated based on the visual feature of the to-be-identified bill image and the text feature of the to-be-identified bill image, and the preset prompt word is also used to instruct the trained bill identification model to determine the target merchant information corresponding to the to-be-identified bill according to the multi-modal feature. Since the multi-modal feature of the to-be-identified bill is a feature obtained by fusing different types of information (for example: visual feature and text feature) of the to-be-identified bill, different types of information can be complementary and aligned, so that the multi-modal feature of the to-be-identified bill can more accurately represent the image of the to-be-identified bill. The target merchant information obtained based on the multi-modal feature of the to-be-identified bill is also more accurate, solving the problem of inaccurate bill identification caused by using a single modality for image identification in the case that the image of the to-be-identified bill contains unclear or blurred character information, or the character information in the bill image is a font icon or a special-shaped character.
[0110] In some embodiments, in order to further improve the recognition result of the image of the to-be-identified bill, the bill recognition method provided in the embodiments of the present application can be combined with other bill recognition processes to determine the merchant information corresponding to the to-be-identified bill, that is, a double recognition mechanism is used to identify the image of the to-be-identified bill, so as to further improve the accuracy of bill recognition.
[0111] The following will be described in combination with Figure 4 The process of identifying the image of the to-be-identified bill by using the double recognition mechanism will be described. Referring to Figure 4 , in the automatic bill accumulation scene of the double recognition mechanism, the merchant information to be identified is taken as an example to be specifically described: first, the image of the to-be-identified bill uploaded by the user is obtained, the text of the to-be-identified bill is recognized by using an OCR model, and the text recognition result of the to-be-identified bill image is output, wherein the text recognition result of the bill image can include bill text, text confidence, text position information, etc.; then the text recognition result of the bill image, the preset standard merchant information, and the first preset prompt word are input into the first bill recognition model to obtain the bill recognition result output by the first bill recognition model, wherein the bill recognition result output by the first bill recognition model can include "merchant name", "small bill number", "consumption time", "consumption amount", etc.; then the target merchant name is found in the preset standard merchant name according to the complete merchant name in the bill recognition result, if no match is found, manual review and accumulation are performed, if the target merchant name is found, the target merchant name found is taken as the value corresponding to "merchant name" in the bill recognition result; then it is judged whether the replaced bill recognition result is modified, if it is modified, manual review and accumulation are performed, if it is not modified, it is judged whether the preset risk control rule is met, if it is met, automatic bill accumulation is performed, if it is not met, manual review and accumulation are performed. Or if no target merchant name is found, the image of the to-be-identified bill, the second preset prompt word, and the preset standard merchant name are input into the second bill recognition model, the multi-modal features corresponding to the image of the to-be-identified bill are extracted, and the second bill recognition result (i.e. the target merchant name corresponding to the to-be-identified bill) is determined according to the multi-modal features of the image of the to-be-identified bill, then the target merchant name is taken as the value corresponding to "merchant name" in the bill recognition result, and other steps are executed, if the target merchant information corresponding to the to-be-identified bill cannot be determined according to the multi-modal features of the image of the to-be-identified bill, it is output as empty, prompting the user that "the bill cannot be accumulated according to the bill".
[0112] Optionally, in the automatic bill accumulation scene, the above-mentioned preset risk control rule can include at least one of the following: whether the same bill is uploaded by different ids multiple times, whether the number of accumulations by the same id reaches the upper limit on the same day, whether the accumulation amount of a single bill reaches the limit, etc.
[0113] It should be noted that the first invoice recognition model is a different invoice recognition model from the second invoice recognition model. The first preset prompt word is used to guide the first invoice recognition model to extract at least the merchant name from the text recognition result of the invoice image, and to guide the first invoice recognition model to match the extracted merchant name with the complete merchant name included in the preset labeled merchant information. The second preset prompt word is used to guide the second invoice recognition model to extract the multimodal features of the image of the invoice to be recognized, and to guide the second invoice recognition model to determine the standard merchant information corresponding to the image of the invoice to be recognized based on the multimodal features. This second invoice recognition model is the trained invoice recognition model provided in the embodiments of this application.
[0114] Understandably, the target merchant information corresponding to the ticket to be identified by the above ticket recognition method can not only be used in the above automatic ticket points scenario, but also in one of the following scenarios: ticket data analysis, user behavior analysis, user profile analysis, etc.
[0115] In this embodiment, different recognition methods were used to conduct invoice image recognition experiments. A total of 815 invoice images were tested. Using a combination of OCR and an invoice recognition model, 71 merchant name recognition errors were found. Using a multimodal model, 13 merchant name recognition errors were found. However, if a dual matching mechanism was used to comprehensively consider both the first invoice recognition model and the invoice recognition model trained in this application for multi-dimensional verification, only 2 merchant name recognition errors remained. This demonstrates that the recognition accuracy is significantly improved under the dual recognition mechanism. Figure 4 The method of identification has the highest accuracy.
[0116] Understandably, in the above Figure 4 Under the dual recognition mechanism shown, when the OCR and ticket recognition model are not matched with the standard merchants in the system, the ticket recognition model trained in this application is used for image recognition. This reduces resource consumption, improves overall efficiency, and ensures a high recognition accuracy.
[0117] The receipt recognition model provided in this application takes merchant name extraction as its core training objective. Its technical solution can effectively address key issues in the automatic points system for shopping receipts, such as large differences in receipt layout, complex heterogeneous fonts and shapes, and blurry or tilted receipts. Currently, it focuses only on the recognition and matching of the key information of the merchant name. This technical solution can be transferred to other key information recognition tasks with similar technical difficulties, such as the accurate extraction of elements like store names and brand names. Its multi-dimensional anti-interference characteristics can systematically solve the complex challenges mentioned in the "problems to be solved" of the technical solution, such as multi-version template conflicts, imaging quality defects, and character morphological ambiguity.
[0118] It should be understood that the size of the serial number of each step in the above embodiments does not mean the order of execution, and the execution order of each process should be determined according to its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0119] The bill recognition method corresponding to the above embodiment, Figure 5 The structural block diagram of the bill recognition device provided by the embodiments of the present application is shown, and only the parts related to the embodiments of the present application are shown for ease of illustration.
[0120] With reference to Figure 5 The bill recognition device 500 comprises an acquisition module 501 and a determination module 502.
[0121] The acquisition module 501 is configured to acquire an image of a bill to be recognized.
[0122] The determination module 502 is configured to process the image of the bill to be recognized, a preset prompt word and preset standard merchant information by using a trained bill recognition model, to determine target merchant information corresponding to the bill to be recognized; wherein the preset prompt word is used to instruct the trained bill recognition model to perform multi-modal feature extraction on the image of the bill to be recognized, the multi-modal feature being generated based on a visual feature of the image of the bill to be recognized and a text feature of the image of the bill to be recognized, and the preset prompt word is also used to instruct the trained bill recognition model to determine the target merchant information corresponding to the bill to be recognized according to the multi-modal feature.
[0123] Optionally, in some embodiments, the determination module 502 is further configured to: perform visual coding on the image of the bill to be recognized by using a visual encoder in the trained bill recognition model, to determine a visual feature sequence; perform text coding on a first text corresponding to the image of the bill to be recognized by using a preset large language model in the trained bill recognition model, to determine a text feature sequence, the first text comprising a text corresponding to the image of the bill to be recognized, the preset prompt word and the preset standard merchant information; and fuse the visual feature sequence and the text feature sequence based on a cross-modal attention mechanism, to generate the multi-modal feature.
[0124] Optionally, in some embodiments, the determination module 502 is further configured to: segment the image of the bill to be recognized to obtain a plurality of segmented image blocks; and determine the visual feature sequence based on the plurality of image blocks obtained by segmenting the image of the bill to be recognized, a position code corresponding to each image block and an image modality identifier of each image block.
[0125] Optionally, in some embodiments, the determination module 502 is further configured to perform word segmentation on the first text by using a word segmentor in a preset large language model to obtain a text word sequence; and determine the text feature sequence according to each text word in the text word sequence, a text position code of each text word, and a text modal identifier of each text word.
[0126] Optionally, in some embodiments, the determination module 502 is further configured to perform autoregressive decoding on the multi-modal feature by using an autoregressive decoder in the trained bill recognition model to generate the target merchant information corresponding to the bill to be recognized.
[0127] Optionally, in some embodiments, the bill recognition apparatus 500 further includes a training module 503 configured to obtain a sample bill image training set, a sample prompt word, and sample merchant information, the sample bill image training set including a plurality of sample bill images; process each sample bill image, the sample prompt word, and the sample merchant information by using an initial bill recognition model to obtain target merchant information of each sample bill image; determine a loss value of each sample bill image according to the target merchant information of each sample bill image, target text corresponding to each sample bill image, and a preset loss function, the target text being real merchant information corresponding to the sample bill image; and update model parameters of the initial bill recognition model by using the loss value of each sample bill image to obtain the trained bill recognition model.
[0128] Optionally, in some embodiments, the determination module 502 is further configured to identify an image of the bill to be recognized by using a preset optical character recognition model in the trained bill recognition model to obtain text corresponding to the image of the bill to be recognized.
[0129] It should be noted that the information interaction and execution process between the above apparatuses / units are based on the same concept as the method embodiments, and the specific functions and technical effects thereof can be referred to the method embodiments part, which will not be repeated here.
[0130] Figure 7 A structural schematic diagram of an electronic device is provided for an embodiment of the present application. As shown in the figure, the electronic device 600 of this embodiment includes at least one processor 60 (only one is shown in the figure), a memory 61, and a computer program 62 stored in the memory 61 and executable on the at least one processor 60. The processor 60 implements the steps in any of the respective method embodiments when executing the computer program 62. Figure 7 Figure 7
[0131] The electronic device 600 can be a desktop computer, a notebook computer, a palm computer, a cloud server, and the like. The electronic device can include, but is not limited to, a processor 60, a memory 61. Those skilled in the art can understand that Figure 6 The electronic device 600 is only an example and does not constitute a limitation on the electronic device 600, and can include more or fewer components than shown, or combine certain components, or different components, for example, the electronic device can also include an input sending device, a network access device, a bus, and the like.
[0132] The processor 60 can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic, discrete hardware components, and the like. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor.
[0133] The memory 61 can be an internal storage unit of the electronic device 600 in some embodiments, for example, a hard disk or a memory of the electronic device 600. The memory 61 can also be an external storage device of the electronic device 600, for example, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, and the like. Further, the memory 61 can include both the internal storage unit and the external storage device of the electronic device 600. The memory 61 is used to store an operating system, application programs, a boot loader, data, and other programs, for example, program codes of the computer program, and the like. The memory 61 can also be used to temporarily store data that has been transmitted or will be transmitted.
[0134] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the functional units and modules is taken as an example, and in actual application, the functions can be completed by different functional units and modules according to needs, that is, the internal structure of the apparatus is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit or module in the embodiment can be integrated in one processing unit, or each unit can exist physically separately, or two or more units can be integrated in one unit, and the integrated unit can be realized in the form of hardware or in the form of a software functional unit. In addition, the specific names of the functional units and modules are only for convenient distinction, and do not limit the protection scope of the present application. The specific working process of the units and modules in the system can refer to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0135] The embodiments of the present application further provide an electronic device, which comprises at least one processor, a memory, and a computer program stored in the memory and executable on the at least one processor, wherein the processor implements the steps in any of the method embodiments described above when executing the computer program.
[0136] The embodiments of the present application further provide a computer readable storage medium, which stores a computer program, wherein the computer program is executable by a processor to implement the steps in any of the method embodiments described above.
[0137] The embodiments of the present application provide a computer program product, which, when running on an electronic device, enables the electronic device to implement the steps in any of the method embodiments described above.
[0138] The integrated unit, if implemented in the form of a software function unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on such understanding, the present application can implement all or part of the processes in the above-mentioned embodiment methods through a computer program to instruct relevant hardware to complete, and the computer program can be stored in a computer readable storage medium. When the computer program is executed by a processor, the steps of each method embodiment described above can be implemented. The computer program includes computer program code, which can be in the form of source code, object code, executable files or some intermediate forms. The computer readable medium can at least include any entity or device capable of carrying the computer program code to the photographing device / terminal equipment, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal and software distribution medium. For example, U disk, mobile hard disk, magnetic disk or optical disk, etc. In some jurisdictions, according to legislation and patent practice, the computer readable medium can not be an electrical carrier signal and a telecommunication signal.
[0139] In the above embodiments, the description of each embodiment has its own focus, and the parts not described or recorded in detail in a certain embodiment can be referred to the relevant description of other embodiments.
[0140] Those skilled in the art can appreciate that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be realized by electronic hardware or a combination of computer software and electronic hardware. Whether the functions are realized in hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0141] In the embodiments provided by the present application, it should be understood that the disclosed apparatus / network device and method can be implemented in other ways. For example, the above-described apparatus / network device embodiments are merely schematic, for example, the division of the modules or units is only a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the displayed or discussed each other can be indirect coupling or communication connection through some interfaces, devices or units, and can be electrical, mechanical or other forms.
[0142] The units described as separate components may or may not be physically separate, and the components displayed as units may or may not be physical units, that is, may be located in one place, or may also be distributed to multiple network units. Part or all of the units can be selected to achieve the purpose of the embodiment scheme according to actual needs.
[0143] The above embodiments are only used to illustrate the technical solutions of the present application, but not limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that the technical solutions recorded in the foregoing embodiments can still be modified, or some technical features can be replaced by equivalents; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the protection scope of the present application.< / end>
Claims
1. A method of document identification, characterized by, The method comprises: acquiring an image of a to-be-identified bill; processing the image of the to-be-identified bill, a preset prompt word and preset standard merchant information by using a trained bill identification model to determine target merchant information corresponding to the to-be-identified bill; the trained bill identification model at least comprises a visual encoder, a preset large language model and a cross-modal fusioner, wherein the preset prompt word is used to instruct the visual encoder to perform visual coding on the image of the to-be-identified bill to determine a visual feature sequence, the preset prompt word is used to instruct the preset large language model to perform text coding on a first text corresponding to the image of the to-be-identified bill to determine a text feature sequence, the preset prompt word is also used to instruct the cross-modal fusioner to perform fusion on the visual feature sequence and the text feature sequence based on a cross-modal attention mechanism to generate a multi-modal feature, and the preset prompt word is also used to instruct the trained bill identification model to determine the target merchant information corresponding to the to-be-identified bill according to the multi-modal feature, wherein the first text comprises a text corresponding to the image of the to-be-identified bill, the preset prompt word and the preset standard merchant information.
2. The method of claim 1, wherein, the visual coding on the image of the to-be-identified bill to determine the visual feature sequence comprises: segmenting the image of the to-be-identified bill to obtain a plurality of segmented image blocks; determining the visual feature sequence according to the plurality of segmented image blocks, a position encoding corresponding to each of the image blocks and an image modality identifier of each of the image blocks.
3. The method of claim 1, wherein, the text coding on the first text corresponding to the image of the to-be-identified bill to determine the text feature sequence further comprises: performing word segmentation on the first text by using a word segmenter in the preset large language model to obtain a text word sequence; determining the text feature sequence according to each text word in the text word sequence, a text position encoding of the text word and a text modality identifier of the text word.
4. The method according to any one of claims 1 to 3, characterized in that, the determination of the target merchant information corresponding to the to-be-identified bill according to the multi-modal feature comprises: performing autoregressive decoding on the multi-modal feature by using an autoregressive decoder in the trained bill identification model to generate the target merchant information corresponding to the to-be-identified bill.
5. The method of claim 1, wherein, Before the processing of the image of the to-be-identified bill, the preset prompt word and the preset standard merchant information by using the trained bill identification model to determine the target merchant information corresponding to the to-be-identified bill, the method further comprises: acquiring a sample bill image training set, a sample prompt word and sample merchant information, the sample bill image training set comprising a plurality of sample bill images; processing each sample bill image, the sample prompt word and the sample merchant information by using an initial bill identification model to obtain target merchant information of the each sample bill image; According to the target merchant information of each sample bill image, the target text corresponding to each sample bill image, and a preset loss function, a loss value of each sample bill image is determined, and the target text is real merchant information corresponding to the sample bill image; The model parameters of the initial bill identification model are updated using the loss value of each sample bill image to obtain the trained bill identification model.
6. The method of claim 1, wherein, The method further comprises: An image of the bill to be identified is identified by a preset optical character recognition model in the trained bill identification model to obtain a text corresponding to the image of the bill to be identified.
7. A document identification apparatus, characterized by comprising: Comprise: An acquisition module is configured to acquire an image of a bill to be identified; A determination module is configured to process the image of the bill to be identified, a preset prompt word, and a preset standard merchant information by using a trained bill identification model to determine target merchant information corresponding to the bill to be identified. The trained bill identification model at least comprises a visual encoder, a preset large language model, and a cross-modal fusioner, wherein the preset prompt word is used to instruct the visual encoder to perform visual coding on the image of the bill to be identified to determine a visual feature sequence, the preset prompt word is used to instruct the preset large language model to perform text coding on a first text corresponding to the image of the bill to be identified to determine a text feature sequence, the preset prompt word is also used to instruct the cross-modal fusioner to perform fusion on the visual feature sequence and the text feature sequence based on a cross-modal attention mechanism to generate a multi-modal feature, and the preset prompt word is also used to instruct the trained bill identification model to determine the target merchant information corresponding to the bill to be identified according to the multi-modal feature, wherein the first text comprises a text corresponding to the image of the bill to be identified, the preset prompt word, and the preset standard merchant information.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, The processor executes the computer program to implement the method of any one of claims 1 to 6.
9. A computer-readable storage medium storing a computer program, the computer program comprising instructions that, when executed by a computer, cause the computer to perform the method of any one of claims 1 to 8. The computer program is executed by the processor to implement the method of any one of claims 1 to 6.
Citation Information
Patent Citations
Bill recognition model training method and bill analysis method
CN119741723A
Image-text harmful information identification method and device, electronic equipment and storage medium
CN120264053A