Bill identification method and device, equipment, storage medium and program product

By fusion of text information and multimodal fusion of visual feature vectors and text content on multiple bill images of the same category, combining bill text sequences and context-aware results, the problem of low bill recognition accuracy in the prior art is solved, and higher recognition accuracy and reliability are achieved.

CN119964180APending Publication Date: 2025-05-09INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510024467.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-07
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

In the prior art, deep learning algorithms have low recognition accuracy in the process of bill image processing, making it difficult to effectively solve the accuracy problem of bill recognition.

Method used

By obtaining multiple bill images of the same category, text information is extracted and fused to form a bill text sequence; at the same time, visual feature vectors are extracted and text content are fused to obtain context-aware results; finally, a comprehensive analysis of the recognition results is carried out in combination with the text sequence and context-aware results to improve the accuracy of bill recognition.

Benefits of technology

It significantly improves the accuracy and reliability of bill recognition. Through multimodal fusion and comprehensive analysis, the error caused by single recognition results is reduced, and the understanding and recognition ability of bill content is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119964180A_ABST
    Figure CN119964180A_ABST
Patent Text Reader

Abstract

The invention provides a bill recognition method and device, equipment, a storage medium and a program product, and relates to the field of artificial intelligence. The method comprises the following steps: acquiring a plurality of bill images; extracting text information in the plurality of bill images, and fusing the extracted text information to obtain a corresponding bill text sequence; for any bill image, extracting a visual feature vector of the bill image, obtaining a corresponding text content according to the visual feature vector, and fusing the text content and the visual feature vector to obtain a context perception result of the bill image; obtaining a first recognition result corresponding to a target field in each bill image according to the bill text sequence, obtaining a second recognition result corresponding to the target field in each bill image according to the context sensing result of each bill image, and obtaining a second recognition result corresponding to the target field in each bill image according to the first recognition result and the second recognition result of each bill image. And obtaining a target identification result of each bill image. According to the method, the technical effect of improving the bill recognition accuracy is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence, and in particular to a bill recognition method, device, equipment, storage medium and program product. Background Art

[0002] In commercial transactions, bills are common vouchers, and accurate extraction of their information is critical to protecting the interests of both parties in the transaction and avoiding economic disputes. With the continuous development and progress of information and automation technology, bill recognition technology realizes the automatic entry and processing of bill information, improving work efficiency.

[0003] In the existing technology, deep learning algorithms can be used to extract features from bill images, and then extract key field information related to bank bill business. However, in the process of processing bill images using deep learning algorithms, there is still a technical problem of low accuracy in bill recognition. Summary of the invention

[0004] The present application provides a bill recognition method, device, equipment, storage medium and program product to solve the technical problem of low bill recognition accuracy.

[0005] In a first aspect, the present application provides a bill recognition method, comprising:

[0006] Acquire multiple bill images, wherein the multiple bill images are multiple bill images belonging to the same category in the bill image set to be predicted;

[0007] Extract text information from multiple bill images, and fuse the extracted text information to obtain the corresponding bill text sequence;

[0008] For any bill image, extract the visual feature vector of the bill image, and obtain the corresponding text content according to the visual feature vector, and fuse the text content and the visual feature vector to obtain the context perception result of the bill image;

[0009] According to the bill text sequence, a first recognition result corresponding to the target field in each bill image is obtained, and according to the context perception result of each bill image, a second recognition result corresponding to the target field in each bill image is obtained, and according to the first recognition result and the second recognition result of each bill image, a target recognition result of each bill image is obtained.

[0010] In a second aspect, the present application provides a bill recognition device, comprising:

[0011] An acquisition module, used for acquiring a plurality of bill images, wherein the plurality of bill images are a plurality of bill images belonging to the same category in a set of bill images to be predicted;

[0012] The first fusion module is used to extract text information from multiple bill images and fuse the extracted text information to obtain a corresponding bill text sequence;

[0013] The second fusion module is used to extract the visual feature vector of any bill image, obtain the corresponding text content according to the visual feature vector, and fuse the text content and the visual feature vector to obtain the context perception result of the bill image;

[0014] A processing module is used to obtain a first recognition result corresponding to a target field in each bill image based on a bill text sequence, obtain a second recognition result corresponding to the target field in each bill image based on a context perception result of each bill image, and obtain a target recognition result for each bill image based on the first recognition result and the second recognition result of each bill image.

[0015] In a third aspect, the present application provides a bill recognition device, including: a memory, a processor;

[0016] Memory stores computer-executable instructions;

[0017] The processor executes the computer-executable instructions stored in the memory, so that the processor executes the above first aspect and / or various possible implementations of the first aspect.

[0018] In a fourth aspect, the present application provides a computer-readable storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, they are used to implement the above first aspect and / or various possible implementations of the first aspect.

[0019] In a fifth aspect, the present application provides a computer program product, including a computer program, which, when executed by a processor, is used to implement the above first aspect and / or various possible implementations of the first aspect.

[0020] The present application provides a bill recognition method, device, equipment, storage medium and program product, which significantly improves the accuracy of the subsequent recognition process by extracting text information from multiple bill images of the same category and fusing this information to form a bill text sequence. This method utilizes the common features between different bills and enhances the ability to understand the content of the bill. At the same time, combining visual feature vectors with text content helps to more comprehensively grasp the context of the bill, thereby more accurately locating and identifying key fields on the bill. Finally, by comparing and analyzing the first recognition result based on the bill text sequence and the second recognition result based on context perception, the error caused by a single result is avoided, and the accuracy and reliability of bill recognition are effectively improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0022] Figure 1 A schematic diagram of a bill recognition method provided in an embodiment of the present application Figure 1 ;

[0023] Figure 2 A schematic diagram of a bill recognition method provided in an embodiment of the present application Figure 2 ;

[0024] Figure 3 A schematic diagram of a bill recognition method provided in an embodiment of the present application Figure 3 ;

[0025] Figure 4 A schematic diagram of a bill recognition method provided in an embodiment of the present application Figure 4 ;

[0026] Figure 5 A schematic diagram of the structure of a bill recognition device provided in an embodiment of the present application;

[0027] Figure 6 A schematic diagram of the structure of a bill recognition device provided in an embodiment of the present application.

[0028] The above drawings have shown clear embodiments of the present application, which will be described in more detail later. These drawings and text descriptions are not intended to limit the scope of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. DETAILED DESCRIPTION

[0029] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The implementations described in the following exemplary embodiments do not represent all implementations consistent with the present application. Instead, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0030] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of the relevant data comply with the relevant laws, regulations and standards of the relevant countries and regions, take necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0031] It should be noted that the bill recognition method, device, equipment, storage medium and program product provided in the present application can be used in the field of artificial intelligence, and can also be used in any field other than artificial intelligence. The application field of the bill recognition method, device, equipment, storage medium and program product in the present application is not limited.

[0032] In modern commercial transactions, bills are an important voucher, and accurate extraction of their information plays a vital role in protecting the interests of both parties to the transaction and avoiding economic disputes. In the financial field, banks can quickly process a large number of checks, bills of exchange and other bills by using deep learning algorithms to extract key field information from bills, automatically enter information and verify it, improve business processing efficiency and accuracy, and reduce manual operation errors.

[0033] However, the recognition accuracy of deep learning algorithms in bill image processing still needs to be improved. Therefore, how to further improve the accuracy of deep learning algorithms in bill recognition is still an urgent problem to be solved.

[0034] The present application provides a bill recognition method, device, equipment, storage medium and program product, which obtain multiple bill images; extract text information from the multiple bill images, and fuse the extracted text information to obtain a corresponding bill text sequence; for any bill image, extract the visual feature vector of the bill image, and obtain the corresponding text content based on the visual feature vector, fuse the text content and the visual feature vector to obtain a context perception result of the bill image; obtain a first recognition result corresponding to a target field in each bill image based on the bill text sequence, obtain a second recognition result corresponding to the target field in each bill image based on the context perception result of each bill image, and obtain a target recognition result for each bill image based on the first recognition result and the second recognition result of each bill image, aiming to solve the above technical problems of the prior art.

[0035] The technical solution of the present application and how the technical solution of the present application solves the above-mentioned technical problems are described in detail below with specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below in conjunction with the accompanying drawings.

[0036] Figure 1 A schematic diagram of a bill recognition method provided in an embodiment of the present application Figure 1 ,like Figure 1 As shown, the method includes:

[0037] S101, acquiring multiple bill images.

[0038] In this embodiment, the multiple bill images are multiple bill images belonging to the same category in the set of bill images to be predicted. First, a set of bill images to be predicted is obtained, and the set of bill images to be predicted includes multiple bill types, such as bills of exchange, promissory notes, and checks. A deep learning model, such as a convolutional neural network model, is used to classify the set of bill images to be predicted to obtain multiple bill images of the same category. Multiple bill images of the same category have high similarity in terms of format, content layout, etc., which makes it easy to extract and analyze common features, thereby improving the accuracy and efficiency of processing.

[0039] In a possible implementation, after obtaining the set of bill images to be predicted, the bill image set may be preprocessed to improve the quality of the bill image set. For example, the bill image set may be enhanced using deep learning technology, such as using data augmentation technology to rotate, translate, and scale the bill images to enhance the diversity of the bill image set. The bill image set may also be denoised using a deep learning denoising model to eliminate noise and interference in the image and improve the clarity and recognizability of the image.

[0040] S102: extract text information from multiple bill images, and fuse the extracted text information to obtain a corresponding bill text sequence.

[0041] In this embodiment, by fusing the text information of multiple bill images, the information in different bills is comprehensively utilized to improve the recognition accuracy of the target field. This comprehensive method helps to make up for the missing or erroneous information in a single bill. During the fusion process, the recognition errors of individual bills can be corrected by the information of other bills, thereby improving the accuracy of the bill recognition results.

[0042] S103, for any bill image, extract the visual feature vector of the bill image, obtain the corresponding text content according to the visual feature vector, fuse the text content and the visual feature vector, and obtain the context perception result of the bill image.

[0043] In this embodiment, for any bill image, not only is its visual feature vector extracted, but also the corresponding text content is obtained based on the visual feature vector. These visual feature vectors may include, for example, text content in the image, such as characters, shapes, and positions. Then, the text content is multimodally fused with the visual feature vector to obtain the context perception result of the bill image. This fusion can improve the ability to understand and recognize the bill content, and can more accurately identify the key information in the bill.

[0044] S104. Obtain a first recognition result corresponding to a target field in each bill image based on a bill text sequence, obtain a second recognition result corresponding to a target field in each bill image based on a context perception result of each bill image, and obtain a target recognition result for each bill image based on the first recognition result and the second recognition result of each bill image.

[0045] In this embodiment, the target recognition result is determined by comprehensively utilizing the information provided by the bill text sequence and the information obtained through the context perception result. On the one hand, the first recognition result obtained based on the bill text sequence can be judged based on the text itself; on the other hand, the second recognition result obtained through the context perception result can be judged in combination with the overall environment and semantic relationship of the image. The combination of the two can verify each other, thereby improving the accuracy and reliability of bill recognition, reducing the possible misjudgment or omission of a single method, and making the final recognition result more comprehensive and accurate.

[0046] The bill recognition method provided by the embodiment of the present application extracts text information by fusing multiple bill images of the same category, and fuses the obtained text information to integrate the features of different bills to obtain the corresponding bill text sequence, so as to improve its accuracy during subsequent recognition processing. By fusing the visual feature vector and the text content, the context information of the bill can be better understood, and the target field in the bill can be more accurately identified. By combining the bill text sequence and the context perception result, the first recognition result and the second recognition result of the target field are obtained respectively. Comprehensive analysis of these two recognition results can effectively improve the accuracy and reliability of recognition. Therefore, the above method achieves the technical effect of improving the accuracy of bill recognition.

[0047] Figure 2 A schematic diagram of a bill recognition method provided in an embodiment of the present application Figure 2 ,like Figure 2 As shown, in this embodiment Figure 1 Based on the embodiment, a bill recognition method is described in detail, and the method includes:

[0048] S201. Acquire multiple bill images.

[0049] Step S201 is similar to step S101 and will not be described in detail here.

[0050] S202: extract text information from multiple bill images, and fuse the extracted text information to obtain a corresponding bill text sequence.

[0051] Step S202 is similar to step S102 and will not be described in detail here.

[0052] S203: Input any bill image into the visual encoder, perform feature extraction on any bill image according to the convolution layer in the visual encoder, and obtain a visual feature vector.

[0053] In this embodiment, any bill image is input into the visual encoder. The bill image can be, for example, in color or grayscale. The visual encoder usually includes multiple convolutional layers, which are responsible for extracting useful visual features from the bill image. Each convolutional layer captures different levels of features in the image, such as edges, textures, and shapes, through a series of convolution operations. After multiple layers of convolution processing, a high-dimensional visual feature vector is finally generated. These vectors contain text information in the image, such as characters, shapes, and positions. These visual feature vectors help to distinguish different bills or identify specific target fields.

[0054] S204: convert the visual feature vector to obtain a first visual feature vector, and input the first visual feature vector into a language model to obtain text content.

[0055] In this embodiment, in order to meet the input requirements of the language model, the visual feature vector extracted from the bill image needs to be transformed to generate the first visual feature vector. These transformations may include, for example, dimensionality reduction, normalization, feature selection and other operations.

[0056] Optionally, principal component analysis or autoencoders may be used, for example, to reduce the feature dimension while retaining the most important information. The transformed first visual feature vector is input into a language model. The language model is typically a deep learning model, such as a recurrent neural network or a long short-term memory network. The language model uses the information in the visual feature vector to predict or generate a text description related to the bill, i.e., text content. This text content may be, for example, the text content on the bill image, a summary or explanation of the bill content, etc.

[0057] S205. Convert the text content into a numerical vector, calculate the similarity between the numerical vector and the visual feature vector according to the attention mechanism, and calculate the similarity with the numerical vector and the visual feature vector respectively to obtain a first numerical vector and a second visual feature vector.

[0058] In this embodiment, the text content generated by the language model is converted into a numerical vector. For example, it can be achieved through word embedding, in which each word or character is mapped to a vector space of fixed dimension. In this way, the entire text can be represented as a numerical vector. The attention mechanism is used to calculate the similarity between the numerical vector and the visual feature vector. The attention mechanism generally includes the following steps: calculating the similarity score between the text numerical vector and the visual feature vector, and normalizing the similarity score to a probability distribution. These probabilities represent the correlation between each visual feature vector and the text numerical vector. After obtaining the similarity between the numerical vector and the visual feature vector, the text numerical vector is calculated according to the above probability distribution to obtain a first numerical vector. This vector represents a weighted representation of the text information in the context of visual information. The visual feature vector is calculated according to the above probability distribution to obtain a second visual feature vector. This vector represents a weighted representation of the visual information in the context of text information.

[0059] S206: Fuse the first numerical vector and the second visual feature vector to obtain a context perception result.

[0060] In this embodiment, the fusion operation may include, for example, direct splicing, weighted summation, attention mechanism, etc. Direct splicing refers to simply splicing the first numerical vector and the second visual feature vector together to form a longer vector; weighted summation refers to giving different weights according to the importance of the first numerical vector and the second visual feature vector, and then performing a summation process; attention mechanism refers to using the attention mechanism to dynamically adjust the fusion method of the first numerical vector and the second visual feature vector. The contribution of each vector can be determined by calculating the attention weight. The first numerical vector and the second visual feature vector are fused to obtain a context-aware result, which contains rich features from text and vision, can better express complex contextual information, and can reduce the error caused by a single modality through multimodal fusion.

[0061] S207: Input the bill text sequence into the sequence labeling model for processing to obtain a first recognition result corresponding to the target field in each bill image and a first confidence level corresponding to the first recognition result.

[0062] In this embodiment, the format of the bill text sequence is converted so as to be suitable for the input format of the sequence annotation model. The formatted bill text sequence is input into the sequence annotation model for processing. The sequence annotation model outputs the recognition result of each target field and the corresponding confidence. It should be noted that before the text information of multiple bill images is fused to obtain the bill text sequence, it is necessary to assign a unique identifier to each bill image or use the file name of each bill image file as an identifier as a source mark. In the process of extracting text information from multiple bill images, each extracted text information is associated with the identifier of the bill image from which it comes. In the process of text information fusion, it is ensured that the identifier is transmitted together with the text information. In the fusion process, the text needs to be merged or modified to ensure that the updated text still retains the correct identifier. Therefore, according to the identifier, the recognition results and corresponding confidences of each target field are mapped back to each bill image to ensure that the recognition results of each bill image are independent, thereby obtaining the first recognition result corresponding to the target field in each bill image and the first confidence corresponding to the first recognition result.

[0063] S208. For any bill image, the context perception result is input into the recurrent neural network model. The input layer, hidden layer and output layer in the recurrent neural network model process the context perception result in turn to obtain a second recognition result corresponding to the target field in the bill image and a second confidence level corresponding to the second recognition result.

[0064] In this embodiment, the input layer in the recurrent neural network model is used to receive context-aware results and transmit the context-aware results to the hidden layer. The hidden layer performs nonlinear changes on the context-aware results, captures complex patterns and dependencies in the data, and transmits the results to the output layer. Among them, the hidden layer can stack multiple hidden layers to enhance the expressiveness of the model. In the output layer, the model generates one or more target fields and their corresponding second confidences, and uses the target field as the second recognition result. For example, if the goal is to identify the amount on the bill, the output is a floating point number representing the amount as the second recognition result, and a confidence score between 0 and 1.

[0065] S209: For any bill image, determine whether the first confidence level is greater than the second confidence level.

[0066] S210: If yes, take the first recognition result as the target recognition result.

[0067] S211: If not, use the second recognition result as the target recognition result.

[0068] In this embodiment, for any bill image, a comparison is made as to whether the first confidence level is greater than the second confidence level. If the first confidence level is greater than the second confidence level, the first recognition result is considered to be more reliable and the first recognition result is used as the target recognition result. If the first confidence level is less than or equal to the second confidence level, the second recognition result is considered to be more reliable and the second recognition result is used as the target recognition result. This method improves the accuracy of the entire bill recognition system by selecting the most reliable recognition result as the target recognition result.

[0069] In a possible implementation, if the target fields to be finally identified are the amount and date, after the above steps S201 to S208 are processed, for any bill image, the first recognition result is the first amount and the first date and the first confidence level corresponding to each. The second recognition result is the second amount and the second date and the second confidence level corresponding to each. For the first amount and the second amount, determine whether the first confidence level corresponding to the first amount is greater than the second confidence level corresponding to the second amount. If so, select the first amount as the target recognition result; if not, select the second amount as the target recognition result. For the first date and the second date, determine whether the first confidence level corresponding to the first date is greater than the second confidence level corresponding to the second date. If so, select the first date as the target recognition result; if not, select the second date as the target recognition result.

[0070] The bill recognition method provided in the embodiment of the present application can mine more valuable information and improve the depth of understanding of bill images by extracting visual features through a visual encoder and performing operations such as conversion and fusion. By using a language model to generate text content and further process it, the ability to understand and analyze semantics is enhanced. By comparing the recognition results and their confidence levels from different sources, the target recognition result is obtained, avoiding the error caused by a single recognition result and improving the accuracy of bill recognition.

[0071] Figure 3 A schematic diagram of a bill recognition method provided in an embodiment of the present application Figure 3 ,like Figure 3 As shown, in this embodiment Figure 2 Based on the embodiment, the method of extracting text information from multiple bill images and fusing the extracted text information to obtain a corresponding bill text sequence is described in detail. The method includes:

[0072] S301. Extract multiple bill images according to optical character recognition technology to obtain corresponding text information, and clean and normalize the text information to obtain standardized text information.

[0073] In this embodiment, optical character recognition technology is used to extract text information from each bill image. Optical character recognition technology can recognize characters in an image and convert them into an editable text format. Since the text information output by optical character recognition technology may contain erroneous or incomplete information, it is necessary to clean the text information, for example, to remove irrelevant characters, blanks, noise, etc. Then, normalization processing is performed, including unifying the date format, address format, etc. After cleaning and normalization processing, multiple standardized text information corresponding to multiple bill images is obtained. Optical character recognition technology can not only extract text information from bill images, but also improve the quality and accuracy of text information through cleaning and normalization processing.

[0074] S302, input the standardized text information into a variational autoencoder, the variational autoencoder includes an encoder and a decoder, the encoder is used to encode the standardized text information, a corresponding plurality of latent space vectors are generated, and the plurality of latent space vectors are fused to obtain a fused latent space vector.

[0075] In this embodiment, multiple standardized text information corresponding to multiple bill images are converted into corresponding digitized text information respectively to adapt to the input format of the variational autoencoder. The digitized text information is input into the encoder part of the variational autoencoder. The encoder converts the digitized text information into a latent space representation and generates multiple latent space vectors. These latent space vectors represent the low-dimensional, continuous features of the text. Multiple latent space vectors are fused to generate a fused latent space vector, which integrates the features of multiple bill text information. The fusion process can use weighted averaging, attention mechanism, etc.

[0076] S303: Acquire prior knowledge related to the bill, and adjust the fused latent space vector according to the prior knowledge to obtain a target latent space vector.

[0077] In this embodiment, the prior knowledge related to the bill refers to the known information, rules, format requirements, etc. when processing bill images and text information. For example, bills usually follow a specific format and structure, such as the location and format of fields such as date, amount, and recipient address. This information can be used to guide the adjustment of the latent space vector to ensure that the generated text information conforms to the expected format. There may be logical relationships and dependencies between the information in the bill, such as the total amount should be the sum of various expenses. These relationships can be used to check and adjust the latent space vector to ensure that the generated text information is logically reasonable.

[0078] The prior knowledge related to the bill is converted into constraints, and an area is defined in the latent space as a feasible domain, where the latent space vector must satisfy the constraints of the prior knowledge. The constraints are used as boundary conditions of the feasible domain, and these conditions will be used to determine whether the latent space vector is within the feasible domain. For latent vectors that are not in the feasible domain, they are mapped to the nearest feasible point through a projection operation. This can be achieved by minimizing the distance to the feasible domain. The projection operation can, for example, use an optimization algorithm (such as the projected gradient method). The adjusted latent space vector is input into the decoder to generate a bill text sequence.

[0079] Applying prior knowledge to the fused latent space vector can ensure that the vector is adjusted to be more consistent with the specifications and requirements of the bill. For example, according to the format and content rules of the bill, certain parts of the latent space vector can be constrained or adjusted to improve the accuracy and compliance of the decoded text sequence.

[0080] A bill recognition method provided in an embodiment of the present application extracts text information through optical character recognition technology, and performs cleaning and normalization processing, thereby improving the accuracy and consistency of data, reducing noise and errors, and providing high-quality input for subsequent encoding and decoding processes. The variational autoencoder can encode complex text information into latent space vectors to capture the potential structure and characteristics of the data. By fusing multiple latent space vectors, it is possible to integrate the information of multiple different bills to generate a more representative feature vector. The latent space vector is adjusted using prior knowledge to ensure that the generated text sequence meets the requirements. The bill text sequence is generated by decoding the adjusted target latent space vector, thereby further achieving the technical effect of improving the accuracy of bill recognition.

[0081] Figure 4 A schematic diagram of a bill recognition method provided in an embodiment of the present application Figure 4 ,like Figure 4 As shown, in this embodiment Figure 2 Based on the embodiment, a possible implementation of the bill recognition method is described in detail, and the method includes:

[0082] A plurality of original bill images are obtained, and the plurality of original bill images are preprocessed to obtain preprocessed bill images. The preprocessing process may include, for example, enhancing the image using deep learning technology, such as rotating, translating, scaling, and the like the original bill images using data augmentation technology to increase the size and diversity of the plurality of bill images. Secondly, the image is denoised using a deep learning denoising model to eliminate noise and interference in the image and improve the clarity and recognizability of the bill image.

[0083] The pre-processed bill image is initially identified to obtain the type and basic structure of the pre-processed bill image. Optionally, in the initial identification stage, a deep learning model, such as a convolutional neural network model, can be used to perform preliminary identification on the pre-processed bill image to determine the type and basic structure of the pre-processed bill image. This initial identification step helps to narrow the scope of subsequent bill identification processing and improve processing efficiency.

[0084] After determining the type and basic structure of the preprocessed bill image, multiple bill images of the same category are screened, and the text information of the multiple bill images of the same category is extracted and fused to obtain the corresponding bill text sequence. Specifically, optical character recognition technology is used to extract the corresponding text information from multiple bill images. The text information is cleaned and normalized to obtain standardized text information. The standardized text information is input into a diffusion model such as a variational autoencoder for processing to generate multiple corresponding latent space vectors, and the multiple latent space vectors are fused to obtain the fused latent space vector. Prior knowledge related to the bill is obtained, and the fused latent space vector is adjusted according to the prior knowledge to obtain the target latent space vector. This process can capture the long-distance dependency in the bill text sequence and improve the accuracy of bill recognition. At the same time, this process can adapt to different bill types and formats and has strong adaptability.

[0085] Any bill image among the multiple bill images of the same category is input into the context-aware model for processing, and any bill image is processed using the visual encoder and language model in the context-aware model to obtain a context-aware result, wherein the context-aware model includes the visual encoder and the language model. Specifically, for example, the output predicted by the visual encoder can be used as the input of the language model, or the output of the language model can be used as the input of the visual encoder. Through the context-aware model, the context information of the bill is better understood, and the accuracy of recognition is further improved.

[0086] According to the bill text sequence, a first recognition result corresponding to the target field in each bill image is obtained, and according to the context perception result of each bill image, a second recognition result corresponding to the target field in each bill image is obtained, and according to the first recognition result and the second recognition result of each bill image, a target recognition result of each bill image is obtained.

[0087] The preprocessing process and the process of adjusting the latent space vector using prior knowledge in the above method reduce the complexity of subsequent processing and improve the processing efficiency. By combining technologies such as deep learning and diffusion models, the accuracy and robustness of bill recognition are improved. This is very important in commercial transactions, which can protect the rights and interests of both parties to the transaction and avoid economic disputes. At the same time, this method can be widely used in various bill recognition scenarios, including but not limited to banks.

[0088] Figure 5 A schematic diagram of a bill recognition device provided in an embodiment of the present application is shown in FIG. Figure 5 As shown, the bill recognition device 500 provided in this embodiment includes:

[0089] An acquisition module 501 is used to acquire a plurality of bill images, wherein the plurality of bill images are a plurality of bill images belonging to the same category in a set of bill images to be predicted;

[0090] The first fusion module 502 is used to extract text information from multiple bill images and fuse the extracted text information to obtain a corresponding bill text sequence;

[0091] The second fusion module 503 is used to extract the visual feature vector of any bill image, obtain the corresponding text content according to the visual feature vector, and fuse the text content and the visual feature vector to obtain the context perception result of the bill image;

[0092] Processing module 504 is used to obtain a first recognition result corresponding to a target field in each bill image based on a bill text sequence, obtain a second recognition result corresponding to a target field in each bill image based on a context perception result of each bill image, and obtain a target recognition result for each bill image based on the first recognition result and the second recognition result of each bill image.

[0093] In a possible implementation, the first fusion module 502 is further configured to:

[0094] Extract multiple bill images according to optical character recognition technology to obtain corresponding text information, and clean and normalize the text information to obtain standardized text information;

[0095] The standardized text information is fused to obtain the bill text sequence.

[0096] In a possible implementation, the first fusion module 502 is further configured to:

[0097] Input the standardized text information into a variational autoencoder, which includes an encoder and a decoder;

[0098] The standardized text information is encoded by using an encoder to generate a corresponding plurality of latent space vectors, and the plurality of latent space vectors are fused to obtain a fused latent space vector;

[0099] According to the fused latent vector space and decoder, the bill text sequence is obtained.

[0100] In a possible implementation, the first fusion module 502 is further configured to:

[0101] Acquire prior knowledge related to the bill, and adjust the fused latent space vector according to the prior knowledge to obtain the target latent space vector;

[0102] The decoder is used to decode the target latent space vector to obtain the bill text sequence.

[0103] In a possible implementation manner, the second fusion module 503 is further configured to:

[0104] Input any bill image into the visual encoder, extract features of any bill image according to the convolution layer in the visual encoder, and obtain a visual feature vector;

[0105] The visual feature vector is converted to obtain a first visual feature vector, and the first visual feature vector is input into a language model to obtain text content.

[0106] In a possible implementation manner, the second fusion module 503 is further configured to:

[0107] Convert the text content into a numerical vector, calculate the similarity between the numerical vector and the visual feature vector according to the attention mechanism, and calculate the similarity with the numerical vector and the visual feature vector respectively to obtain a first numerical vector and a second visual feature vector;

[0108] The first numerical vector and the second visual feature vector are fused to obtain a context-aware result.

[0109] In a possible implementation, the processing module 504 is further configured to:

[0110] Inputting the bill text sequence into the sequence labeling model for processing, obtaining a first recognition result corresponding to the target field in each bill image and a first confidence level corresponding to the first recognition result;

[0111] The context perception result is input into the recurrent neural network model. The input layer, hidden layer and output layer in the recurrent neural network model process the context perception result in turn, and obtain the second recognition result corresponding to the target field in the bill image and the second confidence level corresponding to the second recognition result.

[0112] In a possible implementation, the processing module 504 is further configured to:

[0113] For any bill image, determine whether the first confidence level is greater than the second confidence level, and if so, use the first recognition result as the target recognition result;

[0114] If not, the second recognition result is used as the target recognition result.

[0115] The present embodiment provides a bill recognition device that can execute the method provided in the above method embodiment. Its implementation principle and technical effect are similar, and will not be described in detail in this embodiment.

[0116] Figure 6 This is a schematic diagram of the structure of a bill recognition device provided in an embodiment of the present application. Figure 6 As shown, the bill identification device 600 provided in this embodiment includes: at least one processor 601 and a memory 602. Optionally, the device 600 also includes a communication component 603. The processor 601, the memory 602 and the communication component 603 are connected via a bus 604.

[0117] In a specific implementation process, at least one processor 601 executes the computer execution instructions stored in the memory 602, so that at least one processor 601 executes the above method.

[0118] The specific implementation process of the processor 601 can be found in the above method embodiment, and its implementation principle and technical effect are similar, so this embodiment will not be repeated here.

[0119] In the above embodiments, it should be understood that the processor can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), etc. A general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in the invention can be directly implemented as a hardware processor, or can be implemented by a combination of hardware and software modules in the processor.

[0120] The memory may include a high-speed memory (Random Access Memory, RAM), and may also include a non-volatile memory (NVM), such as at least one disk storage.

[0121] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, the bus in the drawings of this application is not limited to only one bus or one type of bus.

[0122] An embodiment of the present application also provides a computer program product, including a computer program, which implements the above method when executed by a processor.

[0123] An embodiment of the present application also provides a computer-readable storage medium, in which computer-executable instructions are stored. When a processor executes the computer-executable instructions, the above method is implemented.

[0124] The above-mentioned readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk. The readable storage medium can be any available medium that can be accessed by a general or special-purpose computer.

[0125] An exemplary readable storage medium is coupled to a processor so that the processor can read information from the readable storage medium and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can be located in an application specific integrated circuit (Application Specific Integrated Circuits, referred to as: ASIC). Of course, the processor and the readable storage medium can also exist in the device as discrete components.

[0126] It should be noted that, for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the present application is not limited by the described order of actions, because according to the present application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily required by the present application.

[0127] It should be further noted that, although the various steps in the flow chart are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps is not strictly limited in order, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flow chart may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these sub-steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of the sub-steps or stages of other steps.

[0128] It should be noted that the terms "first", "second", etc. in the claims, the specification and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the invention described herein, for example, can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, products or devices.

[0129] It should be understood that the above-mentioned device embodiments are only illustrative, and the device of the present application can also be implemented in other ways. For example, the division of units / modules in the above-mentioned embodiments is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units, modules or components can be combined, or can be integrated into another system, or some features can be ignored or not executed.

[0130] In addition, unless otherwise specified, each functional unit / module in each embodiment of the present application may be integrated into one unit / module, each unit / module may exist physically separately, or two or more units / modules may be integrated together. The above-mentioned integrated unit / module may be implemented in the form of hardware or in the form of a software program module.

[0131] If the integrated unit / module is implemented in the form of hardware, the hardware may be a digital circuit, an analog circuit, etc. The physical implementation of the hardware structure includes but is not limited to transistors, memristors, etc. Unless otherwise specified, the processor may be any appropriate hardware processor, such as a CPU, a GPU, an FPGA, a DSP, an ASIC, etc. Unless otherwise specified, the storage unit may be any appropriate magnetic storage medium or magneto-optical storage medium, such as a resistive random access memory (RRAM), a dynamic random access memory (DRAM), a static random access memory (SRAM), an enhanced dynamic random access memory (EDRAM), a high-bandwidth memory (HBM), a hybrid memory cube (HMC), etc.

[0132] If the integrated unit / module is implemented in the form of a software program module and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a memory, including a number of instructions to enable a computer device (which can be a personal computer, server or network device, etc.) to execute all or part of the steps of the various embodiments of the present application. The aforementioned memory includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, disk or optical disk and other media that can store program codes.

[0133] In the above embodiments, the description of each embodiment has its own emphasis. For the part not described in detail in a certain embodiment, please refer to the relevant description of other embodiments. The technical features of the above embodiments can be combined arbitrarily. In order to make the description concise, all possible combinations of the technical features in the above embodiments are not described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0134] Those skilled in the art will readily appreciate other embodiments of the present application after considering the specification and practicing the invention disclosed herein. The present application is intended to cover any modification, use or adaptation of the present application, which follows the general principles of the present application and includes common knowledge or customary techniques in the art that are not disclosed in the present application. The specification and examples are intended to be exemplary only, and the true scope and spirit of the present application are indicated by the following claims.

[0135] It should be understood that the present application is not limited to the precise structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present application is limited only by the appended claims.

Claims

1. A bill recognition method, characterized in that: include: Acquire multiple bill images, wherein the multiple bill images are multiple bill images belonging to the same category in the bill image set to be predicted; Extracting text information from the multiple bill images, and fusing the extracted text information to obtain a corresponding bill text sequence; For any of the bill images, extract a visual feature vector of the bill image, obtain corresponding text content according to the visual feature vector, and fuse the text content with the visual feature vector to obtain a context perception result of the bill image; Based on the bill text sequence, a first recognition result corresponding to the target field in each of the bill images is obtained, and based on the context perception result of each of the bill images, a second recognition result corresponding to the target field in each of the bill images is obtained, and based on the first recognition result and the second recognition result of each of the bill images, a target recognition result for each of the bill images is obtained.

2. The method according to claim 1, characterized in that: According to the bill text sequence, a first recognition result corresponding to a target field in each of the bill images is obtained, according to the context perception result of each of the bill images, a second recognition result corresponding to the target field in each of the bill images is obtained, and according to the first recognition result and the second recognition result of each of the bill images, a target recognition result of each of the bill images is obtained, including: According to the bill text sequence, obtaining a first recognition result corresponding to a target field in each of the bill images and a first confidence level corresponding to the first recognition result; For any bill image, according to the context perception result of the bill image, obtain a second recognition result corresponding to the target field in the bill image and a second confidence level corresponding to the second recognition result; For any bill image, a target recognition result of the bill image is obtained according to the first recognition result of the bill image, the first confidence level, the second recognition result and the second confidence level.

3. The method according to claim 1, characterized in that Extracting text information from the multiple bill images and fusing the extracted text information to obtain a corresponding bill text sequence, including: Extracting the multiple bill images according to optical character recognition technology to obtain corresponding text information, and cleaning and normalizing the text information to obtain standardized text information; The standardized text information is fused to obtain the bill text sequence.

4. The method according to claim 3, characterized in that The standardized text information is fused to obtain the bill text sequence, including: Inputting the standardized text information into a variational autoencoder, wherein the variational autoencoder includes an encoder and a decoder; Using the encoder to encode the standardized text information to generate a corresponding plurality of latent space vectors, and fusing the plurality of latent space vectors to obtain a fused latent space vector; The bill text sequence is obtained according to the fused latent vector space and the decoder.

5. The method according to claim 4, characterized in that According to the fused latent vector space and the decoder, the bill text sequence is obtained, including: Acquire prior knowledge related to the bill, and adjust the fused latent space vector according to the prior knowledge to obtain a target latent space vector; The target latent space vector is decoded by using the decoder to obtain the bill text sequence.

6. The method according to claim 1, characterized in that For any of the bill images, extracting a visual feature vector of the bill image, and obtaining corresponding text content according to the visual feature vector, including: Inputting any of the bill images into a visual encoder, performing feature extraction on any of the bill images according to a convolutional layer in the visual encoder, and obtaining a visual feature vector; The visual feature vector is converted to obtain a first visual feature vector, and the first visual feature vector is input into a language model to obtain the text content.

7. The method according to claim 1, characterized in that The text content and the visual feature vector are fused to obtain a context perception result of the bill image, including: Convert the text content into a numerical vector, calculate the similarity between the numerical vector and the visual feature vector according to an attention mechanism, and calculate the similarity with the numerical vector and the visual feature vector respectively to obtain a first numerical vector and a second visual feature vector; The first numerical vector and the second visual feature vector are fused to obtain the context perception result.

8. The method according to claim 2, characterized in that: According to the bill text sequence, a first recognition result corresponding to a target field in each of the bill images and a first confidence level corresponding to the first recognition result are obtained; for any bill image, according to a context perception result of the bill image, a second recognition result corresponding to the target field in the bill image and a second confidence level corresponding to the second recognition result are obtained, including: Inputting the bill text sequence into a sequence labeling model for processing, and obtaining a first recognition result corresponding to a target field in each bill image and a first confidence level corresponding to the first recognition result; The context-aware result is input into a recurrent neural network model, and the input layer, hidden layer, and output layer in the recurrent neural network model process the context-aware result in turn to obtain a second recognition result corresponding to the target field in the bill image and a second confidence level corresponding to the second recognition result.

9. The method according to claim 2, characterized in that: For any bill image, obtaining a target recognition result of the bill image according to the first recognition result of the bill image, the first confidence, the second recognition result and the second confidence includes: For any bill image, determine whether the first confidence level is greater than the second confidence level, and if so, use the first recognition result as the target recognition result; If not, the second recognition result is used as the target recognition result.

10. A bill recognition device, characterized in that: include: An acquisition module, used for acquiring a plurality of bill images, wherein the plurality of bill images are a plurality of bill images belonging to the same category in a set of bill images to be predicted; A first fusion module is used to extract text information from the multiple bill images and fuse the extracted text information to obtain a corresponding bill text sequence; A second fusion module is used to extract a visual feature vector of any of the bill images, obtain corresponding text content according to the visual feature vector, and fuse the text content with the visual feature vector to obtain a context perception result of the bill image; A processing module is used to obtain a first recognition result corresponding to a target field in each of the bill images based on the bill text sequence, obtain a second recognition result corresponding to the target field in each of the bill images based on a context perception result of each of the bill images, and obtain a target recognition result for each of the bill images based on the first recognition result and the second recognition result of each of the bill images.

11. A bill recognition device, characterized in that: include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the method according to any one of claims 1 to 9.

12. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, which are used to implement the method according to any one of claims 1 to 9 when executed by a processor.

13. A computer program product, characterized in that The invention comprises a computer program, which implements the method according to any one of claims 1 to 9 when being executed by a processor.

Citation Information

Cited By

  • Bill collaborative management method and system based on multi-source heterogeneous data fusion

    CN120450882A

  • Bill cooperative management method and system based on multi-source heterogeneous data fusion

    CN120450882B

  • Intelligent bill information extraction and analysis method and system fused with deep learning

    CN121564726A