Intelligent bill identification method based on deep neural network model
Through the cross-modal attention mechanism and dynamic weight adjustment mechanism, combined with image and text features, the classification and parsing tasks of the bill intelligent recognition system are optimized, which solves the problem of insufficient capture of image and text correlation and improves the accuracy and efficiency of bill recognition.
Patent Information
- Application Number
- CN202510733918.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-04
- Publication Date
- 2025-09-12
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing intelligent bill recognition systems have difficulty effectively capturing the correlation between image and text information when processing complex bill layouts and semantic information, resulting in a decrease in the accuracy of classification and field parsing. In addition, the optimized resource allocation between classification tasks and parsing tasks in multi-task scenarios is unbalanced, affecting model performance.
A cross-modal attention mechanism and a dynamic weight adjustment mechanism are adopted. The convolutional neural network is used to extract image feature vectors and LSTM is used to generate text feature vectors. The cross-modal attention mechanism is combined to calculate the attention weights between image and text features, dynamically adjust the contribution ratio of image and text modalities in the joint feature representation, and design a joint loss function to optimize classification and parsing tasks.
It improves the accuracy and efficiency of intelligent bill recognition, can significantly reduce manual intervention, is applicable to various bill types, and improves the efficiency and accuracy of financial and tax processing.
Smart Images

Figure CN120635923A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent bill recognition based on a deep neural network model, and specifically to an intelligent bill recognition method based on a deep neural network model. Background Art
[0002] In recent years, with the continuous development of artificial intelligence (AI), intelligent bill recognition technology based on deep neural networks has been widely applied in fields such as financial management, tax processing, and archive digitization. Bills, as a common information carrier in business transactions, contain critical structured information such as invoice number, date, amount, and tax rate, as well as unstructured information such as the bill background and image location. This information is crucial for a company's financial analysis and electronic management. Traditional bill processing relies on manual data entry and verification, which is not only time-consuming and labor-intensive but also prone to human error. Bill recognition based on deep neural networks, through automated learning of image and text modalities, can significantly reduce manual intervention and improve recognition efficiency and accuracy. Furthermore, for different types of bills, such as VAT invoices, receipts, and bank slips, intelligent bill recognition technology can automatically classify bill types and further extract key information from bills, laying the technical foundation for financial automation in enterprises.
[0003] Existing intelligent bill recognition systems typically employ single-modality deep learning techniques, such as image-based convolutional neural networks for bill image classification or natural language processing-based text embedding models for bill field parsing. These methods can, to a certain extent, address the requirements of single-modality scenarios, such as determining bill type based on image features or extracting the invoice number and amount fields from the bill text. However, in real-world applications, the characteristics of bills dictate that image and text information often complement each other. For example, the invoice number is typically presented in numerical form, but its specific location requires information from the image modality to determine; the location and content of the amount field depend on both the image layout and the numerical parsing of the text. Single-modality techniques struggle to effectively capture cross-modal information interactions in these complex scenarios, resulting in insufficient recognition accuracy. Furthermore, existing technologies for multi-task scenarios, such as simultaneous bill classification and field parsing, often employ a simple multi-head output architecture. This fails to effectively address the optimization conflicts between classification and parsing, potentially leading to performance degradation on either task.
[0004] Existing single-modal technologies cannot fully capture the correlation between bill images and text modalities, especially when dealing with complex bill layouts and semantic information, which may lead to a decrease in the accuracy of classification and field parsing. Secondly, although multimodal technologies have been applied, existing methods mostly use simple feature splicing or similarity measurement, and fail to fully utilize cross-modal attention mechanisms to model the deep interactive relationship between images and text. In addition, in terms of multi-task optimization, most existing technologies use fixed weights or manually adjusted weights, and are unable to dynamically adjust the optimization goals according to the real-time needs of the task, resulting in an unbalanced allocation of optimization resources between classification and parsing tasks, thereby limiting the overall performance of the model.
[0005] The above information disclosed in this Background section is only for enhancement of understanding of the background of the present disclosure and therefore it may contain information that does not form the prior art that is already known to a person of ordinary skill in the art. Summary of the Invention
[0006] The purpose of the present invention is to provide a method and device for intelligent bill recognition based on a deep neural network model to solve the problems raised in the above background technology.
[0007] To achieve the above object, the present invention provides the following technical solutions:
[0008] The method for selecting a communication path for transmitting data by a device node includes the following specific steps:
[0009] Step 1: Obtain original images of various types of bills, preprocess them, and use an OCR tool to extract text data from the bill images. The extracted text data includes the invoice number, amount, and date. The image data is preprocessed, including grayscale conversion, binarization, and edge detection. The extracted text data is formatted and validity verified.
[0010] Step 2: Based on the preprocessed image, a convolutional neural network is used to output an image feature vector that represents the visual structure. The verified key field text is input into the LSTM to generate a text feature vector that encodes semantic information. The image feature vector and the text feature vector are then concatenated in the feature dimension to form a preliminary joint feature vector.
[0011] Step 3: The attention weights between the image and text feature vectors are calculated through a cross-modal attention mechanism. Based on the attention weights, a gating mechanism is used to dynamically adjust the contribution ratios of the image and text modalities in the final joint representation. The weighted fusion is then used to generate an optimized cross-modal joint feature representation.
[0012] Step 4: Based on the optimized joint feature representation, the bill type classification task and the key field parsing task are performed simultaneously. A joint loss function is designed, which includes the cross-entropy loss of the classification task and the loss of the parsing task. The total loss is calculated through the joint loss function, and all parameters of the entire model are adaptively adjusted using the back-propagation algorithm. After the model training is completed, a new bill image is input. After the above complete process, the bill type classification result and the parsing value of each key field are finally output.
[0013] Furthermore, the logic for obtaining original images of various types of bills and preprocessing the original images is as follows:
[0014] The various types of receipts obtained include paper receipts and electronic receipts. For paper receipts, a scanner is used to scan the receipts at high resolution. For electronic receipts, electronic samples are obtained from the required business and converted into JPEG / PNG format images.
[0015] Preprocessing the original image includes grayscale conversion, binarization, and edge detection. Grayscale conversion involves converting a color image into a grayscale image to simplify subsequent processing. Binarization involves converting a grayscale image into a black and white binary image using an Otsu algorithm with an adaptive threshold to enhance the contrast between the foreground and background of the bill. Edge detection involves identifying significant edge contours in the image using the Canny edge detection algorithm.
[0016] Based on the edge detection results, the main outer frame of the bill is located using a contour analysis algorithm. According to the detected frame coordinates, the original image is cropped to correct the tilt and extract the region of interest image containing only the main area of the bill.
[0017] Based on the cropped and corrected ROI image, an OCR tool is used to extract the text data on the invoice. The extracted text data contains key field information, namely the invoice number, amount, and date. The extracted raw text data is initially formatted and validated, including checking the date format and amount format.
[0018] Furthermore, based on the detected bill frame coordinates (x min ,y min ,x max ,y max ), and the width of the image W = x max -x min and height H = y max -y min , perform fine preprocessing on the cropped and corrected ROI image, and extract text feature data through the OCR tool. The specific logic is as follows:
[0019] Use the border coordinates to expand and crop the original image to retain the edge information of the bill. The cropping formula is:
[0020] I c =I[ max 0,x min -δ):min(x max +δ,W), max (0,y min -δ):min(y max +δ,H)]
[0021] Among them, I c represents the content of the cropped bill image, δ is the border expansion coefficient, which is set to 5% of the bill border size, that is, δ = 0.05 max (x max -x min ,y max -y min );
[0022] To adapt to the input size of the deep learning model, the cropped bill images are uniformly resized to 224×224, and the scaling ratio is defined as: The scaled size is calculated as:
[0023]
[0024] Where H', W' represent the height and width of the scaled image respectively;
[0025] Next, the image is scaled proportionally to its long side and filled with pixel values 0 to a size of 224×224. The formula is:
[0026] I r =PadToCenter(Resize(I c ,H',W'),224,224);
[0027] Based on the image I after scaling and filling r , use OCR tools to perform text recognition, extract text feature data, including invoice number, amount and date fields, output as a text collection in structured JSON format, and uniformly format and verify the validity of the extracted text collection.
[0028] Furthermore, based on the preprocessed image, the logic of using the convolutional neural network to output the image feature vector representing the visual structure is as follows:
[0029] The cropped and scaled single-channel grayscale image I r The pixel values are normalized to the range [0,1]:
[0030]
[0031] Among them, I n (x,y) represents the normalized pixel value, which serves as the model input;
[0032] Construct a lightweight convolutional neural network for feature extraction. The model structure is as follows:
[0033] Input layer: receives normalized image I n (x,y), size is 224×224×1, convolution layer 1: convolution kernel size 3×3, number of filters 32, stride 1, activation function is ReLU, output size is 222×222×32 feature map; pooling layer 1: maximum pooling 2×2, output size is 111×111×32;
[0034] Convolutional layer 2: convolution kernel size 3×3, number of filters 64, stride 1, activation function is ReLU, output size is 109×109×64; Pooling layer 2: maximum pooling 2×2, output size is 54×54×64;
[0035] Fully connected layer: Flatten the high-dimensional feature map 54×54×64 into a one-dimensional feature vector, and connect two layers of fully connected layers in sequence to generate the image feature vector F g ;
[0036] The attention mechanism is introduced to enhance the feature extraction capability of convolutional neural networks. The formula is as follows:
[0037]
[0038] in, It is the feature vector after attention weighting, and the Attention module is channel attention:
[0039] Attention(F g )=F g ·σ(FC(GlobalPool(F g )))
[0040] Among them, F g Image features extracted by convolutional neural network, size is 128, GlobalPool (F g ) is the eigenvector F g Perform global pooling, FC is the fully connected layer, σ is the Sigmoid activation function, and the weight is limited to the range of [0,1]. g ·σ(·) is an element-wise multiplication that applies the attention weights to the original features.
[0041] Furthermore, the logic of inputting the verified key field text into LSTM to generate a text feature vector encoding semantic information is as follows:
[0042] The extracted text data contains key fields, including invoice number, amount, and date: For the invoice number, replace the 12-digit string T id Decomposed into a character sequence, mapped to a dimension of d through a character-level Embedding layer char =8-dimensional continuous vector, based on the formula:
[0043] CharEmbed(T id )=Embedding(T id ,d char )
[0044] Among them, CAarEmbed(T id ) has an output shape of 12×d char , that is, for each numeric character in the sequence, a d char A vector of dimensions;
[0045] The character embedding sequence is then modeled using a long short-term memory network to generate a fixed-length representation based on the following formula:
[0046] Embedding(T id )=LSTM(CharEmbed(T id ))
[0047] For the amount T am and date T da , which is directly input into the model after normalization;
[0048] The final text feature vector is obtained based on the formula:
[0049]
[0050] in, is the final representation of the text feature vector, Embedding(T id ) Obtain the invoice number T through embedding and sequence modeling id Context feature representation, Normalized(T am ,T da ) Through normalization, the amount T am and date T da Convert to numerical features;
[0051] In order to realize the joint modeling of image and text, a nonlinear interaction mechanism is introduced to transform the image feature vector and text feature vector Splicing into a joint feature index F, the formula is:
[0052]
[0053] Among them, F adopts the following method:
[0054]
[0055] Among them, ⊙ represents element-by-element multiplication, Used to enhance the interaction between image and text features, addition The independent information of single modal features is retained.
[0056] Furthermore, the logic of calculating the attention weight between image and text feature vectors through the cross-modal attention mechanism is as follows:
[0057] The image feature matrix is recorded as where n g represents the number of features output from the convolutional neural network, d g Represents the dimension of each image feature vector; the text feature matrix is recorded as n t represents the number of features output from the text encoding network, d t Represents the dimension of each text feature vector;
[0058] When measuring the similarity between the i-th feature of the image and the j-th feature of the text, first project them to the same hidden dimension d h , specifically: define Among them, W g For each d g The image feature map of dimension d h Wei, W t For each d t The text feature map of dimension d is h dimension;
[0059] For any pair of image features and text features, calculate their similarity in the unified hidden space and obtain the attention weight through Softmax normalization. Specifically, for the i-th image feature and the j-th text feature, the attention weight is defined as:
[0060]
[0061] in, is the i-th row vector of the image feature matrix, i=1,2,…,n g , is the j-th row vector of the text feature matrix j=1,2,…,n t , are respectively the vector representations after being projected to the same hidden layer dimension, A i,jrepresents the attention weight of the i-th image feature and the j-th text feature;
[0062] For all i=1,2,…,n g and j = 1, 2, ..., n t The attention weight {A i,j}Combined into attention matrix
[0063] Use the attention matrix A to analyze the text feature matrix Perform weighted analysis to obtain the attention-enhanced text representation and generate a more semantic cross-modal fusion representation F cross , based on the formula:
[0064]
[0065] in, That is, for each image feature position, the cross-modal representation is formed by concatenating the text context vector captured by attention and the image feature vector. Represents the text features weighted by the attention mechanism, and the Concat operation combines them with the original image features Concatenation, Concat(·) is to connect the attention-enhanced text features with the original image features along the feature dimension;
[0066] Further adaptively determine the contribution ratio of image and text modalities in the final joint feature, and introduce a gating mechanism to F cross Map and generate corresponding weight coefficients; define two sets of learnable parameters: d t +d g It's F cross The dimension of each row, D is the dimension of the gate output, and and The final alignment dimensions are consistent;
[0067] For each image position i, Perform two linear mappings and add Sigmoid to obtain the weight vectors of image modality and text modality: where α g [i]∈[0,1] D The d-th dimension component represents the weight of the image modality at the i-th position and dimension d, α t [i]∈[0,1] D The d-th dimension component represents the weight of the text modality at the i-th position and dimension d;
[0068] The two are weighted and summed in an element-by-element multiplication manner to obtain the fusion vector at the i-th position. The formula is:
[0069]
[0070] in, is the feature obtained by gated fusion at the i-th position, ⊙ represents element-by-element multiplication, that is, multiplication of positions of the same dimension respectively, which means that for each dimension d, there are:
[0071]
[0072] This is obtained For all i=1,2,…,n g The same operation is performed on each row to form the final matrix The F gated That is, the final joint feature representation guided by cross-modal attention and adjusted by gated dynamic weights;
[0073] Based on the optimized joint feature F gated , used for two types of bill intelligent recognition subtasks: bill type classification and key field numerical regression: Pool the rows to obtain a global feature vector of dimension D:
[0074]
[0075] Among them, Pool(·) is average pooling or taking the first row;
[0076] Let the total number of categories be C, and configure a set of classification parameters for each category v = 1, 2, ..., C: Then the predicted probability of the vth class is expressed by Softmax:
[0077]
[0078] Among them, p v represents the predicted probability that the current bill belongs to category v, b v ,b u ,W v ,W u are all learning parameters of the classifier;
[0079] The denominator normalizes the sum of the exponentials of all C categories, and finally selects the index with the highest probability as the output category:
[0080]
[0081] in, The bill category label predicted by the model;
[0082] Use the pooling vector obtained in the classification task Recorded as Assume there are O key fields in total, and set the regression parameter for the oth field as: The predicted value of this field is:
[0083]
[0084] in, is the predicted value of the oth field.
[0085] Furthermore, based on the optimized joint feature representation, the bill type classification task and the key field parsing task are performed simultaneously. The logic of the joint loss function is designed as follows:
[0086] For bill type prediction, the optimized joint feature F gated Predict the category of the bill. This classification task uses cross entropy loss to measure the distance between the true distribution and the predicted distribution. The formula is:
[0087]
[0088] Among them, N is the number of samples in a training batch, C is the total number of bill categories, and y n,v represents the true label of the nth sample, p n,v represents the predicted probability of the nth sample for type v, L cls Represents the loss value of the classification task;
[0089] Perform regression prediction on the key fields in the bill so that the model's output value for each field is as close to the true label as possible. The parsing task uses mean square error as the loss metric, based on the following formula:
[0090]
[0091] Among them, M=3 is the total number of key fields, and y m are the predicted value and true value of the mth field, L parse Represents the loss value of the field parsing task;
[0092] In order to balance the optimization requirements of the classification task and the parsing task, a joint loss function is adopted, based on the formula:
[0093] L=λ cls L cls +λ parse L parse
[0094] Among them, λ cls≥0 is the weight coefficient of the classification task, λ parse ≥0 is the weight coefficient of the parsing task;
[0095] Let the classification loss calculated for the current training iteration / batch be λ cls , the parsing loss is λ parse , the dynamic adjustment formula is:
[0096]
[0097] The task weights are balanced by inverse error. If the classification task performs poorly, that is, λ cls If the model is larger, it will focus more on classification; if the parsing task performs poorly, that is, λ parse Larger, the model will focus more on parsing.
[0098] Compared with the prior art, the present invention has the following beneficial effects:
[0099] The present invention collects various types of bills, models the bill images based on image feature data using a convolutional neural network, and extracts image feature vectors. At the same time, based on text feature data, a deep neural network is used to model the extracted text data to generate text feature vectors. The image and text feature vectors are concatenated to form a joint feature vector, which effectively captures the explicit association between the image modality and the text modality.
[0100] In the process of image and text feature modeling, the image feature vector and the text feature vector are spliced into a joint feature vector, which is then optimized through the cross-modal attention mechanism. The cross-modal attention mechanism can dynamically capture the implicit interaction relationship between the image modality and the text modality, such as the correlation between the semantic content of the text field and its image position, so as to better model the complex information in the bill.
[0101] The present invention utilizes a dynamic weight adjustment mechanism to effectively resolve the conflict between classification tasks and parsing tasks, ensuring an optimal resource allocation balance between the two and improving the overall performance of the model. The present invention is further applicable to a variety of bill types, such as value-added tax invoices, receipts, bank slips, etc., and can significantly reduce manual intervention in actual scenarios, thereby improving the efficiency and accuracy of financial and tax processing. BRIEF DESCRIPTION OF THE DRAWINGS
[0102] Figure 1 Schematic diagram of the overall method of the present invention. DETAILED DESCRIPTION
[0103] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below with reference to specific embodiments.
[0104] It should be noted that, unless otherwise defined, the technical or scientific terms used in the present invention should have the usual meanings understood by people with ordinary skills in the field to which the present invention belongs. The "first", "second" and similar words used in the present invention do not indicate any order, quantity or importance, but are only used to distinguish different components. "Include" or "comprise" and similar words mean that the elements or objects appearing before the word include the elements or objects listed after the word and their equivalents, without excluding other elements or objects. "Connect" or "connected" and similar words are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative position relationships. When the absolute position of the object being described changes, the relative position relationship may also change accordingly.
[0105] Example:
[0106] See also Figure 1 , the present invention provides a technical solution:
[0107] A method for intelligent bill recognition based on a deep neural network model, comprising the following steps:
[0108] Step 1: Obtain original images of various types of bills, preprocess them, and use an OCR tool to extract text data from the bill images. The extracted text data includes the invoice number, amount, and date. The image data is preprocessed, including grayscale conversion, binarization, and edge detection. The extracted text data is formatted and validity verified.
[0109] Obtain original images of various types of bills and preprocess the original images as follows:
[0110] The image border represents the valid area of the bill content after background noise is removed, and the image size represents the height and width of the bill image. A variety of paper bill samples were collected from real-world scenarios, including bills of different sizes and print quality. Handwritten bills were also considered for some complex scenarios. For paper bills, a scanner was used to scan the bills at high resolution, with the resolution set to 300 DPI, and the output was a JPEG / PNG image. For electronic invoices, electronic invoice PDF files were obtained from email attachments, corporate financial systems, or tax platforms and converted into JPEG / PNG images.
[0111] Convert the color image to a grayscale image with a grayscale value range of [0,255]. The value of each pixel is calculated using the following formula:
[0112] C(x,y)=α·R(x,y)+β·G(x,y)+γ·B(x,y)
[0113] Where C(x,y) is the pixel value of the grayscale image at coordinate (x,y), representing the intensity value after conversion from color to grayscale. R(x,y), G(x,y), and B(x,y) are the red, green, and blue channel values of the color image pixel, respectively. The pixel point (x,y) is the pixel in the xth column and yth row of the image. The weights α, β, and γ are adaptively adjusted by minimizing the loss function of the bill classification task and satisfying the constraint α+β+γ=1. The channel values of the input image are normalized to the range [0,1].
[0114] The larger the grayscale value of C(x,y), the brighter the pixel is and the closer it is to white; the smaller the grayscale value, the darker the pixel is and the closer it is to black. When the values of R(x,y), G(x,y), and B(x,y) are equal, the pixel is gray. When a value is significantly greater than the other two, the pixel tends to that color. The weights α, β, and γ represent the contribution ratio of the red, green, and blue channels to the grayscale value.
[0115] In standard grayscale conversion, α = 0.2989, β = 0.5870, and γ = 0.1140 are usually used, which reflects that human vision is more sensitive to green and less sensitive to blue. When the weights remain unchanged, when the pixel value of a certain channel is higher, the grayscale value will be more biased towards the color contribution corresponding to that channel. For example, if α is larger, the influence of the red channel on the grayscale value will be more significant.
[0116] This weight is learned through a fully connected network, and the global pixel mean and global standard deviation are input into the pooling:
[0117] [α,β,γ]=Softmax(FC(Concat(AvgPool(I),StdPool(I))))
[0118] Among them, AvgPool(I) is the global pixel mean, capturing the brightness distribution of the entire image, StdPool(I) is the global standard deviation pooling, supplementing the contrast information of the image with uniform brightness, Concat(AvgPool(I),StdPool(I)) splices the brightness and contrast information into a vector and inputs it into the fully connected network for weight learning; Softmax ensures that the output weights are normalized, that is, the sum of all channel weights is 1;
[0119] [α, β, γ] represents the weights of the red, green, and blue channels, which determine the contribution ratio of each channel to the grayscale value. The weights are dynamically learned by the deep learning network and range from [0, 1], satisfying α + β + γ = 1;
[0120] AvgPool(I) is the global pixel mean, which is used to capture the overall brightness distribution of the image. The higher the brightness, the larger the mean. tdPool(I) is the global pixel standard deviation pooling, which is used to measure the contrast or variation of image brightness. The higher the contrast, the larger the standard deviation.
[0121] The concatenated vector is input into a two-layer fully connected network. The first layer has 128 neurons and the activation function is ReLU. The second layer has 3 neurons for output weights.
[0122] The current formula is suitable for images with uniform brightness and strong contrast regularity. If the bill background is complex, the ability to capture local information is increased:
[0123] Input=Concat(AvgPool(I),StdPool(I),LocalHist(I))
[0124] Among them, LocalHist(I) represents the local histogram, and its calculation window size is set to 16×16 to capture the local brightness distribution;
[0125] The grayscale image is binarized to extract the foreground area of the bill content, that is, the text or bill edge, using the Otsu algorithm adaptive threshold, the formula is:
[0126] Q(z,y)=Otsu(C(x,y))
[0127] If the grayscale value of the pixel C(x,y)>Q(x,y), then the pixel is set to white 255; otherwise, it is set to black 0. The logic is:
[0128]
[0129] The Canny edge detection algorithm is used to accurately detect the edge position of the bill, and the gradient calculation formula is adjusted by adaptive direction weighting:
[0130]
[0131] in, is the gradient of the image in the x direction, that is, the rate of change of pixel intensity in the horizontal direction, is the gradient of the image in the y direction, that is, the rate of change of pixel intensity in the vertical direction. Gradient is the gradient intensity of the image at the coordinate (x, y), which represents the edge information around the pixel. x and w y Learning from the image gradient histogram Gradient(I); To avoid noise interference, a dynamic upper and lower threshold selection strategy is adopted, which automatically adjusts the upper and lower threshold ratio by setting it, such as 1:2;
[0132] w x and w y are the weights of the gradient in the x and y directions, respectively, reflecting the contribution ratio of these two directions to the gradient intensity. The weights are dynamically learned through the image gradient histogram. If w x =1,w y =1, the formula degenerates into the standard Euclidean gradient calculation, indicating that the contributions of all directions are equal; if w x >w y , indicating that the edge features in the x direction are more important, and the formula will be more biased towards the intensity change in the horizontal direction; the larger the gradient value, the more obvious the edge feature at the pixel is, which may be the border of the bill or the edge of the character;
[0133] The above w x and w y You can learn:
[0134] [w x ,w y ]=Softmax(FC(GradHist(I)))
[0135] Among them, GradHist(I) is the image gradient histogram, which is used to capture the global edge distribution. If the edge features of the bill border are regular, w can be directly fixed. x =w y =1 reduces computational complexity;
[0136] Based on the cropped and corrected ROI image, an OCR tool is used to extract the text data on the invoice. The extracted text data contains key field information, namely the invoice number, amount, and date. The extracted raw text data is initially formatted and validated, including checking the date format and amount format.
[0137] Based on the detected bill border coordinates (x min ,y min ,x max ,y max ), corresponding to the upper left corner and lower right corner of the bill content area; the width of the image W = x max -x min and height H = y max -y min , and recorded as numerical features, the input image shape is H×W, that is, height×width;
[0138] The cropped and corrected ROI image is preprocessed and the text feature data is extracted using the OCR tool. The specific logic is as follows:
[0139] Use Tesseract OCR in the OCR tool to perform text recognition on the bill image. The goal is to identify the invoice number, amount, and date data. The processing process is as follows: After inputting the bill image data to be detected, the OCR tool first detects the text area in the image, then segments and recognizes the bill content by line or block, performs text recognition on the segmented areas, and outputs the extracted text set T. The structure of T is in JSON format, including the invoice number, amount, and date data fields;
[0140] Use the border coordinates to expand and crop the original image to retain the edge information of the bill. The cropping formula is:
[0141] I c =I[x min -δ:x max +δ,y min -δ:y max +δ]
[0142] Among them, I c It represents the content of the cropped bill image. δ is the border expansion coefficient, which is used to prevent the edge content from being cropped. Its fixed ratio is set to 5% of the bill border, that is, δ = 0.05 max (x max -xx in ,y max -y min );
[0143] To ensure that the cropping range is within the image boundaries, the above formula needs to limit the coordinate range. The range formula is:
[0144] I c =I[max(0,x min -δ):min(x max +δ,W),max(0,y min -δ):min(y max +δ,H)]
[0145] Among them, I c represents the cropped bill image content, which is the result of extracting the valid area containing the bill content from the original image I and appropriately expanding the border; x min ,x max ,y min ,y max These values represent the bounding coordinates of the minimum bounding rectangle of the bill content, calculated from the edge points detected by Canny. When δ increases, the cropped area will include more background information, which may increase noise in subsequent processing. If the coordinates of the rectangle are close to the image boundary, the formula will limit the range to prevent the index from exceeding the image size.
[0146] To adapt to the input size of the deep learning model, the cropped receipt images are uniformly resized to 224×224. The resizing process is as follows:
[0147] The input image size after setting is H×W, and the scaling ratio is: The scaled size is:
[0148]
[0149] Where H', W' represent the height and width of the scaled image, ensuring that the cropped image is scaled to 224 pixels in proportion to its long side while maintaining the aspect ratio; H, W represent the original height and width of the cropped image;
[0150] The scale determines how the image is resized. If the image's long side is longer, the scale is smaller, indicating a larger scaling margin. Conversely, if the image is smaller, the scaling margin is smaller.
[0151] Scale the image proportionally to its long side and use black, that is, the pixel value is 0, and fill it to 224×224 in the center. The formula is: I r =PadToCenter(Resize(I c ,H',W'),224,224);
[0152] The collected text feature data is uniformly preprocessed, including formatting the amount and date values, removing unnecessary symbols, spaces and possible OCR error output. If a field is missing, the default value is used, that is, the amount defaults to "0.00" and the date defaults to "1970-01-01". The formatting process includes: the invoice number must be a pure numeric string with a length between 8 and 12. If the OCR recognition result contains non-numeric characters, such as "-", it must be removed. The amount must be a positive number greater than 0 and retain two decimal places. For OCR errors, such as recognizing "O" as "0", they are corrected. The date format must comply with the international standard of year, month, and day: "YYYY-MM-DD".
[0153] Step 2: Based on the preprocessed image, a convolutional neural network is used to output an image feature vector that represents the visual structure. The verified key field text is input into the LSTM to generate a text feature vector that encodes semantic information. The image feature vector and the text feature vector are then concatenated in the feature dimension to form a preliminary joint feature vector.
[0154] Based on the preprocessed image, the logic of using the convolutional neural network to output the image feature vector representing the visual structure is:
[0155] The cropped and scaled single-channel grayscale image I r, which is a single-channel grayscale image of size 224×224, with a pixel value range of [0,255]. To adapt to the deep learning model, the input image is normalized and the pixel values are mapped to the interval [0,1]:
[0156]
[0157] Among them, the pixel value is normalized, and I n (x,y)∈[0,1] is used as the input of the model;
[0158] I n (x,y) represents the normalized image pixel value, which is the original image pixel value I r The linear scaling result of (x,y) is used to adapt to the input requirements of the deep learning model; the original image pixel value I r (x,y), the range is [0,255], this is the grayscale image after cropping and adjustment; the normalized pixel value I n (x,y) is directly used as input to the model to reduce the dynamic range of the input data, which helps the model converge faster and avoids instability in gradient calculation caused by excessively large or small values;
[0159] When I r The larger (x,y) is, the higher the normalized I n (x,y) is also larger, if I r (x,y)=255, then I n (x,y)=1, represents the brightest pixel. If I r (x,y)=0, then I n (x,y)=0, indicating the darkest pixel;
[0160] Considering a dataset with thousands to tens of thousands of images, a lightweight convolutional neural network is constructed for feature extraction. The model structure is as follows:
[0161] Input layer: receives normalized image I n (x, y), size is 224×224×1, convolution layer 1: convolution kernel size is 3×3, number of filters is 32, stride is 1, activation function is ReLU, output size is (224-3+1)×(224-3+1)×32=222×222×32;
[0162] The output feature map of 222×222×32 represents the local patterns of the image extracted in the first layer of convolution, such as edges, textures, and other features. These features are captured by 32 filters with different feature distributions. If the convolution kernel size is increased, the size of the output feature map will be reduced more. If the stride is increased, the size of the output feature map will also be reduced more.
[0163] Pooling layer 1: max pooling 2×2, output size is (222 / 2)×(222 / 2)×32=111×111×32;
[0164] The role of pooling is to reduce the dimension of the feature map while retaining the main feature information. The reduced feature map 111×111×32 reduces the redundancy of features while retaining the main feature information. This also reduces the computational cost. If the pooling window is increased, such as 3×3, the size of the feature map will be further reduced. If the stride is reduced, the size of the feature map will be larger.
[0165] Convolutional layer 2: convolution kernel size 3×3, number of filters 64, stride 1, activation function is ReLU, output size is (111-3+1)×(111-3+1)×64=109×109×64;
[0166] The output feature map of 109×109×64 represents higher-level features, such as shape or edge combination information. If the convolution kernel size is increased, the size of the output feature map is further reduced. If the number of filters is increased, the feature dimension extracted will be higher.
[0167] Pooling layer 2: max pooling 2×2, output size is (109 / 2)×(109 / 2)×64=54×54×64;
[0168] Fully connected layer: Flatten the high-dimensional feature map 54×54×64 into a one-dimensional feature vector, and connect two layers of fully connected layers in sequence. The first layer outputs 256 dimensions, and the second layer outputs a fixed 128-dimensional feature vector F g ;
[0169] The flattened feature size is 186624: This is the dimension of the feature map output by the convolutional layer after flattening to one dimension, F g It is a high-dimensional feature vector representation of the image, which condenses all the information extracted by the convolutional layer. The higher the dimension, the richer the representation. The number of neurons in the fully connected layer determines the dimension of the final feature vector. If the output size of the pooling layer increases, the dimension of the flattened feature vector will also increase.
[0170] The model output is the image feature vector F g , represents the high-dimensional features of the bill image, and introduces the attention mechanism to enhance the feature extraction capability of the convolutional neural network. The formula is:
[0171]
[0172] in, It is the feature vector after attention weighting, which has higher discrimination ability and enhances the focus on key features. The Attention module is channel attention:
[0173] Attention(F g )=F g ·σ(FC(GlobalPool(F g )))
[0174] Among them, F g Image features extracted by convolutional neural network, size is 128, GlobalPool (F g ) is the eigenvector F g Perform global pooling to summarize the information of each channel, FC is the fully connected layer, used to map features, σ is the Sigmoid activation function, used to limit the weight to the range of [0,1], F g σ(·) is an element-wise multiplication that applies the attention weights to the original features.
[0175] If the weight of σ(·) output is large, it means that the feature is more important to the task. If the weight of σ(·) output is close to 0, the feature is suppressed. The attention mechanism improves the representation ability of the feature vector and helps the network better focus on the features related to the task.
[0176] Based on the collected text feature data, a deep neural network is used to model the text feature data. The logic for extracting text feature vectors is as follows:
[0177] Use the formatted text feature data from step 1 as input data, including: Invoice number is set to T id , which is a pure numeric string of 8 to 12 characters in length, fixed to 12 characters, and filled with 0 if it is less than 12 characters; the amount is set to T am , is a positive number, retains two decimal places, and is normalized by dividing by the maximum amount; the date is set to T da , which is formatted as "YYYY-MM-DD" and is split into three numerical features: "year", "month", and "day";
[0178] For invoice number T id , string the 12 digits into T id Decomposed into a character sequence, mapped to a dimension of d through a character-level Embedding layer char =8-dimensional continuous vector, based on the formula:
[0179]
[0180] So CharEmbed(T id ) has a shape of (12,8);
[0181] Will Feed it into a single LSTM, the hidden layer dimension of LSTM is d h, get the hidden state of LSTM at each moment {h1,…h 12}, take the last moment h 12 As the overall semantic representation of the invoice number; if a unidirectional LSTM is used, the dimension of the hidden state at the last moment is d h If bidirectional LSTM is used, the size of the concatenated image should be 2d. h ;
[0182] The character embedding sequence is then modeled using a long short-term memory network to generate a fixed-length representation based on the following formula:
[0183] Embedding(T id )=LSTM(CharEmbed(T id ))
[0184] make sure Unify the d text =d h , d h is the hidden layer dimension of LSTM;
[0185] For the amount T am and date T da , and then directly input into the model after min-max normalization;
[0186] The final text feature vector is obtained based on the formula:
[0187]
[0188] in, is the final representation of the text feature vector, which is composed of invoice number, amount and date. Obtain the invoice number T through embedding and sequence modeling id The context feature representation of Through normalization, the amount T am and date T da Convert to numerical features;
[0189] Text feature vector The text information is encoded in a low-dimensional manner to facilitate subsequent joint modeling with image features. id The more complex the character pattern, the more complex its embedding representation. If the distribution of amounts or dates is uneven, normalization helps balance the model's sensitivity to different features.
[0190] Assume that the image feature vector has been obtained through an image encoding network And guarantee and The same in dimension, i.e.
[0191] In order to realize the joint modeling of image and text, a nonlinear interaction mechanism is introduced to transform the image feature vector and text feature vector Splicing into a joint feature index F, the formula is:
[0192]
[0193] Among them, F adopts the following method:
[0194]
[0195] Where ⊙ represents element-by-element multiplication Used to enhance the interaction between image and text features, addition The independent information of the single-modal feature is retained, and the final It is still a D-dimensional vector;
[0196] The joint feature F contains the interactive information of the image and text, which is suitable for multimodal tasks such as bill classification or field parsing. It enhances the interactive relationship between image and text. Joint modeling can comprehensively utilize the image and text information of the bill to improve the model's recognition ability.
[0197] Step 3: The attention weights between the image and text feature vectors are calculated through a cross-modal attention mechanism. Based on the attention weights, a gating mechanism is used to dynamically adjust the contribution ratios of the image and text modalities in the final joint representation. The weighted fusion is then used to generate an optimized cross-modal joint feature representation.
[0198] The logic of calculating the attention weight between image and text feature vectors through the cross-modal attention mechanism is:
[0199] The image feature matrix is recorded as where n g represents the number of features output from the convolutional neural network, d g Represents the dimension of each image feature vector; the text feature matrix is recorded as n t represents the number of features output from the text encoding network, d t Represents the dimension of each text feature vector;
[0200] When measuring the similarity between the i-th feature of the image and the j-th feature of the text, first project them to the same hidden dimension d h , specifically: define Among them, W g For each d g The image feature map of dimension d h Wei, W tFor each d t The text feature map of dimension d is h W g 、W t The cross-modal attention stage projects the image / text features to the hidden dimension d h The matrix parameters used;
[0201] For any pair of image features and text features, calculate their similarity in the unified hidden space and obtain the attention weight through Softmax normalization. Specifically, for the i-th image feature and the j-th text feature, the attention weight is defined as:
[0202]
[0203] in, is the i-th row vector of the image feature matrix, i=1,2,…,n g , is the j-th row vector of the text feature matrix are respectively the vector representations after being projected to the same hidden layer dimension, A i,j Represents the attention weight of the i-th image feature and the j-th text feature, with a value between [0,1] and meeting the normalization condition The denominator accumulates the same image features With all n t The dot product of the projected text features ensures that the attention weights under the same i are normalized to [0,1];
[0204] For all i=1,2,…,n g and j = 1, 2, ..., n t The attention weight {A i,j}Combined into attention matrix
[0205] After the projection is completed, the similarity is measured by the inner product of the two projection vectors, and the similarity scores between the same image feature and all text features are exponentially weighted and normalized to finally obtain an attention weight matrix The element in the i-th row and j-th column of the matrix represents the attention score of the i-th feature of the image and the j-th feature of the text. Since the exponential operation is performed on the same row and divided by the sum of all scores in the row, the sum of the weights of each row is 1;
[0206] Use the attention matrix A to analyze the text feature matrix Perform weighted analysis to obtain the attention-enhanced text representation and generate a more semantic cross-modal fusion representation F cross , based on the formula:
[0207]
[0208] in, It is a cross-modal semantic feature that contains the association information of image layout and text semantics. Represents the text features weighted by the attention mechanism, and the Concat operation combines them with the original image features Concatenation (·) is a concatenation operation along the feature dimension, i.e., the second dimension, which connects the attention-enhanced text features with the original image features. The cross-modal attention mechanism can capture the relationship between the image layout, such as the position of the amount box on the bill, and the text content, such as the amount field value, to improve parsing accuracy.
[0209] F cross Represents cross-modal features, integrating weighted text features and original image features Its size is n g ×(d t ×d g ), which contains rich association information of image layout and text semantics;
[0210] like The value of increases, which means that the text feature contributes more to the joint feature, and the Concat operation retains the image feature. Independent information; cross-modal features can capture the semantic association between bill images and text data, such as the correspondence between the position of the amount field in the image and the amount value in the text;
[0211] Specifically, line i: therefore That is, for each image feature position, the cross-modal representation is formed by concatenating the text context vector captured by attention and the image feature vector;
[0212] Further adaptively determine the contribution ratio of image and text modalities in the final joint feature, and introduce a gating mechanism to F cross Map and generate corresponding weight coefficients; define two sets of learnable parameters: d t +d g It's F cross The dimension of each row, D is the dimension of the gate output, and and The final alignment dimension is consistent; W' g 、W' t The dynamic gating stage will splice the vector F ross [i] The matrix used to map to the weight vector, the output dimension D is consistent with the dimension of the image / text features to be fused;
[0213] For each image position i, Perform two linear mappings and add Sigmoid to obtain the weight vectors of image modality and text modality: where α g [i]∈[0,1] D The d-th dimension component represents the weight of the image modality at the i-th position and dimension d, α t [i]∈[0,1] D The d-th dimension component represents the weight of the text modality at the i-th position and dimension d; for different i, α g [i] and α t [i] are independently learning and dynamic;
[0214] Using the position-by-position weight vector obtained by the above calculation, the original image features With text features (Here it is assumed that some resampling or expansion has been done so that In terms of row number Alignment, if n t ≠n g It is necessary to do position mapping or interpolation in advance) to perform element-by-element weighting to obtain the fused joint feature F gated [i]:
[0215] The two are weighted and summed in an element-by-element multiplication manner to obtain the fusion vector at the i-th position. The formula is:
[0216]
[0217] in, is the feature obtained by gated fusion at the i-th position, ⊙ represents element-by-element multiplication, that is, multiplication of positions of the same dimension respectively, which means that for each dimension d, there are:
[0218]
[0219] This is obtained For all i=1,2,…,n g The same operation is performed on each row to form the final matrix The F gated This is the final joint feature representation guided by cross-modal attention and adjusted by gated dynamic weights. It assigns adaptive and learnable weights to the "image modality" and "text modality" at each position, thereby strengthening the discriminative ability of the fusion expression.
[0220] Based on the optimized joint feature F gated, used for two types of bill intelligent recognition subtasks: bill type classification and key field numerical regression: Pool the rows to obtain a global feature vector of dimension D:
[0221]
[0222] Among them, Pool(·) can use strategies such as average pooling or taking the first row to simplify subsequent calculations; It is the joint feature matrix after final fusion, which can be used for subsequent classification and regression tasks;
[0223] Let the total number of categories be C, and configure a set of classification parameters for each category v = 1, 2, ..., C: Then the predicted probability of the vth class is expressed by Softmax:
[0224]
[0225] Among them, p v represents the predicted probability that the current bill belongs to category v, b v ,b u ,W v ,W u are all learning parameters of the classifier;
[0226] The denominator normalizes the sum of the exponentials of all C categories, and finally selects the index with the highest probability as the output category:
[0227]
[0228] in, The bill category label predicted by the model; is the Softmax parameter used for bill type classification;
[0229] Assume there are three types of bills: VAT invoice, receipt, bank slip, and the category numbers are {1, 2, 3} respectively;
[0230] if It means that the model determines that this bill belongs to "VAT invoice"; if It means that the model determines that this bill belongs to "receipt"; if This means the model determines that this bill belongs to "Bank Receipt";
[0231] Assume that the model output for a bill is [0.8, 0.15, 0.05], which means that the probability that the bill belongs to category 1 VAT invoice, category 2 receipt, and category 3 bank receipt is 80%, 15%, and 5% respectively; the model finally predicts That is, the bill type is: "VAT invoice";
[0232] Use the pooling vector obtained in the classification task Recorded as Assume there are O key fields in total, and set the regression parameter for the oth field as: The predicted value of this field is:
[0233]
[0234] in, is the predicted value of the oth field; is the linear parameter used for numerical regression of key fields;
[0235] If fields such as "Invoice Number" are encoded strings, the regression output can be considered as the probability distribution of the corresponding characters or directly mapped to numerical codes before post-processing. If fields such as "Amount" and "Date" correspond to continuous values, the above linear prediction can be directly applied and combined with losses such as mean squared error for training.
[0236] Assume that the model needs to parse three key fields: invoice number, amount, and date; they are recorded as: is the predicted value of the invoice number, is the predicted value of the amount, is the predicted value of the date; if It means the predicted invoice number is "123456789012"; if It means the predicted amount is "450.50 yuan"; if It means the predicted date is "January 1, 2025";
[0237] Assume that the bill contains "Invoice Number: 123456789012, Amount: 450.50 yuan, Date: 2025-01-01". The model uses regression prediction to output the result [123456789012, 450.50, 2025-01-01]. By parsing the model, the system can extract these key fields and return them to the user.
[0238] Step 4: Based on the optimized joint feature representation, the bill type classification task and the key field parsing task are performed simultaneously. A joint loss function is designed. This joint loss function includes the cross-entropy loss of the classification task and the loss of the parsing task. The total loss is calculated using the joint loss function, and all parameters of the entire model are adaptively adjusted using the backpropagation algorithm. After the model training is completed, a new bill image is input. After the complete process described above, the bill type classification result and the parsed values of each key field are finally output;
[0239] Based on the optimized joint feature representation, the bill type classification task and key field parsing task are performed simultaneously. The logic of designing the joint loss function is as follows:
[0240] For bill type prediction, the optimized joint feature F gated Predict the category of the bill; for the i-th sample, the true category uses the one-hot vector (y n,1 ,y n,2 ,…,y n,C ), where y n,v =1 means that the sample belongs to category v, and the other components are 0; the probability that the nth sample belongs to category v after model prediction is recorded as p n,v ∈(0,1), and for all v=1,2,…,C
[0241] This classification task uses cross entropy loss to measure the distance between the true distribution and the predicted distribution, based on the formula:
[0242]
[0243] Among them, N is the number of samples in a training batch, C is the total number of bill categories, and y n,v represents the true label of the nth sample, p n,v represents the predicted probability of the nth sample for type v, L cls Represents the loss value of the classification task; if the true category of the nth sample is v * ,Right now Then the cross entropy loss of this sample is equivalent to - This formula sums and averages N samples over the entire batch to obtain the average classification loss at the batch level;
[0244] Perform regression prediction on the key fields in the bill so that the model's output value for each field is as close to the true label as possible. The parsing task uses mean square error as the loss metric, based on the following formula:
[0245]
[0246] Among them, M=3 is the total number of key fields, and y m are the predicted value and true value of the mth field, L parse Represents the loss value of the field parsing task, reflecting the error of the model's prediction of the field; L parse is the objective function for optimizing the field parsing task. The smaller it is, the better the performance of the parsing task. Close to y j , that is, the predicted value is close to the true value, the loss is reduced;
[0247] If some fields are identifiers, such as invoice numbers, they need to be discretized. It can be viewed as a continuous real number component and then mapped back to a discrete value. If some fields are pure numeric values, such as amounts and date codes, the regression output is used directly.
[0248] This formula squares and sums the prediction errors of M fields, and then takes the average to obtain a scalar loss. If a field is categorical (but we have encoded it as a numeric value here), we can additionally ensure that the encoding method is compatible with the mean squared error loss; otherwise, we can also use cross entropy for this field alone, but this invention requires that the MSE be used as the unified assumption.
[0249] In order to balance the optimization requirements of the classification task and the parsing task, a joint loss function is adopted, based on the formula:
[0250] L=λ cls L cls +λ parse L parse
[0251] Among them, λ cls ≥0 is the weight coefficient of the classification task, λ parse ≥0 is the weight coefficient of the parsing task; when λ cls =λ parse =1, the two losses are treated equally, indicating that the classification task and the parsing task are given the same optimization weight. When the performance of a task is more critical, the corresponding weight is increased to enhance its optimization effect.
[0252] However, statically setting the λ value is difficult to take into account the real-time performance differences of the two tasks. When the "classification loss" of the model drops rapidly at a certain stage, while the "parsing loss" is still high, it is not optimal to continue to give equal weight to the two. Therefore, a dynamic weight adjustment strategy is introduced to make λ cls and λ parse It can adapt to the current training error and thus balance the optimization speed of the two tasks.
[0253] Let the classification loss calculated for the current training iteration / batch be λ cls , the parsing loss is λ parse , the dynamic adjustment formula is:
[0254]
[0255] The task weights are balanced by inverse error. If the classification task performs poorly, that is, λ cls If the model is large, the model will pay more attention to classification. For example, if the model predicts the bill type inaccurately, the system will give priority to adjusting the parameter weights related to classification. If the parsing task performs poorly, that is, λ parseIf the error is large, the model will focus more on parsing. For example, if the model has a large error in parsing the invoice number, the system will prioritize adjusting the weights of parameters related to parsing.
[0256] λ cls and λ parse Represents the dynamic weights of the classification task and the field parsing task. The larger the weight, the more important the corresponding task. The dynamic weight adjusts the optimization focus of the task inversely proportional to the error to ensure that the model can take into account both classification and parsing tasks.
[0257] If L cls If λ is larger, parse Larger, the model focuses more on field parsing tasks. If L parse If λ is larger, cls The larger the size, the more the model focuses on classification tasks. The dynamic weight adjustment mechanism ensures the balanced optimization of classification and parsing tasks, avoiding the performance improvement of one task leading to the performance degradation of another task, thereby improving the overall accuracy and robustness of the model.
[0258] When the classification loss L cls Much higher than the analytical loss L parse When the sum of the two denominators is also too large, but due to L parse Smaller, resulting in And λ parse ≈1, the model will focus more on reducing the gradient contribution corresponding to the "classification loss" because λ parse Large, the driving parameter update is mainly to reduce L cls direction, on the contrary, when the parsing loss L parse Much higher than the classification loss L cls When λ cls ≈1,λ parse <<1, the model is more inclined to reduce the analytical error. When the two are close, the weights of both are about 0.5, and the model is optimized according to the average contribution of each;
[0259] The final dynamic joint loss is defined as:
[0260]
[0261] This design ensures that during training, tasks with larger losses always receive higher gradient weights, quickly reducing their errors. Tasks with smaller losses temporarily reduce the impact of gradients to avoid over-optimization of already good parts and resulting in uneven overall performance.
[0262] The above formulas are all dimensionless and numerical calculations. The formulas are obtained by collecting a large amount of data and performing software simulation to obtain the most recent real situation. The preset parameters in the formulas are set by technicians in this field according to actual conditions.
[0263] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed by hardware or software depends on the specific application and design constraints of the technical solution.
[0264] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, and may be located in one place or distributed across multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment. The above description is only a specific implementation method of the present application, but the scope of protection of the present application is not limited thereto. Any person skilled in the art can easily think of changes or replacements within the technical scope disclosed in this application, which should be covered by the scope of protection of this application.
Claims
1. A bill intelligent recognition method based on a deep neural network model, characterized in that: The specific steps include: Step 1: Obtain original images of various types of bills, preprocess them, and use an OCR tool to extract text data from the bill images. The extracted text data includes the invoice number, amount, and date. The image data is preprocessed, including grayscale conversion, binarization, and edge detection. The extracted text data is formatted and validity verified. Step 2: Based on the preprocessed image, a convolutional neural network is used to output an image feature vector that represents the visual structure. The verified key field text is input into the LSTM to generate a text feature vector that encodes semantic information. The image feature vector and the text feature vector are then concatenated in the feature dimension to form a preliminary joint feature vector. Step 3: The attention weights between the image and text feature vectors are calculated through a cross-modal attention mechanism. Based on the attention weights, a gating mechanism is used to dynamically adjust the contribution ratios of the image and text modalities in the final joint representation. The weighted fusion is then used to generate an optimized cross-modal joint feature representation. Step 4: Based on the optimized joint feature representation, the bill type classification task and the key field parsing task are performed simultaneously. A joint loss function is designed, which includes the cross-entropy loss of the classification task and the loss of the parsing task. The total loss is calculated through the joint loss function, and all parameters of the entire model are adaptively adjusted using the back-propagation algorithm. After the model training is completed, a new bill image is input. After the above complete process, the bill type classification result and the parsing value of each key field are finally output.
2. The method for intelligent bill recognition based on a deep neural network model according to claim 1, characterized in that: Obtain original images of various types of bills and preprocess the original images as follows: The various types of receipts obtained include paper receipts and electronic receipts. For paper receipts, a scanner is used to scan the receipts at high resolution. For electronic receipts, electronic samples are obtained from the required business and converted into JPEG / PNG format images. Preprocessing the original image includes grayscale conversion, binarization, and edge detection. Grayscale conversion involves converting a color image into a grayscale image to simplify subsequent processing. Binarization involves converting a grayscale image into a black and white binary image using an Otsu algorithm with an adaptive threshold to enhance the contrast between the foreground and background of the bill. Edge detection involves identifying significant edge contours in the image using the Canny edge detection algorithm. Based on the edge detection results, the main outer frame of the bill is located using a contour analysis algorithm. According to the detected frame coordinates, the original image is cropped to correct the tilt and extract the region of interest image containing only the main area of the bill. Based on the cropped and corrected ROI image, the OCR tool is used to extract the text data on the bill. The extracted text data contains key field information, namely the invoice number, amount, and date. The extracted original text data is preliminarily formatted and validated, including checking the date format and amount number format.
3. The method for intelligent bill recognition based on a deep neural network model according to claim 2, characterized in that: Based on the detected bill border coordinates (x min ,y min ,x max ,y max ), and the width of the image W = x max -x min and height H = y max -y min , perform fine preprocessing on the cropped and corrected ROI image, and extract text feature data through the OCR tool. The specific logic is as follows: Use the border coordinates to expand and crop the original image to retain the edge information of the bill. The cropping formula is: I c =I[max0,x min -δ):min(x max +δ,W),max(0,y min -δ):min(y max +δ,H)] Among them, I c represents the content of the cropped bill image, δ is the border expansion coefficient, which is set to 5% of the bill border size, that is, δ = 0.05 max (x max -x min ,y max -y min ); To adapt to the input size of the deep learning model, the cropped bill images are uniformly resized to 224×224, and the scaling ratio is defined as: The scaled size is calculated as: Where H', W' represent the height and width of the scaled image respectively; Next, the image is scaled proportionally to its long side and filled with pixel values 0 to a size of 224×224. The formula is: I r =PadToCenter(Resize(I c ,H',W'),224,224); Based on the image I after scaling and filling r , use OCR tools to perform text recognition, extract text feature data, including invoice number, amount and date fields, output as a text collection in structured JSON format, and uniformly format and verify the validity of the extracted text collection.
4. The method for intelligent bill recognition based on a deep neural network model according to claim 3, characterized in that: Based on the preprocessed image, the logic of using the convolutional neural network to output the image feature vector representing the visual structure is: The cropped and scaled single-channel grayscale image I r The pixel values are normalized to the range [0,1]: Among them, I n (x,y) represents the normalized pixel value, which serves as the model input; Construct a lightweight convolutional neural network for feature extraction. The model structure is as follows: Input layer: receives normalized image I n (x,y), size is 224×224×1, convolution layer 1: convolution kernel size 3×3, number of filters 32, stride 1, activation function is ReLU, output size is 222×222×32 feature map; pooling layer 1: maximum pooling 2×2, output size is 111×111×32; Convolutional layer 2: convolution kernel size 3×3, number of filters 64, stride 1, activation function is ReLU, output size is 109×109×64; Pooling layer 2: maximum pooling 2×2, output size is 54×54×64; Fully connected layer: Flatten the high-dimensional feature map 54×54×64 into a one-dimensional feature vector, and connect two layers of fully connected layers in sequence to generate the image feature vector F g ; The attention mechanism is introduced to enhance the feature extraction capability of convolutional neural networks. The formula is as follows: in, It is the feature vector after attention weighting, and the Attention module is channel attention: Attention(F g )=F g ·σ(FC(GlobalPool(F g ))) Among them, F g Image features extracted by convolutional neural network, size is 128, GlobalPool (F g ) is the eigenvector F g Perform global pooling, FC is the fully connected layer, σ is the Sigmoid activation function, and the weight is limited to the range of [0,1]. g ·σ(·) is an element-wise multiplication that applies the attention weights to the original features.
5. The method for intelligent bill recognition based on a deep neural network model according to claim 4, characterized in that: The logic of inputting the verified key field text into LSTM to generate a text feature vector encoding semantic information is as follows: The extracted text data contains key fields, including invoice number, amount, and date: For the invoice number, replace the 12-digit string T id Decomposed into a character sequence, mapped to a dimension of d through a character-level Embedding layer ch ar =8-dimensional continuous vector, based on the formula: Ch arEmbed(T id )=Embedding(T id ,d ch ar ) Among them, CharEmbed(T id ) has an output shape of 12×d ch ar , that is, for each numeric character in the sequence, a d ch ar A vector of dimensions; The character embedding sequence is then modeled using a long short-term memory network to generate a fixed-length representation based on the following formula: Embedding(T id )=LSTM(Ch arEmbed(T id )) For the amount T am and date T da , is directly input into the model after normalization; The final text feature vector is obtained based on the formula: in, is the final representation of the text feature vector, Embedding(T id ) Obtain the invoice number T through embedding and sequence modeling id Context feature representation, Normalized(T am ,T da ) Through normalization, the amount T am and date T da Convert to numerical features; In order to realize the joint modeling of image and text, a nonlinear interaction mechanism is introduced to transform the image feature vector and text feature vector Splicing into a joint feature index F, the formula is: Among them, F adopts the following method: Among them, ⊙ represents element-by-element multiplication, Used to enhance the interaction between image and text features, addition The independent information of single modal features is retained.
6. The method for intelligent bill recognition based on a deep neural network model according to claim 5, characterized in that: The logic of calculating the attention weight between image and text feature vectors through the cross-modal attention mechanism is: The image feature matrix is recorded as where n g represents the number of features output from the convolutional neural network, d g Represents the dimension of each image feature vector; the text feature matrix is recorded as n t represents the number of features output from the text encoding network, d t Represents the dimension of each text feature vector; When measuring the similarity between the i-th feature of the image and the j-th feature of the text, first project them to the same hidden dimension d h , specifically: define Among them, W g For each d g The image feature map of dimension d h Wei, W t For each d t The text feature map of dimension d is h dimension; For any pair of image features and text features, calculate their similarity in the unified hidden space and obtain the attention weight through Softmax normalization. Specifically, for the i-th image feature and the j-th text feature, the attention weight is defined as: in, is the i-th row vector of the image feature matrix, i=1,2,…,n g , is the j-th row vector of the text feature matrix are respectively the vector representations after being projected to the same hidden layer dimension, A i,j represents the attention weight of the i-th image feature and the j-th text feature; For all i=1,2,…,n g and j = 1, 2, ..., n t The attention weight {A i,j }Combined into attention matrix Use the attention matrix A to analyze the text feature matrix Perform weighted analysis to obtain the attention-enhanced text representation and generate a more semantic cross-modal fusion representation F cross , based on the formula: in, That is, for each image feature position, the cross-modal representation is formed by concatenating the text context vector captured by attention and the image feature vector. Represents the text features weighted by the attention mechanism, and the Concat operation combines them with the original image features Concatenation, Concat(·) is to connect the attention-enhanced text features with the original image features along the feature dimension; Further adaptively determine the contribution ratio of image and text modalities in the final joint feature, and introduce a gating mechanism to F cross Map and generate corresponding weight coefficients; define two sets of learnable parameters: d t +d g It's F cross The dimension of each row, D is the dimension of the gate output, and and The final alignment dimensions are consistent; For each image position i, Perform two linear mappings and add Sigmoid to obtain the weight vectors of image modality and text modality: where α g [i]∈[0,1] D The d-th dimension component represents the weight of the image modality at the i-th position and dimension d, α t [i]∈[0,1] D The d-th dimension component represents the weight of the text modality at the i-th position and dimension d; The two are weighted and summed in an element-by-element multiplication manner to obtain the fusion vector at the i-th position. The formula is: in, is the feature obtained by gated fusion at the i-th position, ⊙ represents element-by-element multiplication, that is, multiplication of positions of the same dimension respectively, which means that for each dimension d, there are: This is obtained For all i=1,2,…,n g The same operation is performed on each row to form the final matrix The F gated That is, the final joint feature representation guided by cross-modal attention and adjusted by gated dynamic weights; Based on the optimized joint feature F gated , used for two types of bill intelligent recognition subtasks: bill type classification and key field numerical regression: Pool the rows to obtain a global feature vector of dimension D: Among them, Pool(·) is average pooling or taking the first row; Let the total number of categories be C, and configure a set of classification parameters for each category v = 1, 2, ..., C: Then the predicted probability of the vth class is expressed by Softmax: Among them, p v represents the predicted probability that the current bill belongs to category v, b v ,b u ,W v ,W u are all learning parameters of the classifier; The denominator normalizes the sum of the exponentials of all C categories, and finally selects the index with the highest probability as the output category: in, The bill category label predicted by the model; Use the pooling vector obtained in the classification task Recorded as Assume there are O key fields in total, and set the regression parameter for the oth field as: The predicted value of this field is: in, is the predicted value of the oth field.
7. The method for intelligent bill recognition based on a deep neural network model according to claim 6, characterized in that: Based on the optimized joint feature representation, the bill type classification task and key field parsing task are performed simultaneously. The logic of designing the joint loss function is as follows: For bill type prediction, the optimized joint feature F gated Predict the category of the bill. This classification task uses cross entropy loss to measure the distance between the true distribution and the predicted distribution. The formula is: Among them, N is the number of samples in a training batch, C is the total number of bill categories, and y n,v represents the true label of the nth sample, p n,v represents the predicted probability of the nth sample for type v, λ cls Represents the loss value of the classification task; Perform regression prediction on the key fields in the bill so that the model's output value for each field is as close to the true label as possible. The parsing task uses mean square error as the loss metric, based on the following formula: Among them, M=3 is the total number of key fields, and y m are the predicted value and true value of the mth field, L parse Represents the loss value of the field parsing task; In order to balance the optimization requirements of the classification task and the parsing task, a joint loss function is adopted, based on the formula: L=λ cls L cls +λ narse L parse Among them, λ cls ≥0 is the weight coefficient of the classification task, λ parse ≥0 is the weight coefficient of the parsing task; Let the classification loss calculated for the current training iteration / batch be λ cls , the parsing loss is λ parse , the dynamic adjustment formula is: The task weights are balanced by inverse error. If the classification task performs poorly, that is, λ cls If the model is larger, it will focus more on classification; if the parsing task performs poorly, that is, λ parse Larger, the model will focus more on parsing.
Citation Information
Cited By
Multi-modal large model-based worker safety protection equipment missing detection method
CN121190756A
Financial document intelligent verification method and system
CN121505641A