Vision-based bill content identification method and system

Through the visual-based bill content recognition method, preprocessing the bill image, detecting text areas, identifying text content and converting the format, the problems of inaccurate invoice recognition and inconsistent format in the prior art are solved, and invoice information recognition with high accuracy and format consistency are achieved.

CN120107986APending Publication Date: 2025-06-06WUHAN LINGYU INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510060735.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

When the prior art recognizes complex and diverse invoice formats, fuzzy handwriting or changes in layout, the recognition effect is inaccurate and the recognition result format is inconsistent, which affects the efficiency of data statistical analysis.

Method used

The visual-based bill content recognition method is adopted to preprocess the bill image, determine the text area using the preset detection model, and perform text detection to extract the text content. Then, text recognition is performed based on the preset recognition model, invoice key information is obtained, format conversion is performed, and structured invoice information is generated.

Benefits of technology

It improves the accuracy of invoice information identification, ensures the consistency of the identification data format, and enhances the efficiency of subsequent data statistical analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107986A_ABST
    Figure CN120107986A_ABST
Patent Text Reader

Abstract

The invention discloses a bill content identification method and system based on vision, and the method comprises the steps: obtaining a bill image of a to-be-identified bill, and carrying out the preprocessing of the bill image, and obtaining a preprocessed image; determining a character area in the preprocessed image based on a preset detection model, performing line text detection on the character area, and extracting text content of a line text in the character area; and performing text recognition on the text content based on a preset recognition model to obtain invoice key information, and performing format conversion on the invoice key information to obtain structured invoice information for storage. According to the invention, accurate identification of invoices and consistency of identification result formats are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a vision-based bill content recognition method and system. Background Art

[0002] Invoices are not only the original basis for accounting, but also an important basis for law enforcement inspections by auditing agencies and tax authorities. In the past, people often needed to enter invoice information data into the corresponding system for reimbursement, auditing, certification, and archiving. However, with the advancement of digital transformation, a large number of paper invoices and electronic invoices will be scanned or directly generated. The resulting invoice recognition technology can automatically extract key information from these invoices, such as invoice number, amount, date, etc., to improve work efficiency.

[0003] At present, the extraction of key invoice information mainly uses optical character recognition (OCR) technology to directly recognize invoice images. However, when faced with complex and diverse invoice formats, blurred handwriting or layout changes, the recognition effect is often unsatisfactory and there is a problem of inaccurate recognition. At the same time, there may also be problems with the format of the recognition results being inconsistent, which affects the efficiency of subsequent data statistical analysis. Therefore, how to improve the accuracy of invoice information recognition and ensure the consistency of the recognition data format is one of the technical problems that need to be solved at present. Summary of the invention

[0004] The main purpose of the present invention is to provide a vision-based bill content recognition method and system, aiming to solve the technical problem of how to improve the accuracy of invoice information recognition and ensure the consistency of recognition data format in the prior art.

[0005] To achieve the above object, the present invention provides a method for bill content recognition based on vision, and the method for bill content recognition based on vision includes:

[0006] Acquire a bill image of a bill to be identified, and preprocess the bill image to obtain a preprocessed image;

[0007] Determine a text area in the preprocessed image based on a preset detection model, perform line text detection on the text area, and extract text content of the line text in the text area;

[0008] The text content is recognized based on a preset recognition model to obtain key invoice information, and the key invoice information is formatted to obtain structured invoice information for storage.

[0009] Preferably, the step of performing text recognition on the line text content based on a preset recognition model to obtain key invoice information includes:

[0010] The text content is segmented based on a preset recognition model to obtain a plurality of words, and the following steps are performed for each of the words:

[0011] The first word vector corresponding to the word and the second word vector corresponding to the preset invoice keyword are respectively calculated for similarity, and the target keyword corresponding to the word in the preset invoice keyword is determined, wherein the formula for calculating the similarity is:

[0012]

[0013] Where S represents similarity, n represents the larger value between the dimension of the first word vector and the dimension of the second word vector, ai represents the first word vector, and bi represents the second word vector;

[0014] The key information corresponding to the target keyword in the line text content is obtained, and the target keyword and the key information are generated together as the invoice key information.

[0015] Preferably, the step of performing text recognition on the text content based on a preset recognition model to obtain key invoice information includes:

[0016] Randomly selecting target key information from each of the invoice key information, and identifying the information type of the target key information;

[0017] Verify the target key information according to the preset information format corresponding to the information type, and determine whether the target key information passes the verification

[0018] If the target key information passes the verification, the step of converting the format of the invoice key information is executed;

[0019] If the target key information fails to pass the verification, feedback information is generated, and based on the feedback information, a step of performing word segmentation processing on the text content based on a preset recognition model is executed.

[0020] Preferably, the step of determining the text area in the preprocessed image based on a preset detection model comprises:

[0021] Detecting the four corner positioning points of the text contained in the preprocessed image based on a preset detection model, and constructing a cutting template according to the four corner positioning points;

[0022] Determine an initial text area in the preprocessed image according to the cutting template and the preset template, and pre-segment the initial text area to obtain a pre-segmented area;

[0023] The pre-segmented area is identified to determine whether the text in the pre-segmented area is complete; if the text in the pre-segmented area is complete, the pre-segmented area is determined as a text area in the pre-processed image.

[0024] Preferably, before the step of detecting the four corner positioning points of the text contained in the preprocessed image based on a preset detection model, the step includes:

[0025] Acquire a large number of image samples, and divide the large number of image samples into training samples and test samples;

[0026] Training a preset initial model based on the training samples, and when the training time reaches a preset time, testing the preset initial model based on the test samples to obtain a test result;

[0027] The loss function value of the preset initial model is calculated based on the test results, and the calculation formula is:

[0028]

[0029] Where L represents the loss function value, m represents the number of test samples, pj represents the test result corresponding to the j-th test sample, and qj represents the reference result corresponding to the j-th test sample;

[0030] Compare the loss function value with a preset loss threshold to determine whether the loss function value is less than the preset loss threshold, and if so, generate the preset initial model as a preset detection model;

[0031] If the loss function value is greater than the preset loss threshold, the model parameters of the preset initial model are updated based on a preset update formula, and the preset update formula is:

[0032]

[0033] Among them, θ t+1 represents the updated model parameters, θ t represents the model parameters before updating, is the learning rate of the preset initial model, It is the gradient of the loss function of the preset initial model before the model parameters are updated;

[0034] For the updated preset initial model, a step of training the preset initial model based on the training sample is performed until the loss function value is less than a preset loss threshold.

[0035] Furthermore, to achieve the above-mentioned purpose, the present invention also provides a vision-based bill content recognition system, the vision-based bill content recognition system comprising:

[0036] A first acquisition module is used to acquire a bill image of a bill to be identified, and preprocess the bill image to obtain a preprocessed image;

[0037] A detection module, used to determine a text area in the preprocessed image based on a preset detection model, perform line text detection on the text area, and extract text content of the line text in the text area;

[0038] The recognition module is used to perform text recognition on the text content based on a preset recognition model to obtain key invoice information, and to perform format conversion on the key invoice information to obtain structured invoice information for storage.

[0039] Preferably, the identification module further includes:

[0040] The word segmentation unit is used to perform word segmentation processing on the text content based on a preset recognition model to obtain multiple words, and perform the following steps for each of the words:

[0041] A calculation unit is used to perform similarity calculation on the first word vector corresponding to the word and the second word vector corresponding to the preset invoice keyword, respectively, to determine the target keyword corresponding to the word in the preset invoice keyword, wherein the formula for similarity calculation is:

[0042]

[0043] Where S represents similarity, n represents the larger value between the dimension of the first word vector and the dimension of the second word vector, ai represents the first word vector, and bi represents the second word vector;

[0044] The acquisition unit is used to acquire key information corresponding to the target keyword in the line text content, and generate the target keyword and the key information together as the invoice key information.

[0045] Preferably, the bill content recognition system further comprises:

[0046] A screening module, used for randomly screening out target key information from each of the invoice key information, and identifying the information type of the target key information;

[0047] A verification module is used to verify the target key information according to a preset information format corresponding to the information type, and determine whether the target key information passes the verification.

[0048] The identification module is also used to execute the step of format conversion of the invoice key information if the target key information passes the verification;

[0049] If the target key information fails to pass the verification, feedback information is generated, and based on the feedback information, a step of performing word segmentation processing on the text content based on a preset recognition model is executed.

[0050] Preferably, the detection module comprises:

[0051] A detection unit, configured to detect the four corner positioning points of the text contained in the preprocessed image based on a preset detection model, and to construct a cutting template according to the four corner positioning points;

[0052] A preprocessing unit, configured to determine an initial text region in the preprocessed image according to the cutting template and the preset template, and pre-segment the initial text region to obtain a pre-segmented region;

[0053] The determination unit is used to identify the pre-segmented area and determine whether the text in the pre-segmented area is complete. If the text in the pre-segmented area is complete, the pre-segmented area is determined as the text area in the pre-processed image.

[0054] Preferably, the vision-based bill content recognition system further includes:

[0055] A second acquisition module is used to acquire a large number of image samples and divide the large number of image samples into training samples and test samples;

[0056] A training module, used to train a preset initial model based on the training samples, and when the training time reaches a preset time, test the preset initial model based on the test samples to obtain a test result;

[0057] The calculation module is used to calculate the loss function value of the preset initial model based on the test result, and the calculation formula is:

[0058]

[0059] Where L represents the loss function value, m represents the number of test samples, pj represents the test result corresponding to the j-th test sample, and qj represents the reference result corresponding to the j-th test sample;

[0060] A judgment module, used for comparing the loss function value with a preset loss threshold, judging whether the loss function value is less than the preset loss threshold, and if it is less than the preset loss threshold, generating the preset initial model as a preset detection model;

[0061] An updating module is used to update the model parameters of the preset initial model based on a preset updating formula if the loss function value is greater than a preset loss threshold, and the preset updating formula is:

[0062]

[0063] Among them, θ t+1 represents the updated model parameters, θ t represents the model parameters before updating, is the learning rate of the preset initial model, It is the gradient of the loss function of the preset initial model before the model parameters are updated;

[0064] The training module is also used to execute the step of training the preset initial model based on the training sample for the updated preset initial model until the loss function value is less than a preset loss threshold.

[0065] The visual-based bill content recognition method and system of the present invention, after obtaining the bill image of the bill to be recognized, pre-processes the bill image to obtain a pre-processed image; then determines the text area in the pre-processed image according to the pre-trained preset detection model, and performs line text detection on the text area to extract the text content of the line text in the text area; in addition, a preset recognition model is also pre-trained, and the extracted line text content is recognized by the preset recognition model to obtain the key information of the invoice, and then the key information of the invoice is formatted to obtain structured invoice information for storage. In this way, the text area where various types of text in the invoice are located is accurately detected by the pre-trained preset detection model and the preset recognition model to extract the text content, and the extracted text content is accurately recognized, thereby improving the accuracy of obtaining the key information of the invoice. At the same time, by formatting the obtained key information of the invoice, the resultant text information is generated and stored, so that each extracted key information of the invoice has a unified structured format. Accurate recognition of the invoice and consistency of the recognition result format are achieved. BRIEF DESCRIPTION OF THE DRAWINGS

[0066] Figure 1 It is a module diagram of the first embodiment of the bill content recognition method based on vision of the present invention;

[0067] Figure 2 It is a module diagram of the second embodiment of the bill content recognition method based on vision of the present invention;

[0068] Figure 3 It is a flow chart of the first embodiment of the bill content recognition system based on vision of the present invention;

[0069] Figure 4 It is a flow chart of the second embodiment of the vision-based bill content recognition system of the present invention.

[0070] The realization of the purpose, functional features and advantages of the present invention will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0071] It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention.

[0072] The present invention provides a method for identifying bill contents based on vision. Figure 1 , Figure 1 The flowchart of the first embodiment of the method for identifying bill contents based on vision of the present invention is shown in the flowchart. It should be noted that although the logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than that shown here. Specifically, the method for identifying bill contents based on vision in this embodiment includes:

[0073] Step S10, obtaining a bill image of the bill to be identified, and preprocessing the bill image to obtain a preprocessed image;

[0074] The vision-based bill content recognition method of this embodiment is applied to a vision-based bill content recognition system. When the bill content recognition is required, a recognition request is initiated to the system based on the image of the bill to be recognized. After receiving the recognition request, the system takes the bill to be recognized as the bill to be recognized, and obtains the image carried in the recognition request as the bill image of the bill to be recognized. Then, the image to be recognized is preprocessed to obtain a preprocessed image. Preprocessing includes but is not limited to image denoising, binarization, edge detection and other operations to improve image quality. Image denoising is used to reduce or eliminate noise in the bill image while retaining the details and features of the image as much as possible; the denoising method can be any one of Gaussian filtering, median filtering, and wavelet transform. The binarization of the image is to convert the image into a pixel value containing only black and white pixels, thereby simplifying the image data and reducing the complexity of subsequent processing. Specifically, the binarization can be performed based on a threshold. The edge detection of the image is to identify the contour and boundary of the object in the image, and the edge detection can be performed by the Canny operator. In this embodiment, after the bill image is denoised, binarized and edge detected, the preprocessed image is obtained.

[0075] Step S20, determining a text area in the preprocessed image based on a preset detection model, performing line text detection on the text area, and extracting text content of the line text in the text area;

[0076] Furthermore, a preset detection model is formed in advance through training with a large number of image samples, and the pre-processed image is identified through the preset detection model to determine the text area containing the text. The identification of the text area is mainly achieved through four-corner point positioning, and the position of the text area is determined by identifying the positions of the four corners of the text area in the pre-processed image. Specifically, the step of determining the text area in the pre-processed image based on the preset detection model includes:

[0077] Step S21, detecting the four corner positioning points of the text contained in the preprocessed image based on a preset detection model, and constructing a cutting template according to the four corner positioning points;

[0078] Step S22, determining an initial text area in the preprocessed image according to the cutting template and the preset template, and pre-segmenting the initial text area to obtain a pre-segmented area;

[0079] Step S23, identifying the pre-segmented area to determine whether the text in the pre-segmented area is complete; if the text in the pre-segmented area is complete, determining the pre-segmented area as the text area in the pre-processed image.

[0080] Furthermore, a two-dimensional coordinate system of the preprocessed image is established by a preset detection model, for example, the upper left corner of the preprocessed image is used as the coordinate origin, the horizontal rightward direction of the preprocessed image is used as the X coordinate axis, and the vertical downward direction of the preprocessed image is used as the Y coordinate axis. Then, the coordinate point positions of each character contained in the preprocessed image in the two-dimensional coordinate system are determined, and the four diagonal coordinate points with the farthest distance are determined by the positions of each coordinate point. The characters corresponding to the four diagonal coordinate points are the four corner positioning points of the characters contained in the preprocessed image.

[0081] Furthermore, a preset template for representing a text area is pre-set for the preprocessed image. The quadrilateral formed by the four corner positioning points is constructed as a cutting template, and the cutting template is compared with the preset template, and the initial text area in the preprocessed image is determined by the comparison. Among them, the comparison is to determine whether the length, width and area between the cutting template and the preset template are all matched, and the match means that the length, width and area of ​​the cutting template are all greater than the length, width and area of ​​the preset template within a certain range. If the comparison determines that they all match, it means that the cutting area corresponding to the cutting template matches the text area reflected by the preset template, so that the cutting area corresponding to the cutting template is determined as the initial text area in the pre-segmented area, and the initial text area is pre-segmented to obtain the corresponding pre-segmented area.

[0082] Furthermore, considering the matching of the cutting template and the preset template, there may be only a certain area in the preprocessed image that matches the text area reflected by the preset template in size, and the cutting area of ​​the cutting template does not accurately divide the text area. In this regard, the present embodiment identifies the pre-segmented area to identify whether the text contained therein is complete. If it is complete, it means that the division of the pre-segmented area is accurate, and the pre-segmented area is determined as the text area in the preprocessed image. On the contrary, if it is identified that the text contained in the pre-segmented area is incomplete, it means that the division of the pre-segmented area is inaccurate. At this time, the pre-segmented area is adjusted to contain complete text, and the adjusted pre-segmented area is determined as the text area.

[0083] Furthermore, the text area may contain multiple lines of text. In order to recognize more accurately, the text area is detected by line text detection, the text area is recognized and divided into multiple lines of text, and the text content contained in each line of text is extracted. The text content extraction is the operation of distinguishing each line of text from the text area.

[0084] Step S30, performing text recognition on the text content based on a preset recognition model to obtain key invoice information, and performing format conversion on the key invoice information to obtain structured invoice information for storage.

[0085] Furthermore, a preset recognition model is formed in advance through training, and text recognition is performed on each extracted text content according to the preset recognition model, that is, the specific content corresponding to each line of text is recognized, and the key invoice information is obtained through recognition. The key invoice information includes but is not limited to the invoice date, invoice number, invoice amount, etc. Specifically, the step of performing text recognition on the line of text content based on the preset recognition model to obtain the key invoice information includes:

[0086] Step S31, performing word segmentation processing on the text content based on a preset recognition model to obtain multiple words, and performing the following steps for each of the words:

[0087] Step S32, respectively calculating the similarity between the first word vector corresponding to the word and the second word vector corresponding to the preset invoice keyword, and determining the target keyword corresponding to the word in the preset invoice keyword;

[0088] Step S33, obtaining key information corresponding to the target keyword in the line text content, and generating the target keyword and the key information together as the invoice key information.

[0089] Furthermore, the preset recognition model performs word segmentation processing on the text content to obtain multiple words. Then, for each divided word, the corresponding word vector is searched from the preset dictionary as the first word vector. At the same time, preset invoice keywords and their corresponding second word vectors are pre-set according to the key information of the invoice. The first word vector and each second word vector are respectively calculated for similarity, and the similarity between the word and each preset invoice keyword is determined by calculation. The similarity calculation formula can be referred to as the following formula (1).

[0090]

[0091] Among them, S represents similarity, n represents the larger value of the dimension of the first word vector and the dimension of the second word vector, ai represents the first word vector, and bi represents the second word vector. The calculated value is between 0 and 1. The closer it is to 1, the higher the similarity, and the closer it is to 0, the lower the similarity.

[0092] Furthermore, the calculated similarities are compared to determine the maximum value, i.e., the value closest to 1. Then, the preset invoice keyword that generates the maximum value is searched as the target keyword corresponding to the word. The key information corresponding to the target keyword is identified from the text content. For example, if the target keyword is the invoice number, the specific invoice number is identified from the text content as the key information. Moreover, the key information is usually located in the same line of text as the target keyword, so the key information can be identified from the same line of text. Then, the target keyword and its corresponding key information are generated as invoice key information together, for example, the two are formed into invoice key information in the form of a key-value pair.

[0093] Furthermore, in order to ensure the consistency of the format of the key invoice information obtained from various forms of invoices, a structured format is pre-set. After the key invoice information is obtained, the key invoice information is formatted and converted into a unified structured format, and the structured invoice information is obtained and stored for subsequent statistical analysis and other processing.

[0094] In order to ensure the accuracy of the generated invoice key information, this embodiment also provides a verification mechanism. Specifically, the step of performing text recognition on the text content based on a preset recognition model to obtain the invoice key information includes:

[0095] Step a1, randomly selecting target key information from each of the invoice key information, and identifying the information type of the target key information;

[0096] Step a2: verify the target key information according to the preset information format corresponding to the information type, and determine whether the target key information passes the verification.

[0097] Step a3, if the target key information passes the verification, then executing the step of converting the format of the invoice key information;

[0098] Step a4, if the target key information fails to pass the verification, feedback information is generated, and based on the feedback information, a step of performing word segmentation processing on the text content based on a preset recognition model is executed;

[0099] Furthermore, from the invoice key information obtained by identifying the various text contents, any one item is randomly selected as the target key information, and the information type of the target key information is identified. The information type includes at least the invoice number, the invoice amount, and the invoicing time. Different information formats are pre-set for different information types. For the information type of the target key information, the corresponding preset information format is searched, and the target key information is verified according to the preset information format. It is determined whether the information format of the target key information is consistent with the preset information format. If they are consistent, a verification result of passing the verification is generated, and if they are inconsistent, a verification result of failing the verification is generated.

[0100] Furthermore, after the verification result is generated, it is determined whether the target key information has passed the verification based on the verification result. If the target key information has passed the verification, it means that the generated invoice key information is accurate, and the invoice key information is formatted and converted into a unified structured format for storage. On the contrary, if the target key information has not passed the verification, it means that the generated invoice key information is inaccurate. At this time, feedback information is generated, and the text content is re-segmented by the preset recognition model based on the feedback information. The feedback information contains incorrect target key information, and the preset recognition model combines it to re-segment the text content, which can correct the incorrect target key information and obtain accurate target key information.

[0101] The vision-based bill content recognition method implemented in this embodiment, after obtaining the bill image of the bill to be recognized, pre-processes the bill image to obtain a pre-processed image; then determines the text area in the pre-processed image according to the pre-trained preset detection model, and performs line text detection on the text area to extract the text content of the line text in the text area; in addition, a preset recognition model is also pre-trained, and the extracted line text content is recognized by the preset recognition model to obtain the key information of the invoice, and then the key information of the invoice is formatted to obtain structured invoice information for storage. In this way, the text area where various types of text in the invoice are located is accurately detected by the pre-trained preset detection model and the preset recognition model to extract the text content, and the extracted text content is accurately recognized, thereby improving the accuracy of obtaining the key information of the invoice. At the same time, by formatting the obtained key information of the invoice, the resultant text information is generated for storage, so that each extracted key information of the invoice has a unified structured format. Accurate recognition of the invoice and consistency of the recognition result format are achieved.

[0102] For further information, please refer to Figure 2 Based on the first embodiment of the vision-based bill content recognition method of the present invention, a second embodiment of the vision-based bill content recognition method of the present invention is proposed.

[0103] The difference between the second embodiment of the vision-based bill content recognition method and the first embodiment of the vision-based bill content recognition method is that:

[0104] The step of detecting the four corner positioning points of the text contained in the preprocessed image based on the preset detection model includes:

[0105] Step S40, obtaining a large number of image samples, and dividing the large number of image samples into training samples and test samples;

[0106] Step S50, training a preset initial model based on the training sample, and when the training time reaches a preset time, testing the preset initial model based on the test sample to obtain a test result;

[0107] Step S60, calculating the loss function value of the preset initial model based on the test result;

[0108] Step S70, comparing the loss function value with a preset loss threshold to determine whether the loss function value is less than the preset loss threshold, and if so, generating the preset initial model as a preset detection model;

[0109] Step S80, if the loss function value is greater than a preset loss threshold, updating the model parameters of the preset initial model based on a preset update formula;

[0110] Step S90, for the updated preset initial model, executing the step of training the preset initial model based on the training sample until the loss function value is less than a preset loss threshold.

[0111] Furthermore, the present embodiment forms a preset detection model by training a large number of image samples. Specifically, a large number of image samples are obtained, and each image sample is divided into training samples and test samples according to a preset ratio, such as 8 to 2, 7 to 3, etc., with more training samples than test samples. Then, the preset initial model is trained based on the training samples, and a preset duration is preset. When the statistical training duration reaches the preset duration, the preset initial model is tested by the test sample. The preset initial model performs detection and analysis on the test sample to generate a corresponding test result.

[0112] Furthermore, in order to reflect the performance of the preset initial model, the loss function value of the preset initial model is calculated based on the test results. The loss function value represents the difference between the test result obtained by the preset initial model for detecting and analyzing the test sample and the reference result corresponding to the test sample. The specific calculation formula can be found in the following formula (2).

[0113]

[0114] Among them, L represents the loss function value, m represents the number of test samples, pj represents the test result corresponding to the j-th test sample, and qj represents the reference result corresponding to the j-th test sample.

[0115] The larger the loss function value, the larger the difference between the test result and the reference result, the poorer the performance of the preset initial model, and the need to continue iterative training of the preset initial model. Conversely, the smaller the loss function value, the smaller the difference between the test result and the reference result, the better the performance of the preset initial model, and the training of the preset initial model can be stopped. In order to indicate the size of the loss function value, a preset loss threshold is pre-set, and the loss function value is compared with the preset loss threshold to determine whether the loss function value is less than the preset loss threshold. If it is less than, it means that the preset initial model has a good detection and analysis performance for the test sample. At this time, the training of the preset initial model is stopped and it is generated as the preset detection model. Conversely, if the loss function value is determined to be greater than the preset loss threshold through comparison, it means that the performance of the preset initial model is poor, and the preset initial model needs to be iteratively trained. A preset update formula for updating the model parameters is pre-set, and the model parameters of the preset initial model are updated by the preset update formula. After the update, the preset initial model is iteratively trained, and the loss function value is re-tested and calculated until the loss function value obtained by the test is less than the preset loss threshold. The preset update formula can be specifically referred to in the following formula (3).

[0116]

[0117] Among them, θ t+1 represents the updated model parameters, θ t represents the model parameters before updating, is the learning rate of the preset initial model, It is the gradient of the loss function of the preset initial model before the model parameters are updated. The gradient refers to the contribution of the model parameters to the loss function calculated by backpropagating the error from the output layer to the input layer.

[0118] In this embodiment, a large number of image samples are used to train the preset initial model, and the test samples are used to test it. The model parameters are updated in combination with the gradient until the loss function value calculated by the test is less than the preset loss threshold. The preset initial model is then generated as a preset detection model, ensuring the detection accuracy of the preset detection model, thereby making the recognition of the bill content more accurate.

[0119] In addition, the present invention provides a bill content recognition system based on vision, please refer to Figure 3 , Figure 3 This is a module diagram of the first embodiment of the vision-based bill content recognition system of the present invention.

[0120] Specifically, the vision-based bill content recognition system in this embodiment includes a first acquisition module 10, a detection module 20, and a recognition module 30, wherein:

[0121] The first acquisition module 10 is used to acquire a bill image of a bill to be identified, and pre-process the bill image to obtain a pre-processed image;

[0122] The vision-based bill content recognition method of this embodiment is applied to a vision-based bill content recognition system. When the bill content recognition is required, a recognition request is initiated to the system based on the image of the bill to be recognized. After receiving the recognition request, the system takes the bill to be recognized as the bill to be recognized, and obtains the image carried in the recognition request as the bill image of the bill to be recognized. Then, the image to be recognized is preprocessed to obtain a preprocessed image. Preprocessing includes but is not limited to image denoising, binarization, edge detection and other operations to improve image quality. Image denoising is used to reduce or eliminate noise in the bill image while retaining the details and features of the image as much as possible; the denoising method can be any one of Gaussian filtering, median filtering, and wavelet transform. The binarization of the image is to convert the image into a pixel value containing only black and white pixels, thereby simplifying the image data and reducing the complexity of subsequent processing. Specifically, the binarization can be performed based on a threshold. The edge detection of the image is to identify the contour and boundary of the object in the image, and the edge detection can be performed by the Canny operator. In this embodiment, after the bill image is denoised, binarized and edge detected, the preprocessed image is obtained.

[0123] A detection module 20, configured to determine a text region in the preprocessed image based on a preset detection model, perform line text detection on the text region, and extract text content of the line text in the text region;

[0124] Furthermore, a preset detection model is formed by training a large number of image samples in advance, and the pre-processed image is identified by the preset detection model to determine the text area containing the text. The identification of the text area is mainly achieved by four-corner point positioning, and the position of the text area is determined by identifying the positions of the four corners of the text area in the pre-processed image. Specifically, the detection module 20 includes:

[0125] A detection unit 21, configured to detect the four corner positioning points of the text contained in the preprocessed image based on a preset detection model, and construct a cutting template according to the four corner positioning points;

[0126] A preprocessing unit 22, configured to determine an initial text region in the preprocessed image according to the cutting template and the preset template, and pre-segment the initial text region to obtain a pre-segmented region;

[0127] The determination unit 23 is used to identify the pre-segmented area and determine whether the text in the pre-segmented area is complete. If the text in the pre-segmented area is complete, the pre-segmented area is determined as the text area in the pre-processed image.

[0128] Furthermore, a two-dimensional coordinate system of the preprocessed image is established by a preset detection model, for example, the upper left corner of the preprocessed image is used as the coordinate origin, the horizontal rightward direction of the preprocessed image is used as the X coordinate axis, and the vertical downward direction of the preprocessed image is used as the Y coordinate axis. Then, the coordinate point positions of each character contained in the preprocessed image in the two-dimensional coordinate system are determined, and the four diagonal coordinate points with the farthest distance are determined by the positions of each coordinate point. The characters corresponding to the four diagonal coordinate points are the four corner positioning points of the characters contained in the preprocessed image.

[0129] Furthermore, a preset template for representing a text area is pre-set for the preprocessed image. The quadrilateral formed by the four corner positioning points is constructed as a cutting template, and the cutting template is compared with the preset template, and the initial text area in the preprocessed image is determined by the comparison. Among them, the comparison is to determine whether the length, width and area between the cutting template and the preset template are all matched, and the match means that the length, width and area of ​​the cutting template are all greater than the length, width and area of ​​the preset template within a certain range. If the comparison determines that they all match, it means that the cutting area corresponding to the cutting template matches the text area reflected by the preset template, so that the cutting area corresponding to the cutting template is determined as the initial text area in the pre-segmented area, and the initial text area is pre-segmented to obtain the corresponding pre-segmented area.

[0130] Furthermore, considering the matching of the cutting template and the preset template, there may be only a certain area in the preprocessed image that matches the text area reflected by the preset template in size, and the cutting area of ​​the cutting template does not accurately divide the text area. In this regard, the present embodiment identifies the pre-segmented area to identify whether the text contained therein is complete. If it is complete, it means that the division of the pre-segmented area is accurate, and the pre-segmented area is determined as the text area in the preprocessed image. On the contrary, if it is identified that the text contained in the pre-segmented area is incomplete, it means that the division of the pre-segmented area is inaccurate. At this time, the pre-segmented area is adjusted to contain complete text, and the adjusted pre-segmented area is determined as the text area.

[0131] Furthermore, the text area may contain multiple lines of text. In order to recognize more accurately, the text area is detected by line text detection, the text area is recognized and divided into multiple lines of text, and the text content contained in each line of text is extracted. The text content extraction is the operation of distinguishing each line of text from the text area.

[0132] The recognition module 30 is used to perform text recognition on the text content based on a preset recognition model to obtain key invoice information, and to perform format conversion on the key invoice information to obtain structured invoice information for storage.

[0133] Furthermore, a preset recognition model is formed in advance through training, and text recognition is performed on each extracted text content according to the preset recognition model, that is, the specific content corresponding to each line of text is recognized, and the key invoice information is obtained through recognition. The key invoice information includes but is not limited to the invoice date, invoice number, invoice amount, etc. Specifically, the recognition module 30 includes:

[0134] The word segmentation unit 31 is used to perform word segmentation processing on the text content based on a preset recognition model to obtain multiple words, and perform the following steps for each of the words:

[0135] A calculation unit 32, configured to respectively calculate similarity between a first word vector corresponding to the word and a second word vector corresponding to a preset invoice keyword, and determine a target keyword corresponding to the word in the preset invoice keyword;

[0136] The acquisition unit 33 is used to acquire key information corresponding to the target keyword in the line text content, and generate the target keyword and the key information together as the invoice key information.

[0137] Furthermore, the preset recognition model performs word segmentation processing on the text content to obtain multiple words. Then, for each divided word, the corresponding word vector is searched from the preset dictionary as the first word vector. At the same time, preset invoice keywords and their corresponding second word vectors are pre-set according to the key invoice information. The first word vector and each second word vector are respectively calculated for similarity, and the similarity between the word and each preset invoice keyword is determined by calculation. The similarity calculation formula can be found in the above formula (1), which will not be repeated here.

[0138] Furthermore, the calculated similarities are compared to determine the maximum value, i.e., the value closest to 1. Then, the preset invoice keyword that generates the maximum value is searched as the target keyword corresponding to the word. The key information corresponding to the target keyword is identified from the text content. For example, if the target keyword is the invoice number, the specific invoice number is identified from the text content as the key information. Moreover, the key information is usually located in the same line of text as the target keyword, so the key information can be identified from the same line of text. Then, the target keyword and its corresponding key information are generated as invoice key information together, for example, the two are formed into invoice key information in the form of a key-value pair.

[0139] Furthermore, in order to ensure the consistency of the format of the key invoice information obtained from various forms of invoices, a structured format is pre-set. After the key invoice information is obtained, the key invoice information is formatted and converted into a unified structured format, and the structured invoice information is obtained and stored for subsequent statistical analysis and other processing.

[0140] In order to ensure the accuracy of the key information of the generated invoice, this embodiment also provides a verification mechanism.

[0141] Specifically, the bill content recognition system further includes:

[0142] The screening module b1 is used to randomly screen out target key information from each of the invoice key information and identify the information type of the target key information;

[0143] Verification module b2, used to verify the target key information according to the preset information format corresponding to the information type, and determine whether the target key information passes the verification

[0144] The identification module 30 is also used to execute the step of format conversion of the invoice key information if the target key information passes the verification;

[0145] If the target key information fails to pass the verification, feedback information is generated, and based on the feedback information, a step of performing word segmentation processing on the text content based on a preset recognition model is executed;

[0146] Furthermore, from the invoice key information obtained by identifying the various text contents, any one item is randomly selected as the target key information, and the information type of the target key information is identified. The information type includes at least the invoice number, the invoice amount, and the invoicing time. Different information formats are pre-set for different information types. For the information type of the target key information, the corresponding preset information format is searched, and the target key information is verified according to the preset information format. It is determined whether the information format of the target key information is consistent with the preset information format. If they are consistent, a verification result of passing the verification is generated, and if they are inconsistent, a verification result of failing the verification is generated.

[0147] Furthermore, after the verification result is generated, it is determined whether the target key information has passed the verification based on the verification result. If the target key information has passed the verification, it means that the generated invoice key information is accurate, and the invoice key information is formatted and converted into a unified structured format for storage. On the contrary, if the target key information has not passed the verification, it means that the generated invoice key information is inaccurate. At this time, feedback information is generated, and the text content is re-segmented by the preset recognition model based on the feedback information. The feedback information contains incorrect target key information, and the preset recognition model combines it to re-segment the text content, which can correct the incorrect target key information and obtain accurate target key information.

[0148] In the vision-based bill content recognition system implemented in this embodiment, after the first acquisition module obtains the bill image of the bill to be recognized, the bill image is preprocessed to obtain the preprocessed image; then the detection module determines the text area in the preprocessed image according to the pre-trained preset detection model, and performs line text detection on the text area to extract the text content of the line text in the text area; in addition, a preset recognition model is also pre-trained, and the recognition module performs text recognition on the extracted line text content through the preset recognition model to obtain the key information of the invoice, and then performs format conversion on the key information of the invoice to obtain structured invoice information for storage. In this way, the text area where various types of text in the invoice are located is accurately detected by the pre-trained preset detection model and the preset recognition model to extract the text content, and the extracted text content is accurately recognized, thereby improving the accuracy of obtaining the key information of the invoice. At the same time, by performing format conversion on the obtained key information of the invoice, the resultant text information is generated for storage, so that each extracted key information of the invoice has a unified structured format. Accurate recognition of invoices and consistency of the recognition result format are achieved.

[0149] For further information, please refer to Figure 4 Based on the first embodiment of the vision-based bill content recognition system of the present invention, a second embodiment of the vision-based bill content recognition system of the present invention is proposed.

[0150] The difference between the second embodiment of the vision-based bill content recognition system and the first embodiment of the vision-based bill content recognition system is that the vision-based bill content recognition system further includes:

[0151] A second acquisition module 40 is used to acquire a large number of image samples and divide the large number of image samples into training samples and test samples;

[0152] A training module 50 is used to train a preset initial model based on the training sample, and when the training time reaches a preset time, test the preset initial model based on the test sample to obtain a test result;

[0153] A calculation module 60, configured to calculate a loss function value of the preset initial model based on the test result;

[0154] A judgment module 70, used to compare the loss function value with a preset loss threshold, to judge whether the loss function value is less than the preset loss threshold, and if it is less than the preset loss threshold, to generate the preset initial model as a preset detection model;

[0155] An updating module 80, configured to update the model parameters of the preset initial model based on a preset updating formula if the loss function value is greater than a preset loss threshold;

[0156] The training module 50 is further used to execute the step of training the preset initial model based on the training sample for the updated preset initial model until the loss function value is less than the preset loss threshold.

[0157] Furthermore, the present embodiment forms a preset detection model by training a large number of image samples. Specifically, a large number of image samples are obtained, and each image sample is divided into training samples and test samples according to a preset ratio, such as 8 to 2, 7 to 3, etc., with more training samples than test samples. Then, the preset initial model is trained based on the training samples, and a preset duration is preset. When the statistical training duration reaches the preset duration, the preset initial model is tested by the test sample. The preset initial model performs detection and analysis on the test sample to generate a corresponding test result.

[0158] Furthermore, in order to reflect the performance of the preset initial model, the loss function value of the preset initial model is calculated based on the test results. The loss function value represents the difference between the test result obtained by the preset initial model for detecting and analyzing the test sample and the reference result corresponding to the test sample. The specific calculation formula can be found in the above formula (2), which will not be repeated here.

[0159] Among them, the larger the loss function value, the larger the difference between the test result and the reference result, the poorer the performance of the preset initial model, and the need to continue iterative training of the preset initial model. Conversely, the smaller the loss function value, the smaller the difference between the test result and the reference result, the better the performance of the preset initial model, and the training of the preset initial model can be stopped. In order to indicate the size of the loss function value, a preset loss threshold is pre-set, and the loss function value is compared with the preset loss threshold to determine whether the loss function value is less than the preset loss threshold. If it is less than, it means that the preset initial model has a good detection and analysis performance for the test sample. At this time, the training of the preset initial model is stopped and it is generated as the preset detection model. Conversely, if the loss function value is determined to be greater than the preset loss threshold through comparison, it means that the performance of the preset initial model is poor, and the preset initial model needs to be iteratively trained. A preset update formula for updating the model parameters is pre-set, and the model parameters of the preset initial model are updated by the preset update formula. After the update, the preset initial model is iteratively trained, and the loss function value is re-tested and calculated until the loss function value obtained by the test is less than the preset loss threshold. The preset update formula can be specifically referred to in the above formula (3), which will not be elaborated here.

[0160] In this embodiment, a large number of image samples are used to train the preset initial model, and the test samples are used to test it. The model parameters are updated in combination with the gradient until the loss function value calculated by the test is less than the preset loss threshold. The preset initial model is then generated as a preset detection model, ensuring the detection accuracy of the preset detection model, thereby making the recognition of the bill content more accurate.

[0161] The embodiments of the present invention are described above in conjunction with the accompanying drawings, but the present invention is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the enlightenment of the present invention, ordinary technicians in this field can also make many forms without departing from the scope of protection of the present invention and the claims. All equivalent structures or equivalent process changes made using the contents of the specification and drawings of the present invention, or directly or indirectly used in other related technical fields, are protected by the present invention.

Claims

1. A method for identifying bill content based on vision, characterized in that: The bill content identification method comprises: Acquire a bill image of a bill to be identified, and preprocess the bill image to obtain a preprocessed image; Determine a text area in the preprocessed image based on a preset detection model, perform line text detection on the text area, and extract text content of the line text in the text area; The text content is recognized based on a preset recognition model to obtain key invoice information, and the key invoice information is formatted to obtain structured invoice information for storage.

2. The method for identifying bill contents according to claim 1, characterized in that: The step of performing text recognition on the line text content based on a preset recognition model to obtain key invoice information includes: The text content is segmented based on a preset recognition model to obtain a plurality of words, and the following steps are performed for each of the words: The first word vector corresponding to the word and the second word vector corresponding to the preset invoice keyword are respectively calculated for similarity, and the target keyword corresponding to the word in the preset invoice keyword is determined, wherein the formula for calculating the similarity is: Where S represents similarity, n represents the larger value between the dimension of the first word vector and the dimension of the second word vector, ai represents the first word vector, and bi represents the second word vector; The key information corresponding to the target keyword in the line text content is obtained, and the target keyword and the key information are generated together as the invoice key information.

3. The method for identifying bill contents as claimed in claim 2, characterized in that: The step of performing text recognition on the text content based on a preset recognition model to obtain key invoice information includes: Randomly selecting target key information from each of the invoice key information, and identifying the information type of the target key information; Verify the target key information according to the preset information format corresponding to the information type, and determine whether the target key information passes the verification If the target key information passes the verification, the step of converting the format of the invoice key information is executed; If the target key information fails to pass the verification, feedback information is generated, and based on the feedback information, a step of performing word segmentation processing on the text content based on a preset recognition model is executed.

4. The method for identifying bill contents according to any one of claims 1 to 3, characterized in that: The step of determining the text area in the preprocessed image based on a preset detection model comprises: Detecting the four corner positioning points of the text contained in the preprocessed image based on a preset detection model, and constructing a cutting template according to the four corner positioning points; Determine an initial text area in the preprocessed image according to the cutting template and the preset template, and pre-segment the initial text area to obtain a pre-segmented area; The pre-segmented area is identified to determine whether the text in the pre-segmented area is complete; if the text in the pre-segmented area is complete, the pre-segmented area is determined as a text area in the pre-processed image.

5. The method for identifying bill contents according to any one of claims 1 to 3, characterized in that: The step of detecting the four corner positioning points of the text contained in the preprocessed image based on the preset detection model includes: Acquire a large number of image samples, and divide the large number of image samples into training samples and test samples; Training a preset initial model based on the training samples, and when the training time reaches a preset time, testing the preset initial model based on the test samples to obtain a test result; The loss function value of the preset initial model is calculated based on the test results, and the calculation formula is: Where L represents the loss function value, m represents the number of test samples, pj represents the test result corresponding to the j-th test sample, and qj represents the reference result corresponding to the j-th test sample; Compare the loss function value with a preset loss threshold to determine whether the loss function value is less than the preset loss threshold, and if so, generate the preset initial model as a preset detection model; If the loss function value is greater than the preset loss threshold, the model parameters of the preset initial model are updated based on a preset update formula, and the preset update formula is: Among them, θ t+1 represents the updated model parameters, θ t represents the model parameters before updating, is the learning rate of the preset initial model, It is the gradient of the loss function of the preset initial model before the model parameters are updated; For the updated preset initial model, a step of training the preset initial model based on the training sample is performed until the loss function value is less than a preset loss threshold.

6. A vision-based bill content recognition system, characterized in that: The bill content recognition system comprises: A first acquisition module is used to acquire a bill image of a bill to be identified, and preprocess the bill image to obtain a preprocessed image; A detection module, used to determine a text area in the preprocessed image based on a preset detection model, perform line text detection on the text area, and extract text content of the line text in the text area; The recognition module is used to perform text recognition on the text content based on a preset recognition model to obtain key invoice information, and to perform format conversion on the key invoice information to obtain structured invoice information for storage.

7. The visual-based bill content recognition system according to claim 6, characterized in that: The identification module also includes: The word segmentation unit is used to perform word segmentation processing on the text content based on a preset recognition model to obtain multiple words, and perform the following steps for each of the words: A calculation unit is used to perform similarity calculation on the first word vector corresponding to the word and the second word vector corresponding to the preset invoice keyword, respectively, to determine the target keyword corresponding to the word in the preset invoice keyword, wherein the formula for similarity calculation is: Where S represents similarity, n represents the larger value between the dimension of the first word vector and the dimension of the second word vector, ai represents the first word vector, and bi represents the second word vector; The acquisition unit is used to acquire key information corresponding to the target keyword in the line text content, and generate the target keyword and the key information together as the invoice key information.

8. The visual-based bill content recognition system according to claim 7, characterized in that: The bill content recognition system also includes: A screening module, used for randomly screening out target key information from each of the invoice key information, and identifying the information type of the target key information; A verification module is used to verify the target key information according to a preset information format corresponding to the information type, and determine whether the target key information passes the verification. The identification module is also used to execute the step of format conversion of the invoice key information if the target key information passes the verification; If the target key information fails to pass the verification, feedback information is generated, and based on the feedback information, a step of performing word segmentation processing on the text content based on a preset recognition model is executed.

9. The visual-based bill content recognition system according to any one of claims 6 to 8, characterized in that: The detection module comprises: A detection unit, configured to detect the four corner positioning points of the text contained in the preprocessed image based on a preset detection model, and construct a cutting template according to the four corner positioning points; A preprocessing unit, configured to determine an initial text region in the preprocessed image according to the cutting template and a preset template, and pre-segment the initial text region to obtain a pre-segmented region; The determination unit is used to identify the pre-segmented area and determine whether the text in the pre-segmented area is complete. If the text in the pre-segmented area is complete, the pre-segmented area is determined as the text area in the pre-processed image.

10. The vision-based bill content recognition system according to any one of claims 6 to 8, characterized in that: The vision-based bill content recognition system also includes: A second acquisition module is used to acquire a large number of image samples and divide the large number of image samples into training samples and test samples; A training module, used to train a preset initial model based on the training samples, and when the training time reaches a preset time, test the preset initial model based on the test samples to obtain a test result; The calculation module is used to calculate the loss function value of the preset initial model based on the test result, and the calculation formula is: Where L represents the loss function value, m represents the number of test samples, pj represents the test result corresponding to the j-th test sample, and qj represents the reference result corresponding to the j-th test sample; A judgment module, used for comparing the loss function value with a preset loss threshold, judging whether the loss function value is less than the preset loss threshold, and if it is less than the preset loss threshold, generating the preset initial model as a preset detection model; An updating module is used to update the model parameters of the preset initial model based on a preset updating formula if the loss function value is greater than a preset loss threshold, and the preset updating formula is: Among them, θ t+1 represents the updated model parameters, θ t represents the model parameters before updating, is the learning rate of the preset initial model, It is the gradient of the loss function of the preset initial model before the model parameters are updated; The training module is also used to execute the step of training the preset initial model based on the training sample for the updated preset initial model until the loss function value is less than a preset loss threshold.