Invoice content authenticity identification method and storage medium

Through deep learning technology and multi-step processing methods, the area, element location and text content of the invoice are identified, and combined with QR code analysis and handwritten data detection, the problem of insufficient ability to identify invoices in the existing technology is solved, and efficient identification of massive invoices is achieved.

CN120014660AInactive Publication Date: 2025-05-16SENSCAPE TECH BEIJING CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510490147.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-05-16
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the prior art, the ability to identify the authenticity of invoices is low, and it is difficult to effectively identify the authenticity of invoice content.

Method used

Through a method that includes multiple steps, including image acquisition, detection segmentation, image correction, field recognition, feature recognition, QR code analysis and invoice judgment steps, deep learning technology is used to judge the area of ​​the invoice, detect the location and text content of the identified elements, and judge the authenticity of the invoice through QR code analysis and handwritten data detection.

Benefits of technology

It realizes the authenticity of massive invoice content, improves the ability to identify the authenticity of invoices, and effectively solves the problem of insufficient identification capabilities in the existing technology.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014660A_ABST
    Figure CN120014660A_ABST
Patent Text Reader

Abstract

The invention provides an invoice content authenticity identification method and a storage medium, and the method comprises an image acquisition step, a detection and segmentation step, an image correction step, a field identification step, an element identification step, a two-dimensional code analysis step and an invoice judgment step, through a deep learning method, an invoice region is judged, the position of an identification element is detected, and the authenticity of the invoice content is identified. The method comprises the following steps of: obtaining an element result of an invoice through the character content of the element, detecting the position of a two-dimensional code, identifying the content of the two-dimensional code to obtain two-dimensional code analysis data, detecting handwritten data, identifying the element content of the invoice, matching the element content of the invoice with the two-dimensional code data, and judging the authenticity of the invoice according to the overlapping ratio of the element identification result and the handwritten area. And true and false identification of massive invoice contents is realized. The method has the advantage that the invoice authenticity identification capability of the deep learning method is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to a method for identifying the authenticity of invoice content and a storage medium. Background Art

[0002] In the process of invoice reimbursement recognition, invoices with falsified contents often appear. The main causes are Photoshop editing of invoice images, overlaying data such as the name, specification, model, and unit of goods from other invoices onto the image to be recognized, and modifying the date of the invoice. For example, invoices submitted several years ago are modified with the invoice date and reused for reimbursement, or key elements of the invoice are modified by hand, such as modifying the price and tax amount, total amount and other information of the invoice.

[0003] With the deepening of digital transformation, automated invoice processing has become a key need for many companies. Behind the development of invoice recognition technology, there are many technical means and methods. First of all, optical character recognition (OCR) technology is the basis of invoice recognition. OCR technology converts the text in the image into an editable text format by scanning and recognizing the characters on the invoice. However, OCR technology still has some challenges. For example, the recognition accuracy may be reduced when processing handwritten text, blurred images or invoices with complex typesetting. Secondly, machine learning and deep learning technologies play an important role in invoice recognition. By training models to recognize different types of invoices and learning invoice layout and structure, machine learning algorithms can help improve recognition accuracy. However, current machine learning models and optical character recognition methods have certain problems in identifying the authenticity of invoices. Summary of the invention

[0004] The present application provides a method and storage medium for identifying the authenticity of invoice contents, so as to solve the problem of low ability to identify the authenticity of invoices in the prior art.

[0005] The present application provides a method for identifying the authenticity of invoice contents, which specifically includes an image acquisition step, a detection and segmentation step, an image correction step, a field recognition step, an element recognition step, a QR code parsing step, and an invoice judgment step.

[0006] The image acquisition step is used to obtain a first invoice image in a target format, wherein the first invoice image includes image information and text information, and the image information includes an invoice image and a QR code image; the detection and segmentation step is to input the image information into a first model, obtain the first position information and first segmentation information of the invoice image and the second position information and second segmentation information of the QR code image; the image correction step is to crop the invoice image and the QR code image based on the obtained first position information and first segmentation information and the second position information and second segmentation information, and perform rotation correction processing to obtain a corrected second invoice image; the field recognition step is to convert the second invoice image into a second model, obtain the The field position and field category information in the second invoice image; the element recognition step is to input the second invoice image into the third model, obtain the text recognition result of the second invoice image, and obtain the element recognition result by corresponding the text recognition result and the preset element type, and record the position of each element at the same time; the QR code parsing step is based on the second position information and the second segmentation information of the QR code in the invoice image, and parses the QR code content through the pyzbar library to obtain the QR code recognition result; the invoice judgment step is to obtain the invoice recognition result by corresponding the text recognition result with the field category information, and judge whether the invoice in the invoice image is a real invoice based on the invoice recognition result and the QR code recognition result.

[0007] Furthermore, the detection and segmentation step specifically includes a first model building step, a data labeling step, a data processing step, a model training step and a model reasoning step.

[0008] The first model building step is used to build a maskRcnn target detection segmentation model implemented by the pytorch framework; the data labeling step is used to label the position points of the invoice box in the invoice image and the position points of the segmented area of ​​the invoice; the position points of the QR code box in the invoice image and the position points of the segmented area of ​​the QR code are labeled to obtain labeled data; the data processing step is to preprocess the labeled data to obtain preprocessed data, and divide the preprocessed data into a training set, a validation set and a test set; the model training step is used to set different learning rates lr, batch sizes batch_size and epoch numbers, use the training set to train the first model, and verify it through the validation set; the model inference step is to use the test set to infer the first model, obtain the optimal mask threshold and the optimal position threshold, and obtain the trained first model.

[0009] Furthermore, the data processing step specifically includes a data enhancement step, a mask matrix acquisition step and a normalization step.

[0010] The data enhancement step is used to count the distribution of the labeled data and perform data enhancement processing on the data whose amount is less than 1 / 2 of the maximum data category; the mask matrix acquisition step is to obtain the mask matrix by annotating the segmented area in the labeled data, and its formula is: ; Among them, segment represents the marked segmentation area; i xy Represents the corresponding element in the mask matrix; The normalization step is to normalize the frame of the marked area in the marked data, and its formula is: ; Among them, box represents the box, b x , b y , b w , b h Indicates the location information of the box, m w , m h Indicates the width and height of the box.

[0011] Furthermore, the model reasoning step includes a first intersection-over-union calculation step, a first intersection-over-union judgment step, a first accuracy acquisition step, a first recall acquisition step, a first indicator acquisition step, a first threshold judgment step and a first optimal threshold selection step.

[0012] The first intersection-and-union ratio calculation step is to input the test set into the first model, obtain at least one predicted mask matrix, and obtain a first intersection-and-union ratio between the labeled mask matrix and the predicted mask matrix. The calculation formula of the first intersection-and-union ratio is: ; Among them, IOU1 represents the first intersection-over-union ratio, mask p Represents the predicted mask matrix, mask r Represents the label mask matrix; The first intersection-and-union ratio judgment step is used to set an initial threshold set, which includes more than two initial thresholds, and to determine whether there is a first intersection-and-union ratio greater than the minimum initial threshold in the initial threshold set. If so, execute the next step; the first accuracy acquisition step is to obtain the accuracy of the predicted segmentation area based on the number of obtained prediction mask matrices, and its formula is: ; Wherein, pre1 represents the accuracy of the predicted mask matrix, TP1 represents the number of correct predictions of the predicted mask matrix, and P1 represents the number of predicted mask matrices; the first recall rate acquisition step is to obtain the recall rate of the predicted mask matrix based on the obtained accuracy of the predicted mask matrix, and its formula is: ; Wherein, recall1 represents the recall rate of the predicted mask matrix, and R1 represents the number of labeled mask matrices; the first indicator acquisition step is to obtain a first position indicator based on the obtained recall rate of the predicted segmented area, and the calculation formula of the first position indicator is: ; Among them, F1 represents the first position indicator; the first threshold judgment step is used to determine whether the initial threshold corresponding to the first position indicator is the maximum initial threshold in the initial threshold set. If not, take the initial threshold next to the initial threshold and repeat the first intersection-and-union calculation step, the first intersection-and-union judgment step, the first accuracy acquisition step, the first recall acquisition step and the first indicator acquisition step, until the initial threshold corresponding to the first position indicator is the maximum initial threshold in the initial threshold set, and a first position indicator set is obtained; the first optimal threshold selection step is based on the obtained first position indicator set, and the initial threshold corresponding to the largest first position indicator in the first position indicator set is selected as the optimal mask threshold of the first model.

[0013] Furthermore, the model reasoning step also includes a second intersection-over-union calculation step, a second intersection-over-union judgment step, a second accuracy acquisition step, a second recall acquisition step, a second indicator acquisition step, a second threshold judgment step and a second optimal threshold selection step.

[0014] The second IoU calculation step is to input the test set into the first model, obtain at least one prediction box, and obtain the second IoU ratio of the marked box and the prediction box. The calculation formula of the second IoU ratio is: ; Among them, IOU2 second intersection and union ratio, box p Represents the prediction box, box r The second intersection-and-union ratio judgment step is used to set an initial threshold set, which includes more than two initial thresholds, and to determine whether there is a second intersection-and-union ratio greater than the minimum initial threshold in the initial threshold set. If so, execute the next step; the second accuracy acquisition step is to obtain the accuracy of the predicted position point based on the number of predicted frames obtained, and its formula is: ; Among them, pre2 represents the accuracy of the prediction box, TP2 represents the number of correct predictions of the prediction box, and P2 represents the number of prediction boxes; the second recall rate acquisition step is to obtain the recall rate of the prediction box based on the accuracy of the obtained prediction box, and its formula is: ; Wherein, recall2 represents the recall rate of the predicted box, and R2 represents the number of annotated boxes; the second indicator acquisition step is to obtain a second position indicator based on the obtained recall rate of the predicted position, and the calculation formula of the second position indicator is: ; Among them, F2 represents the second position indicator; the second threshold judgment step is used to determine whether the initial threshold corresponding to the second position indicator is the maximum initial threshold in the initial threshold set. If not, take the initial threshold next to the initial threshold and repeat the second intersection and union calculation step, the second intersection and union judgment step, the second accuracy acquisition step, the second recall acquisition step and the second indicator acquisition step until the initial threshold corresponding to the second position indicator is the maximum initial threshold in the initial threshold set, and a second position indicator set is obtained; the second optimal threshold selection step is based on the obtained second position indicator set, and the initial threshold corresponding to the largest second position indicator in the second position indicator set is selected as the optimal position threshold of the first model.

[0015] Furthermore, the image correction step includes an invoice cropping step, an invoice rotating step, a contour acquiring step, a contour mapping step and an invoice correcting step.

[0016] The invoice cropping step is to crop the invoice image based on the first position information of the invoice image and the data category of the invoice image using the Crop function of opencv; the invoice rotation step is to rotate the invoice image based on the first position information of the invoice image and the data category of the invoice image using the getRotationMatrix2D method; the contour acquisition step is to transform the mask matrix between 0 and 1 according to the optimal mask threshold, and use the findContours function to traverse the contour points of the mask matrix to obtain the contour of the mask matrix; the contour mapping step is to map the contour of the mask matrix to the positions of the four corners of the invoice image to obtain the mapping mask matrix; the invoice correction step is to encapsulate the mask matrix and the mapping mask matrix using the vstack function in opencv, and perform target transformation on the encapsulated mask matrix and the mapping mask matrix through the warpAffine function in opencv to obtain the corrected invoice image.

[0017] Furthermore, the image correction step also includes a QR code clipping step, a rotation angle acquisition step and a QR code correction step.

[0018] The two-dimensional code cropping step is based on the second position information of the two-dimensional code image, and the two-dimensional code image is cropped using the Crop function of opencv, and its formula is: img qr =crop(img, box qr ) Among them, img qr Indicates the QR code image area, box qr represents a two-dimensional code frame; the rotation angle acquisition step is to obtain the position of the contour point of the two-dimensional code image based on the second segmentation information of the two-dimensional code image, and then calculate the rotation angle of the two-dimensional code image, and its formula is: p i =(p ix , p iy ), i=1,2,3,4; ; Among them, p i Represents the contour points of the QR code image, (p ix , p iy ) represents the position of the contour point and represents the rotation angle of the two-dimensional code image; the two-dimensional code correction step is based on the obtained two-dimensional code rotation angle, and uses the getRotationMatrix2D method to rotate the two-dimensional code image to obtain a corrected two-dimensional code image.

[0019] Furthermore, the invoice judgment step includes a first judgment step, a second judgment step, a third judgment step and a fourth judgment step.

[0020] The first judgment step is used to judge whether the QR code recognition result is consistent with the invoice recognition result. If not, the invoice is judged to be a fake invoice. If so, the next step is executed; the second judgment step is used to judge whether there is a handwritten field in the identified field. If so, the third judgment step is executed. If not, the fourth judgment step is executed; the third judgment step is used to judge whether the overlap rate between the handwritten field position and the invoice element position is less than a preset threshold. If not, the invoice is judged to be a fake invoice. If so, the next step is executed; the fourth judgment step is used to judge whether the handwritten field belongs to the key field category. If not, the invoice is judged to be a real invoice.

[0021] The present application also provides a storage medium storing computer-readable instructions. When the computer-readable instructions are read by at least one processor, the at least one processor executes at least one step of the method for identifying the authenticity of invoice contents.

[0022] The present application provides a method and storage medium for identifying the authenticity of invoice content. Through a deep learning method, the area of ​​the invoice is judged, the position of the identification element is detected, and the element result of the invoice is obtained by the text content of the element. The QR code position is detected and the QR code content is identified to obtain QR code parsing data, and handwriting data is detected. The authenticity of the invoice is judged by matching the invoice element content recognition with the QR code data and the overlap rate between the element recognition result and the handwriting area, thereby realizing the authenticity identification of massive invoice contents, and effectively improving the ability of the deep learning method to identify the authenticity of the invoice. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.

[0024] Figure 1 It is a flow chart of the method for identifying the authenticity of invoice contents described in the embodiment of the present application; Figure 2 is a flow chart of the detection and segmentation steps described in an embodiment of the present application; Figure 3 is a flow chart of the data processing steps described in the embodiment of the present application; Figure 4 This is the model reasoning step process described in the embodiment of this application Figure 1 ; Figure 5 This is the model reasoning step process described in the embodiment of this application Figure 2 ; Figure 6 The image correction process described in the embodiment of the present application is Figure 1 ; Figure 7 The image correction process described in the embodiment of the present application is Figure 2 ; Figure 8 is a flowchart of the invoice judgment steps described in the embodiment of the present application; Fig. 9 It is a schematic diagram of the storage medium described in an embodiment of the present application.

[0025] Description of reference numerals: 100 storage medium, 110 processor, 120 memory. DETAILED DESCRIPTION

[0026] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of this application.

[0027] like Figure 1 As shown, the present application provides a method for identifying the authenticity of invoice content, which specifically includes step S1) image acquisition step, step S2) detection and segmentation step, step S3) image correction step, step S4) field recognition step, step S5) element recognition step, step S6) QR code parsing step and step S7) invoice judgment step.

[0028] Step S1) An image acquisition step is to acquire a first invoice image in a target format, wherein the first invoice image includes image information and text information, and the image information includes an invoice image and a QR code image.

[0029] In this embodiment, the data formats of the first invoice image include jpeg, jpg, png, pdf, and OFD, wherein the first invoice image in the jpeg, jpg, png, pdf, and OFD formats contains image data, i.e., the image information, and the first invoice image in the pdf and OFD formats contains text data, i.e., the text information. When the text information in the image in the pdf and OFD formats cannot be extracted, the image information in the first invoice image in the pdf and OFD formats is extracted.

[0030] Step S2) Detection and segmentation step: input the image information into a first model to obtain the first position information and first segmentation information of the invoice image and the second position information and second segmentation information of the two-dimensional code image.

[0031] In this embodiment, since the invoice and the QR code are both quadrilaterals, the first position information is the position information of the four corners of the invoice, the first segmentation information is the segmentation information of the four corners of the invoice, the second position information is the position information of the four corners of the QR code, and the second segmentation information is the segmentation information of the four corners of the QR code.

[0032] like Figure 2 As shown, step S2) the detection and segmentation step specifically includes step S21) a first model building step, step S22) a data labeling step, step S23) a data processing step, step S24) a model training step and step S25) a model inference step.

[0033] Step S21) The first model building step is to build a maskRcnn target detection and segmentation model implemented by the pytorch framework.

[0034] Step S22) Data labeling step, labeling the position points of the invoice frame in the invoice image and the position points of the segmented area of ​​the invoice; labeling the position points of the QR code frame in the invoice image and the position points of the segmented area of ​​the QR code to obtain labeled data.

[0035] In this embodiment, the position of the invoice frame in the invoice image is marked, including the upper left corner in_x1 and in_y1 of the invoice, the width and height in_w and in_h of the invoice, and the position points of the segmented areas of the invoice are inp_x1, inp_y1, inp_x2, inp_y2....inp_xn, inp_yn; the position of the QR code frame in the invoice image is marked, including the upper left corner qr_x1 and qr_y1 of the QR code, the width and height qr_w and qr_h of the invoice, and the position points of the segmented areas of the QR code are qrp_x1, qrp_y1, qrp_x2, qrp_y2....qrp_xn, qrp_yn.

[0036] Step S23) Data processing step, preprocessing the labeled data to obtain preprocessed data, dividing the preprocessed data into a training set, a validation set and a test set, using 3 / 4 of the preprocessed data for the training set, 1 / 8 of the preprocessed data for the validation set, and 1 / 8 of the preprocessed data for the test set.

[0037] like Figure 3 As shown, step S23) the data processing step specifically includes step S231) a data enhancement step, step S232) a mask matrix acquisition step and step S233) a normalization step.

[0038] Step S231) Data enhancement step, statistics the distribution of the labeled data, and perform data enhancement processing on the data whose amount is less than 1 / 2 of the maximum data category.

[0039] In this embodiment, the acquired marked data is processed, the distribution of statistical data is analyzed, and data with a data volume less than 1 / 2 of the maximum data category is enhanced, including changes in lighting, changes in grayscale images, rotation of invoice images, splicing and combination of non-QR code images, etc.

[0040] Step S232) A mask matrix acquisition step is performed to acquire a mask matrix by marking the segmented area in the labeled data. The formula is: ; Among them, segment represents the marked segmentation area; i xyRepresents the corresponding element in the mask matrix; Step S233) Normalization step, normalizing the frame of the marked area in the marked data, the formula is: ; Among them, box represents the box, b x , b y , b w , b h Indicates the location information of the box, m w , m h Indicates the width and height of the box.

[0041] Step S24) Model training step, setting different learning rates lr, batch sizes batch_size and epoch numbers, training the first model using the training set, and verifying it using the verification set.

[0042] Step S25) Model inference step, using the test set to infer the first model, obtain the optimal mask threshold and the optimal position threshold, and obtain the trained first model.

[0043] like Figure 4 As shown, step S25) the model reasoning step includes step S251) a first intersection-over-union calculation step, step S252) a first intersection-over-union judgment step, step S253) a first accuracy acquisition step, step S254) a first recall acquisition step, step S255) a first indicator acquisition step, step S256) a first threshold judgment step and step S257) a first optimal threshold selection step.

[0044] Step S251) A first intersection-and-union ratio calculation step is performed, wherein the test set is input into the first model, and at least one predicted mask matrix is ​​obtained, and a first intersection-and-union ratio between the labeled mask matrix and the predicted mask matrix is ​​obtained. The calculation formula of the first intersection-and-union ratio is: ; Among them, IOU1 represents the first intersection-over-union ratio, mask p Represents the predicted mask matrix, mask r Represents the labeled mask matrix, that is, the mask matrix obtained by labeling the data.

[0045] Step S252) A first intersection-and-union ratio judgment step is to set an initial threshold set, wherein the initial threshold set includes more than two initial thresholds, and to judge whether there is a first intersection-and-union ratio greater than the minimum initial threshold in the initial threshold set. If so, proceed to the next step.

[0046] Step S253) The first accuracy acquisition step is to obtain the accuracy of the predicted segmented area based on the obtained number of predicted mask matrices, and the formula is: ; Among them, pre1 represents the accuracy of the predicted mask matrix, TP1 represents the number of correct predictions of the predicted mask matrix, and P1 represents the number of predicted mask matrices.

[0047] Step S254) The first recall rate acquisition step is to obtain the recall rate of the predicted mask matrix based on the obtained accuracy of the predicted mask matrix, and the formula is: ; Among them, recall1 represents the recall rate of the predicted mask matrix, and R1 represents the number of labeled mask matrices.

[0048] Step S255) A first indicator acquisition step is to obtain a first position indicator based on the obtained recall rate of the predicted segmented area. The calculation formula of the first position indicator is: ; Among them, F1 represents the first position index; Step S256) a first threshold judgment step, judging whether the initial threshold corresponding to the first position indicator is the maximum initial threshold in the initial threshold set, if not, taking the initial threshold next to the initial threshold and repeating step S251) a first intersection-and-union ratio calculation step, step S252) a first intersection-and-union ratio judgment step, step S253) a first accuracy acquisition step, step S254) a first recall rate acquisition step, step S255) a first indicator acquisition step, until the initial threshold corresponding to the first position indicator is the maximum initial threshold in the initial threshold set, and a first position indicator set is obtained.

[0049] Step S257) A first optimal threshold selection step, based on the obtained first position indicator set, selecting an initial threshold corresponding to the largest first position indicator in the first position indicator set as the optimal mask threshold of the first model.

[0050] In this embodiment, the initial threshold set F is set b [0.5, 0.55, 0.6, 0.65...0.9], the minimum initial threshold is 0.5, the maximum initial threshold is 0.9, in step S252) the first intersection-over-union judgment step, according to F bThe initial thresholds are taken in order from small to large. When there is a first intersection-and-union ratio greater than an initial threshold, the prediction mask matrix corresponding to the first intersection-and-union ratio is a qualified matrix. At this time, the first position index corresponding to the prediction mask matrix is ​​calculated. Therefore, there is a first position index corresponding to each initial threshold, and a first position index set is obtained. The initial threshold corresponding to the largest first position index in the first position index set is selected as the optimal mask threshold of the first model.

[0051] like Figure 5 As shown, step S25) the model reasoning step also includes step S261) a second intersection-over-union calculation step, step S262) a second intersection-over-union judgment step, step S263) a second accuracy acquisition step, step S264) a second recall acquisition step, step S265) a second indicator acquisition step, step S266) a second threshold judgment step and step S267) a second optimal threshold selection step.

[0052] Step S261) A second IoU calculation step is to input the test set into the first model, obtain at least one predicted box, and obtain a second IoU ratio between the marked box and the predicted box. The calculation formula of the second IoU ratio is: ; Among them, IOU2 represents the second intersection-and-union ratio, box p Represents the prediction box, box r Represents the annotated box, that is, the box obtained by labeling the data.

[0053] Step S262) A second intersection-and-union ratio judgment step is to set an initial threshold set, wherein the initial threshold set includes more than two initial thresholds, and to judge whether there is a second intersection-and-union ratio greater than the minimum initial threshold in the initial threshold set. If so, proceed to the next step.

[0054] Step S263) The second accuracy acquisition step is to obtain the accuracy of the predicted position point based on the number of predicted boxes obtained, and the formula is: ; Among them, pre2 represents the accuracy of the prediction box, TP2 represents the number of correct predictions of the prediction box, and P2 represents the number of prediction boxes.

[0055] Step S264) The second recall rate acquisition step is to obtain the recall rate of the prediction frame based on the obtained accuracy of the prediction frame, and the formula is: ; Among them, recall2 represents the recall rate of the predicted box, and R2 represents the number of labeled boxes; Step S265) A second indicator acquisition step is performed to obtain a second position indicator based on the obtained recall rate of the predicted position. The calculation formula of the second position indicator is: ; Wherein, F2 represents the second position index.

[0056] Step S266) a second threshold judgment step, judging whether the initial threshold corresponding to the second position indicator is the maximum initial threshold in the initial threshold set, if not, taking the initial threshold next to the initial threshold and repeating step S261) a second intersection-and-union ratio calculation step, step S262) a second intersection-and-union ratio judgment step, step S263) a second accuracy acquisition step, step S264) a second recall rate acquisition step, step S265) a second indicator acquisition step, until the initial threshold corresponding to the second position indicator is the maximum initial threshold in the initial threshold set, and a second position indicator set is obtained.

[0057] Step S267) A second optimal threshold selection step, based on the obtained second position indicator set, selecting an initial threshold corresponding to the largest second position indicator in the second position indicator set as the optimal position threshold of the first model.

[0058] Step S3) an image correction step, based on the obtained first position information and first segmentation information and the second position information and second segmentation information, the invoice image and the two-dimensional code image are cropped, rotated and corrected to obtain a corrected second invoice image.

[0059] like Figure 6 As shown, step S3) the image correction step includes step S31) an invoice cropping step, step S32) an invoice rotation step, step S33) a contour acquisition step, step S34) a contour mapping step and step S35) an invoice correction step.

[0060] Step S31) Invoice cropping step, based on the first position information of the invoice image and the data category of the invoice image, the invoice image is cropped using the Crop function of opencv, and the data categories of the invoice image include: 0 degree invoice, 90 degree invoice, 180 degree invoice, and 270 degree invoice.

[0061] Step S32) Invoice rotation step: based on the first position information of the invoice image and the data category of the invoice image, the invoice image is rotated using the getRotationMatrix2D method.

[0062] Step S33) Contour acquisition step: transform the mask matrix between 0 and 1 according to the optimal mask threshold, and use the findContours function to traverse the contour points of the mask matrix to obtain the contour of the mask matrix.

[0063] Step S34) The contour mapping step is to map the contour of the mask matrix to the positions of the four corners of the invoice image to obtain a mapped mask matrix.

[0064] In this embodiment, if the 0 and 1 results of the mask matrix are ; Then use the findContours function to find the contour of the mask as m=[[1, 2], [2, 1], [3, 2], [3, 3]]. Since the first position information is the position information of the four corners of the invoice, there are four free points. These four points are the positions of the upper left, upper right, lower right and lower left points of the mask contour. Map the four points m=[[2, 1], [3, 2], [3, 3], [1, 2]] to dm =[[1, 1], [3, 1], [3, 3], [1, 3]] to obtain the mapping mask matrix.

[0065] Step S35) Invoice correction step, the mask matrix and the mapping mask matrix are encapsulated by using the vstack function in opencv, and the encapsulated mask matrix and the mapping mask matrix are subjected to target transformation by using the warpAffine function in opencv to obtain the invoice image.

[0066] like Figure 7 As shown, step S3) the image correction step also includes step S36) a two-dimensional code cutting step, step S37) a rotation angle acquisition step and step S38) a two-dimensional code correction step.

[0067] Step S36) A two-dimensional code cropping step, based on the second position information of the two-dimensional code image, the two-dimensional code image is cropped using the Crop function of opencv, and the formula is: img qr =crop(img, box qr ) Among them, img qr Indicates the QR code image area, box qr Indicates a QR code box.

[0068] Step S37) A rotation angle acquisition step is performed to obtain the positions of the contour points of the two-dimensional code image based on the second segmentation information of the two-dimensional code image, and then calculate the rotation angle of the two-dimensional code image. The formula is: p i =(p ix , p iy ), i=1,2,3,4; ; Among them, p i Represents the contour points of the QR code image, (p ix , p iy ) represents the position of the contour point, and a represents the rotation angle of the QR code image.

[0069] Step S38) A two-dimensional code correction step, based on the obtained two-dimensional code rotation angle, the two-dimensional code image is rotated using the getRotationMatrix2D method to obtain a corrected two-dimensional code image.

[0070] In this embodiment, the data of the two-dimensional code image cannot be deformed, and only the angle can be calculated for rotation, because the two-dimensional code information cannot be easily stretched and deformed, which will lead to unrecognizable results.

[0071] Step S4) Field recognition step, converting the second invoice image into a second model, obtaining the field position and field category information in the second invoice image, when there are handwritten words in the invoice image, the field position includes the handwritten field position, and the field category information includes the handwritten field category information.

[0072] Step S5) element recognition step, input the second invoice image into the third model, obtain the text recognition result of the second invoice image, obtain the element recognition result by corresponding the text recognition result to the preset element type, and record the position of each element, that is, whether each character in the text recognition result can correspond to the text segment in the preset element type.

[0073] In this embodiment, the second model is a DBNet (Resnet18) model, and the third model is a CRNN (resnet18) model. Both the second model and the third model are open source models and can be retrained using the training set in the present invention.

[0074] In this embodiment, the invoice element recognition result of the invoice includes "VAT invoice type", "invoice code", "invoice number", "invoice date", "purchaser name", "purchaser taxpayer identification number", "purchaser address and telephone number", "purchaser bank and account number", "total amount", "total tax amount", "total price and tax", "seller name", "seller taxpayer identification number", "seller address and telephone number", "seller bank and account number", "remarks", "payee", "review", "invoice issuer", "invoice special seal (presence or absence)", "special seal for invoice issuance on behalf of others (presence or absence)", "print invoice code", "print invoice number", "agent issuance mark", "goods or taxable labor or service name", "specification model", "unit", "quantity", "unit price", "amount", "tax rate", "tax amount", and "verification code"; the invoice element recognition result output data is in a dictionary format, including the invoice element name, recognition content, and recognition position.

[0075] Step S6) A QR code parsing step, based on the second position information and the second segmentation information of the QR code in the invoice image, the QR code content is parsed by using the pyzbar library to obtain a QR code recognition result.

[0076] Step S7) Invoice determination step, the text recognition result is matched with the field category information to obtain an invoice recognition result, and based on the invoice recognition result and the QR code recognition result, it is determined whether the invoice in the invoice image is a real invoice.

[0077] like Figure 8 As shown, step S7) the invoice judgment step includes step S71) a first judgment step, step S72) a second judgment step, step S73) a third judgment step and step S74) a fourth judgment step.

[0078] Step S71) The first judgment step is to judge whether the QR code recognition result is consistent with the invoice recognition result. If not, the invoice is determined to be a fake invoice. If so, proceed to the next step.

[0079] In this embodiment, the inconsistency between the QR code recognition result and the invoice recognition result is as follows: there are elements that are not recognized in the invoice element recognition result, but the element is recognized in the QR code result; the first element of the invoice element recognition result and the QR code recognition result is missing; the invoice element recognition result and the QR code recognition result are both complete, and there is one content mismatch.

[0080] Step S72) The second judgment step is to judge whether there is a handwritten field in the identified field. If so, the third judgment step is executed; if not, the fourth judgment step is executed.

[0081] Step S73) The third judgment step is to judge whether the overlap rate between the position of the handwritten field and the position of the invoice element is less than a preset threshold. If not, the invoice is determined to be a fake invoice. If so, the next step is executed.

[0082] In this embodiment, if the first judgment step of step S71) determines that the invoice is a genuine invoice, it is further determined whether the invoice is artificially modified. When there is a handwritten field in the identified field, the overlap rate between the position of the handwriting recognition and the position of the element in the invoice is determined. If the overlap rate is min_iou<0.9, it is determined that the field has not been artificially modified and is a genuine invoice; when there is no handwritten field in the identified field, the third judgment step of step S73) is skipped and the fourth judgment step of step S74) is directly executed.

[0083] Step S74) The fourth judgment step is to judge whether the handwritten field belongs to the key field category. If not, the invoice is judged to be a real invoice.

[0084] In this embodiment, the key fields include "invoice code", "invoice number", "invoice date", "purchaser's name", "purchaser's taxpayer identification number", "purchaser's address and telephone number", "purchaser's bank account and account number", "total amount", "total tax amount", "total price and tax", "seller's name", "seller's taxpayer identification number", "seller's address and telephone number", "seller's bank account and account number", "print invoice code", "print invoice number", "name of goods or taxable services", "specification model", "unit", "quantity", "unit price", "amount", "tax rate", "tax amount", and "verification code".

[0085] like Fig. 9 As shown, the present application also proposes a storage medium 100 storing computer-readable instructions. When the computer-readable instructions are read by at least one processor 110, at least one processor 120 executes at least one step of the method for identifying the authenticity of the contents of multiple invoices.

[0086] The present application provides a method and storage medium for identifying the authenticity of invoice content. Through a deep learning method, the area of ​​the invoice is judged, the position of the identification element is detected, and the element result of the invoice is obtained by the text content of the element. The QR code position is detected and the QR code content is identified to obtain QR code parsing data, and handwriting data is detected. The authenticity of the invoice is judged by matching the invoice element content recognition with the QR code data and the overlap rate between the element recognition result and the handwriting area, thereby realizing the authenticity identification of massive invoice contents and effectively improving the ability to identify the authenticity of invoices.

[0087] The above is a detailed introduction to the authenticity identification method and storage medium for invoice contents provided by the present application. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. At the same time, for general technicians in this field, according to the idea of ​​the present application, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.

Claims

1. A method for identifying the authenticity of invoice content, characterized in that: The specific steps include: An image acquisition step, acquiring a first invoice image in a target format, wherein the first invoice image includes image information and text information, and the image information includes an invoice image and a QR code image; A detection and segmentation step, inputting the image information into a first model, obtaining first position information and first segmentation information of the invoice image and second position information and second segmentation information of the QR code image; An image correction step, based on the obtained first position information and first segmentation information and the second position information and second segmentation information, the invoice image and the two-dimensional code image are cropped, rotated and corrected to obtain a corrected second invoice image; a field recognition step, converting the second invoice image into a second model, and obtaining field positions and field category information in the second invoice image; An element recognition step, inputting the second invoice image into the third model, obtaining a text recognition result of the second invoice image, obtaining an element recognition result by corresponding the text recognition result to a preset element type, and recording the position of each element; A QR code parsing step, based on the second position information and the second segmentation information of the QR code in the invoice image, the QR code content is parsed by using the pyzbar library to obtain a QR code recognition result; as well as In the invoice judgment step, the text recognition result is matched with the field category information to obtain an invoice recognition result, and based on the invoice recognition result and the QR code recognition result, it is judged whether the invoice in the invoice image is a real invoice.

2. The method for identifying the authenticity of invoice contents according to claim 1, characterized in that: The detection and segmentation step specifically includes the following steps: The first model building step is to build a maskRcnn target detection and segmentation model implemented by the pytorch framework; A data labeling step, labeling the position points of the invoice frame and the position points of the segmented area of ​​the invoice in the invoice image; Marking the position points of the frame of the two-dimensional code in the invoice image and the position points of the segmented area of ​​the two-dimensional code to obtain marking data; A data processing step of preprocessing the labeled data to obtain preprocessed data, and dividing the preprocessed data into a training set, a validation set, and a test set; A model training step, setting different learning rates lr, batch sizes batch_size and epoch numbers, training the first model using the training set, and verifying it using the verification set; as well as The model inference step uses the test set to infer the first model, obtain the optimal mask threshold and the optimal position threshold, and obtain the trained first model.

3. The method for identifying the authenticity of invoice contents as claimed in claim 2, characterized in that: The data processing step specifically includes the following steps: A data enhancement step is to count the distribution of the labeled data and perform data enhancement processing on data whose amount is less than 1 / 2 of the maximum data category; The mask matrix acquisition step is to obtain the mask matrix by marking the segmented area in the labeled data, and its formula is: ; Among them, segment represents the marked segmentation area; i xy represents the corresponding element in the mask matrix; and The normalization step is to normalize the frame of the marked area in the labeled data, and the formula is: ; Among them, box represents the box, b x , b y , b w , b h Indicates the location information of the box, m w , m h Indicates the width and height of the box.

4. The method for identifying the authenticity of invoice contents as claimed in claim 2, characterized in that: The model reasoning step includes the following steps: In the first intersection-and-union ratio calculation step, the test set is input into the first model to obtain at least one predicted mask matrix, and a first intersection-and-union ratio between the labeled mask matrix and the predicted mask matrix is ​​obtained. The calculation formula of the first intersection-and-union ratio is: ; Among them, IOU1 represents the first intersection-over-union ratio, mask p Represents the predicted mask matrix, mask r Represents the label mask matrix; A first intersection-and-union ratio judgment step, setting an initial threshold set, wherein the initial threshold set includes more than two initial thresholds, and judging whether there is a first intersection-and-union ratio greater than the minimum initial threshold in the initial threshold set, and if so, executing the next step; The first accuracy acquisition step is to obtain the accuracy of the predicted segmentation area based on the number of predicted mask matrices obtained. The formula is: ; Among them, pre1 represents the accuracy of the predicted mask matrix, TP1 represents the number of correct predictions of the predicted mask matrix, and P1 represents the number of predicted mask matrices; In the first recall rate acquisition step, the recall rate of the predicted mask matrix is ​​obtained based on the accuracy of the obtained predicted mask matrix. The formula is: ; Among them, recall1 represents the recall rate of the predicted mask matrix, and R1 represents the number of labeled mask matrices; The first indicator acquisition step is to obtain a first position indicator based on the obtained recall rate of the predicted segmented area. The calculation formula of the first position indicator is: ; Among them, F1 represents the first position index; A first threshold determination step, determining whether the initial threshold corresponding to the first position indicator is the maximum initial threshold in the initial threshold set, and if not, taking the initial threshold next to the initial threshold to repeatedly perform the first intersection-and-union ratio calculation step, the first intersection-and-union ratio determination step, the first accuracy acquisition step, the first recall rate acquisition step and the first indicator acquisition step, until the initial threshold corresponding to the first position indicator is the maximum initial threshold in the initial threshold set, thereby obtaining a first position indicator set; and The first optimal threshold selection step selects, based on the obtained first position indicator set, an initial threshold corresponding to the largest first position indicator in the first position indicator set as the optimal mask threshold of the first model.

5. The method for identifying the authenticity of invoice contents as claimed in claim 2, characterized in that: The model reasoning step also includes the following steps: The second IoU calculation step is to input the test set into the first model, obtain at least one prediction box, and obtain the second IoU ratio between the marked box and the prediction box. The calculation formula of the second IoU ratio is: ; Among them, IOU2 represents the second intersection-and-union ratio, box p Represents the prediction box, box r A box representing a callout; A second intersection-and-union ratio judgment step, setting an initial threshold set, wherein the initial threshold set includes more than two initial thresholds, and judging whether there is a second intersection-and-union ratio greater than the minimum initial threshold in the initial threshold set, and if so, executing the next step; The second accuracy acquisition step is to obtain the accuracy of the predicted position point based on the number of prediction boxes obtained. The formula is: ; Among them, pre2 represents the accuracy of the prediction box, TP2 represents the number of correct predictions of the prediction box, and P2 represents the number of prediction boxes; The second recall rate acquisition step is to obtain the recall rate of the prediction box based on the accuracy of the obtained prediction box. The formula is: ; Among them, recall2 represents the recall rate of the predicted box, and R2 represents the number of labeled boxes; The second indicator acquisition step is to obtain a second position indicator based on the obtained recall rate of the predicted position. The calculation formula of the second position indicator is: ; Wherein, F2 represents the second position index; a second threshold determination step, determining whether the initial threshold corresponding to the second position indicator is the maximum initial threshold in the initial threshold set; if not, taking the initial threshold next to the initial threshold to repeatedly perform the second intersection-over-union calculation step, the second intersection-over-union determination step, the second accuracy acquisition step, the second recall acquisition step and the second indicator acquisition step, until the initial threshold corresponding to the second position indicator is the maximum initial threshold in the initial threshold set, thereby obtaining a second position indicator set; and The second optimal threshold selection step is to select, based on the obtained second position indicator set, an initial threshold corresponding to the largest second position indicator in the second position indicator set as the optimal position threshold of the first model.

6. The method for identifying the authenticity of invoice contents according to claim 1, characterized in that: The image correction step comprises the following steps: an invoice cropping step, cropping the invoice using the Crop function of opencv based on the first position information of the invoice image and the data category of the invoice image; an invoice rotation step, rotating the invoice using a getRotationMatrix2D method based on the first position information of the invoice image and the data category of the invoice image; Contour acquisition step: transform the mask matrix into 0 and 1 according to the optimal mask threshold, and use the findContours function to traverse the contour points of the mask matrix to obtain the contour of the mask matrix; a contour mapping step, mapping the contour of the mask matrix to the position in the first position information to obtain a mapping mask matrix; as well as In the invoice correction step, the mask matrix and the mapping mask matrix are encapsulated by using the vstack function in opencv, and the encapsulated mask matrix and the mapping mask matrix are subjected to target transformation by using the warpAffine function in opencv to obtain a corrected invoice image.

7. The method for identifying the authenticity of invoice contents according to claim 1, characterized in that: The image correction step also includes the following steps: The two-dimensional code cropping step is to crop the two-dimensional code image based on the second position information of the two-dimensional code image using the Crop function of opencv, and the formula is: img qr =crop(img, box qr ); Among them, img qr Indicates the QR code image area, box qr Indicates a QR code frame; The rotation angle acquisition step obtains the position of the contour point of the two-dimensional code image based on the second segmentation information of the two-dimensional code image, and then calculates the rotation angle of the two-dimensional code image, and its formula is: p i =(p ix , p iy ), i=1,2,3,4; ; Among them, p i Represents the contour points of the QR code image, (p ix , p iy ) represents the position of the contour point, a represents the rotation angle of the two-dimensional code image; and The two-dimensional code correction step is to rotate the two-dimensional code image using the getRotationMatrix2D method based on the acquired two-dimensional code rotation angle to obtain a corrected two-dimensional code image.

8. The method for identifying the authenticity of invoice contents according to claim 1, characterized in that: The invoice judgment step includes the following steps: The first judgment step is to judge whether the QR code recognition result is consistent with the invoice recognition result. If not, the invoice is determined to be a fake invoice. If so, proceed to the next step; A second judgment step, judging whether there is a handwritten field in the recognized field, if yes, executing the third judgment step, if no, executing the fourth judgment step; The third judgment step is to judge whether the overlap rate between the position of the handwritten field and the position of the invoice element is less than a preset threshold value. If not, the invoice is determined to be a fake invoice. If so, the next step is executed; as well as The fourth judgment step is to judge whether the handwritten field belongs to the key field category. If not, the invoice is judged to be a genuine invoice.

9. A storage medium storing computer-readable instructions, which, when read by at least one processor, causes at least one processor to execute at least one step of the method for identifying the authenticity of invoice contents as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Bill inspection method, device, terminal apparatus and storage medium

    CN111462388A

  • Text graph offset angle prediction and correction method

    CN112784836A

  • Alluvial-proluvial fan deposition remote sensing intelligent identification method and device

    CN114387501A

  • Invoice image recognition method and device, equipment and storage medium

    CN114611541A

  • Multi-mode invoice automatic classification identification method, verification method and system

    CN116052186A