False invoice making identification method, device and equipment
By extracting and aligning the image, text and structured data feature vectors of invoices and combining them with an unsupervised anomaly detection model, the shortcomings of existing false invoice identification methods are solved and efficient recognition of digital invoices is achieved.
Patent Information
- Application Number
- CN202511024762.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-24
- Publication Date
- 2025-10-17
AI Technical Summary
Existing methods for identifying false invoices rely on manual audits or simple rules, which cannot effectively identify digital electronic invoices. They also lack joint modeling of multi-source information, resulting in unsatisfactory identification results.
By extracting the joint feature vectors of the invoice's image, text and structured data, using an unsupervised anomaly detection model for identification, and combining hierarchical density clustering and independent forest algorithm for weighted sum scoring, accurate identification of false invoices can be achieved.
It improves the accuracy and applicability of false invoice identification, supports automatic identification of paper, electronic and digital invoices, reduces dependence on labeled data, and adapts to multi-field applications.
Smart Images

Figure CN120808363A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computers, and in particular to a virtual invoice identification method, device and equipment. BACKGROUND
[0002] Virtual invoice opening is an irregular behavior in the financial field, which mainly manifests as behaviors such as fictitious transaction background, fictitious transaction amount or frequent association of abnormal accounts.
[0003] At present, the traditional virtual invoice identification method relies on manual audit or simple rule-based verification mode, and has problems such as low efficiency, rule update lag, and difficulty in covering new fraud modes. The tax system constructs risk indicators based on numerical fields, and cannot capture graphic-text contradictions, and relies too much on single-modal data. Although the existing Optical Character Recognition (OCR) technology can extract text, it does not model jointly with semantics and transaction network features, resulting in weak graphic-text consistency detection capability. In addition, with the popularization and application of digital electronic invoices, compared with traditional paper invoices and electronic invoices, the style has been greatly revised, the content and format have been adjusted, and dynamic two-dimensional codes have been added, so that the existing identification method cannot support automatic identification of virtual digital electronic invoices. It can be seen that the existing virtual invoice identification method cannot achieve ideal identification effect. SUMMARY
[0004] The main purpose of the embodiments of the present application is to provide a virtual invoice identification method, device and equipment, which can jointly analyze image data, text data and structured data corresponding to an invoice, so as to more accurately identify whether a virtual invoice situation occurs, and achieve ideal identification effect.
[0005] The embodiments of the present application provide a virtual invoice identification method, comprising:
[0006] Obtaining a target image in which a target invoice to be identified is located; and extracting text data and structured data of the target invoice from the target image;
[0007] Extracting an image feature vector of a first dimension corresponding to the target image;
[0008] Extracting a semantic feature vector of a second dimension corresponding to the text data of the target invoice;
[0009] Extracting a structured feature vector of a third dimension corresponding to the structured data of the target invoice;
[0010] Aligning the image feature vector of the first dimension, the semantic feature vector of the second dimension and the structured feature vector of the third dimension to obtain an aligned joint feature vector.
[0011] inputting the joint feature vector into an unsupervised anomaly detection model to identify the target invoice, to obtain a virtual invoice identification result corresponding to the target invoice.
[0012] In an optional implementation, the extracting the image feature vector of the first dimension corresponding to the target image comprises:
[0013] correcting the target image to obtain a corrected target image;
[0014] inputting the corrected target image into a pre-trained residual network model to extract an image feature, to obtain an image feature vector of the first dimension.
[0015] In an optional implementation, the extracting the semantic feature vector of the second dimension corresponding to the text data of the target invoice comprises:
[0016] extracting a cargo name field information in the text data of the target invoice by using an optical character recognition (OCR) algorithm;
[0017] generating a dynamic word vector of the second dimension corresponding to the cargo name field information by using a pre-trained language model and a word segmenter, as a semantic feature vector of the second dimension corresponding to the text data of the target invoice.
[0018] In an optional implementation, the image feature vector of the first dimension is a 512-dimensional image feature vector; the semantic feature vector of the second dimension is a 768-dimensional semantic feature vector; the structured feature vector of the third dimension is a 128-dimensional structured feature vector; and the aligning the image feature vector of the first dimension, the semantic feature vector of the second dimension, and the structured feature vector of the third dimension to obtain a joint feature vector after alignment comprises:
[0019] mapping the 512-dimensional image feature vector and the 768-dimensional semantic feature vector to a unified 256-dimensional vector space to obtain a joint image-text feature vector after alignment;
[0020] projecting and aligning the joint image-text feature vector after alignment and the 128-dimensional structured feature vector to obtain a joint feature vector after alignment.
[0021] In an optional implementation, the inputting the joint feature vector into an unsupervised anomaly detection model to identify the target invoice, to obtain a virtual invoice identification result corresponding to the target invoice comprises:
[0022] inputting the joint feature vector into an unsupervised anomaly detection model to obtain an anomaly score of hierarchical density clustering and an anomaly score of independent forest;
[0023] The weighted sum score is obtained by performing a weighted sum processing on the hierarchical density clustering anomaly score and the independent forest anomaly score; and the target invoice is identified by using the weighted sum score, so as to obtain a virtual invoice identification result corresponding to the target invoice.
[0024] Corresponding to the virtual invoice identification method, the application provides a virtual invoice identification device, comprising:
[0025] An acquisition unit is configured to acquire a target image in which a target invoice to be identified is located, and extract text data and structured data of the target invoice from the target image;
[0026] A first extraction unit is configured to extract a first-dimensional image feature vector corresponding to the target image;
[0027] A second extraction unit is configured to extract a second-dimensional semantic feature vector corresponding to the text data of the target invoice;
[0028] A third extraction unit is configured to extract a third-dimensional structured feature vector corresponding to the structured data of the target invoice;
[0029] An alignment unit is configured to perform alignment processing on the first-dimensional image feature vector, the second-dimensional semantic feature vector and the third-dimensional structured feature vector, so as to obtain an aligned joint feature vector;
[0030] An identification unit is configured to input the joint feature vector into an unsupervised anomaly detection model, identify the target invoice, and obtain a virtual invoice identification result corresponding to the target invoice.
[0031] In an optional implementation, the first extraction unit comprises:
[0032] A correction subunit is configured to correct the target image, so as to obtain a corrected target image;
[0033] A first extraction subunit is configured to input the corrected target image into a pre-trained residual network model to extract image features, so as to obtain a first-dimensional image feature vector.
[0034] In an optional implementation, the second extraction unit comprises:
[0035] A second extraction subunit is configured to extract information of a goods name field in the text data of the target invoice by using an optical character recognition (OCR) algorithm;
[0036] The generating subunit is configured to generate a second-dimension dynamic word vector corresponding to the goods name field information by using a pre-trained language model and a word segmenter, as a second-dimension semantic feature vector corresponding to the text data of the target invoice.
[0037] In an optional implementation, the first-dimension image feature vector is a 512-dimension image feature vector; the second-dimension semantic feature vector is a 768-dimension semantic feature vector; the third-dimension structured feature vector is a 128-dimension structured feature vector; and the aligning unit includes:
[0038] The mapping subunit is configured to map the 512-dimension image feature vector and the 768-dimension semantic feature vector to a unified 256-dimension vector space to obtain an aligned image-text feature vector.
[0039] The aligning subunit is configured to perform projection alignment processing on the aligned image-text feature vector and the 128-dimension structured feature vector to obtain an aligned joint feature vector.
[0040] In an optional implementation, the identifying unit includes:
[0041] The input subunit is configured to input the joint feature vector into an unsupervised anomaly detection model to obtain an anomaly score of hierarchical density clustering and an anomaly score of independent forest;
[0042] The identifying subunit is configured to perform weighted summation processing on the anomaly score of hierarchical density clustering and the anomaly score of independent forest to obtain a weighted summation score, and identify the target invoice by using the weighted summation score to obtain a virtual invoice identification result corresponding to the target invoice.
[0043] Embodiments of the present application also provide a virtual invoice identification device, which includes a processor, a memory, and a system bus.
[0044] The processor and the memory are connected through the system bus.
[0045] The memory is configured to store one or more programs, and the one or more programs include instructions which, when executed by the processor, cause the processor to perform any one of the implementation manners of the virtual invoice identification method.
[0046] Embodiments of the present application also provide a computer readable storage medium, which stores instructions, and when the instructions run on a terminal device, cause the terminal device to perform any one of the implementation manners of the virtual invoice identification method.
[0047] Therefore, embodiments of the present application have the following beneficial effects:
[0048] The embodiment of the present application provides a kind of virtual invoice identification method, device and equipment, first, the target image where target invoice to be identified is obtained;And the text data and structured data of target invoice are extracted from target image.Then the first dimension image feature vector corresponding to target image of target invoice, the second dimension semantic feature vector corresponding to the text data of target invoice, and the third dimension structured feature vector corresponding to the structured data of target invoice are extracted.Image feature vector of first dimension, second dimension semantic feature vector and third dimension structured feature vector are aligned, and the joint feature vector after alignment is obtained;Joint feature vector is input into unsupervised anomaly detection model, and target invoice is identified, and the virtual invoice identification result corresponding to target invoice is obtained.Thereby, the joint analysis of the image data, text data and structured data corresponding to target invoice can be carried out, and whether the situation of virtual invoice appears can be more accurately identified, and ideal identification effect is achieved. BRIEF DESCRIPTION OF DRAWINGS
[0049] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or prior art description will be briefly introduced. Obviously, the drawings in the following description are some embodiments of the present application, and those skilled in the art can also obtain other drawings according to these drawings without creative labor.
[0050] Figure 1 A flow chart of a virtual invoice identification method provided by the embodiment of the present application is provided.
[0051] Figure 2 A composition schematic diagram of a virtual invoice identification device provided by the embodiment of the present application is provided. DETAILED DESCRIPTION
[0052] In order to make the purpose, technical scheme and advantages of the embodiments of the present application more clear, the technical scheme in the embodiments of the present application will be described clearly and completely with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0053] Traditional false invoice identification methods only rely on structured data (such as amount, tax rate) or a single modality (such as text or image), and it is difficult to capture the inconsistency of multi-source information in false invoices. Currently, a supervised method is usually used to train the model for false invoice identification. However, there is no publicly available annotated invoice dataset, and manual annotation of data is time-consuming and labor-intensive. In addition, the scarcity of false invoice samples also leads to poor model training results. The generalization ability of invoices with various formats and cross-industry invoices is poor, and only the identification of enterprises or taxpayers is realized, and the processing and judgment of the invoice itself are lacking.
[0054] In addition, with the popularization and application of digital electronic invoices, compared with traditional paper invoices and electronic invoices, the style has been greatly revised, and the content and format of the fill-in have been adjusted. For example, the invoice number of the digital electronic invoice is 20 digits, the password area is cancelled, the invoice code, machine code, invoice special seal and enterprise digital front, opening bank, bank account, payee, reviewer, etc. can be displayed in the remarks column, and a dynamic two-dimensional code is added, which makes the existing identification method unable to support the automatic identification of false digital electronic invoices.
[0055] Therefore, the existing false invoice identification method cannot achieve ideal identification results.
[0056] Based on this, the present application provides a false invoice identification method, device and equipment, which combines image, text and structured data joint analysis to improve identification accuracy and reduce dependence on labeled data, avoids the problem of overfitting caused by low sample proportion and sample imbalance of false invoices, and relies on oversampling or cost-sensitive learning, so as to achieve ideal identification results.
[0057] The false invoice identification method provided by the embodiments of the present application will be described in detail below with reference to the accompanying drawings. Referring to Figure 1 As shown in the figure, a flowchart of a false invoice identification method embodiment provided by the present application is shown, and the present embodiment can include the following steps:
[0058] S101: Obtain a target image in which a target invoice to be identified is located; and extract text data and structured data of the target invoice from the target image.
[0059] In this embodiment, any invoice that needs to be judged whether it is false opening is defined as a target invoice to be segmented and recognized, and an image containing the target invoice is defined as a target image. It should be noted that this embodiment does not limit the type of target image, for example, the target image can be a color image composed of red (R), green (G), and blue (B) three primary colors, or a grayscale image, etc. Moreover, this application also does not limit the type and content of the target invoice in the target image, for example, the target image can be a value-added tax invoice scan, a PDF electronic invoice screenshot, an invoice image taken by a mobile terminal (such as a mobile phone), etc.
[0060] Moreover, in order to improve the recognition accuracy and realize joint analysis of image data, text data and structured data of the target invoice, after obtaining the target image in which the target invoice to be recognized is located, the text data and structured data of the target invoice can be extracted from the target image by using existing or future image recognition methods (such as OCR) to perform subsequent steps S102-S104. The text data can include but is not limited to commodity information (i.e. goods name field information), buyer and seller information, note field, etc.; the structured data can include but is not limited to invoice amount, tax rate, transaction timestamp, invoice date, etc. numerical field data and invoice type, transaction area, commodity (goods) category, etc. category type field data.
[0061] S102: Extracting an image feature vector of a first dimension corresponding to the target image.
[0062] In this embodiment, after obtaining the target image in which the target invoice to be recognized is located by step S101, since the image (such as an invoice image taken by a mobile terminal) can have problems such as uneven illumination, blurred seal, image blur, picture tilt or background noise, in order to remove these interferences and ensure the clarity and uniformity of the image, the original invoice picture is preprocessed by standardization and normalization operation to improve the recognition effect of the target invoice. Specifically, one optional implementation manner is that first, the target image can be corrected for tilt or distortion by using existing or future image correction methods (such as using perspective transformation), to obtain a corrected target image. Then, the corrected target image is input into a pre-trained residual network model (such as ResNet-50 model) for image feature extraction, to obtain an image feature vector of a first dimension (the specific value is not limited, which can be set according to actual situation and experience value, such as 512 dimensions, etc.) and represented as f i .
[0063] Specifically, first, the target image in which the target invoice is located can be corrected using perspective transformation, that is, the transformation matrix is calculated by detecting the four corner key points of the invoice, such as the two-dimensional code area and the invoice code frame, and then the binary processing is performed, the adaptive threshold algorithm is adopted, the histogram is calculated, the histogram data is smoothed according to a certain radius, and the threshold value is determined according to the distance between the peak value and the minimum value in proportion (the specific value is not limited), and the image after processing can complete the separation of the text and the background. Then, non-local mean denoising is used to effectively retain the text edge and the image details of the target image in which the target invoice is located, and the corrected target image is obtained.
[0064] Then, the corrected target image can be subjected to text detection and recognition, such as inputting the aforementioned corrected target image as a training set to train a model to improve the recognition accuracy of each field of the target invoice. However, it should be noted that since the invoice has a fixed format, the OCR recognition result is mapped to the structured field, including the name, taxpayer identification number, address and telephone number, bank and account number, project name, specification and model, amount, tax rate, tax amount, total price including tax, remarks, etc. For the two key fields of invoice code and taxpayer identification number containing numbers and characters, regular expressions are used for verification after the OCR extraction mapping is completed to ensure the accuracy of the field.
[0065] In addition, for the invoice seal area, table area and text block area, the seal bounding box coordinates can be represented as (x min , y min , x max , y max ), the positional relationship with the key field is calculated, and the table structure analysis can use the Hough transform to extract the table line, and the specific formula is as follows:
[0066]
[0067] Where (x, y) is the rectangular coordinate representing a point in the plane, x is the horizontal axis coordinate, and y is the vertical axis coordinate; (p, q) is the polar coordinate, which also corresponds to a point in the plane, p is the distance (modulus) from the origin, and q is the angle (polar angle) between the point and the positive direction of the x-axis.
[0068] During the detection process, when the peak point in the p-q parameter space is detected, a straight line is detected, and a grid is generated based on the focal point of the straight line, and the number of rows and columns and the size of the cells are counted. The coefficient of variation of the area of all cells is calculated, and if it is greater than a threshold value (the specific value is not limited), it is marked as abnormal.
[0069] In addition, the detected seal coordinates can be converted into a proportion relative to the image size, and the specific formula is as follows:
[0070]
[0071] wherein, W and H represent the width and height of the target image (corrected) respectively; if the distance between the seal center and the preset key region exceeds a threshold value (the specific value is not limited), it is marked as an anomaly.
[0072] Further, for the stroke width, the edge E(x, y) can be extracted using the Canny operator or the like, a ray is tracked along the edge gradient direction, the stroke width s is calculated, and the variance of the stroke width within the same text region is calculated , and the specific formula is as follows:
[0073]
[0074] wherein, s represents the total stroke width; μ represents the average value of the stroke width, and if the value of is greater than a preset threshold value (the specific value is not limited), it is determined that there is multi-font mixing, and it is marked as an anomaly.
[0075] In this way, by using a pre-trained residual network model (such as a ResNet-50 model) with a removed full connection layer, a 512-dimensional global feature vector corresponding to the target image can be extracted as the first-dimensional image feature vector, which is used to perform the subsequent step S105.
[0076] S103: Extracting a second-dimensional semantic feature vector corresponding to the text data of the target invoice.
[0077] In the present embodiment, after the text data of the target invoice is extracted through step S101, in order to avoid the interference of redundant information on the recognition result, the text data can be cleaned and filtered, and general stop words such as “Limited Company” and “invoice number” can be removed, and special symbols such as “【】”, “*”, and “#” can be deleted using a regular expression. Due to the possibility of semantic errors during OCR extraction, the aforementioned results can also be corrected based on context similarity, using edit distance and TF-IDF. Specifically, one optional implementation is that the cargo name field information in the text data of the target invoice can be extracted using an OCR algorithm. Then, a pre-trained language model and a tokenizer are used to generate a second-dimensional (the specific value is not limited and can be set according to actual conditions and experience values, such as 768 dimensions) dynamic word vector corresponding to the cargo name field information, as the second-dimensional semantic feature vector corresponding to the text data of the target invoice, and denoted as f t , which is used to perform the subsequent step S105.
[0078] Specifically, first, a 768-dimensional dynamic word vector can be generated for the cargo name field extracted by OCR using a pre-trained language model and a word segmenter, and then entity linking is performed to map the cargo name to a standard commodity classification, eliminating differences in the expression of commodity names, such as “notebook” corresponding to “stationery” or “electronic stationery”. This step requires a pre-constructed knowledge graph, but the construction method is not limited, such as division by industry category -> medium category -> small category, and each classification node is associated with a standard name, code, and a list of typical commodity names. The cosine similarity between the dynamic word vector of the cargo (denoted as v item ) and the static vector of the standard commodity node in the knowledge graph (denoted as v std ) is calculated, and the specific calculation formula is as follows:
[0079]
[0080] Among them, for the standard commodity name, a static vector can also be generated using a pre-trained language model and a word segmenter, and then the top 5 (or other values, only for example) standard classifications with the highest similarity are selected to enter the verification stage. Detect conflicts between cargo descriptions and enterprise industry types, verify whether the statutory tax rate corresponding to the cargo classification is consistent with the actual tax rate on the invoice, and count the co-occurrence frequency of goods and buyers in invoices issued by the same enterprise to detect irregular combinations. If the cargo name cannot be mapped to the existing classification, i.e., the similarity is less than a threshold value (the specific value is not limited), the manual review process is triggered, and after passing the review, it is added to the knowledge graph. Periodically fine-tune the pre-trained language model and word segmenter with new commodity names and their classification labels to enhance semantic representation capabilities. The specific fine-tuning formula is as follows:
[0081]
[0082] Where L represents the cross-entropy loss; represents the gradient descent algorithm function to find the minimum value of the function by iteration, i.e., to minimize the loss function; represents the learning rate of the loop, and the specific value is not limited and can be determined according to actual conditions and experience values, such as taking a value of 0.01 or the like; represents the standard classification label.
[0083] S104: Extract a third-dimensional structured feature vector corresponding to the structured data of the target invoice.
[0084] In this embodiment, after the structured data of the target invoice is extracted by step S101, the missing values of the numerical fields such as invoice amount, transaction timestamp, transaction frequency, etc. are filled with the historical invoicing mean or median, the missing values of the categorical fields such as commodity category, transaction area, invoice type, etc. are filled with the "unknown" label, the unexpected values of the non-standard tax rate are filtered out, and whether the tax rate is compliant is verified.
[0085] 56-dimensional basic features are constructed by calculating the single-day invoicing amount standard deviation, the 7-day (other numerical values can also be used, which are only examples) transaction frequency coefficient of variation, the non-working time transaction proportion, the amount kurtosis, the transaction interval time skewness, the transaction counterparty concentration, the commodity category distribution entropy value, the Fourier transform to extract the main frequency amplitude, the week transaction volume autocorrelation coefficient, the cross-provincial transaction proportion, and the number of related accounts.
[0086] A transaction graph is constructed by the buyer, seller, and bank account information, and 32-dimensional graph structure features are extracted.
[0087] According to the invoicing time, the proportion of weekend invoicing is calculated according to the week field, the seasonal fluctuations are analyzed according to the month field, the 7-day moving average is calculated by a sliding window, the recent invoicing frequency trend of the enterprise is calculated, and a hybrid deep learning model combining Long Short-Term Memory (LSTM) and Autoencoder is used to extract 24-dimensional time series features. Then, the above features are concatenated to obtain a 112-dimensional feature vector. Finally, 95% of the variance is retained by dimension reduction, and about 100 dimensions are output. Then, a fully connected layer is used for adjustment to output 128-dimensional structured modal features as the third-dimensional structured feature vector corresponding to the structured data of the target invoice, and is represented as f s to perform the subsequent step S105.
[0088] S105: Aligning the first-dimensional image feature vector, the second-dimensional semantic feature vector, and the third-dimensional structured feature vector to obtain an aligned joint feature vector.
[0089] In this embodiment, after the first-dimensional image feature vector f i (such as a 512-dimensional image feature vector), the second-dimensional semantic feature vector f t (such as a 768-dimensional semantic feature vector), and the third-dimensional structured feature vector f s (such as a 128-dimensional structured feature vector) corresponding to the target image are extracted by steps S102-S104, respectively, the first-dimensional image feature vector f i , the second-dimensional semantic feature vector f tstructured feature vector f s The alignment processing is performed to obtain an aligned joint feature vector. For example, the 512-dimensional image feature vector and the 768-dimensional semantic feature vector can be mapped to a unified 256-dimensional vector space to obtain an aligned image-text feature vector, and then the aligned image-text feature vector and the 128-dimensional structured feature vector are subjected to projection alignment processing to obtain an aligned joint feature vector, which is used to perform the subsequent step S106.
[0090] Specifically, when the image, text and structured data of the same target invoice are associated through a unique identifier, invoice code + invoice number, first, the image-text features can be aligned to obtain a 512-dimensional image feature vector f i and a 768-dimensional semantic feature vector f t Then, the image-text features can be mapped to a unified 256-dimensional space through a fully connected layer, and the projection layer formula is as follows:
[0091]
[0092]
[0093] wherein, W i ∈R 256*512 and W t ∈R 256*512 represent learnable parameters; b i and b t represent bias terms.
[0094] To narrow the distance of the image-text features of the same invoice and push away the features of different invoices, the loss function is as follows:
[0095]
[0096] wherein, represents a cosine similarity; τ represents a temperature coefficient, which can be taken as 0.07 by default, and controls the smoothing value of the similarity distribution; N represents the number of negative samples in a batch (all other samples in the same batch are negative samples).
[0097] Then, the third-dimensional structured feature vector f t is aligned with the image-text feature vector. Similarly, the multi-modal features can be mapped to a unified 256-dimensional space through a fully connected layer, and the projection layer formula is as follows:
[0098]
[0099]
[0100] wherein, “;” represents a splicing operation; and represents a projection matrix.
[0101] Further, the aligned joint feature vector (denoted as ) is:
[0102]
[0103] wherein, α, β and γ represent learnable weights, initialized as 1, and can be updated by the change of the following loss function:
[0104]
[0105] S106: input the joint feature vector structured feature vector into the unsupervised anomaly detection model, identify the target invoice, and obtain the virtual invoice identification result corresponding to the target invoice.
[0106] In this embodiment, the aligned joint feature vector Further, the joint feature vector structured feature vector can be input into the unsupervised anomaly detection model first, to obtain the hierarchical density clustering anomaly score and the independent forest anomaly score, and then the hierarchical density clustering anomaly score and the independent forest anomaly score are processed by weighted summation to obtain a weighted summation score; and the target invoice is identified by using the weighted summation score, to obtain a more accurate virtual invoice identification result corresponding to the target invoice.
[0107] Specifically, in actual application, the hierarchical density clustering (Hierarchical Density-Based Spatial Clustering of Applications with Noise, HDBSCAN) algorithm can be used, which is based on the density connection and uses a hierarchical method to construct a clustering structure. The advantage of this algorithm is that it can automatically determine the number of clusters, adapt to non-convex distribution, and use the distance of the sample from the nearest cluster center as the abnormality degree. The specific calculation formula of the hierarchical density clustering anomaly score is as follows:
[0108]
[0109] wherein, C represents a cluster center set.
[0110] During the training of the model, actual experience can be combined for parameter optimization, such as taking min_cluster_size=15 as the minimum virtual gang scale, taking min_samples=5 as the minimum number of neighbors of the core point, and taking cluster_selection_epsilon=0.3 to merge clusters with a distance of <0.3, which is used to capture associated gangs, cluster_selection_method='eom', using the excess maximum method to balance the size of the cluster, alpha=1.0 to control the boundary sensitivity.
[0111] On this basis, the aligned joint feature vector obtained through the above steps Input the training model, calculate the core distance of each point, generate a mutual reachable distance matrix, arrange the MST edges in descending order of distance, gradually merge the clusters, calculate the stability of each cluster, and select the cluster set with the maximum stability sum.
[0112] Combined with the Isolation Forest algorithm, the specific formula of the independent forest anomaly score is as follows:
[0113]
[0114] Among them, E(h(x)) represents the average sample path length; c(n) represents the standardization factor.
[0115] Further define the parameter space n_estimators=[50,100,200], max_samples=[0.5,0.8,'auto'], estimate the abnormal proportion contamination=[0.01,0.05,0.1].max_features=[0.7,0.9,1.0], automatically optimize the training model, and evaluate the unsupervised model based on the path length. Select the parameter with the largest separation degree.
[0116] Further to the anomaly score of hierarchical density clustering And the independent forest anomaly score Weighted sum processing is performed to obtain the weighted sum score The specific formula is as follows:
[0117]
[0118] Among them, And represent the weight value, and the specific value is not limited, which can be optimized by unsupervised indicators.
[0119] When the weighted sum score when the value of the weighted sum score is greater than a preset score threshold (specific value is not limited), it is determined that the target invoice is not a fake invoice, otherwise, when the value of the weighted sum score is not greater than the preset score threshold (specific value is not limited), it is determined that the target invoice is not a fake invoice.
[0120] In this way, when identifying whether the target invoice is a fake invoice by executing the steps S101-S106, the following beneficial effects can be achieved:
[0121] (1) The three kinds of multi-modal data of image, text and structured are adopted, which not only solves the problem of single-modal information limitation, but also makes the detection coverage more comprehensive due to the complementarity of multi-modal, and improves the recognition accuracy.
[0122] (2) The unsupervised method does not need prior labels, the core of which is to explore the internal structure and law of data and discover hidden patterns. Since there is no label, the algorithm needs to find the relevance and distribution characteristics in the data by itself, which has wider applicability and solves the problem of insufficient number of fake invoices in the training data set.
[0123] (3) Scanned copies, electronic invoices and the like can be supported as target invoices, which can handle complex scenes such as inclination and blur, and the knowledge graph and semantic analysis can be combined with industry classification standards to verify the rationality of the goods, which can be adapted to multiple fields.
[0124] (4) The judgment result is more accurate by combining the image information of the invoice itself with the identification of enterprises and taxpayers.
[0125] (5) Not only paper invoices and electronic invoices are supported, but also digital electronic invoices are supported for automatic identification, and the recognition accuracy is effectively improved by integrating multiple model algorithms.
[0126] In summary, the virtual invoice identification method provided by the embodiment first acquires a target image in which a target invoice to be identified is located; and extracts text data and structured data of the target invoice from the target image. Then, a first-dimensional image feature vector corresponding to the target image, a second-dimensional semantic feature vector corresponding to the text data of the target invoice, and a third-dimensional structured feature vector corresponding to the structured data of the target invoice are extracted. Next, the first-dimensional image feature vector, the second-dimensional semantic feature vector and the third-dimensional structured feature vector are aligned to obtain an aligned joint feature vector; and then the joint feature vector is input into an unsupervised anomaly detection model to identify the target invoice, and a virtual invoice identification result corresponding to the target invoice is obtained. Thus, the joint analysis of the image data, the text data and the structured data corresponding to the target invoice can more accurately identify whether a virtual invoice occurs, achieving an ideal identification effect.
[0127] Referring toFigure 2 As shown, the application also provides a virtual invoice identification device embodiment, which can include:
[0128] The acquisition unit 201 is configured to acquire a target image in which a target invoice to be identified is located, and extract text data and structured data of the target invoice from the target image.
[0129] The first extraction unit 202 is configured to extract a first-dimension image feature vector corresponding to the target image.
[0130] The second extraction unit 203 is configured to extract a second-dimension semantic feature vector corresponding to the text data of the target invoice.
[0131] The third extraction unit 204 is configured to extract a third-dimension structured feature vector corresponding to the structured data of the target invoice.
[0132] The alignment unit 205 is configured to perform alignment processing on the first-dimension image feature vector, the second-dimension semantic feature vector, and the third-dimension structured feature vector, to obtain an aligned joint feature vector.
[0133] The identification unit 206 is configured to input the joint feature vector into an unsupervised anomaly detection model, identify the target invoice, and obtain a virtual invoice identification result corresponding to the target invoice.
[0134] In some possible implementation manners of the application, the first extraction unit 202 includes:
[0135] The correction sub-unit is configured to correct the target image to obtain a corrected target image.
[0136] The first extraction sub-unit is configured to input the corrected target image into a pre-trained residual network model to extract image features, and obtain a first-dimension image feature vector.
[0137] In some possible implementation manners of the application, the second extraction unit 203 includes:
[0138] The second extraction sub-unit is configured to extract goods name field information in the text data of the target invoice by using an optical character recognition (OCR) algorithm.
[0139] The generation sub-unit is configured to generate a second-dimension dynamic word vector corresponding to the goods name field information by using a pre-trained language model and a word segmenter, as a second-dimension semantic feature vector corresponding to the text data of the target invoice.
[0140] In some possible implementation manners of the present application, the image feature vector of the first dimension is a 512-dimensional image feature vector; the semantic feature vector of the second dimension is a 768-dimensional semantic feature vector; the structured feature vector of the third dimension is a 128-dimensional structured feature vector; the alignment unit 205 comprises:
[0141] a mapping subunit configured to map the 512-dimensional image feature vector and the 768-dimensional semantic feature vector to a unified 256-dimensional vector space to obtain an aligned image-text feature vector;
[0142] an alignment subunit configured to perform projection alignment processing on the aligned image-text feature vector and the 128-dimensional structured feature vector to obtain an aligned joint feature vector.
[0143] In some possible implementation manners of the present application, the recognition unit 206 comprises:
[0144] an input subunit configured to input the joint feature vector into an unsupervised anomaly detection model to obtain an anomaly score of hierarchical density clustering and an anomaly score of independent forest;
[0145] a recognition subunit configured to perform weighted summation processing on the anomaly score of hierarchical density clustering and the anomaly score of independent forest to obtain a weighted summation score; and perform recognition on the target invoice by using the weighted summation score to obtain a virtual invoice recognition result corresponding to the target invoice.
[0146] As can be seen from the above embodiments, the virtual invoice recognition apparatus provided by the embodiments of the present application first acquires a target image in which a target invoice to be recognized is located; and extracts text data and structured data of the target invoice from the target image. Then, an image feature vector of a first dimension corresponding to the target image, a semantic feature vector of a second dimension corresponding to the text data of the target invoice, and a structured feature vector of a third dimension corresponding to the structured data of the target invoice are extracted. Next, the image feature vector of the first dimension, the semantic feature vector of the second dimension, and the structured feature vector of the third dimension are aligned to obtain an aligned joint feature vector; and the joint feature vector is input into an unsupervised anomaly detection model to recognize the target invoice, thereby obtaining a virtual invoice recognition result corresponding to the target invoice. Thus, the joint analysis on the image data, the text data, and the structured data corresponding to the target invoice can be performed to more accurately recognize whether the virtual invoice appears or not, thereby achieving an ideal recognition effect.
[0147] Further, the embodiments of the present application also provide a virtual invoice recognition device, comprising: a processor, a memory, a system bus;
[0148] The processor and the memory are connected through the system bus;
[0149] The memory is configured to store one or more programs, the one or more programs including instructions that, when executed by the processor, cause the processor to perform any of the above-mentioned virtual invoice identification method implementation methods.
[0150] Further, the embodiments of the present application also provide a computer readable storage medium, wherein instructions are stored, and when the instructions run on a terminal device, the terminal device executes any of the above-mentioned virtual invoice identification method implementation methods.
[0151] From the above description of the embodiments, those skilled in the art can clearly understand that all or part of the steps of the above-mentioned embodiment methods can be implemented by means of software plus necessary universal hardware platforms. Based on such an understanding, the technical solutions of the present application can be embodied in the form of a software product, and the computer software product can be stored in a storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions to make a computer device (which can be a personal computer, a server, or a network communication device such as a media gateway, etc.) execute the methods described in the various embodiments or some parts of the embodiments of the present application.
[0152] It should be noted that the various embodiments in the specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments, and the same or similar parts of each embodiment can be mutually referred to. For the device disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and the related parts are referred to the method part.
[0153] It should also be noted that in this document, the relationship terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply that there is any such actual relationship or order between these entities or operations. Moreover, the terms "include", "contain" or any other variants thereof are intended to cover non-exclusive inclusion, so that the processes, methods, articles or devices including a series of elements not only include those elements, but also include other elements not explicitly listed or inherent to such processes, methods, articles or devices. Without more limitations, the element defined by the statement "including a" does not exclude the presence of another identical element in the process, method, article or device including the element.
[0154] The foregoing description of the disclosed embodiments enables a person skilled in the art to make or use the application. Modifications of these embodiments will occur to persons of skill in the art, and that the generic principles defined herein can be applied to other embodiments without departing from the spirit or scope of the application. Therefore, the present application is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for identifying false invoices, characterized in that: include: Obtain a target image containing a target invoice to be identified; and extracting text data and structured data of the target invoice from the target image; Extracting an image feature vector of a first dimension corresponding to the target image; Extracting a semantic feature vector of a second dimension corresponding to the text data of the target invoice; Extracting a structured feature vector of a third dimension corresponding to the structured data of the target invoice; Aligning the image feature vector of the first dimension, the semantic feature vector of the second dimension, and the structural feature vector of the third dimension to obtain an aligned joint feature vector; The joint feature vector is input into an unsupervised anomaly detection model to identify the target invoice, and a false invoice identification result corresponding to the target invoice is obtained.
2. The method according to claim 1, characterized in that The extracting the image feature vector of the first dimension corresponding to the target image includes: Correcting the target image to obtain a corrected target image; The corrected target image is input into a pre-trained residual network model to extract image features to obtain an image feature vector of the first dimension.
3. The method according to claim 1, characterized in that The step of extracting the semantic feature vector of the second dimension corresponding to the text data of the target invoice includes: Extracting the goods name field information from the text data of the target invoice using an optical character recognition (OCR) algorithm; Using a pre-trained language model and a word segmenter, a dynamic word vector of the second dimension corresponding to the commodity name field information is generated as a semantic feature vector of the second dimension corresponding to the text data of the target invoice.
4. The method according to claim 1, wherein The image feature vector of the first dimension is a 512-dimensional image feature vector; the semantic feature vector of the second dimension is a 768-dimensional semantic feature vector; the structured feature vector of the third dimension is a 128-dimensional structured feature vector; and the image feature vector of the first dimension, the semantic feature vector of the second dimension, and the structured feature vector of the third dimension are aligned to obtain an aligned joint feature vector, including: Mapping the 512-dimensional image feature vector and the 768-dimensional semantic feature vector to a unified 256-dimensional vector space to obtain an aligned image-text feature vector; The aligned image-text feature vector and the 128-dimensional structured feature vector are subjected to projection alignment processing to obtain an aligned joint feature vector.
5. The method according to any one of claims 1 to 4, characterized in that Inputting the joint feature vector into an unsupervised anomaly detection model to identify the target invoice and obtain a false invoice identification result corresponding to the target invoice includes: Inputting the joint feature vector into an unsupervised anomaly detection model to obtain an anomaly score of hierarchical density clustering and an anomaly score of independent forests; The anomaly score of the hierarchical density clustering and the independent forest anomaly score are weighted and summed to obtain a weighted sum score; and the weighted sum score is used to identify the target invoice to obtain a false invoice identification result corresponding to the target invoice.
6. A device for identifying false invoices, characterized in that: include: An acquisition unit, configured to acquire a target image containing a target invoice to be identified; and extracting text data and structured data of the target invoice from the target image; A first extraction unit, configured to extract an image feature vector of a first dimension corresponding to the target image; A second extraction unit, configured to extract a semantic feature vector of a second dimension corresponding to the text data of the target invoice; a third extraction unit, configured to extract a structured feature vector of a third dimension corresponding to the structured data of the target invoice; an alignment unit, configured to align the image feature vector of the first dimension, the semantic feature vector of the second dimension, and the structural feature vector of the third dimension to obtain an aligned joint feature vector; The identification unit is used to input the joint feature vector into an unsupervised anomaly detection model, identify the target invoice, and obtain a false invoice identification result corresponding to the target invoice.
7. The device according to claim 6, characterized in that The first extraction unit includes: a correction subunit, configured to correct the target image to obtain a corrected target image; The first extraction subunit is used to input the corrected target image into a pre-trained residual network model to extract image features and obtain an image feature vector of a first dimension.
8. The device according to claim 6, characterized in that The second extraction unit includes: A second extraction subunit is configured to extract the commodity name field information from the text data of the target invoice using an optical character recognition (OCR) algorithm; The generating subunit is used to generate a dynamic word vector of the second dimension corresponding to the commodity name field information by using a pre-trained language model and a word segmenter, as a semantic feature vector of the second dimension corresponding to the text data of the target invoice.
9. A device for identifying false invoices, characterized in that: include: Processor, memory, system bus; The processor and the memory are connected via the system bus; The memory is configured to store one or more programs, wherein the one or more programs include instructions, and when the instructions are executed by the processor, the processor is enabled to perform the method according to any one of claims 1 to 5.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores instructions, and when the instructions are executed on a terminal device, the terminal device executes the method according to any one of claims 1 to 5.