Systems and method for generating labelled datasets

By using pre-trained models to generate numerical representations and cluster documents based on similarity, the method effectively categorizes diverse document types and enhances model performance by identifying areas for retraining, addressing classification challenges and vendor diversity.

AU2021428224B2Pending Publication Date: 2026-07-23XERO
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
AU2021428224
Authority / Receiving Office
AU · AU
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-02-18
Filing Date
2021-08-19
Publication Date
2026-07-23
Estimated Expiration
2041-08-19

Smart Images

  • Figure 00000001_0000
    Figure 00000001_0000
  • Figure 00000034_0000
    Figure 00000034_0000
  • Figure 00000035_0000
    Figure 00000035_0000
Patent Text Reader

Abstract

A method comprises determining a plurality of documents and for each document of the plurality of documents: (i) providing the document to a numerical representation generation model; (ii) generating, by the numerical representation generation model, a numerical representation of the document; and (iii) determining a document score for the document based on the numerical representation. The method further comprises c) providing the document scores to a clustering module; d) determining, by the clustering module, one or more clusters, each cluster being associated with a class of the documents; e) outputting, by the clustering module, a cluster identifier indicative of the class of each document; and f) associating each document with its respective cluster identifier.
Need to check novelty before this filing date? Find Prior Art

Description

[87] Consider a dataset of example documents which may include one or more different types of documents, such as invoices, debit notes, credit notes, receipts etc. In some cases, the dataset of example documents may be associated or originate from vendor #A, but in other cases, the dataset may include example documents from different vendors.

[88] For each document, use a pre-trained model to determine a numerical representation of the document. For example, an input to the model may be an image of the document, character or text data from the document, or both. Where image data is provided to the model, in some embodiments, Gaussian blurring may be performed prior to blur the image data prior to providing it to the model.

[89] Determine a document score for each document. The document score may be the numerical representation (which may be a multi-dimensional vector), or may be a dimensionally reduced numerical representation and / or may be a single value. 2021428224   04 Jun 2026

[90] Cluster the document scores to identify document groups within the dataset with similar document scores. The documents with similar document scores are considered as being documents with common or similar attributes, such as class of document. Allocate each cluster or document group a cluster identifier, which is indicative of the attributes, and accordingly of the class or type of document.

[91] A second use case relates to performing the method 300 of Figure 3 to identify subsets of training data (for example, document classes) with which a model has low confidence in predicting attributes of the training data.

[92] Consider a dataset of example documents which may include one or more different types of documents, such as invoices, debit notes, credit notes, receipts etc. In some cases, the dataset of example documents may be associated or originate from vendor #A, but in other cases, the dataset may include example documents from different vendors.

[93] For each document, use a pre-trained model to determine confidence scores in making predictions about attributes of the document. For example, the pre-trained model may be configured to receive, as an input, the document and to provide, as an output, a vector of confidence scores associated with the confidence or certainty the model has in being able to accurately extract respective attributes from the document.

[94] Determine a document score for each document. The document score may be a vector of the confidence scores, or may be a dimensionally reduced vector of confidence scores or may be a single value.

[95] Cluster the document scores to identify document groups within the dataset with similar document scores, and accordingly, where confidence about the predictions is similar. Allocate each cluster or document group a cluster identifier, which is indicative of an attribute such as the class of the document. Documents for which the model has similar confidence in predicting its attributes are likely to have similar documents scores, and accordingly to be grouped together in the same cluster. As the 2021428224   04 Jun 2026 model is likely to determine or generate similar confidence scores and accordingly document scores for similar classes of document, the document groups or clusters will be indicative of documents with similar attributes, such as types of documents (e.g., invoices, receipts, credit notes etc.) or categories of documents (e.g. financial documents, advertisements, personal correspondence). Furthermore, clusters with relatively low document scores, which are based on the confidence scores, are indicative of documents that the model is underperforming or struggling with. Accordingly, these clusters can inform which documents the model needs to be trained or retrained on to improve the overall performance of the model.

[96] For either use case, and depending on the nature of the documents to be clustered, and / or the features used to generate the numerical representation, and accordingly, the document score, the described processes may be capable of distinguishing between documents of different entities, documents of different class, type, documents associated with a broader or more general classification or category, such as financial documents, advertisements, personal correspondence etc. For example, each category may comprise one or more types of document associated with the category; financial documents may include documents types receipt, invoice, credit note etc. In some embodiments, the class may be invariant to the vendor or entity who issued the document, or a country of origin etc.

[97] It will be appreciated by persons skilled in the art that numerous variations and / or modifications may be made to the above-described embodiments, without departing from the broad general scope of the present disclosure. The present embodiments are, therefore, to be considered in all respects as illustrative and not restrictive.

Claims

1. A method comprising:a) determining a plurality of documents, wherein each document comprises image data;b) for each document:i) applying a filter to the document to blur image data of the document and produce a filtered document;ii) providing the filtered document to a numerical representation generation model;iii) generating, by the numerical representation generation model,a numerical representation of the filtered document; andiv) determining a document score for the filtered document based on the numerical representation,wherein the document score comprises a confidence score associated with a confidence of the numerical representation generation model;c) providing the document scores for the documents to a clustering module;d) determining, by the clustering module and based on the document scores, one or more clusters, each cluster being associated with a class of the documents,wherein the clustering model is configured to cluster documents with a similar confidence score;e) outputting, by the clustering module, a cluster identifier indicative of the class of each document; andf) associating each document with its respective cluster identifier.

2. The method of claim 1, wherein each document is associated with acorresponding label of a plurality of labels, and the method further comprises:determining a dataset for each label of the plurality of labels, the dataset comprising the documents associated with the label; andperforming steps b) to e) for each dataset separately.2021428224   04 Jun 20263.      The method of claim 1 or claim 2, wherein the numerical representation is amulti-dimensional vector, and wherein the method further comprises:providing the numerical representation to a dimensionality reduction model to determine the document score.

4. The method of claim 3, wherein the dimensionality reduction model performsPrincipal Component Analysis (PCA) to generate a dimensionally reduced numerical representation of the document, and the method further comprises:multiplying the dimensionally reduced numerical representation by the variance-ratio to determine the document score.

5. The method of any one of the preceding claims, wherein the document scorecomprises one or more confidence scores for corresponding one of more attributes of the document, wherein the one or more attributes comprise: (i) amount; (ii) entity; (iii) due date; (iv) bill date; (v) invoice number; (vi) tax amount; and / or (vii) currency.

6. The method of any one of the preceding claims, wherein one or more of thecluster identifiers are indicative of a low confidence score class of document for which one or more low confidence scores have been allocated and the method further comprises:selecting the documents of the one or more low confidence score classes for label review.

7. The method of any one of the preceding claims, wherein one or more of thecluster identifiers are indicative of a low confidence score class of document for which one or more low confidence scores have been allocated and the method further comprises:retraining a model used to generate the confidence scores using documents from the one or more low confidence score classes of document.2021428224   04 Jun 20268.      The method of any one of the preceding claims, wherein determining, by theclustering module, one or more clusters, comprises:determining a plurality of histogram bins based on the document scores; determining a bin score for each document; andgrouping histogram bins into clusters.

9. The method of claim 8, wherein grouping histogram bins into clusterscomprises determining the clusters as respective local minima of the histogram bins.

10. The method of any one of claims 1 to 8, wherein determining, by theclustering module, one or more clusters, comprises performing k-means clustering, density-based spatial clustering (DBSCAN) or hierarchical clustering.

11. The method of any one of the preceding claims, wherein each documentcomprises character data, and the method further comprising:for each document:providing the character data to an attribute determination model;determining an attribute associated with the document based on the character data; andassociating the document with the determined attribute as the label.

12. The method of any one of the preceding claims, wherein the method furthercomprises:for each document:extracting the image data from the document;providing the image data to an attribute determination model;determining, by the attribute determination module, an attribute associated with the document based on the image data; andassociating the document with the determined attribute as the label.2021428224   04 Jun 202613.     The method of any one of claims 1 to 11, wherein the method furthercomprises:for each document:extracting the image data from the document;providing the image data to an image-based numerical representation generation module;determining by the image-based numerical representation generation module, an image-based numerical representation of the document;providing the character data to a character-based numerical representation generation module;determining by the character-based numerical representation generation module, a character-based numerical representation of the document;providing, to a consolidated numerical representation generation module, the image-based numerical representation of the document and the characterbased numerical representation of the document; andgenerating, by the consolidated numerical representation generation module, a combined numerical representation of the character data and the image data of the document;providing the combined numerical representation to an attribute prediction module; anddetermining, by the attribute prediction module, the attribute associated with the document; andassociating the document with the determined attribute as the label.

14. The method of any one of claims 11 to 13, wherein the attribute is an entityassociated with the document.

15. The method of any one of any one of the preceding claims, wherein thedocuments are derived from previously reconciled accounting documents of an accounting system, each of which has been associated with a respective entity, and wherein the label of each document is indicative of the respective entity.2021428224   04 Jun 202616.     The method of any one of the preceding claims, wherein the document is anaccounting document and the class of documents includes one or more of: (i) an invoice; (ii) a credit note; (iii) a receipt; (iv) a purchase order; and (v) a quote.

17. A system comprising:one or more processors; andmemory comprising computer executable instructions, which when executed by the one or more processors, cause the system to perform the method of any one of claims 1 to 16.

18. A computer-readable storage medium storing instructions that, when executedby a computer, cause the computer to perform the method of any one of claims 1 to 17.

Citation Information

Patent Citations

  • Adaptive Automatic Defect Classification

    US20170082555A1

  • Dynamic Document Clustering and Keyword Extraction

    US20200311414A1