Identify key-value pairs in a document

Through neural network model and optical character recognition technology, key-value pairs in unstructured documents are automatically identified, solving the inefficiency and error rate problems of manually extracting structured data in the prior art, and achieving high-accuracy automatic document conversion.

CN114072857BActive Publication Date: 2025-07-18GOOGLE LLC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202080016688.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-02-27
Filing Date
2020-02-26
Publication Date
2025-07-18
Estimated Expiration
2040-02-26

AI Technical Summary

Technical Problem

In the prior art, manually extracting structured data from unstructured documents is expensive, time-consuming and error-prone, and it is difficult to efficiently convert it into structured key-value pairs automatically.

Method used

Using neural network model and optical character recognition technology, the bounding box in the document image is generated by the detection model, the key-value pair is identified and determined, and the filtering engine is used to verify whether the text data surrounded by the bounding box defines the key-value pair.

Benefits of technology

It realizes automatic conversion of a large number of unstructured documents into structured key-value pairs, with high accuracy (such as more than 99%), suitable for financial document processing that requires high accuracy, and is independent of document style and content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114072857B_ABST
    Figure CN114072857B_ABST
Patent Text Reader

Abstract

Methods, systems, and apparatuses for converting unstructured documents into structured key-value pairs, including computer programs encoded on computer storage media. In one aspect, a method includes: providing an image of a document to a detection model, where: the detection model is configured to process the image to generate an output defined as one or more bounding boxes generated for the image; and each bounding box generated for the image is predicted to enclose a key-value pair including key text data and value text data, where the key text data defines a label characterizing the value text data; and for each of the one or more bounding boxes generated for the image: using optical character recognition technology to identify the text data enclosed by the bounding box; and determining whether the text data enclosed by the bounding box defines a key-value pair.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to document processing. Background Art

[0002] Understanding documents (such as invoices, payment stubs, purchase receipts, etc.) is an important business requirement for many modern enterprises. Most of the enterprise data (such as 90% or more) is stored and represented in the form of unstructured documents. Manually extracting structured data from documents can be expensive, time-consuming, and error-prone. Summary of the Invention

[0003] This specification generally describes a parsing system and a parsing method implemented as a computer program on one or more computers in one or more locations, which automatically convert an unstructured document into structured key-value pairs. More specifically, the parsing system is configured to process a document to identify "key" text data and corresponding "value" text data in the document. Broadly speaking, a key defines a label that characterizes (i.e., describes) the corresponding value. For example, the key "Date" can correspond to the value "2-23-2019".

[0004] According to a first aspect, there is provided a method executed by one or more data processing devices, the method comprising: providing an image of a document to a detection model, wherein: the detection model is configured to process the image according to values of a plurality of detection model parameters to generate an output defined as one or more bounding boxes generated for the image; and each bounding box generated for the image is predicted to enclose a key-value pair including key text data and value text data, wherein the key text data defines a label characterizing the value text data; and for each of the one or more bounding boxes generated for the image: using optical character recognition technology to identify the text data enclosed by the bounding box; determining whether the text data enclosed by the bounding box defines a key-value pair; and in response to determining that the text data enclosed by the bounding box defines a key-value pair, providing the key-value pair for use in characterizing the document.

[0005] In some implementations, the detection model is a neural network model.

[0006] In some implementations, the neural network model includes a convolutional neural network.

[0007] In some implementations, the neural network model is trained on a set of training examples, each training example including a training input and a target output, the training input including a training image of a training document, and the target output including data defining one or more bounding boxes that respectively enclose corresponding key-value pairs in the training image.

[0008] In some implementations, the document is an invoice.

[0009] In some implementations, providing an image of a document to a detection model includes: identifying a specific category of the document; and providing the image of the document to the detection model, which is trained to process documents of the specific category.

[0010] In some implementations, determining whether text data enclosed by a bounding box defines a key-value pair includes: determining that the text data enclosed by the bounding box includes a key from a predetermined set of valid keys; identifying the type of a portion of the text data enclosed by the bounding box that does not include a key; identifying a set of one or more valid types of values corresponding to the key; and determining that the type of the portion of the text data enclosed by the bounding box that does not include a key is included in the set of one or more valid types of values corresponding to the key.

[0011] In some implementations, identifying a set of one or more valid types of values corresponding to a key includes: using a predetermined mapping to map the key to the set of one or more valid types of values corresponding to the key.

[0012] In some implementations, the set of valid keys and the mapping of the corresponding set from keys to valid types of values corresponding to the keys are provided by a user.

[0013] In some implementations, the bounding box has a rectangular shape.

[0014] In some implementations, the method further includes: receiving a document from a user; and converting the document into an image, where the image depicts the document.

[0015] According to another aspect, there is provided a method performed by one or more data processing devices, the method including: providing an image of a document to a detection model configured to process the image to identify, in the image, one or more bounding boxes predicted to enclose key-value pairs including key text data and value text data, where the key defines a label characterizing the value corresponding to the key; for each of the one or more bounding boxes generated for the image: using optical character recognition technology to identify the text data enclosed by the bounding box and determining whether the text data enclosed by the bounding box defines a key-value pair; and outputting the one or more key-value pairs for use in characterizing the document.

[0016] In some implementations, the detection model is a machine learning model having a set of parameters that can be trained on a training data set.

[0017] In some implementations, the machine learning model includes a neural network model, particularly a convolutional neural network.

[0018] In some implementations, a machine learning model is trained on a set of training examples, each training example including a training input and a target output, the training input including a training image of a training document, and the target output including data defining one or more bounding boxes that respectively enclose corresponding key-value pairs in the training image.

[0019] In some implementations, the document is an invoice.

[0020] In some implementations, providing an image of a document to a detection model includes: identifying a specific category of the document; and providing the image of the document to the detection model, which is trained to process documents of the specific category.

[0021] In some implementations, determining whether text data enclosed by a bounding box defines a key-value pair includes: determining that the text data enclosed by the bounding box includes a key from a predetermined set of valid keys; identifying the type of a portion of the text data enclosed by the bounding box that does not include a key; identifying a set of one or more valid types of values corresponding to the key; and determining that the type of the portion of the text data enclosed by the bounding box that does not include a key is included in the set of one or more valid types of values corresponding to the key.

[0022] In some implementations, identifying a set of one or more valid types of values corresponding to a key includes: using a predetermined mapping to map the key to the set of one or more valid types of values corresponding to the key.

[0023] In some implementations, the set of valid keys and the mapping of the set corresponding to valid types of values from the key are provided by a user.

[0024] In some implementations, the bounding box has a rectangular shape.

[0025] In some implementations, the method further includes: receiving a document from a user; and converting the document into an image, where the image depicts the document.

[0026] According to another aspect, there is provided a system including: one or more computers; and one or more storage devices communicatively coupled to the one or more computers, where the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform operations including the operations of the previously described method.

[0027] According to another aspect, there is provided one or more non-transitory computer storage media storing instructions that, when executed by one or more computers, cause the one or more computers to perform operations including the operations of the previously described method.

[0028] Specific embodiments of the subject matter described in this specification can be implemented to realize one or more of the following advantages.

[0029] The systems described in this specification can be used to automatically convert a large number of unstructured documents into structured key-value pairs. Thus, the systems eliminate the need for manual extraction of structured data from unstructured documents, which can be expensive, time-consuming, and error-prone.

[0030] The systems described in this specification can identify key-value pairs in documents at a high level of accuracy (e.g., greater than 99% accuracy for some types of documents). Thus, the systems can be suitable for deployment in applications that require a high level of accuracy (e.g., processing financial documents).

[0031] The systems described in this specification can generalize better than some conventional systems, i.e., have improved generalization ability compared to some conventional systems. In particular, by using a machine learning detection model trained to identify visual signals that distinguish key-value pairs in documents, the systems can accurately identify key-value pairs independent of the specific style, structure, or content of the documents.

[0032] Details of one or more embodiments of the subject matter of this specification are set forth in the accompanying drawings and the following description. Other features, aspects, and advantages of the subject matter will become apparent from the description, the drawings, and the claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 Shows an example parsing system.

[0034] Figure 2A Illustrates an example of an invoice document that can be provided to the parsing system.

[0035] Figure 2B Illustrates the bounding boxes generated by the detection model of the parsing system.

[0036] Figure 2C Illustrates the keys and values enclosed by the bounding boxes.

[0037] Figure 2D Illustrates the key-value pairs identified by the parsing system in the invoice.

[0038] Figure 3 Is a flowchart of an example process for identifying key-value pairs in a document.

[0039] Figure 4 Is a block diagram of an example computer system.

[0040] Like reference numerals and names in the different drawings indicate like elements. DETAILED DESCRIPTION

[0041] Figure 1Disclosed is an example parsing system 100. The parsing system 100 is an example of a system implemented as a computer program on one or more computers in one or more locations, in which the following systems, components, and technologies are implemented.

[0042] The parsing system 100 is configured to process a document 102 (e.g., an invoice, a payment stub, or a purchase receipt) to identify one or more key-value pairs 104 in the document 102. A "key-value pair" refers to a key and a corresponding value, both of which are typically text data. "Text data" should be understood to refer to at least: alphabetic characters, numbers, and special symbols. As described earlier, the key defines a label characterizing the corresponding value. Figures 2A to 2D An example of a key-value pair in an invoice document is shown.

[0043] The system 100 can receive the document 102 in any of a variety of ways. For example, the system 100 can receive the document 102 as an upload from a remote user of the system 100 via a data communication network (e.g., using an application programming interface (API) provided by the system 100). The document 102 can be represented in any suitable unstructured data format, e.g., as a Portable Document Format (PDF) document or as an image document (e.g., a Portable Network Graphics (PNG) or Joint Photographic Experts Group (JPEG) document).

[0044] The system 100 uses a detection model 106, an Optical Character Recognition (OCR) engine 108, and a filtering engine 110 to identify the key-value pairs 104 in the document 102.

[0045] The detection model 106 is configured to process an image 112 of the document 102 to generate an output defining one or more bounding boxes 114 in the image 112, each bounding box being predicted to enclose text data defining a corresponding key-value pair. That is, each bounding box 114 is predicted to enclose text data defining: (i) a key, and (ii) a value corresponding to the key. For example, a bounding box can enclose the text data "Name: John Smith", which defines the key "Name" and the corresponding value "John Smith". The detection model 106 can be configured to generate bounding boxes 114 each enclosing a single key-value pair (i.e., rather than multiple key-value pairs).

[0046] The image 112 of document 102 is an ordered set of values representing the visual appearance of document 102. For example, image 112 can be a black and white image of the document. In this example, image 112 can be represented as a two-dimensional array of numerical intensity values. As another example, the image can be a color image of the document. In this example, image 112 can be represented as a multi-channel image where each channel corresponds to a respective color (e.g., red, green, or blue), and image 112 can be represented as a two-dimensional array of numerical intensity values.

[0047] The bounding box 114 can be a rectangular bounding box. The rectangular bounding box can be represented by the coordinates of specific corners of the bounding box and the corresponding width and height of the bounding box. More generally, other bounding box shapes and other ways of representing the bounding box are possible.

[0048] Although the detection model 106 can implicitly identify and use any frame or boundary present in document 102 as a visual signal, the bounding box 114 is not limited to aligning (i.e., coinciding) with any existing frame of the boundaries present in document 102. Additionally, system 100 can generate the bounding box 114 without visually displaying the bounding box 114 in the image 112 of document 102. That is, system 100 can generate the data defining the bounding box without displaying a visual indication of the position of the bounding box to the user of system 100.

[0049] The detection model 106 is generally a machine learning model, i.e., a model having a set of parameters that can be trained on a training data set. The training data includes multiple training examples, each of which includes: (i) a training image depicting a training document, and (ii) a target output defining one or more bounding boxes each enclosing a respective key-value pair in the training image. The training data can be generated by manual annotation, i.e., by a person manually identifying the bounding boxes around the key-value pairs in the training document (e.g., using appropriate annotation software).

[0050] Using machine learning techniques to train the detection model 106 on the training data set enables the detection model 106 to implicitly identify the visual signals that enable it to recognize the key-value pairs in the document. For example, the detection model 106 can be trained to implicitly identify both local signals (e.g., the text style and relative spatial position of words) and global signals (e.g., the presence of boundaries in the document) that enable it to recognize the key-value pairs. The visual signals that enable the detection model to recognize the key-value pairs in the document generally do not include signals representing the explicit meaning of the words in the document.

[0051] The training detection model 106 is trained to implicitly identify visual signals that distinguish key-value pairs in a document such that the detection model can "generalize" beyond the training data used to train the detection model. That is, if a document is not included in the training data used to train the detection model 106, the trained detection model 106 can also process an image depicting the document to accurately generate bounding boxes that enclose the key-value pairs in the document.

[0052] In one example, the detection model 106 can be a neural network object detection model (e.g., including one or more convolutional neural networks), where the "objects" correspond to key-value pairs in a document. The trainable parameters of the neural network model include the weights of the neural network model, e.g., the weights that define the convolutional filters in the neural network model.

[0053] A suitable machine learning training procedure such as stochastic gradient descent can be used to train the neural network model on a training data set. In particular, at each iteration of multiple training iterations, the neural network model can process a "batch" (i.e., a set) of training images from a training example to generate bounding boxes that are predicted to enclose the corresponding key-value pairs in the training images. The system 100 can evaluate an objective function that represents a measure of the similarity between the bounding boxes generated by the neural network model and the bounding boxes specified by the target output of the corresponding training example. A measure of the similarity between two bounding boxes can be, for example, the sum of the squared distances between the corresponding vertices of the bounding boxes. The system can determine the gradient of the objective function with respect to the neural network parameter values (e.g., using backpropagation), and thereafter use the gradient to adjust the current neural network parameter values. In particular, the system 100 can use the parameter update rule from any suitable gradient descent optimization algorithm (e.g., Adam or RMSprop) to use the gradient to adjust the current neural network parameter values. The system 100 trains the neural network model until a training termination criterion is met (e.g., until a predetermined number of training iterations have been performed, or until the change in the value of the objective function between training iterations drops below a predetermined threshold).

[0054] Before using the detection model 106, the system 100 can identify the "category" of the document 102 (e.g., invoice, payment stub, or purchase receipt). For example, a user of the system 100 can identify the category of the document 102 when providing the document to the system 100. As another example, the system 100 can use a classification neural network to automatically classify the category of the document 102. As another example, the system 100 can use OCR technology to identify the text in the document 102 and then identify the category of the document 102 based on the text in the document 102. In a specific example, in response to identifying the phrase "Net Pay", the system 100 can identify the category of the document 102 as "pay stub". In another specific example, in response to identifying the phrase "Sales tax", the system 100 can identify the category of the document 102 as "invoice". After identifying a specific category of the document 102, the system 100 can use the detection model 106 that is trained to process documents of the specific category. That is, the system 100 can use the detection model 106 that is trained on training data that includes only documents of the same specific category as the document 102. Using the detection model 106 that is specifically trained to process documents of the same category as the document 102 can improve the performance of the detection model (e.g., by enabling the detection model to generate bounding boxes around key-value pairs with greater accuracy).

[0055] For each of the bounding boxes 114, the system 100 uses the OCR engine 108 to process the portion of the image 112 enclosed by the bounding box to identify the text data (i.e., the text 116) enclosed by the bounding box. In particular, the OCR engine 108 identifies the text 116 enclosed by the bounding box by identifying each letter, numerical value, or special character enclosed by the bounding box. The OCR engine 108 is capable of using any suitable OCR technology to identify the text 116 enclosed by the bounding box 114.

[0056] The filtering engine 110 is configured to determine whether the text 116 enclosed by the bounding box 114 represents a key-value pair. The filtering engine can determine whether the text 116 enclosed by the bounding box 114 represents a key-value pair in any suitable manner. For example, for a given bounding box, the filtering engine 110 can determine whether the text enclosed by the bounding box includes a valid key from a predetermined set of valid keys. For example, the set of valid keys can include: "Date", "Time", "Invoice#", "Amount Due", etc. When comparing different parts of the text to determine whether the text enclosed by the bounding box includes a valid key, the filtering engine 110 can determine that they "match" even if the two parts of the text are not identical. For example, even if the two parts of the text include different capitalizations or punctuation, the filtering engine 110 can determine that they match (e.g., the filtering system 100 can determine that "Date", "Date:", "date", and "date:" all match).

[0057] In response to determining that the text enclosed by the bounding box does not include a valid key from the set of valid keys, the filtering engine 110 determines that the text enclosed by the bounding box does not represent a key-value pair.

[0058] In response to determining that the text enclosed by the bounding box includes a valid key, the filtering engine 110 identifies the "type" (e.g., alphabetical, numerical, temporal) of the portion of the text enclosed by the bounding box that is not recognized as a key (i.e., "non-key" text). For example, for a bounding box enclosing the text: "Date: 2-23-2019", in the case where the filtering engine 110 identifies "Date:" as a key (as described earlier), the filtering engine 110 can identify the type of the non-key text "2-23-2019" as "temporal".

[0059] In addition to identifying the type of the non-key text, the filtering engine 110 also identifies a set of one or more valid types of the value corresponding to the key. In particular, the filtering engine 110 can map the key to a set of valid data types of the value corresponding to the key according to a predetermined mapping. For example, the filtering engine 110 can map the key "Name" to the corresponding value data type "alphabetical", indicating that the value corresponding to the key should have an alphabetical data type (e.g., "JohnSmith"). As another example, the filtering engine 110 can map the key "Date" to the corresponding value data type "temporal", indicating that the value corresponding to the key should have a temporal data type (e.g., "2-23-2019" or "17:30:22").

[0060] The filtering engine 110 determines whether the type of the non-key text is included in the set of valid types of the values corresponding to the key. In response to determining that the type of the non-key text is included in the set of valid types of the values corresponding to the key, the filtering engine 110 determines that the text surrounded by the bounding box represents a key-value pair. In particular, the filtering engine 110 identifies the non-key text as the value corresponding to the key. Otherwise, the filtering engine 110 determines that the text surrounded by the bounding box does not represent a key-value pair.

[0061] The set of valid keys and the mapping from the valid keys to the set of valid data types of the values corresponding to the valid keys can be provided by the user of the system 100 (e.g., via an API provided by the system 100).

[0062] After using the filtering engine 110 to identify the key-value pairs 104 from the text 116 surrounded by the respective bounding boxes 114, the system 100 outputs the identified key-value pairs 104. For example, the system 100 can provide the key-value pairs 104 to a remote user of the system 100 via a data communication network (e.g., using an API provided by the system 100). As another example, the system 100 can store the data defining the identified key-value pairs in a database (or other data structure) accessible to the user of the system 100.

[0063] In some cases, the user of the system 100 can request the system 100 to identify the value corresponding to a specific key (e.g., "Invoice#") in a document. In these cases, instead of identifying and providing every key-value pair in the document, the system 100 can process the text 116 identified in the respective bounding boxes 114 until the requested key-value pair is identified, and thereafter output the requested key-value pair.

[0064] As described above, the detection model 106 can be trained to generate bounding boxes each surrounding a respective key-value pair. Alternatively, instead of using a single detection model 106, the system 100 can include: (i) a "key detection model" trained to generate bounding boxes surrounding the respective keys, and (ii) a "value detection model" trained to generate bounding boxes surrounding the respective values. The system 100 can identify key-value pairs from the key bounding boxes and the value bounding boxes in any suitable manner. For example, for each pair of a key bounding box and a value bounding box, the system 100 can generate a "matching score" based on: (i) the spatial proximity of the bounding boxes, (ii) whether the key bounding box surrounds a valid key, and (iii) whether the type of the value surrounded by the value bounding box is included in the set of valid types of the values corresponding to the key. If the matching score between the key bounding box and the value bounding box exceeds a threshold, the system 100 can identify the key surrounded by the key bounding box and the value surrounded by the value bounding box as a key-value pair.

[0065] Figure 2AAn example of an invoice document 200 is shown. The parsing system 100 (refer to Figure 1 A user (described in FIG. 1 ) may provide an invoice 200 (eg, as a scanned image or PDF file) to the parsing system 100 .

[0066] Figure 2B The bounding boxes (e.g., 202, 204, 206, 208, 210, 212, 214, and 216) generated by the detection model 106 of the parsing system 100 are illustrated. Each of the bounding boxes is predicted to enclose text data that defines a key-value pair. The detection model 106 does not generate a bounding box to enclose text 218 (i.e., "Thank you for your business!") because this text does not represent a key-value pair. As shown in FIG. Figure 1 As described, the parsing system 100 uses OCR technology to identify the text inside each bounding box, and then identifies the valid key-value pairs enclosed by the bounding box.

[0067] Figure 2C Illustrated is a key 220 (ie, “Date:”) and a value 222 (ie, “2-23-2019”) enclosed by a bounding box 202 .

[0068] Figure 2D The key-value pairs identified by the parsing system 100 in the invoice 200 are illustrated.

[0069] Figure 3 is a flow chart of an example process 300 for identifying key-value pairs in a document. For convenience, process 300 will be described as being performed by a system of one or more computers located in one or more locations. For example, a parsing system appropriately programmed according to the present specification, such as Figure 1 The parsing system 100 is capable of executing the process 300.

[0070] The system receives a document (302). For example, the system can receive the document as an upload from a remote user of the system via a data communications network (e.g., using an API provided by the system). The document can be represented in any suitable unstructured data format, for example, as a PDF document or an image document (e.g., a PNG or JPEG document).

[0071] The system converts the document into an image, ie, an ordered collection of numerical values representing the visual appearance of the document (304). For example, the image may be a black and white image of the document represented as a two-dimensional array of numerical intensity values.

[0072] The system provides an image of a document to a detection model that is configured to process the image according to a set of detection model parameters to generate an output (306) that defines one or more bounding boxes in the image of the document. Each bounding box is predicted to enclose a key-value pair that includes key text data and value text data, where the key defines a label characterizing the value. The detection model can be a neural network object detection model that includes one or more convolutional neural networks.

[0073] Steps 308 - 310 are performed for each bounding box in the image of the document. For convenience, steps 308 - 310 are described with reference to a given bounding box.

[0074] The system uses optical character recognition (OCR) technology to identify the text data enclosed by the bounding box (308). In particular, the system uses OCR technology to identify each letter, numerical value, or special character enclosed by the bounding box.

[0075] The system determines whether the text data enclosed by the bounding box defines a key-value pair (310). For example, the system can determine whether the text enclosed by the bounding box includes a valid key from a predetermined set of valid keys. In response to determining that the text enclosed by the bounding box does not include a valid key from the set of valid keys, the system determines that the text enclosed by the bounding box does not represent a key-value pair. In response to determining that the text enclosed by the bounding box includes a valid key, the system identifies the "type" (e.g., letter, numerical value, time, or a combination thereof) of the portion of the text enclosed by the bounding box that is not recognized as a key (i.e., "non-key" text). In addition to identifying the type of the non-key text, the system also identifies a set of one or more valid types of the value corresponding to the key. The system determines whether the type of the non-key text is included in the set of valid types of the value corresponding to the key. In response to determining that the type of the non-key text is included in the set of valid types of the value corresponding to the key, the system determines that the text enclosed by the bounding box represents a key-value pair. In particular, the system identifies the non-key text as the value corresponding to the key. Otherwise, the system determines that the text enclosed by the bounding box does not represent a key-value pair.

[0076] The system provides the identified key-value pairs for use in characterizing the document (312). For example, the system can provide the key-value pairs to a remote user of the system via a data communication network (e.g., using an API provided by the system).

[0077] Figure 4FIG. 400 is a block diagram of an example computer system 400 that can be used to perform the operations described previously. System 400 includes a processor 410, a memory 420, a storage device 430, and an input / output device 440. Each of the components 410, 420, 430, and 440 can be interconnected, for example, using a system bus 450. The processor 410 is capable of processing instructions executed within the system 400. In one embodiment, the processor 410 is a single-threaded processor. In another embodiment, the processor 410 is a multi-threaded processor. The processor 410 is capable of processing instructions stored in the memory 420 or on the storage device 430.

[0078] The memory 420 stores information within the system 400. In one embodiment, the memory 420 is a computer-readable medium. In one embodiment, the memory 420 is a volatile memory unit. In another embodiment, the memory 420 is a non-volatile memory unit.

[0079] The storage device 430 is capable of providing mass storage for the system 400. In one embodiment, the storage device 430 is a computer-readable medium. In various different embodiments, the storage device 430 can include, for example, a hard disk device, an optical disk device, a storage device shared by multiple computing devices (e.g., a cloud storage device) over a network, or some other mass storage device.

[0080] The input / output device 440 provides input / output operations for the system 400. In one embodiment, the input / output device 440 can include one or more network interface devices, e.g., an Ethernet card, a serial communication device such as an RS-232 port, and / or a wireless interface device such as an 802.11 card. In another embodiment, the input / output device 440 can include a driver device configured to receive input data and send output data to other input / output devices such as a keyboard, a printer, and a display device 460. However, other embodiments can also be used, such as mobile computing devices, mobile communication devices, and set-top box TV client devices, etc.

[0081] Although an example processing system has been described in Figure 4 , embodiments of the subject matter and functional operations described in this specification can be implemented using other types of digital electronic circuitry or using computer software, firmware, or hardware, including the structures disclosed in this specification and their structural equivalents, or using a combination of one or more of them.

[0082] This specification uses the term "configured" in connection with systems and computer program components. For a system of one or more computers to be configured to perform particular operations or actions means that the system has installed thereon software, firmware, hardware, or a combination of software, firmware, and hardware that in operation cause the system to perform those operations or actions. For one or more computer programs to be configured to perform particular operations or actions means that the one or more programs include instructions that, when executed by a data processing apparatus, cause the apparatus to perform the operations or actions.

[0083] Embodiments of the subject matter and the functional operations described in this specification can be implemented in digital electronic circuitry, in tangibly embodied computer software or firmware, in computer hardware including the structures disclosed in this specification and their structural equivalents, or in one or more of them in combination. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible non-transitory storage medium for execution by, or to control the operation of, a data processing apparatus. The computer storage medium can be a machine-readable storage device, a machine-readable storage substrate, a random or serial access storage device, or a combination of one or more of them. Alternatively or in addition, the program instructions can be encoded on an artificially generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to a suitable receiver apparatus for execution by the data processing apparatus.

[0084] The term "data processing apparatus" refers to data processing hardware and includes all kinds of devices, equipment, and machines for processing data, including, by way of example, a programmable processor, a computer, or multiple processors or computers. The apparatus can also be, or further include, special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit). The apparatus can optionally include, in addition to hardware, code that creates an execution environment for computer programs, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them.

[0085] A computer program, which may also be referred to as or described as a program, software, software application, app, module, software module, script, or code, can be written in any form of programming language, including compiled or interpreted languages or declarative or procedural languages; and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. The program may, but need not, correspond to a file in a file system. The program may be stored as part of a file that holds other programs or data, such as one or more scripts in a markup language document; in a single file dedicated to the program in question or in multiple coordinated files, such as files that store one or more modules, subroutines, or portions of code. The computer program may be deployed to execute on one computer or on multiple computers located at one site or distributed across multiple sites and interconnected by a data communication network.

[0086] In this specification, the term "engine" is used broadly to refer to a software-based system, subsystem, or process that is programmed to perform one or more specific functions. Typically, an engine will be implemented as one or more software modules or components installed on one or more computers in one or more locations. In some cases, one or more computers will be dedicated to a particular engine; in other cases, multiple engines may be installed and run on the same computer or on multiple computers.

[0087] The processes and logical flows described in this specification can be performed by one or more programmable computers executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logical flows can also be performed by, for example, a special-purpose logic circuit such as an FPGA or ASIC, or by a combination of a special-purpose logic circuit and one or more programmed computers.

[0088] A computer suitable for executing a computer program may be based on a general-purpose microprocessor, a special-purpose microprocessor, or both, or any other kind of central processing unit. Generally, the central processing unit will receive instructions and data from a read-only memory, a random access memory, or both. Essential elements of a computer are a central processing unit for executing or implementing instructions and one or more storage devices for storing instructions and data. The central processing unit and the memory may be supplemented by, or incorporated in, special-purpose logic circuitry. Generally, a computer will also include one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, or is operatively coupled to receive data from, or transfer data to, the one or more mass storage devices, or both, for storing data. However, a computer need not have such devices. In addition, a computer may be embedded in another device, such as a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game controller, a global positioning system (GPS) receiver, or a portable storage device, such as a universal serial bus (USB) flash drive, etc.

[0089] Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and storage devices, by way of example including semiconductor storage devices, such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD ROM and DVD-ROM disks.

[0090] To provide interaction with a user, embodiments of the subject matter described in this specification may be implemented on a computer having a display device for displaying information to the user and a keyboard and a pointing device by which the user may provide input to the computer, the display device being, for example, a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, and the pointing device being, for example, a mouse or a trackball. Other kinds of devices may also be used to provide interaction with the user; for example, feedback provided to the user may be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback; and input from the user may be received in any form, including acoustic, speech, or tactile input. In addition, a computer may interact with a user by sending documents to and receiving documents from a device used by the user; for example, by sending a web page to a web browser on the user's device in response to a request received from the web browser. Additionally, a computer may interact with a user by sending a text message or other form of message to a personal device and then receiving a response message from the user, the personal device being, for example, a smart phone running a messaging application.

[0091] The data processing apparatus for implementing a machine learning model may also include, for example, a dedicated hardware accelerator unit for processing the common and computationally intensive parts of machine learning training or production, i.e., inference, workloads.

[0092] A machine learning framework can be used to implement and deploy a machine learning model, such as the TensorFlow framework, the Microsoft Cognitive Toolkit framework, the Apache Singa framework, or the Apache MXNet framework.

[0093] Embodiments of the subject matter described in this specification can be implemented in a computing system that includes backend components, such as a data server; or includes middleware components, such as an application server; or includes frontend components, such as a client computer having a graphical user interface, a web browser, or an app that a user can use to interact with an implementation of the subject matter described in this specification; or includes any combination of one or more such backend, middleware, or frontend components. The components of the system can be interconnected by digital data communication in any form or medium, such as a communication network. Examples of communication networks include local area networks (LANs) and wide area networks (WANs), such as the Internet.

[0094] The computing system can include a client and a server. The client and the server are generally far apart from each other and typically interact via a communication network. The relationship between the client and the server arises by virtue of computer programs running on the respective computers and having a client - server relationship with each other. In some embodiments, the server transmits data, such as an HTML page, to a user device, for example, for the purpose of displaying data to a user interacting with the device acting as a client and receiving user input from the user. Data generated at the user device can be received at the server, for example, the result of a user interaction.

[0095] Although this specification contains many details of specific embodiments, these should not be construed as limitations on the scope of any invention or of what may be claimed, but rather as descriptions of features specific to particular embodiments of a particular invention. Certain features described in the context of separate embodiments in this specification can also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment can also be implemented separately or in any suitable sub - combination in multiple embodiments. Additionally, although features may be described above as acting in certain combinations and even initially claimed as such, one or more features from a claimed combination can in some cases be excluded from the combination, and the claimed combination can be directed to a sub - combination or variation of a sub - combination.

[0096] Similarly, although operations are depicted in the drawings and recited in the claims in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous. Additionally, the separation of various system modules and components in the above-described embodiments should not be construed as requiring such separation in all embodiments, and it should be understood that the described program components and systems generally may be integrated together in a single software product or packaged into multiple software products.

[0097] Particular embodiments of the subject matter have been described. Other embodiments are within the scope of the appended claims. For example, the acts recited in the claims may be performed in a different order and still achieve the desired result. As one example, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous.

Claims

1. A method performed by one or more data processing devices, the method comprising: Providing an image of a document to a detection model, wherein: The detection model is configured to process the image according to values of a plurality of detection model parameters to generate an output defining one or more bounding boxes generated for the image; and Each bounding box generated for the image is predicted to enclose a key-value pair including key text data and value text data, wherein the key text data defines a label characterizing the value text data; and For each of the one or more bounding boxes generated for the image: Using optical character recognition technology to identify the text data enclosed by the bounding box; Determining whether the text data enclosed by the bounding box defines a key-value pair by: Determining that the text data enclosed by the bounding box includes a key from a predetermined set of valid keys; Identifying the type of a portion of the text data enclosed by the bounding box that does not include the key; Identifying a set of one or more valid types of values corresponding to the key; and Determining that the type of the portion of the text data enclosed by the bounding box that does not include the key is included in the set of one or more valid types of values corresponding to the key; and In response to determining that the text data enclosed by the bounding box defines a key-value pair, providing the key-value pair for use in characterizing the document.

2. The method according to claim 1, wherein, The detection model is a neural network model.

3. The method according to claim 2, wherein The neural network model includes a convolutional neural network.

4. The method according to claim 2, wherein, Training the neural network model on a set of training examples, each training example including a training input and a target output, the training input including a training image of a training document, and the target output including data defining one or more bounding boxes each enclosing a corresponding key-value pair in the training image.

5. The method according to claim 1, wherein, The document is an invoice.

6. The method according to claim 1, wherein Providing an image of a document to a detection model includes: identifying a specific category of the document; and providing the image of the document to a detection model trained to process documents of the specific category.

7. The method according to claim 1, wherein Identifying a set of one or more valid types of values corresponding to the key includes: Using a predetermined mapping to map the key to the set of one or more valid types of values corresponding to the key.

8. The method according to claim 7, wherein The set of valid keys and the mapping of the correspondence from keys to valid types of values corresponding to the keys are provided by a user.

9. The method according to claim 1, wherein, The bounding box has a rectangular shape.

10. The method according to claim 1, further comprising: Receiving the document from a user; and converting the document into the image, wherein the image depicts the document.

11. A system for identifying key-value pairs in a document, comprising: One or more computers; And One or more storage devices communicatively coupled to the one or more computers, wherein the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform operations including the following: Providing an image of a document to a detection model, wherein: The detection model is configured to process the image according to values of a plurality of detection model parameters to generate an output defining one or more bounding boxes generated for the image; and Each bounding box generated for the image is predicted to enclose a key-value pair including key text data and value text data, wherein the key text data defines a label characterizing the value text data; and For each of the one or more bounding boxes generated for the image: Use optical character recognition technology to identify the text data enclosed by the bounding box; Determine whether the text data enclosed by the bounding box defines a key-value pair by: Determining that the text data enclosed by the bounding box includes a key from a predetermined set of valid keys; Identifying the type of a portion of the text data enclosed by the bounding box that does not include the key; Identifying a set of one or more valid types of values corresponding to the key; and Determining that the type of the portion of the text data enclosed by the bounding box that does not include the key is included in the set of one or more valid types of values corresponding to the key; and In response to determining that the text data enclosed by the bounding box defines a key-value pair, provide the key-value pair for use in characterizing the document.

12. The system according to claim 11, wherein, The detection model is a neural network model.

13. The system according to claim 12, wherein, The neural network model includes a convolutional neural network.

14. The system according to claim 12, wherein, The neural network model is trained on a set of training examples, each training example including a training input and a target output, the training input including a training image of a training document, and the target output including data defining one or more bounding boxes each enclosing a respective key-value pair in the training image.

15. The system according to claim 11, wherein The document is an invoice.

16. A non-transitory computer storage medium storing instructions that, when executed by one or more computers, cause the one or more computers to perform operations including the following: Provide an image of a document to a detection model, wherein: The detection model is configured to process the image according to values of a plurality of detection model parameters to generate an output defining one or more bounding boxes generated for the image; and Each bounding box generated for the image is predicted to enclose a key-value pair including key text data and value text data, wherein the key text data defines a label characterizing the value text data; and For each of the one or more bounding boxes generated for the image: Use optical character recognition technology to identify the text data enclosed by the bounding box; Determine whether the text data enclosed by the bounding box defines a key-value pair by: Determining that the text data enclosed by the bounding box includes a key from a predetermined set of valid keys; Identifying the type of a portion of the text data enclosed by the bounding box that does not include the key; Identifying a set of one or more valid types of values corresponding to the key; and Determining that the type of the portion of the text data enclosed by the bounding box that does not include the key is included in the set of one or more valid types of values corresponding to the key; and In response to determining that the text data enclosed by the bounding box defines a key-value pair, provide the key-value pair for use in characterizing the document.

17. The non-transitory computer storage medium according to claim 16, wherein, The detection model is a neural network model.

18. The non-transitory computer storage medium according to claim 17, wherein, The neural network model includes a convolutional neural network.

19. The non-transitory computer storage medium according to claim 17, wherein, Train the neural network model on a set of training examples, each training example including a training input and a target output, the training input including a training image of a training document, and the target output including data defining one or more bounding boxes that respectively enclose corresponding key-value pairs in the training image.

Citation Information

Patent Citations

  • Systems and methods for generating and using semantic images in deep learning for classification and data extraction

    US20190050639A1