Domain-adaptive graph networks for visually rich documents

The method employs a graph neural network with pre-trained models to classify key-value pairs in visually rich documents, addressing the need for extensive labeled data and improving classification accuracy.

JP2026509741APending Publication Date: 2026-03-25ORACLE INT CORP
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-02-22
Publication Date
2026-03-25

AI Technical Summary

Technical Problem

Existing machine learning models struggle to effectively classify key-value pairs in visually rich documents due to the need for large amounts of labeled data and the challenge of generalizing across various domains, leading to inefficiencies in training and classification accuracy.

Method used

A method involving the use of a graph neural network trained on visually rich documents, utilizing pre-trained language and visual models, along with positional features, to classify text as key-value pairs, requiring only a few labeled documents for effective training.

Benefits of technology

Enables efficient training and accurate classification of key-value pairs in visually rich documents, reducing the need for extensive labeled data and improving model performance across different domains.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026509741000001_ABST
    Figure 2026509741000001_ABST
Patent Text Reader

Abstract

In some embodiments, the techniques described herein may include identifying text in a visually rich document and determining a sequence of identified text. These techniques may include selecting a language model based at least in part on the identified text and the determined sequence. Furthermore, these techniques may include generating textual features corresponding to the identified text by assigning each word of the identified text to its respective token. These techniques may include extracting visual features corresponding to the identified text. These techniques may include determining the spatial features of each word of the identified text. These techniques may also include generating a graph representing the visually rich document, where each node in the graph represents a visual feature, textual feature, and spatial feature, respectively, of each word of the identified text. These techniques may include training a classifier on the graph to classify each of the words of the identified text.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - Reference to Related Applications This application is a regular application of Indian Provisional Patent Application No. 202341013172 filed on February 27, 2023 and US Patent Application No. 18 / 240,480 filed on August 31, 2023 under 35 U.S.C. § 119(e), claiming the benefit and priority thereof, and all of its contents are incorporated herein by reference for all purposes.

[0002] Field This disclosure relates to machine learning techniques. Specifically, it relates to a machine learning model trained to perform key - value extraction.

Background Art

[0003] Background Training a machine learning model to perform key - value extraction from physical documents may involve a large amount of labeled data. In supervised machine learning algorithms, learning patterns may require a significant amount of labeled data with sufficient variation to generalize and extract key - value pairs from a new set of documents. Since the available training data may be from a general domain, models trained on this data may struggle to classify data from various target domains. Therefore, it is desirable to improve the training of models for labeling key - value labeled documents.

Summary of the Invention

Means for Solving the Problems

[0004] Brief Summary In a general embodiment, the techniques of the present disclosure may include identifying text in a visually rich document. These techniques may also include determining a sequence of identified text, the sequence having a numerical order for each word of the identified text. The methods may further include selecting a language model based at least in part on the identified text and the determined sequence. These techniques may also include generating textual features corresponding to the identified text by assigning each word of the identified text to a respective token, each token having a string of one or more words, using the selected language model and the determined sequence. These techniques may further include extracting visual features corresponding to the identified text, the visual features having information about multiple pixels representing each word of the identified text. These techniques may also include determining a positional feature for each word of the identified text, the positional feature having the respective coordinates of each word of the identified text within the visually rich document. These techniques may further include generating a graph representing the visually rich document, where each node in the graph represents a visual feature, textual feature, and positional feature, respectively, for each word of the identified text. Furthermore, these techniques may include training a classifier to classify each individual word in the identified text, the classifier being trained on a graph representing a visually rich document. Other embodiments of this aspect include corresponding computer systems, devices, systems, one or more non-temporary computer-readable media, devices, and computer programs recorded in one or more computer storage devices, each configured to perform the operations of the above techniques.

[0005] Embodiments may include one or more of the following features: The technology may include classifying each word of an identified text by a classifier. In the technology, each word is classified as a key or value of a key-value pair. In the technology, the graph is a graph neural network. In the technology, the language model is selected based at least in part on the domain of the identified text. In the technology, the domain may include the language or subject of the identified text. In the technology, the visually rich document may include at least one of an invoice, receipt, insurance form, boarding pass, or identification document. Embodiments of the described technology may include hardware, methods or processes, systems, non-temporary computer-readable media, or computer tangible media. Embodiments of the described technology may include embodiments of classifying text in an incoming visually rich document by deployment of a trained classifier. Embodiments of the described technology may include embodiments realized by using a computer program product which, when executed by a processor, includes a computer program / instruction that causes a processor to execute any of the technologies described herein. [Brief explanation of the drawing]

[0006] [Figure 1] This is an illustrative diagram showing a process for extracting textual features from a visually rich document (VRD) according to at least one embodiment. [Figure 2] This is a simplified illustrative diagram of a model training service according to at least one embodiment. [Figure 3] This figure shows an exemplary machine learning model that follows at least one embodiment. [Figure 4A] This figure shows an exemplary machine learning model of a neural network, following at least one embodiment. [Figure 4B] This figure shows an exemplary machine learning model of a Support Vector Machine (SVM) according to at least one embodiment. [Figure 5]This is an exemplary simplified diagram illustrating a method for training a model that performs key-value extraction, according to at least one embodiment. [Figure 6] This is a simplified diagram illustrating an exemplary system architecture for a model training service according to one embodiment. [Figure 7] This figure shows an exemplary architecture for prompt-enhanced services, comprising one or more service provider computers, user devices, and one or more facility computers, according to at least one embodiment. [Figure 8] This is an exemplary block diagram illustrating a pattern for implementing a service-based cloud infrastructure system, following at least one embodiment. [Figure 9] This is an exemplary block diagram illustrating another pattern for implementing a service-based cloud infrastructure system, following at least one embodiment. [Figure 10] This is an exemplary block diagram illustrating another pattern for implementing a service-based cloud infrastructure system, following at least one embodiment. [Figure 11] This is an exemplary block diagram illustrating another pattern for implementing a service-based cloud infrastructure system, following at least one embodiment. [Figure 12] This is an exemplary block diagram showing an exemplary computer system according to at least one embodiment. [Modes for carrying out the invention]

[0007] Detailed explanation Various embodiments are described in this specification. For illustrative purposes, and to enable a complete understanding of the embodiments, specific configurations and details are described. However, it will be apparent to those skilled in the art that embodiments can be realized without these specific details. Furthermore, well-known features may be omitted or simplified so as not to obscure the description of the embodiments.

[0008] Embodiments of this disclosure provide a technique for training a model that labels key-value pairs in visually rich documents (VRDs). Training a machine learning model that performs key-value extraction can require a large amount of training data. The technique of the disclosure offers a technical advantage by enabling the training of a model that performs key-value extraction with only five labeled documents. Thus, the technique of the disclosure can reduce the storage and processing requirements of the model training system. A VRD can be a document that conveys information beyond text, and a VRD can convey data through location information, text information, and visual information. For example, a VRD can convey that the field "Jones" is a surname because "Jones" is close to the "Name" field. VRDs may include driver's licenses, passports, identification cards, checks, receipts, invoices, medical forms, insurance forms, tax forms, transaction statements, etc.

[0009] Information in a VRD can be structured as key-value pairs (e.g., name-value pairs, attribute-value pairs, field-value pairs, semantic classes, etc.). A pair may contain a key that defines the dataset and one or more values ​​belonging to the dataset. For example, a key could be "country," defining the dataset as containing a list of countries. The values ​​associated with the key could include one or more countries, such as "Mexico" and "Ukraine." To convey information, it may be necessary to concatenate values ​​into pairs, because isolated values ​​may not have enough context to provide meaningful information. For example, the value "Ukraine" may be difficult to understand without its corresponding key. "Ukraine" could represent a country, but it could also represent a business, a person, etc. For example, the name of a business could be "Ukraine Imports."

[0010] Location information can be conveyed through a specific document layout, including the position and relative arrangement of elements such as words, images, and graphs within the document. Location information may also include the relative arrangement or relative size of elements or fields within the document. Location information can be generated as location embeddings through training with neural networks. Location information can be determined using pre-trained models such as PICK (Processing Key Information Extraction from Documents using Improved Graph Learning-Convolutional Networks) (e.g., Yu, Wenwen et al. "PICK: processing key information extraction from documents using improved graph learning-convolutional networks." 2020 25th International Conference on Pattern Recognition (ICPR). IEEE, 2021)) and SDMG-R (Spatial Dual-Modality Graph Reasoning for Key Information Extraction) (e.g., Sun, Hongbin et al. "Spatial Dual-Modality Graph Reasoning for Key Information Extraction." arXiv preprint arXiv:2103.14470 (2021)).

[0011] Text information may include characters extracted from VRDs by optical character recognition. These characters can be tokenized and converted into textual features using text embeddings generated by language models, including deep learning-based language models. Text embeddings can be vectors that encode the meaning of words such that the distance between two words in a vector space represents the similarity between those two words. In the vector space, similar words can be closer than dissimilar words. These text embeddings can be one-hot coded vectors or sparse matrices usable for syntactic or textual matching. Such models can be difficult to train and may require large amounts of training data. Alternatively, instead of training a language model, it is also possible to identify textual information in visually rich documents by selecting and using pre-trained language models. Using pre-trained language models can enable the identification of textual features from documents in various languages.

[0012] The language model may be selected based on the language identified in the visually rich document, and the language model can be a general domain model for that language or a domain-specific language model. A domain can be a set of data, and a domain for a language model can be a specific set of descriptive text. This text can be a general descriptive text for a particular language (e.g., a general domain), a descriptive text on a particular topic, text written for a particular audience, text written for a particular industry, or any other set of descriptive text.

[0013] Visual information may include font design, color, background, or image styling within the VRD. Styling can convey information within the VRD; for example, important words may be bold or in a color that stands out against the background. Keys may have a uniform style across the document, while values ​​may have different styles across documents. For example, in medical records, keys may be printed in a uniform font, while values ​​may be handwritten. Visual information may be determined by trained models, including convolutional deep learning models such as U-Net (Ronneberger, Olaf & Fischer, Philipp & Brox, Thomas. (2015). U-Net: Convolutional Networks for Biomedical Image Segmentation. LNCS. 9351. 234-241. 10.1007 / 978-3-319-24574-4_28.).

[0014] Machine learning models can be trained to classify text from visually rich documents as key-value pairs. These models can receive text, visual, and spatial information from visually rich documents as input. The model can be a graph neural network initialized with visual and textual features related to a given word. The nodes in the graph neural network can be initialized by fusing and using these visual and textual features. After training, the classification layer of the graph neural network can be used to classify text from visually rich documents as either a key corresponding to a specific value or a value corresponding to a specific key.

[0015] In one example, an insurance company may desire to train a model that performs key-value extraction on medical records. The insurance company supplies a set of labeled VRDs with the same layout to a model training service. The labeled documents are English documents that contain text from the medical domain. After selecting a suitable pre-trained language model, the model training service can use the supplied VRDs to extract visual features, text features, and location features from visually rich documents. Visually rich documents can be time-consuming and costly to label. The disclosed model training service can train a model that performs key-value extraction using only one labeled document and four unlabeled documents.

[0016] Once visual, text, and location features are extracted from the visually rich document, the visual and text features can be used to initialize the nodes of the model being trained. During training, the features of the nodes can be propagated and aggregated by the model training service in the graph neural network being trained until the model learns the features of the edges. The features of the edges can be learned by message passing using position embeddings. After learning the features of the edges and training the model, key-value pairs can be extracted from the insurance company's medical records using the model.

[0017] FIG. 1 is a diagram 100 showing a process for extracting text features from a visually rich document (VRD) according to at least one embodiment. In this case, the visually rich document is a medical record, but other visually rich documents including identity documents (ID cards), driver's licenses, passports, receipts, advertisements, checks, etc. are also possible.

[0018] Text information can be extracted from visually rich documents. To extract text information, the text in the document is identified by using an optical character recognition model, and a vector (referred to as text embedding) that records information about the identified text is generated by using a language model. Through text embedding, a computer can assign meaning to the identified text, and by comparing the similarity of text embeddings, the relationships between words can be understood. For example, the text embedding can be a multi-dimensional vector, and in the vector space, similar words can be arranged close to each other, while dissimilar words are more distant.

[0019] Describing process 101 in more detail, in block 125, text can be identified in the visually rich document shown in FIG. 100. This text can be identified by a machine learning model that uses optical character recognition (OCR) to identify shapes that match the characters in the visually rich document. Optical character recognition can include identifying the area of the document containing the text, determining the orientation of the text, and determining the characters in the text. Identifying the area of the document containing the text can include assigning a bounding box that surrounds the identified text. For example, in FIG. 100, the bounding box is shown as a rounded rectangle surrounding the identified text 102-124. In some embodiments, identifying the text in the visually rich document can be performed by an individual who manually identifies the area of the document containing the text and converts the text into a computer-readable format.

[0020] In block 130, sequences can be assigned to discriminant text. During training, a particular visually rich document may be shown to the model being trained multiple times. If the model learns too many features of the training data, it may become overfitted to each set of training data. Such an overfitted model may be able to correctly classify the training data, but it may struggle to correctly classify new data that is different from the training data because it cannot generalize. Overfitting can be mitigated by test time extension, in which the document is randomly changed each time the visually rich document is shown to the model being trained. For example, selective blurring of areas of the document, changes in the orientation of the document, or changes in the color of the document are possible. Extension techniques may include rotation (±z degrees), perspective transformation, affine transformation, as well as scaling and padding. Test time extension allows the model to learn the training data without becoming overfitted or struggling to classify new data.

[0021] When the reading order of text changes, learning how to extract textual features from visually rich documents can become difficult. Identifying key-value pairs, such as the key "1a.Country" in identified text 110 and the value "Poland" in identified text 112, can become more difficult if the key and value are read in a different order from document to document. Recognition of key-value pairs can be based at least partially on the relative positions of the identified texts, and various augmentation techniques, i.e., mismatched text extraction, can cause the order and position of text in visually rich documents to change randomly. This potential problem can be mitigated by assigning a consistent order to text reading from visually rich documents. This order or sequence can be assigned to the visually rich document by the model, or it can be assigned manually to the text within the visually rich document. Furthermore, after running test-time augmentation techniques on the visually rich document, this order can be assigned to and maintained in the original document. For example, an order can be assigned to a document, the document can be rotated, and the rotated document can be presented to a model with the above sequence. Continuing with this example, even if the text is in a different position in the two images, the model can read the text in the original visually rich document and the rotated document in the order specified in the sequence.

[0022] Sequences can reflect the order in which readers read text within a document. Key-value pairs may be placed within a visually rich document in a way that aligns with the reader's intuition, but models may struggle to identify these pairs because they are unaware of this placement. Sequences assigned to text within a visually rich document can present information to the model in a format that can help the model process the identified text in the same order as a reader. For example, since English text is read from left to right, a language model can process the text in this order. This is because text-based documents are read from left to right when information is not conveyed by visual or spatial features.

[0023] However, people do not necessarily read the identification text in the document shown in 100 from left to right. For example, if read from left to right, the identification texts might be read in the order of identification text 106 "1. Patient's initials (first name, last name)", identification text 110 "1a. Country", identification text 108 "Vincent", and identification text 112 "Poland". People might also read the text in the order of identification text 106 "1. Patient's initials (first name, last name)", identification text 108 "Vincent", identification text 110 "1a. Country", and identification text 112 "Poland". Because people read keys and values ​​sequentially, the order in which people read them can make it easier to identify two key-value pairs (key 1 "1. Patient's initials (first name, last name)" and value 1 "Vincent", key 2 "1a. Country" and value 2 "Poland"). Therefore, assigning sequences to the text in the document can improve the model's key-value pair identification.

[0024] In block 135, textual features can be determined by processing the identified text in the assigned sequence. Processing the text may include providing the text to the language model in the order specified in the sequence. The language model can perform various natural language processing techniques on the input text. For example, the model can generate feature vectors for the input text. Natural language techniques may include part-of-speech tagging, parsing, grammatical guidance, and information retrieval.

[0025] While a language model can perform natural language processing tasks on the text it is input to, it may struggle to perform these tasks on texts different from its training data. For example, a model trained on Spanish text may struggle to process English text. A model can be trained on general texts of a language, but some models may be trained on a specific set of texts in a particular language. Performing some tasks may require training the model on a specific text corpus of a particular language. For example, a model trained on a specific language may struggle to process texts intended for a specific audience. For instance, texts may be targeted at members of a particular industry (e.g., medical texts) or specifically for a particular audience (e.g., fans of a particular musician). Group-specific jargon, slang, and code can make it difficult for models trained on general datasets to process texts intended for a specific audience. For example, a general English language model may not be able to accurately classify medical terminology (e.g., subdural hematoma) or fan-based slang (e.g., Swiftie).

[0026] Training language models can require large amounts of training data. Instead of training models to perform natural language processing, one can train models for key-value extraction, or select pre-trained language models to use for text processing. Pre-trained language models can have their weights fixed so that they do not learn from the text of the visually rich documents input to them. If a model fixes its weights, these weights may be fixed for all layers. Alternatively, these weights may be fixed for a subset of layers, while some unfixed layers are trained on new datasets. Pre-trained models can be selected manually by a human, or they can be selected by the model itself.

[0027] Figure 2 is a simplified diagram 200 of a model training framework 205 according to at least one embodiment. This model is trainable to perform key-value extraction. Key-value extraction can be a process of identifying constants (referred to as keys) that define a dataset and concatenating them with variables (e.g., values) that belong to the dataset. The model can identify key-value pairs in a visually rich document using textual, visual, and spatial features. Each service within the model training framework 205 and any other services in this disclosure include software, hardware, or any combination of software and hardware components.

[0028] Key-value extraction from a VRD may include text extraction from a document. Text extraction can be performed in such a way that textual features can be generated for a visually rich document. Text extraction can be a process of converting typed or handwritten text into a machine-readable format, and text detection can be performed by optical character recognition (OCR). The model training framework 205 can receive a visually rich document 210. The received document may be provided to an OCR service 215 which may include at least one of a text detection service 220, an orientation classification service 225, or a text recognition service 230.

[0029] The text detection service 220 can detect areas in the VRD 210 that contain text. Because OCR can be computationally intensive, the use of the text detection service 220 reduces the search space by segmenting the VRD 210 into areas containing text that can be recognized and areas that do not contain text that can be excluded from text detection. These segments can correspond to bounding boxes, and in some embodiments, the bounding boxes can be shared with other elements in the model training framework 205.

[0030] The orientation classification service 225 can detect the orientation of words in the VRD 210. Since information can be conveyed through word orientation, for example, "smug" could become "gums" depending on the word orientation and the order in which the letters are read. The text within the words can be detected by the text recognition service 230. The text can be recognized using various techniques, including feature extraction and matrix matching.

[0031] After recognition, the text may be provided to the sequence service 235, which can determine the order of the recognized text. The text may also be provided from the sequence service 235 to the language model service 240. The language model, which runs within the language model service 240, can tokenize the text and decode the sequence to generate a set of subtoken embeddings for the recognized text. Tokenization of text may mean dividing the text into a set of tokens that represent words, phrases, sentences, paragraphs, etc. Subtoken embeddings can be vectors representing the features of words within a token. A token may have one or more subtoken embeddings. For example, a token representing a single word may have one subtoken embedding, while a token representing a phrase may have multiple subtoken embeddings (for example, for multiple words within that token).

[0032] The language model in the language model service 240 may be selected based on the visually rich document 210. Training a language model to perform natural language processing may require a large amount of training data. The number of labeled visually rich documents required to train a model that performs key-value extraction can be reduced by using a pre-trained language model that has been trained to tokenize and generate embeddings for the input text. This language model can be trained on a general domain of a language (e.g., Cantonese) or a specific domain within that language (e.g., computer science text in Cantonese).

[0033] The visual features service 250 within the model training framework 205 can identify visual features from the visually rich document 210. Visually rich documents can use a variety of visual stimuli to convey information. A reader of a visually rich document might understand, for example, bold text as a key and handwritten text as a value. In another example, the key might be text in a large font and the value in a small font. In yet another example, each key-value pair might share similar visual characteristics, and a specific font and font size of text might belong to the key-value pair. Because visual features can convey these visual stimuli in a machine-readable format, it is possible to train a model to learn text classification based at least partially on the visual characteristics of the text.

[0034] In the feature extraction service 255 of the visual feature service 250, vectors representing the features of each pixel in the visual rich document 210 can be generated. These features can be extracted by a visual feature extractor, such as a UNET model with a residual network (ResNet) backbone or any other model capable of extracting visual features from images or documents. The feature extractor can be a convolutional neural network pre-trained for the domain corresponding to the visual rich document, and the feature extractor can be replaced with any other visual feature extractor based on the type of visual rich document being analyzed. The features extracted by the feature extractor can be document-specific, and the weights of the feature extractor do not have to be fixed in the disclosed training technique. Models with fixed weights may not be able to continue training and learning the features of the visual rich document input to the model. However, models with non-fixed weights can continue training.

[0035] The visually rich document 210 can be passed from the feature extraction service 255 to the cropping service 260. Cropping can be used to isolate visual features around a specific selected text. The cropping service 260 can identify areas of the visually rich document corresponding to text using bounding boxes identified by the text detection service 220. In some embodiments, the cropping service 260 can identify text and generate bounding boxes. These bounding boxes can be shared by the cropping service with the text detection service 220 or any other service in the model training framework 205. The cropping service 260 can identify feature vectors of pixels corresponding to areas within the bounding boxes. These identified features can be visual features corresponding to text within the bounding boxes.

[0036] Textual and visual features can be combined by the deep fusion model service 265. The deep fusion of the deep fusion model can consist of a set of small neural networks capable of reweighting textual and visual features. These small neural networks can be pre-trained models, or they can learn weights during training. Textual features identified by the language model service 240 and visual features identified by the visual feature service 250 can be correlated by the deep fusion model service 265 to generate vectors or matrices representing combined visual and textual features for a particular word, phrase, sentence, or other group of text. Tokens or subtoken embeddings may be associated with the visual features of pixels representing the text corresponding to the token or subtoken embedding.

[0037] The model training service 270 can initialize a machine learning model that fuses the visual and textual features output by the deep fusion model service 265. The machine learning model trained by the model training service 270 may include a graph neural network, a transformer model, a multilayer perceptron (MLP), or a statistical model. The model can be a neural network, and each token or subtoken can be a node in the initialized machine learning model. For example, the machine learning model can be a graph neural network such as a convolutional neural network. A graph neural network may include nodes connected by edges. The model training service 270 can initialize the edges between nodes by the distance between bounding boxes corresponding to at least each node. For example, the distance between two bounding boxes can be the distance between the center coordinates of each bounding box. These distances may be provided to the model training service 270 by the positional feature service 275.

[0038] In the model training service 270, positional features determined by the positional feature service 275 can be assigned to nodes. Positional features may include representations of the absolute position of text within the visually rich document 210 and can be converted into feature vectors by a neural network in the model training service 270. Positional features can be coordinates in any appropriate coordinate service, such as orthogonal coordinates (e.g., xy coordinates). In the model training service 270, node embeddings (e.g., feature vectors corresponding to the textual and visual features of a node) can be linked to positional embeddings (e.g., feature vectors corresponding to the feature vectors representing the positional features of a node). In the model training service 270, a linkage vector for a node can be linked to an initial edge embedding corresponding to that node (e.g., a vector representing the features of an edge). Edge embeddings corresponding to a node can be edge embeddings of edges connected to the node.

[0039] The positional feature service 275 allows a visually rich document to be divided into multiple regions. For example, a visually rich document 210 may be divided into a 5x5 grid consisting of 25 regions of equal size. In some embodiments, the size of each region may differ, and the number of regions may be more or less than 25. The positional feature of a particular text region (e.g., text corresponding to a token or subtoken embedding) may be the absolute coordinates of the text region (e.g., the xy coordinates corresponding to the bounding box surrounding the text) or the region coordinates corresponding to the region containing the text. For example, for all text within a particular region, a positional embedding corresponding to the entire region or the centroid of the region may be assigned.

[0040] Model training service 270 allows training a model to classify text in a visually rich document as either a key or a value. Classifying text as a key means identifying the corresponding value for that key, and classifying text as a value may mean identifying the corresponding key for that value. This model can classify text as belonging to a key-value pair by using the visual, textual, and spatial features of the text in the visually rich document 210. A key may have multiple corresponding values, and a value may have multiple corresponding keys.

[0041] The model training service 270 can use various techniques for training the model. For example, by using multi-head attention techniques, the model can be focused on different areas of the visually rich document 210. The model training service can use various filters to focus the model on different areas of the visually rich document 210. Attention techniques are described in more detail in the paper "Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Proceedings of the 31st International Conference on Neural Information Processing Services (NIPS'17). Curran Associates Inc., Red Hook, NY, USA, 6000-6010". The model training framework 205 can use the aforementioned extension techniques, such as test time extension. Once model training is complete using the model training service 270, the trained model 280 may be output from the model training framework 205.

[0042] Figure 3 shows a machine learning model according to an embodiment of the present disclosure. The training vectors 305 are shown with service properties 310 and known classifications 315. As an example, the service could be a machine learning model deployable on a cloud network. The service properties 310 may include various fields. For ease of illustration, only two training vectors are shown, but the number of training vectors may be much larger, for example, 10, 30, 100, 1,000, 10,000, 100,000 or more. The training vectors may also be configured for different services or for the same service over different time periods.

[0043] The service property 310 has property fields that can correspond to the properties of a machine learning model or cloud service, and a person skilled in the art will recognize the various ways in which such a service or model can be constructed. Known classifications 315 include hardware or software characteristics such as the number of nodes, the number of central processing unit (CPU) cores, the number of graphics processing units (GPUs), the number of CPUs, the type of CPU, the type of GPU, and the amount of memory. The classification can have arbitrary backing (e.g., real numbers) or elements of a small finite set. The classification can also be ordinal, and the backing can be provided as an integer. Thus, the classification can be categorical, ordinal, or real, and can be related to a single measurement or multiple measurements, and may be high-dimensional.

[0044] Training 320 can be performed by the learning service 325 using the training vector 305. A service such as the learning service 325 is one or more computing devices configured to execute computer code that performs one or more operations that constitute the service. The learning service 325 can optimize the parameters of model 335 so that a quality criterion (e.g., the accuracy of model 335) is achieved by one or more specific criteria. Accuracy may be measured by comparing a known classification 315 with a predicted classification. The parameters of model 335 can be iteratively modified to improve accuracy. Determining the quality criterion can be done for any function that includes all risk, loss, utility, and decision functions.

[0045] In some embodiments of training, it is possible to determine the gradient of how parameter changes affect the cost function, which can provide a measure of how accurate the current state of the machine learning model is. This gradient can be used in conjunction with the training steps (e.g., a measure of how much the model's parameters should be updated for a given time step in the optimization process). Thus, optimization of parameters (which may include weights, matrix transformations, and probability distributions) can provide an optimal value for the cost function, which can be measured, for example, either above or below a threshold (i.e., above the threshold), or as long as the cost function does not change significantly over multiple time steps. In other embodiments, training can be performed by methods that do not require Hessian computation or gradient computation, such as dynamic programming or evolutionary algorithms.

[0046] In prediction stage 330, a predicted entity classification 355 can be provided for the entity signature vector 340 of a new entity based on a new service property 345. The new service property can be of a similar type to the service property 310. If the new service property is of a different type, data in a format similar to the service property 310 can be obtained by performing a transformation on the data. Ideally, the predicted service classification 355 corresponds to the true service classification of the input vector 340.

[0047] Examples of machine learning models include deep learning models, neural networks (e.g., deep learning neural networks), kernel-based regression, adaptive basic regression or classification, Bayesian methods, ensemble methods, logistic regression and extension, Gaussian processes, support vector machines (SVMs), probabilistic models, and probabilistic graphical models. Embodiments using neural networks may employ wide tensorized deep architectures, convolutional layers, dropout, various neural activations, and regularization steps.

[0048] Figure 4A shows an exemplary machine learning model of a neural network. For example, model 435 can be a neural network containing multiple neurons (e.g., adaptive basis functions) organized as layers. For example, neuron 405 can be a part of layer 410. Neurons can be connected by edges between them. For example, neuron 405 can be connected to neuron 415 by edge 420. Neurons can be connected to any number of different neurons in any number of layers. For example, neuron 405 can be connected to neuron 415 and also to neuron 425 by edge 430.

[0049] During neural network training, the best configuration of neural network parameters can be iteratively explored for feature recognition and classification performance. Various numbers of layers and nodes may be used. Those skilled in the art will readily recognize changes in the design of neural networks and other machine learning models. For example, artificial neural networks may include graph neural networks configured to operate on unstructured data. A graph neural network can receive a graph (e.g., nodes connected by edges) as input to the model and can learn the features of this input through paired message passing. In paired message passing, nodes exchange information, and each node iteratively updates its representation based on the information received. Further details on graph neural networks can be found in the reference Wu, Zonghan et al. "A comprehensive survey on graph neural networks." IEEE Transactions on Neural Networks and Learning Systems 32.1 (2020): 4-24. The artificial neural networks described in the background of this application can be implemented in hardware or emulated using software.

[0050] Figure 4B shows an exemplary machine learning model of a Support Vector Machine (SVM). As another example, Model 435 is a possible Support Vector Machine. Features can be treated as coordinates in coordinate space. A sample of training data points (e.g., multidimensional data points consisting of measurement data). The training data points are distributed in space, and the Support Vector Machine can identify boundaries between classifications. For example, points 435 and 440 may be separated by boundary 445.

[0051] Figure 5 is a simplified diagram illustrating Method 500 for training a model that performs key-value extraction, according to at least one embodiment. The Method is presented as a logical flow diagram, where each operation may be implemented by hardware, computer instructions, or a combination thereof. In the context of computer instructions, these operations may represent computer-executable instructions, stored in one or more computer-readable storage media and executed by one or more processors to perform the described operations. Generally, computer-executable instructions include routines, programs, objects, components, data structures, etc., that perform specific functions or implement specific data types. The order in which the operations are described is not intended to have any restrictive interpretation, and any number of described operations in any order and / or in parallel combinations may be used to implement this process or method.

[0052] To elaborate on Method 500, in Block 505, text in a visually rich document can be identified. The text can be identified by a computer service. A visually rich document can be a physical or digital document that conveys information through visual stimuli, in addition to the text of the document. For example, the position, font, size, and color of the text in the document can characterize the meaning of the text. As an example, key-value pairs in a document can be identified if the key is bold text while the corresponding value is standard text. The text can be identified by using a model trained to perform optical character recognition in the document. A visually rich document can be an invoice, receipt, insurance form, boarding pass, identification card, or any other document that conveys information using both text and visual characteristics.

[0053] In block 510, a sequence of identified text can be determined. This sequence can be determined by a computer service and can be the numerical order of each word in the identified text. Alternatively, it can be the order in which the model processes the text from the visually rich document. The model can process text as input to the model, and this sequence can be the order in which the text is input to the model. The determination of the sequence can be achieved by clustering and sorting the text identified in block 505, as described above with reference to Figure 1.

[0054] In block 515, a language model may be selected based at least in part on the identified text and decision sequence. The language model can be a pre-trained model trained for a specific language. The language model may be selected in response to user input. In some embodiments, the model may be trained for a specific domain. Some domains may include text in a particular language that a general natural language processing model might struggle to classify and understand. For example, a general English language model might struggle to classify medical records because these records contain text from the medical domain. This medical domain may include the selection of medical-specific words that have different meanings from medical terminology and common usage that a general model would not have encountered during training. For example, "the patient is coding" could mean that the patient is in cardiac arrest, while a general language model might misinterpret the expression as meaning that the patient is writing software.

[0055] In block 520, the selected language model may assign each word of the identified text to a token. The identified text may be assigned by the model in an order based on the sequence determined in 510. Tokens can be groups of one or more words, or single words, phrases, sentences, groups of sentences, etc. For example, the sentence "The dog barked, and I told him to be quiet" can be broken down into tokens "the dog barked" and "and I told him to be quiet," or the sentence can be tokenized as individual words. The words in a token may be consecutive words, and the token represents a block of text (for example, a token may contain adjacent words).

[0056] The use of tokens allows textual features to be assigned to words within those tokens. For example, a textual feature might include the order of tokens derived from the sequence determined in 510. Textual features can be represented as numerical vectors representing the properties of the corresponding words. Two words with similar meanings can be placed close to each other in the vector space, while two words with different meanings can be placed far apart.

[0057] In block 525, visual features corresponding to the identified text may be extracted. The visual features of a word may include information about the pixels that represent that word in a visually rich document. A word can be enclosed in a bounding box, and the multiple pixels that represent a word can be the pixels within the bounding box. The visual information may include the color of each pixel and one or more aggregated statistics calculated from the colors of each pixel in the bounding box (e.g., the total number of pixels, the average color of the pixels in the bounding box, etc.). The visual information may also include information from the region of interest, such as edge-corner interactions.

[0058] In block 530, the positional features of each word in the identified text can be determined. The positional features of a word can include the coordinates within the corresponding visual rich document. For example, the coordinates could be the coordinates of the center of the bounding box surrounding the word. Alternatively, the coordinates could be any information indicating the position of the word within the visual rich document.

[0059] In block 535, a document model representing a visually rich document may be generated. Any type of machine learning model is possible as the document model, for example, a graph neural network where nodes are connected by edges. A particular node may represent the visual, textual, and spatial features of a particular word in the identified text. In some embodiments, a particular node may represent the visual, textual, and spatial features of a particular token. Each word or token identified in a visually rich document may have a corresponding node.

[0060] In block 540, a classifier may be trained to classify each individual word in the identified text. The classifier may be trained on a graph representing a visually rich document. These techniques may include classifying words in the identified text by the trained classifier. Words may be classified as key or value, and classifying a word may include identifying the word as a key or value for a key-value pair of a particular class. For example, a word may be classified as a key or value for the key-value pair "surname". In some embodiments, words may be classified into an unknown category. An unknown category may mean that the model does not have enough information to classify the word as a key or value. In some embodiments, multiple words may be classified as constituting a single key, and multiple words may be classified as belonging to a particular pair. Once trained, the classifier may be deployed to identify and extract text in incoming visually rich documents that have not been analyzed previously. Based on fixed weights of the trained model, the classifier is able to identify each word in the document.

[0061] Figure 6 is a simplified diagram 600 showing a service architecture for a model training service according to one embodiment. Each service in Figure 600 and any other services in this disclosure include software, hardware, or any combination of software and hardware components. The pseudo-labeling service 605 may be hosted on a computing device 610. The optical character recognition (OCR) service 615 in the pseudo-labeling service 605 can identify characters in a visually rich document (VRD). Because OCR can be computationally intensive, the text detection service 620 can reduce the amount of OCR processing by segmenting the VRD into areas containing characters to be recognized and areas not containing characters to be excluded from recognition. The orientation classification service 625 can reduce the amount of OCR by identifying the orientation of the text so that the text recognition is performed in the correct orientation. The text recognition service 630 can recognize and extract text from the VRD.

[0062] The sequence service 635 can assign sequences to words identified in a visually rich document. These sequences can be an order assigned to some or all of the words in the visually rich document. The language model service 640 can select an appropriate pre-trained language model for the text extracted by the OCR service 615. The model can be selected based on the user's identification of the appropriate language (for example, via a graphical user interface), or the language model service 640 can identify and select an appropriate language model. Once a model is selected, the language model service 640 can tokenize the text in the visually rich document and generate word embeddings for the text.

[0063] In training the final key-value extraction model, the weights of the pre-trained language model may be fixed so that the pre-trained language model does not continue training itself. Fixing weights may include fixing some of the weights in the model and allowing the training of other weights. Word embeddings generated by the language model in the language model service 640 may be projected onto a linear layer by the linear projection service as described above. Positional features of a visually rich document may be extracted by the positional features service 650. Visual features may be extracted by the visual features service 655. In the feature extraction service 660 of the visual features service 655, feature vectors can be assigned to some or all of the pixels in the visually rich document. In the cropping service 655 of the visual features service 655, pixels corresponding to specific words can be identified by cropping the visually rich document, and the feature vectors of these pixels may be associated with specific words. For example, pixels corresponding to a specific word may be pixels within the bounding box corresponding to that specific word.

[0064] The deep fusion model service 670 can initialize the node features of a graph neural network trained by the model training service 675 by fusing visual and word embeddings (e.g., visual and word features). In some embodiments, these features may be fused using Kronecker fusion. The model training service 675 can train a graph neural network to identify and classify key-value pairs. The model training service can train a graph neural network using location features from the location feature service 650. The model training service 675 can propagate and aggregate node features in the GNN during message passing, learning edge features by using location embeddings and multi-head attention. The graph neural network may include a classification layer that can be used to classify words from visually rich documents as key-value pairs. Models trained by the model training service may include any family of machine learning models, including deep learning models and statistical models.

[0065] Figure 7 shows an architecture 700 for a model training service (e.g., model training framework 205) comprising one or more service provider computers, user devices, and one or more facility computers, according to at least one embodiment. In architecture 700, one or more users 702, such as customers, requesting key-value extraction from a visually rich document may receive the visually rich document via one or more networks 708 by using user computing devices 704A-704N (collectively referred to as user devices 704) to access a browser application 706 or a user interface (UI) accessible through the browser application 706, which is present and interactable via the browser application 706 or the UI accessible through the browser application 706. The “browser application” 706 may include any browser control or native application that is capable of accessing and / or displaying information such as network pages. Native applications may include applications or programs developed for use on a specific platform such as an operating service or on a specific device such as a specific type of mobile device.

[0066] According to at least one embodiment, the user device 704 may be configured to communicate with a service provider computer 714 and a facility computer 730 via a network 708. The user device 704 may include at least one memory, such as memory 710, and one or more processing units or one or more processors 712. The memory 710 may store program instructions that can be loaded and executed by one or more processors 712, as well as data generated when these programs are executed. Depending on the configuration and type of the user device 704, the memory 710 may be volatile, such as random access memory (RAM), and / or read-only memory (ROM), or non-volatile, such as flash memory. The user device 704 may also include additional removable storage and / or non-removable storage, including, but not limited to, magnetic storage, optical disks, and / or tape storage. Disk drives and their respective associated non-temporary computer-readable media may provide non-volatile storage for computer-readable instructions, data structures, program services, and other data to the user device 704. In some embodiments, the memory 710 may include several different types of memory, such as static random access memory (SRAM), dynamic random access memory (DRAM), and ROM.

[0067] More specifically, the contents of memory 710 may include operating services and one or more application programs or services for implementing the features disclosed herein. Alternatively, memory 710 may include one or more services for implementing the features described herein, such as the model training framework 205.

[0068] Architecture 700 may additionally include one or more service provider computers 714 that can provide computing resources in some examples, such as client entities, low-latency data storage, durable data storage, data access, management, virtualization, host computing environments or "cloud-based" solutions, prompt enhancement or engineering function implementations, etc. A service provider computer 714 may implement, or may not implement, one or more machine learning models or one or more service provider computers described herein by reference in Figures 1 to 6 and / or throughout this disclosure. Furthermore, one or more service provider computers 714 may be capable of operating to provide site hosting, computer application development, and / or implementation platforms, or combinations thereof, to one or more users 702 via a user device 704.

[0069] In some examples, network 708 may include one or a combination of many different types of networks, such as cable networks, the Internet, wireless networks, cellular networks, and other private and / or public networks. The illustrated example shows a case where user 702 communicates with service provider computer 714 via network 708, but the described technique may equally apply to a case where user 702 interacts with one or more service provider computers 714 via one or more user devices 704 by a landline telephone, kiosk, or any other method. The described technique may also apply to other client / server configurations such as set-top boxes, as well as non-client / server configurations such as locally stored applications and peer-to-peer configurations. In embodiments, user 702 may communicate with facility computer 730 via network 708, and facility computer 730 may communicate with service provider computer 714 via network 708. In some embodiments, the service provider computer 714 may obtain data inputs to various algorithms of the generation function described herein by communicating with one or more third-party computers (not shown) via the network 708. According to at least one embodiment, the service provider computer 714 may receive text data, video data, image data, one or more prompts, aggregate inputs generated therefrom, etc., in order to enhance prompts for at least the generation model.

[0070] One or more service provider computers 714 may be any type of computing device, or may include any type of computing device, such as, but are not limited to, mobile phones, smartphones, personal digital assistants (PDAs), laptop computers, desktop computers, server computers, thin client devices, and tablet PCs. It should also be noted that in some embodiments, one or more service provider computers 714 may be run by one or more virtual machines implemented in a host computing environment. The host computing environment may include one or more computing resources that are rapidly provisioned and released, and these computing resources may include computing, networking, and / or storage devices. The host computing environment may also be referred to as a cloud computing environment or a distributed computing environment. In some examples, one or more service provider computers 714 may communicate with user devices 704 via a network 708 or other network connection. One or more service provider computers 714 may include one or more servers that can be deployed in a cluster configuration or as individual, unrelated servers. In one embodiment, the service provider computer 714 may communicate with one or more third-party computers (not shown) via the network 708 to receive or acquire data including text data, video data, image data, one or more prompts, aggregate inputs generated therefrom, etc., in order to enhance prompts for the generative model.

[0071] In an exemplary configuration, one or more service provider computers 714 may include at least one memory, such as memory 716, and one or more processing units or one or more processors 718. One or more processors 718 may be implemented as hardware, computer executable instructions, firmware, or any combination thereof, as needed. Embodiments of computer executable instructions or firmware for one or more processors 718 may include computer executable or machine executable instructions written in a optionally preferred programming language to perform the various functions described above when executed by a hardware computing device such as a processor. Memory 716 may store program instructions that can be loaded and executed by one or more processors 718, as well as data generated when these programs are executed. Depending on the configuration and type of one or more service provider computers 714, memory 716 may be volatile, such as RAM, and / or non-volatile, such as ROM or flash memory. One or more service provider computers 714 or servers may also include additional storage 720, which may include removable storage and / or non-removable storage. The additional storage 720 may include, but is not limited to, magnetic storage, optical disks, and / or tape storage. Disk drives and their respective associated computer-readable media may provide non-volatile storage of computer-readable instructions, data structures, program services, and other data to the computing device. In some embodiments, the memory 716 may include several different types of memory, such as SRAM, DRAM, and ROM.

[0072] Memory 716 and additional removable and / or non-removable storage 720 are examples of non-temporary computer-readable storage media. For example, computer-readable storage media may include volatile or non-volatile, removable or non-removable media implemented in any method or technique for storing information such as computer-readable instructions, data structures, program services, or other data. Memory 716 and additional storage 720 are examples of non-temporary computer storage media. Additional types of non-temporary computer storage media that may be present in one or more service provider computers 714 may include, but are not limited to, PRAM, SRAM, DRAM, RAM, ROM, EEPROM, flash memory, or other memory technologies, CD-ROM, DVD, or other optical storage, magnetic cassette, magnetic tape, magnetic disk storage, or other magnetic storage devices, or any other media that can be used to store desired information and are accessible by one or more service provider computers 714. Any combination of the above is also included in the scope of non-temporary computer-readable media.

[0073] Furthermore, one or more service provider computers 714 may include one or more communication connection interfaces 722 that can enable communication with data stores, other computing devices or servers, user terminals, and / or other devices on the network 708. Also, one or more service provider computers 714 may include one or more I / O devices 724 such as keyboards, mice, pens, voice input devices, touch input devices, displays, speakers, and printers.

[0074] More specifically, the contents of memory 716 may include an operating system 726 and one or more data stores 728 and / or one or more application programs or services for implementing the features disclosed herein (including the model training framework 205). Architecture 700 includes a facility computer 730. In embodiments, the service provider computer 714 and the model training framework 205 may be configured to generate instructions and transmit them to a component 736 communicating with or associated with the facility computer 730 via a network 708. For example, the instructions may be configured to transmit a trained model or a visually rich document in accordance with the operation of the model training framework 205, upon activation or triggering of component 736. The facility computer 730 may include at least one memory, such as memory 732, and one or more processing units or one or more processors 734. Memory 732 may store program instructions (which may include one or more machine learning models as disclosed herein) that can be loaded and executed by one or more processors 734, as well as data generated when these programs are executed. Depending on the configuration and type of the facility computer 730, the memory 732 may be volatile, such as random access memory (RAM), and / or non-volatile, such as read-only memory (ROM), or flash memory. The facility computer 730 may also include additional removable and / or non-removable storage (including, but not limited to, magnetic storage, optical disks, and / or tape storage). Disk drives and their respective associated non-temporary computer-readable media may provide non-volatile storage for computer-readable instructions, data structures, program services, and other data to the facility computer 730.In some embodiments, the memory 732 may include several different types of memory, such as static random access memory (SRAM), dynamic random access memory (DRAM), and ROM.

[0075] More specifically, the contents of memory 732 may include an operating system and one or more application programs or services for implementing the features disclosed herein. Memory 732 may also include one or more services for implementing the features described herein (which may include the model training framework 205). In some embodiments, the service provider computer 714 and the model training framework 205 can train a model to perform key-value extraction based at least partially on a visually rich document provided to the model training framework 205. The user device 704 and the browser application 706 may be configured to send output to the user 702. According to at least one embodiment, the model training framework 205 may be configured to receive a visually rich document, a pre-trained language model, etc. In some embodiments, some or all of this input data may be stored and transmitted as a text file or other file (which may include text data). In some embodiments, the model training framework 205 may be configured to select a specific pre-trained language model based on the input visually rich document, etc., by implementing one or more machine learning models, computer models, computer algorithms, etc.

[0076] The model training framework 205 may be presented via a browser application 706 and a user device 704 and may be configured to generate and transmit a user interface or data objects for updating the user interface to present labeled visual rich documents, key-value pair identified visual rich documents, aggregate statistics based on identified key-value pairs, or any of these components or any components associated therewith to the user 702. Furthermore, data object generation associated with other graphical updates, feedback mechanisms, and prompt enhancement features described herein may be implemented by the service provider computer 714 and / or the model training framework 205.

[0077] As mentioned above, Infrastructure as a Service (IaaS) is a specific type of cloud computing. IaaS can be configured to provide virtualized computing resources over a public network (e.g., the internet). In the IaaS model, a cloud computing provider can host infrastructure components (e.g., servers, storage devices, network nodes (e.g., hardware), deployment software, platform virtualization (e.g., hypervisor layer), etc.). In some cases, the IaaS provider may also supply various services associated with these infrastructure components (exemplary services include billing software, monitoring software, logging software, load balancing software, and clustering software, etc.). Therefore, since these services can be policy-driven, IaaS users can maintain application availability and performance by implementing policies that promote load balancing.

[0078] In some cases, IaaS customers can access resources and services over a wide area network (WAN), such as the internet, and use the cloud provider's services to install other elements of their application stack. For example, a user can log into the IaaS platform and create virtual machines (VMs), install operating systems (OS) on each VM, deploy middleware such as databases, create storage buckets for workloads and backups, and even install enterprise software on those VMs. The customer can then use the provider's services to perform various functions, such as distributing network traffic, troubleshooting application issues, monitoring performance, and managing disaster recovery.

[0079] In most cases, the cloud computing model requires the participation of a cloud provider. This cloud provider may, but does not have to be, a third-party service specializing in providing IaaS (e.g., granting, leasing, or selling). Alternatively, an entity could deploy a private cloud and become its own infrastructure service provider.

[0080] In some cases, IaaS deployment is the process of placing a new application or a new version of an application on a prepared application server, etc. It may also include the process of preparing the server (e.g., installing libraries, daemons, etc.). This is often managed by the cloud provider under the hypervisor layer (e.g., servers, storage, network hardware, and virtualization). Therefore, the customer may be responsible for handling (OS), middleware, and / or application deployment (e.g., on self-service virtual machines that can be spun up on demand).

[0081] In some cases, IaaS provisioning represents acquiring the computers or virtual hosts to be used, and may even represent installing the necessary libraries or services on them. In most cases, a deployment does not include provisioning, and provisioning may need to be performed first.

[0082] In some cases, IaaS provisioning presents two distinct challenges. Firstly, there is the initial challenge of provisioning the initial set of infrastructure before anything is operational. Secondly, there is the challenge of evolving the existing infrastructure after all provisioning is complete (e.g., adding new services, modifying services, removing services, etc.). In some cases, these two challenges can be addressed by enabling the declarative definition of infrastructure configuration. In other words, the infrastructure (e.g., the required components and the way those components interact) can be defined by one or more configuration files. Thus, the overall topology of the infrastructure (e.g., resource dependencies and how each resource works together) can be described declaratively. In some cases, once the topology is defined, a workflow can be generated to form and / or manage the various components described in the configuration files.

[0083] In some examples, infrastructure can consist of many interconnected elements. For instance, there may be one or more virtual private clouds (VPCs), also known as core networks (e.g., potential on-demand pools of configurable and / or shared computing resources). In some examples, there may also be one or more inbound / outbound traffic group rules provisioning that define how inbound and / or outbound network traffic is configured and one or more virtual machines (VMs). Other infrastructure elements such as load balancers and databases may also be provisioned. As there is a demand for and / or addition of more infrastructure elements, the infrastructure can gradually evolve.

[0084] In some cases, the adoption of sequential deployment techniques can enable the deployment of infrastructure code across various virtual computing environments. Furthermore, the techniques described can enable infrastructure management within these environments. In some examples, a service team may write code that is desirable to be deployed to one or more (but often many) different generation environments (e.g., geographically diverse locations, sometimes even worldwide). However, in some examples, the infrastructure to which the code is deployed must be configured first. In some cases, manual provisioning, the use of provisioning tools for provisioning resources, and / or the use of deployment tools for deploying code after infrastructure provisioning are also possible.

[0085] Figure 8 is a block diagram 800 showing an exemplary pattern of an IaaS architecture according to at least one embodiment. A service operator 802 may be communicated with a secure host tenancy 804 which may include a virtual cloud network (VCN) 806 and a secure host subnet 808. In some examples, the service operator 802 may use one or more client computing devices, which may be portable handheld devices (e.g., iPhone®, mobile phones, iPad®, computing tablets, personal digital assistants (PDAs)) or wearable devices (e.g., Google® Glasses Head-Mounted Display) that run software such as Microsoft Windows Mobile® and / or various mobile operating systems such as iOS, Windows Phone, Android, BlackBerry 8, Palm OS, and are capable of using the Internet, email, short message service (SMS), BlackBerry®, or other communication protocols. Alternatively, the client computing device may be a general-purpose personal computer, such as a PC and / or laptop computer running various versions of the Microsoft Windows®, Apple Macintosh®, and / or Linux® operating systems. Client computing devices can be workstation computers running any of the various commercially available UNIX® or UNIX-like operating systems, including, but are not limited to, various GNU / Linux® operating systems such as Google® Chrome OS.Alternatively or in addition, the client computing device may be any other electronic device, such as a thin client computer that can communicate via a network that can access the VCN806 and / or the Internet, an Internet-enabled gaming system (e.g., a Microsoft Xbox game console with or without a Kinect® gesture input device), and / or a personal messaging device.

[0086] VCN806 may include an LPG810 that can be connected to SSH VCN812 via a local peering gateway (LPG)810 included in Secure Shell (SSH) VCN812. SSH VCN812 may include an SSH subnet 814, and SSH VCN812 may be connected to control plane VCN816 via an LPG810 included in control plane VCN816. Furthermore, SSH VCN812 may be connected to data plane VCN818 via LPG810. Control plane VCN816 and data plane VCN818 may be included in a service tenancy 819 that may be owned and / or operated by an IaaS provider.

[0087] The control plane VCN 816 may include a control plane buffer zone (DMZ) layer 820 that functions as a perimeter network (e.g., part of the corporate network between the corporate intranet and the external network). DMZ-based servers may have limited liability and help mitigate breaches. The DMZ layer 820 may also include a control plane application layer 824 that may include one or more load balancer (LB) subnets 822, an application subnet 826, and a control plane data layer 828 that may include a database (DB) subnet 830 (e.g., a front-end DB subnet and / or a back-end DB subnet). The LB subnet 822 included in the control plane DMZ layer 820 may be coupled to the application subnet 826 included in the control plane application layer 824 and an internet gateway 834 that may be included in the control plane VCN 816, and the application subnet 826 may be coupled to the DB subnet 830, a service gateway 836, and a network address translation (NAT) gateway 838 included in the control plane data layer 828. The control plane VCN816 may include a service gateway 836 and a NAT gateway 838.

[0088] The control plane VCN 816 may include a data plane mirror application layer 840 which may include an application subnet 826. The application subnet 826 included in the data plane mirror application layer 840 may include a virtual network interface controller (VNIC) 842 which may run a compute instance 844. The compute instance 844 may communicate-couple the application subnet 826 of the data plane mirror application layer 840 to an application subnet 826 which may be included in the data plane application layer 846.

[0089] The data plane VCN818 may include a data plane application layer 846, a data plane DMZ layer 848, and a data plane data layer 850. The data plane DMZ layer 848 may include an LB subnet 822 that can be connected to the application subnet 826 of the data plane application layer 846 and the internet gateway 834 of the data plane VCN818. The application subnet 826 may be connected to the service gateway 836 of the data plane VCN818 and the NAT gateway 838 of the data plane VCN818. The data plane data layer 850 may also include a DB subnet 830 that can be connected to the application subnet 826 of the data plane application layer 846.

[0090] The Internet gateway 834 of the control plane VCN816 and data plane VCN818 can be connected to a metadata management service 852, which can be connected to the public internet 854. The public internet 854 can be connected to the NAT gateway 838 of the control plane VCN816 and data plane VCN818. The service gateway 836 of the control plane VCN816 and data plane VCN818 can be connected to a cloud service 856.

[0091] In some cases, a service gateway 836 of the control plane VCN816 or data plane VCN818 can make application programming interface (API) calls to a cloud service 856 without going through the public internet 854. API calls from the service gateway 836 to the cloud service 856 can be unidirectional. The service gateway 836 can make API calls to the cloud service 856, and the cloud service 856 can send the requested data to the service gateway 836. However, the cloud service 856 does not have to initiate an API call to the service gateway 836.

[0092] In some examples, secure host tenancy 804 may be directly connected to service tenancy 819, or otherwise isolated. Secure host subnet 808 can communicate with SSH subnet 814 via LPG 810, which can enable bidirectional communication through systems that would otherwise be isolated. By connecting secure host subnet 808 to SSH subnet 814, secure host subnet 808 becomes able to access other entities within service tenancy 819.

[0093] The control plane VCN816 may enable users of the service tenancy 819 to configure or provision desired resources. Desired resources provisioned in the control plane VCN816 may be deployed or used in the data plane VCN818. In some examples, the control plane VCN816 can be isolated from the data plane VCN818, and the data plane mirror application layer 840 of the control plane VCN816 can communicate with the data plane application layer 846 of the data plane VCN818 via a VNIC 842 which may be included in the data plane mirror application layer 840 and the data plane application layer 846.

[0094] In some examples, a system user, or customer, can make a request (for example, create, read, update, or erase an operation (CRUD)) through the public internet 854, and the public internet 854 can send the request to the metadata management service 852. The metadata management service 852 can send the request to the control plane VCN 816 through the internet gateway 834. The request may be received by the LB subnet 822, which is included in the control plane DMZ layer 820. The LB subnet 822 may determine that the request is valid, and in response to this determination, the LB subnet 822 may send the request to the application subnet 826, which is included in the control plane application layer 824. If the request is validated and a call to the public internet 854 is required, the call to the public internet 854 may be sent to a NAT gateway 838, which can make a call to the public internet 854. Metadata that is deemed desirable to be stored by the request may be stored in the DB subnet 830.

[0095] In some examples, the data plane mirror application layer 840 may facilitate direct communication between the control plane VCN816 and the data plane VCN818. For example, it may be desirable that configuration changes, updates, or other preferred modifications be applied to resources contained in the data plane VCN818. The control plane VCN816 can perform configuration changes, updates, or other preferred modifications to resources by communicating directly with the resources contained in the data plane VCN818 via VNIC842.

[0096] In some embodiments, the control plane VCN816 and the data plane VCN818 may be included in the service tenancy 819. In this case, the system user, i.e., the customer, does not have to own either the control plane VCN816 or the data plane VCN818, or does not have either of them running. Alternatively, the IaaS provider may own both the control plane VCN816 and the data plane VCN818, or have both running, or both may be included in the service tenancy 819. This embodiment may enable network isolation that can prevent interaction between the user, i.e., the customer, and other users, i.e., other customers' resources. Furthermore, this embodiment may enable private storage of databases by the system user, i.e., the customer, without having to rely on the public internet 854, which may not have the desired level of threat prevention for storage.

[0097] In another embodiment, the LB subnet 822 included in the control plane VCN 816 can be configured to receive signals from the service gateway 836. In this embodiment, the control plane VCN 816 and the data plane VCN 818 may be configured to be invoked by the IaaS provider's customer without calling the public internet 854. The IaaS provider's customer is likely to prefer this embodiment because the database used by the customer may be stored in a service tenancy 819 that is controlled by the IaaS provider and can be isolated from the public internet 854.

[0098] Figure 9 is a block diagram 900 showing another exemplary pattern of an IaaS architecture according to at least one embodiment. A service operator 902 (e.g., service operator 802 in Figure 8) may be connected to a secure host tenancy 904 (e.g., secure host tenancy 804 in Figure 8), which may include a virtual cloud network (VCN) 906 (e.g., VCN806 in Figure 8) and a secure host subnet 908 (e.g., secure host subnet 808 in Figure 8). The VCN 906 may include an LPG 910, which may be connected to the SSH VCN 912 via a local peering gateway (LPG) 910 (e.g., LPG810 in Figure 8), which is included in the Secure Shell (SSH) VCN 912 (e.g., SSH VCN812 in Figure 8). SSH VCN912 may include SSH subnet 914 (for example, SSH subnet 814 in Figure 8), and SSH VCN912 may be communicated with control plane VCN916 via LPG910 included in control plane VCN916 (for example, control plane VCN816 in Figure 8). Control plane VCN916 may be included in service tenancy 919 (for example, service tenancy 819 in Figure 8), and data plane VCN918 (for example, data plane VCN818 in Figure 8) may be included in customer tenancy 921, which may be owned or operated by the system's users, i.e., customers.

[0099] The control plane VCN916 may include a control plane DMZ layer 920 (for example, the control plane DMZ layer 820 in Figure 8) which may include an LB subnet 922 (for example, the LB subnet 822 in Figure 8), a control plane application layer 924 (for example, the control plane application layer 824 in Figure 8) which may include an application subnet 926 (for example, the application subnet 826 in Figure 8), and a control plane data layer 928 (for example, the control plane data layer 828 in Figure 8) which may include a DB subnet 930 (similar to the database (DB) subnet 830 in Figure 8). The LB subnet 922 included in the control plane DMZ layer 920 is connected to the application subnet 926 included in the control plane application layer 924 and the Internet gateway 934 (for example, Internet gateway 834 in Figure 8) which may be included in the control plane VCN 916. The application subnet 926 may be connected to the DB subnet 930, the service gateway 936 (for example, service gateway 836 in Figure 8), and the Network Address Translation (NAT) gateway 938 (for example, NAT gateway 838 in Figure 8) included in the control plane data layer 928. The control plane VCN 916 may include the service gateway 936 and the NAT gateway 938.

[0100] The control plane VCN916 may include a data plane mirror application layer 940 (for example, the data plane mirror application layer 840 in Figure 8) which may include an application subnet 926. The application subnet 926 included in the data plane mirror application layer 940 may include a virtual network interface controller (VNIC) 942 (for example, VNIC842) which may run a compute instance 944 (for example, similar to compute instance 844 in Figure 8). The compute instance 944 may facilitate communication between the application subnet 926 of the data plane mirror application layer 940 and the application subnet 926 that may be included in the data plane application layer 946 via the VNIC942 included in the data plane mirror application layer 940 and the VNIC942 included in the data plane application layer 946 (for example, the data plane application layer 846 in Figure 8).

[0101] The Internet gateway 934 included in the control plane VCN916 can be connected to a metadata management service 952 (e.g., metadata management service 852 in Figure 8), which can be connected to the public internet 954 (e.g., public internet 854 in Figure 8). The public internet 954 can be connected to a NAT gateway 938 included in the control plane VCN916. The service gateway 936 included in the control plane VCN916 can be connected to a cloud service 956 (e.g., cloud service 856 in Figure 8).

[0102] In some examples, the data plane VCN918 may be included in a customer tenancy 921. In this case, the IaaS provider may provide a control plane VCN916 for each customer, or the IaaS provider may configure a unique compute instance 944 included in a service tenancy 919 for each customer. Each compute instance 944 may enable communication between the control plane VCN916 included in the service tenancy 919 and the data plane VCN918 included in the customer tenancy 921. The compute instance 944 may enable the deployment or use of resources provisioned in the control plane VCN916 included in the service tenancy 919 in the data plane VCN918 included in the customer tenancy 921.

[0103] In another example, the IaaS provider's customer may have a database located in customer tenancy 921. In this example, the control plane VCN916 may include a data plane mirror application layer 940, which may include an application subnet 926. The data plane mirror application layer 940 may reside in data plane VCN918, but may not. That is, the data plane mirror application layer 940 may be accessible to customer tenancy 921, but may not reside in data plane VCN918, nor may it be owned or operated by the IaaS provider's customer. The data plane mirror application layer 940 may be configured to make calls to data plane VCN918, but may not be configured to make calls to any entities included in control plane VCN916. The customer is expected to want to deploy or use resources in the data plane VCN918 that are provisioned in the control plane VCN916, and the data plane mirror application layer 940 can facilitate the deployment or other use of the resources desired by the customer.

[0104] In some embodiments, a customer of the IaaS provider can apply filters to the data plane VCN918. In this embodiment, the customer can determine what the data plane VCN918 can access, and may also restrict access from the data plane VCN918 to the public internet 954. The IaaS provider may not be able to apply filters or control the data plane VCN918's access to any external network or database. The application of filters and controls by the customer to the data plane VCN918 included in the customer tenancy 921 may help isolate the data plane VCN918 from other customers and the public internet 954.

[0105] In some embodiments, the cloud service 956 becomes accessible by a call from the service gateway 936 to a service that could not exist on the public internet 954, the control plane VCN916, or the data plane VCN918. The connection between the cloud service 956 and the control plane VCN916 or data plane VCN918 does not have to be live or continuous. The cloud service 956 may reside on different networks owned or operated by the IaaS provider. The cloud service 956 may be configured to receive calls from the service gateway 936 and may be configured not to receive calls from the public internet 954. Some cloud services 956 may be isolated from other cloud services 956, and the control plane VCN916 may be isolated from cloud services 956 that could not be in the same region as the control plane VCN916. For example, the control plane VCN916 may be located in "Region 1", and the cloud service "Deployment 8" may be located in Region 1 and "Region 2". If a call to deployment 8 is made by a service gateway 936 included in the control plane VCN916 located in region 1, the call may be sent to deployment 8 in region 1. In this example, the control plane VCN916 or deployment 8 in region 1 does not have to be communication-coupled to or in communication with deployment 8 in region 2.

[0106] Figure 10 is a block diagram 1000 showing another exemplary pattern of an IaaS architecture according to at least one embodiment. A service operator 1002 (e.g., service operator 802 in Figure 8) may be connected to a secure host tenancy 1004 (e.g., secure host tenancy 804 in Figure 8), which may include a virtual cloud network (VCN) 1006 (e.g., VCN806 in Figure 8) and a secure host subnet 1008 (e.g., secure host subnet 808 in Figure 8). VCN 1006 may include an LPG 1010 (e.g., LPG810 in Figure 8), which may be connected to SSH VCN 1012 (e.g., SSH VCN812 in Figure 8) via an LPG 1010 (e.g., LPG810 in Figure 8). SSH VCN1012 may include SSH subnet 1014 (for example, SSH subnet 814 in Figure 8), and SSH VCN1012 may be connected to control plane VCN1016 via LPG1010 included in control plane VCN1016 (for example, control plane VCN816 in Figure 8), and may be connected to data plane VCN1018 via LPG1010 included in data plane VCN1018 (for example, data plane VCN818 in Figure 8). Control plane VCN1016 and data plane VCN1018 may be included in service tenancy 1019 (for example, service tenancy 819 in Figure 8).

[0107] The control plane VCN1016 may include a control plane DMZ layer 1020 (e.g., control plane DMZ layer 820 in Figure 8) which may include a load balancer (LB) subnet 1022 (e.g., LB subnet 822 in Figure 8), a control plane application layer 1024 (e.g., control plane application layer 824 in Figure 8) which may include an application subnet 1026 (e.g., similar to application subnet 826 in Figure 8), and a control plane data layer 1028 (e.g., control plane data layer 828 in Figure 8) which may include a DB subnet 1030. The LB subnet 1022 included in the control plane DMZ layer 1020 is connected to the application subnet 1026 included in the control plane application layer 1024 and the Internet gateway 1034 (for example, Internet gateway 834 in Figure 8), which may be included in the control plane VCN 1016. The application subnet 1026 may be connected to the DB subnet 1030 included in the control plane data layer 1028, the service gateway 1036 (for example, the service gateway in Figure 8), and the Network Address Translation (NAT) gateway 1038 (for example, the NAT gateway 838 in Figure 8). The control plane VCN 1016 may include the service gateway 1036 and the NAT gateway 1038.

[0108] The data plane VCN 1018 may include a data plane application layer 1046 (for example, the data plane application layer 846 in Figure 8), a data plane DMZ layer 1048 (for example, the data plane DMZ layer 848 in Figure 8), and a data plane data layer 1050 (for example, the data plane data layer 850 in Figure 8). The data plane DMZ layer 1048 may include a trusted application subnet 1060 and an untrusted application subnet 1062 of the data plane application layer 1046, as well as an LB subnet 1022 that can be connected to the internet gateway 1034 included in the data plane VCN 1018. The trusted application subnet 1060 may be connected to a service gateway 1036 included in the data plane VCN 1018, a NAT gateway 1038 included in the data plane VCN 1018, and a DB subnet 1030 included in the data plane data layer 1050. The non-trusted application subnet 1062 may be connected to the service gateway 1036 included in the data plane VCN 1018 and the DB subnet 1030 included in the data plane data layer 1050. The data plane data layer 1050 may include the DB subnet 1030, which may be connected to the service gateway 1036 included in the data plane VCN 1018.

[0109] The untrusted application subnet 1062 may contain one or more primary VNICs 1064(1) to 1064(N) that can be connected to tenant virtual machines (VMs) 1066(1) to 1066(N). Each tenant VM 1066(1) to 1066(N) may be connected to each application subnet 1067(1) to 1067(N) that may be included in each container output VCN 1068(1) to 1068(N) that may be included in each customer tenancy 1070(1) to 1070(N). Each secondary VNIC 1072(1) to 1072(N) may facilitate communication between the untrusted application subnet 1062 included in the data plane VCN 1018 and the application subnets included in the container output VCNs 1068(1) to 1068(N). Each container output VCN 1068(1) to 1068(N) may include a NAT gateway 1038 that can be connected to the public internet 1054 (for example, the public internet 854 in Figure 8).

[0110] The Internet gateway 1034, included in the control plane VCN1016 and data plane VCN1018, can be connected to a metadata management service 1052 (for example, the metadata management system 852 in Figure 8), which can be connected to the public internet 1054. The public internet 1054 can be connected to a NAT gateway 1038, included in the control plane VCN1016 and data plane VCN1018. The service gateway 1036, included in the control plane VCN1016 and data plane VCN1018, can be connected to a cloud service 1056.

[0111] In some embodiments, the data plane VCN 1018 may be integrated with a customer tenancy 1070. This integration may be useful or desirable for the IaaS provider's customer, for example, if they may want support for executing code. The customer may provide code to be executed, which may be destructive, communicate with other customer resources, or have undesirable effects. In response, the IaaS provider may determine whether to execute the code provided to the IaaS provider by the customer.

[0112] In some examples, an IaaS provider's customer may grant temporary network access to the IaaS provider and request functionality to be granted to the data plane application layer 1046. The code that performs this functionality may run in VMs 1066(1) to 1066(N), and may not be configured to run elsewhere on the data plane VCN 1018. Each VM 1066(1) to 1066(N) may be connected to a single customer tenancy 1070. Each container 1071(1) to 1071(N) contained within VMs 1066(1) to 1066(N) may be configured to execute code. In this case, a double isolation may exist (containers 1071(1)-1071(N) executing the code may be contained in VMs 1066(1)-1066(N) that are at least in the non-trusted app subnet 1062), which may help prevent damage to the IaaS provider's network or a different customer's network by erroneous or undesirable code. Containers 1071(1)-1071(N) may be communication coupled to customer tenancy 1070 and may be configured to send or receive data to or from customer tenancy 1070. Containers 1071(1)-1071(N) may also be configured not to send or receive data to or from any other entity in the data plane VCN 1018. Upon completion of code execution, the IaaS provider may disable or discard containers 1071(1)-1071(N).

[0113] In some embodiments, the trusted application subnet 1060 may be configured to execute code owned or operated by the IaaS provider. In this embodiment, the trusted application subnet 1060 may be connected to the DB subnet 1030 and may be configured to perform CRUD operations in the DB subnet 1030. The non-trusted application subnet 1062 may be connected to the DB subnet 1030, but in this embodiment, may be configured to perform read operations in the DB subnet 1030. Containers 1071(1) to 1071(N) contained in each customer's VMs 1066(1) to 1066(N) and capable of executing code from the customer may not be connected to the DB subnet 1030.

[0114] In other embodiments, the control plane VCN1016 and the data plane VCN1018 do not have to be directly connected. In this embodiment, direct communication between the control plane VCN1016 and the data plane VCN1018 is not required. However, communication can be performed indirectly by at least one method. An LPG1010 that can facilitate communication between the control plane VCN1016 and the data plane VCN1018 may be established by the IaaS provider. In another example, the control plane VCN1016 or the data plane VCN1018 can make a call to the cloud service 1056 via the service gateway 1036. For example, a call from the control plane VCN1016 to the cloud service 1056 may include a request for a service that can communicate with the data plane VCN1018.

[0115] Figure 11 is a block diagram 1100 showing another exemplary pattern of an IaaS architecture according to at least one embodiment. A service operator 1102 (e.g., service operator 802 in Figure 8) may be connected to a secure host tenancy 1104 (e.g., secure host tenancy 804 in Figure 8), which may include a virtual cloud network (VCN) 1106 (e.g., VCN806 in Figure 8) and a secure host subnet 1108 (e.g., secure host subnet 808 in Figure 8). VCN 1106 may include an LPG 1110 (e.g., LPG810 in Figure 8), which may be connected to SSH VCN 1112 (e.g., SSH VCN812 in Figure 8). SSH VCN1112 may include SSH subnet 1114 (for example, SSH subnet 814 in Figure 8), and SSH VCN1112 may be connected to control plane VCN1116 via LPG1110 included in control plane VCN1116 (for example, control plane VCN816 in Figure 8), and may be connected to data plane VCN1118 via LPG1110 included in data plane VCN1118 (for example, data plane VCN818 in Figure 8). Control plane VCN1116 and data plane VCN1118 may be included in service tenancy 1119 (for example, service tenancy 819 in Figure 8).

[0116] The control plane VCN1116 may include a control plane DMZ layer 1120 (for example, the control plane DMZ layer 820 in Figure 8) which may include an LB subnet 1122 (for example, the LB subnet 822 in Figure 8), a control plane application layer 1124 (for example, the control plane application layer 824 in Figure 8) which may include an application subnet 1126 (for example, the application subnet 826 in Figure 8), and a control plane data layer 1128 (for example, the control plane data layer 828 in Figure 8) which may include a DB subnet 1130 (for example, the DB subnet 1030 in Figure 10). The LB subnet 1122 included in the control plane DMZ layer 1120 is connected to the application subnet 1126 included in the control plane application layer 1124 and the Internet gateway 1134 (for example, Internet gateway 834 in Figure 8), which may be included in the control plane VCN 1116. The application subnet 1126 may be connected to the DB subnet 1130, the service gateway 1136 (for example, the service gateway in Figure 8), and the Network Address Translation (NAT) gateway 1138 (for example, the NAT gateway 838 in Figure 8), which are included in the control plane data layer 1128. The control plane VCN 1116 may include the service gateway 1136 and the NAT gateway 1138.

[0117] The data plane VCN 1118 may include a data plane application layer 1146 (for example, the data plane application layer 846 in Figure 8), a data plane DMZ layer 1148 (for example, the data plane DMZ layer 848 in Figure 8), and a data plane data layer 1150 (for example, the data plane data layer 850 in Figure 8). The data plane DMZ layer 1148 may include the trusted application subnet 1160 (for example, the trusted application subnet 1060 in Figure 10) and the untrusted application subnet 1162 (for example, the untrusted application subnet 1062 in Figure 10) of the data plane application layer 1146, as well as an LB subnet 1122 that can be connected to the internet gateway 1134 included in the data plane VCN 1118. Trusted application subnet 1160 can be connected to service gateway 1136 included in data plane VCN 1118, NAT gateway 1138 included in data plane VCN 1118, and DB subnet 1130 included in data plane data layer 1150. Non-trusted application subnet 1162 can be connected to service gateway 1136 included in data plane VCN 1118 and DB subnet 1130 included in data plane data layer 1150. Data plane data layer 1150 may include DB subnet 1130 which can be connected to service gateway 1136 included in data plane VCN 1118.

[0118] The untrusted application subnet 1162 may include primary VNICs 1164(1) to 1164(N) that can be connected to tenant virtual machines (VMs) 1166(1) to 1166(N) residing within the untrusted application subnet 1162. Each tenant VM 1166(1) to 1166(N) can execute code in each of the containers 1167(1) to 1167(N) and can be connected to an application subnet 1126 that may be included in the data plane application layer 1146, which may be included in the container output VCN 1168. The secondary VNICs 1172(1) to 1172(N) can each facilitate communication between the untrusted application subnet 1162 included in the data plane VCN 1118 and the application subnet included in the container output VCN 1168. The container output VCN may include a NAT gateway 1138 that can be connected to the public internet 1154 (for example, the public internet 854 in Figure 8).

[0119] The Internet gateway 1134, included in the control plane VCN1116 and data plane VCN1118, can be connected to a metadata management service 1152 (for example, the metadata management system 852 in Figure 8), which can be connected to the public internet 1154. The public internet 1154 can be connected to a NAT gateway 1138, included in the control plane VCN1116 and data plane VCN1118. The service gateway 1136, included in the control plane VCN1116 and data plane VCN1118, can be connected to a cloud service 1156.

[0120] In some examples, the architecture pattern shown in block diagram 1100 of Figure 11 is considered an exception to the architecture pattern shown in block diagram 1000 of Figure 10, and is considered desirable for the IaaS provider's customers when the IaaS provider cannot communicate directly with the customer (for example, in an unconnected region). Each customer has containers 1167(1) to 1167(N) contained within VMs 1166(1) to 1166(N), each of which is accessible to the customer in real time. Each container 1167(1) to 1167(N) may be configured to make calls to each of the secondary VNICs 1172(1) to 1172(N) contained within the application subnet 1126 of the data plane application layer 1146, which may be contained within the container output VCN 1168. Secondary VNICs 1172(1) to 1172(N) can send calls to the NAT gateway 1138, which may then send calls to the public internet 1154. In this example, the containers 1167(1) to 1167(N), which are accessible to customers in real time, can be isolated from the control plane VCN 1116, as well as from other entities included in the data plane VCN 1118. Furthermore, the containers 1167(1) to 1167(N) can also be isolated from other customers' resources.

[0121] In another example, a customer can use containers 1167(1) to 1167(N) to invoke cloud service 1156. In this example, a customer can execute code in containers 1167(1) to 1167(N) to request services from cloud service 1156. Containers 1167(1) to 1167(N) can send this request to secondary VNICs 1172(1) to 1172(N), which can send the request to a NAT gateway, which can send the request to the public internet 1154. The public internet 1154 can send the request to LB subnet 1122, which is included in control plane VCN 1116, via internet gateway 1134. In response to the determination that the request is valid, the LB subnet can send the request to the application subnet 1126, and the application subnet 1126 can send the request to the cloud service 1156 via the service gateway 1136.

[0122] Naturally, the IaaS architectures 800, 900, 1000, and 1100 shown in the drawings may have components other than those shown. Furthermore, the embodiments shown in the drawings are only a few examples of cloud infrastructure systems that may encompass one embodiment of this disclosure. In some other embodiments, the IaaS system may have more or fewer components than shown, may combine two or more components, or may have different component configurations or arrangements.

[0123] In one embodiment, the IaaS system described herein may include a set of applications, middleware, and database services delivered to customers in a self-service, subscription-based, elastically scalable, reliable, highly available, and secure manner. Oracle Cloud Infrastructure (OCI), offered by the assignee, is an example of such an IaaS system.

[0124] Figure 12 shows an exemplary computer system 1200 in which various embodiments can be realized. System 1200 may be used to realize any of the computer systems described above. As shown in the drawing, computer system 1200 includes a processing unit 1204 that communicates with a number of peripheral subsystems via a bus subsystem 1202. Peripheral subsystems may include a processing acceleration unit 1206, an I / O subsystem 1208, a storage subsystem 1218, and a communication subsystem 1224. The storage subsystem 1218 includes a tangible computer-readable storage medium 1222 and system memory 1210.

[0125] The bus subsystem 1202 provides a mechanism for various components and subsystems of the computer system 1200 to communicate with each other as intended. While the bus subsystem 1202 is schematically shown as a single bus, alternative embodiments of the bus subsystem may utilize multiple buses. The bus subsystem 1202 may be one of several types of bus structures, including a memory bus or memory controller, a peripheral bus, and a local bus, using any of various bus architectures. For example, such architectures may include industry standard architecture (ISA) buses, microchannel architecture (MCA) buses, extended ISA (EISA) buses, video electronics standards (VESA) local buses, and peripheral interconnect (PCI) buses, which can be implemented as mezzanine buses manufactured according to the IEEE P1386.1 standard.

[0126] A processing unit 1204, which can be implemented as one or more integrated circuits (e.g., conventional microprocessors or microcontrollers), controls the operation of the computer system 1200. One or more processors may be included in the processing unit 1204. These processors may include single-core or multi-core processors. In some embodiments, the processing unit 1204 may be implemented as one or more independent processing units 1232 and / or 1234, each containing a single-core or multi-core processor. In other embodiments, the processing unit 1204 may also be implemented as a quad-core processing unit formed by integrating two dual-core processors onto a single chip.

[0127] In various embodiments, the processing unit 1204 can execute various programs in response to program code and maintain multiple concurrently running programs or processes. At any given time, some or all of the program code to be executed may reside in the processor 1204 and / or the storage subsystem 1218. With suitable programming, the processor 1204 can provide the various functions described above. The computer system 1200 may also include a processing acceleration unit 1206 which may include a digital signal processor (DSP), a dedicated processor, and / or similar.

[0128] The I / O subsystem 1208 may include user interface input devices and user interface output devices. User interface input devices may include pointing devices such as keyboards, mice or trackballs, touchpads or touchscreens integrated into displays, scroll wheels, click wheels, dials, buttons, switches, keypads, audio input devices with voice command recognition systems, microphones, and other types of input devices. User interface input devices may include motion detection and / or gesture recognition devices such as Microsoft Kinect® motion sensors that enable user control and interaction with input devices such as Microsoft Xbox® 360 game controllers through a natural user interface using gestures and voice commands. User interface input devices may also include eye gesture recognition devices such as Google Glass® blink detectors that detect the user's eye activity (e.g., blinking while taking photos and / or selecting menus) and convert eye gestures into input to an input device (e.g., Google Glass®). Furthermore, the user interface input device may include a voice recognition detection device that enables user interaction with a voice recognition system (e.g., Siri® Navigator) via voice commands.

[0129] Furthermore, user interface input devices may include, but are not limited to, three-dimensional (3D) mice, joysticks or pointing sticks, gamepads and graphic tablets, as well as auditory / visual devices such as speakers, digital cameras, digital video cameras, portable media players, webcams, image scanners, fingerprint scanners, barcode readers, 3D scanners, 3D printers, laser rangefinders, and eye-tracking devices. User interface input devices may also include, for example, medical imaging input devices such as computed tomography, magnetic resonance imaging, positron emission tomography, and medical ultrasound imaging devices. Additionally, user interface input devices may include, for example, audio input devices such as MIDI keyboards and digital musical instruments.

[0130] User interface output devices may include non-visual displays such as display subsystems, indicator lights, or audio output devices. Display subsystems may also include flat panel devices such as cathode ray tubes (CRTs), liquid crystal displays (LCDs), or plasma displays, projection devices, touchscreens, etc. Generally, the use of the term “output device” is intended to include all conceivable types of devices and mechanisms for outputting information from the computer system 1200 to a user or another computer. For example, user interface output devices may include, but are not limited to, various display devices that visually convey text, graphics, and audio / video information, such as monitors, printers, speakers, headphones, car navigation systems, plotters, audio output devices, and modems.

[0131] The computer system 1200 may include a storage subsystem 1218 that provides a tangible, non-temporary, computer-readable storage medium for storing software and data constructs that provide the functionality of the embodiments described in this disclosure. The software may include programs, code modules, instructions, scripts, etc., that provide the above-described functionality when executed by one or more cores or processors of the processing unit 1204. The storage subsystem 1218 may also provide a repository for storing data used in accordance with this disclosure.

[0132] As shown in the example in Figure 12, the storage subsystem 1218 may include various components, including system memory 1210, computer-readable storage medium 1222, and computer-readable storage medium reader 1220. System memory 1210 may store program instructions that can be loaded and executed by the processing unit 1204. System memory 1210 may also store data used when instructions are executed and / or data generated when program instructions are executed. Various different types of programs may be loaded into system memory 1210, including, but not limited to, client applications, web browsers, middle-tier applications, relational database management systems (RDBMS), virtual machines, containers, etc.

[0133] Furthermore, the system memory 1210 may be configured to store the operating system 1216. Examples of operating systems 1216 include various versions of Microsoft Windows®, Apple Macintosh®, and / or Linux operating systems, various commercially available UNIX® or UNIX-like operating systems (including, but not limited to, various GNU / Linux operating systems, Google Chrome® OS, etc.), and / or mobile operating systems such as iOS, Windows® Phone, Android® OS, BlackBerry® OS, and Palm® OS. In some embodiments in which the computer system 1200 runs one or more virtual machines, the virtual machines, along with their respective guest operating systems (GOS), may be loaded into the system memory 1210 and executed by one or more processors or cores of the processing unit 1204.

[0134] The system memory 1210 can have various configurations depending on the type of computer system 1200. For example, the system memory 1210 may be volatile memory (such as random access memory (RAM)) and / or non-volatile memory (such as read-only memory (ROM) or flash memory). It may also be provided with different types of RAM configurations, including static random access memory (SRAM), dynamic random access memory (DRAM), and others. In some embodiments, the system memory 1210 may include a basic input / output system (BIOS) that includes basic routines useful for communicating information between elements within the computer system 1200, such as during startup.

[0135] The computer-readable storage medium 1222 may represent a remote, local, fixed, and / or removable storage device, as well as a storage medium for temporarily and / or permanently containing and storing computer-readable information used by the computer system 1200 (including instructions that can be executed by the processing unit 1204 of the computer system 1200).

[0136] The computer-readable storage medium 1222 may include, but is not limited to, any suitable medium known or used in the art (including storage and communication media), and includes volatile and non-volatile, removable and non-removable media implemented in any method or technique for storing and / or transmitting information. This may include memory technologies such as RAM, ROM, electronically erasable programmable ROM (EEPROM), and flash memory; optical storage such as CD-ROM and digital multipurpose discs (DVDs); magnetic storage devices such as magnetic cassettes, magnetic tapes, and magnetic disk storage; or other tangible computer-readable storage media.

[0137] For example, the computer-readable storage medium 1222 may include a hard disk drive that reads and writes to a non-removable non-volatile magnetic medium, a magnetic disk drive that reads and writes to a removable non-volatile magnetic disk, and an optical disk drive that reads and writes to a removable non-volatile optical disk such as a CD-ROM, DVD, Blu-ray® disc, or other optical medium. The computer-readable storage medium 1222 may include, but is not limited to, Zip® drives, flash memory cards, Universal Serial Bus (USB) flash drives, Secure Digital (SD) cards, DVD discs, digital videotapes, etc. Furthermore, the computer-readable storage medium 1222 may include solid-state drives (SSDs) based on non-volatile memory (flash memory-based SSDs, enterprise flash drives, solid-state ROMs, etc.), SSDs based on volatile memory (solid-state RAM, dynamic RAM, static RAM, DRAM-based SSDs, magnetoresistive RAM (MRAM) SSDs), and hybrid SSDs (using a combination of DRAM and flash memory-based SSDs). The disk drives and their respective associated computer-readable media may provide non-volatile storage for computer-readable instructions, data structures, program modules, and other data to the computer system 1200.

[0138] Machine-readable instructions executable by one or more processors or cores of the processing unit 1204 may be stored in a non-temporary computer-readable storage medium. The non-temporary computer-readable storage medium may include physically tangible memory or storage devices, including volatile memory devices and / or non-volatile storage devices. Examples of non-temporary computer-readable storage media include magnetic storage media (e.g., disks or tapes), optical storage media (e.g., DVDs, CDs), various types of RAM, ROM, or flash memory, hard drives, floppy disk drives, removable memory drives (e.g., USB drives), or other types of storage devices.

[0139] The communication subsystem 1224 provides an interface to other computer systems and networks. The communication subsystem 1224 functions as an interface for sending and receiving data between the computer system 1200 and other systems. For example, the communication subsystem 1224 may enable the computer system 1200 to connect to one or more devices via the Internet. In some embodiments, the communication subsystem 1224 may include a wireless voice and / or data network (using, for example, cellular technology, 3G, 4G, or EDGE (Enhanced Data Rates for Global Evolution), advanced data network technologies such as WiFi (IEEE 802.11 family standards), or other mobile communication technologies, or any combination thereof), a Global Positioning System (GPS) receiver component, and / or a radio frequency (RF) transceiver component for accessing other components. In some embodiments, the communication subsystem 1224 may provide a wired network connection (e.g., Ethernet) as an addition to or alternative to the wireless interface.

[0140] In addition, in some embodiments, the communication subsystem 1224 is capable of receiving incoming communications in the form of structured and / or unstructured data feeds 1226, event streams 1228, event updates 1230, etc., on behalf of one or more users who may be using the computer system 1200.

[0141] For example, the communication subsystem 1224 may be configured to receive data feeds 1226 in real time from users of social networks, and / or from users of other communication services, such as Twitter® feeds, Facebook® updates, RSS (Rich Site Summary) feeds, and / or real-time updates from one or more third-party sources.

[0142] Furthermore, the communication subsystem 1224 may be configured to receive data in the form of a continuous data stream, which may include an event stream 1228 and / or event update 1230 of real-time events, which may be continuous or virtually infinite with no explicit end. Examples of applications that generate continuous data include, for example, sensor data applications, financial tickers, network performance measurement tools (e.g., network monitoring and traffic management applications), clickstream analysis tools, and automotive traffic monitoring.

[0143] Furthermore, the communication subsystem 1224 may be configured to output structured and / or unstructured data feeds 1226, event streams 1228, event updates 1230, etc., to one or more databases that can communicate with one or more streaming data source computers connected to the computer system 1200.

[0144] The computer system 1200 can be one of a variety of types, including portable handheld devices (e.g., iPhone® mobile phones, iPad® computing tablets, PDAs, etc.), wearable devices (e.g., Google Glass® head-mounted displays), PCs, workstations, mainframes, kiosks, server racks, or any other data processing systems.

[0145] Due to the constantly changing nature of computers and networks, the description of the illustrated computer system 1200 is intended only as an example. Many other configurations with more or fewer components than the illustrated system are possible. For example, the use of customized hardware and / or the implementation of specific elements in hardware, firmware, software (including applets), or combinations thereof are also possible. Furthermore, connections to other computing devices such as network input / output devices may be employed. Based on the disclosures and teachings contained herein, those skilled in the art will recognize other methods and / or ways of implementing various embodiments.

[0146] While specific embodiments have been described above, various improvements, modifications, alternative configurations, and equivalents are also included in the scope of this disclosure. The embodiments are not limited to operation within a few specific data processing environments, but can freely operate within multiple data processing environments. Furthermore, while the embodiments have been described using a specific set of transactions and steps, it will be clear to those skilled in the art that the scope of this disclosure is not limited to the set of transactions and steps described. The various features and aspects of the embodiments described above may be used individually or in combination.

[0147] Furthermore, while embodiments have been described using specific combinations of hardware and software, it will be recognized that other combinations of hardware and software are also included in the scope of this disclosure. Embodiments may be implemented in hardware only, in software only, or by using a combination of these. The various processes described herein may be implemented on the same processor or on any combination of different processors. Thus, where a component or service is described as being configured to perform some operation, such configuration may be achieved, for example, by designing electronic circuits to perform this operation, by programming programmable electronic circuits (such as a microprocessor) to perform this operation, or by any combination thereof. Processes may be communicative using various techniques, including but not limited to conventional techniques for inter-process communication. Also, different pairs of processes may use different techniques, or the same pair of processes may use different techniques at different times.

[0148] Therefore, this specification and the drawings are intended to be illustrative and not limiting in any way. However, it is clear that additions, reductions, deletions, and other improvements and modifications can be made without departing from the broad idea and scope set forth in the claims. For this reason, specific embodiments of this disclosure have been described, but these are not intended to be limiting in any way. Various improvements and equivalents are included in the following claims.

[0149] In the context describing embodiments of the disclosure (in particular, in the context of the following claims), the terms “a,” “an,” and “the,” and similar reference subjects, shall be interpreted as referring to both singular and plural forms unless otherwise indicated herein or there is a clear contextual inconsistency. The terms “comprising,” “having,” “including,” and “containing” shall be interpreted as open-ended terms (i.e., “including, but not limited to,”) unless otherwise specified. The term “connected” shall be interpreted as including, attaching, or integrally combining, in whole or in part, even if something is intervening. Unless otherwise indicated herein, descriptions of ranges of values ​​are intended merely as a concise way of referring individually to each distinct value contained within that range, and each distinct value is incorporated herein as if it were individually described herein. Unless otherwise indicated herein or there is a clear contextual inconsistency, all methods described herein may be performed in any preferred order. The use of any examples or illustrative expressions (e.g., "such as") described herein is intended solely to facilitate the understanding of the embodiments and, unless otherwise claimed, does not limit the scope of this disclosure. Nothing described herein shall be construed as indicating that any non-claimed element is essential to the implementation of this disclosure.

[0150] Unless otherwise specified, disjunctive expressions such as "at least one of X, Y, or Z" are intended to be understood in contexts where they are commonly used to indicate that an item, term, etc., can be any one of X, Y, or Z, or any combination thereof (e.g., X, Y, and / or Z). Therefore, such disjunctive expressions are not intended, nor should they be used, to imply that certain embodiments require the presence of at least one of X, at least one of Y, or at least one of Z.

[0151] This specification describes preferred embodiments of the Disclosure, including the best known modes for the execution of the Disclosure. Those skilled in the art will be able to see, by reading the above description, variations of these preferred embodiments. Those skilled in the art may adopt such variations as appropriate, and the Disclosure may be executed in a manner different from the specific description herein. Therefore, to the extent permitted by applicable law, the Disclosure includes all improvements and equivalents to the subject matter described in the claims appended herein. Furthermore, unless otherwise indicated herein, any combination of all conceivable variations of the elements described above is included in the Disclosure.

[0152] All references cited herein, including publications, patent applications, and patents, are incorporated herein by reference to the same extent as all their contents are included herein, with each reference individually and specifically indicated as being incorporated by reference.

[0153] While the above specification describes aspects of the disclosure with reference to specific embodiments, those skilled in the art will recognize that the disclosure is not limited thereto. The various features and aspects of the disclosure described above may be used individually or in combination. Furthermore, embodiments may be used in any number of environments and applications beyond those described herein, without departing from the broader concept and scope of this specification. Accordingly, this specification and the drawings should be considered illustrative and not limiting.

Claims

1. A method by which a computer performs an action, A computing system identifies text in a visually rich document, The computing system includes determining the sequence of the identification text, The sequence includes the number order of each word in the identification text, and the method further includes The computing system selects a language model based at least partially on the identified text and the decision sequence. The computing system includes generating textual features corresponding to the identified text by assigning each word of the identified text to its respective token using the selected language model and the decision sequence, Each token contains a string of one or more words, and the method further, The computing system includes extracting visual features corresponding to the identified text, The visual feature includes information about multiple pixels representing each word of the identification text, and the method further, The computing system includes determining the positional characteristics of each word in the identified text. The positional feature includes the respective coordinates of each word of the identified text within the visual rich document, wherein the respective coordinates of the word are the coordinates corresponding to the position of the word or the coordinates corresponding to the area of ​​the visual rich document containing the word, and the method further, The computing system includes generating a document model representing the visually rich document, The document model includes nodes connected by edges, each node in the document model representing the visual, textual, and spatial features of each word of the identified text, and the method further, The computing system further includes training a classifier to classify each of the words in the identified text, A method wherein the classifier is trained on the document model representing the visually rich document.

2. The method according to claim 1, further comprising the computing system classifying each of the words of the identification text by the classifier.

3. The method according to claim 1 or claim 2, wherein each of the aforementioned words is classified as a key or value of a key-value pair.

4. The method according to any one of claims 1 to 3, wherein the document model is a graph neural network.

5. The method according to any one of claims 1 to 4, wherein the language model is selected at least partially based on the domain of the identified text.

6. The method according to claim 5, wherein the domain includes the language or subject of the identifying text.

7. The method according to any one of claims 1 to 6, wherein the visually rich document includes at least one of an invoice, receipt, insurance form, boarding pass, or identification document.

8. The method according to any one of claims 1 to 7, further comprising classifying text in an incoming visually rich document by deploying the classifier.

9. A non-temporary computer-readable medium that stores multiple instructions, which, when executed by a computer system, perform an action, The aforementioned operation is, A computing system identifies text in a visually rich document, The computing system includes determining a sequence of the identified text, wherein the sequence includes the numerical order of each word in the identified text, and the operation further includes: The computing system selects a language model based at least partially on the identified text and the decision sequence. The computing system includes generating textual features corresponding to the identified text by assigning each word of the identified text to a token using the selected language model and the decision sequence, wherein each token comprises a string of one or more words, and the operation further includes: The computing system includes extracting visual features corresponding to the identified text, wherein the visual features include information about a plurality of pixels representing each word of the identified text, and the operation further includes: The computing system includes determining the positional features of each word in the identified text, the positional features including the respective coordinates of each word in the identified text within the visual rich document, the respective coordinates of each word being the coordinates corresponding to the location of the word or the coordinates corresponding to the area of ​​the visual rich document containing the word, and the operation further includes The computing system includes generating a document model representing the visually rich document, the document model including nodes connected by edges, each node in the document model representing the visual, textual, and spatial features of each word of the identified text, and the operation further includes: The computing system includes training a classifier for classifying each word of the identified text, the classifier being trained against a document model representing the visually rich document, in a non-temporal, computer-readable medium.

10. The non-temporary computer-readable medium according to claim 9, further comprising the computing system classifying each of the words of the identification text by the classifier.

11. The non-temporary computer-readable medium according to claim 9 or 10, wherein each of the aforementioned words is classified as a key or value of a key-value pair.

12. The non-temporary computer-readable medium according to any one of claims 9 to 11, wherein the document model is a graph neural network.

13. The non-temporary computer-readable medium according to any one of claims 9 to 12, wherein the language model is selected at least in part on the domain of the identified text.

14. The non-temporary computer-readable medium according to claim 13, wherein the domain includes the language or subject of the identification text.

15. The non-temporary computer-readable medium according to any one of claims 9 to 14, wherein the visual-rich document includes at least one of an invoice, receipt, insurance form, boarding pass, or identification document.

16. The non-temporary computer-readable medium according to any one of claims 9 to 15, further comprising the computing system classifying text in an incoming visually rich document by deploying the classifier.

17. It is a system, Computer-readable media and, The system comprises one or more processors for executing at least one operation by executing instructions stored in the computer-readable medium, The aforementioned operation is, A computing system identifies text in a visually rich document, The computing system includes determining a sequence of the identified text, wherein the sequence includes the numerical order of each word in the identified text, and the operation further includes: The computing system selects a language model based at least partially on the identified text and the decision sequence. The computing system includes generating textual features corresponding to the identified text by assigning each word of the identified text to a token using the selected language model and the decision sequence, wherein each token comprises a string of one or more words, and the operation further includes: The computing system includes extracting visual features corresponding to the identified text, wherein the visual features include information about a plurality of pixels representing each word of the identified text, and the operation further includes: The computing system includes determining the positional features of each word in the identified text, the positional features including the respective coordinates of each word in the identified text within the visual rich document, the respective coordinates of each word being the coordinates corresponding to the location of the word or the coordinates corresponding to the area of ​​the visual rich document containing the word, and the operation further includes The computing system includes generating a document model representing the visually rich document, the document model including nodes connected by edges, each node in the document model representing the visual, textual, and spatial features of each word of the identified text, and the operation further includes: The computing system includes training a classifier to classify each word of the identified text, the classifier being trained against a document model representing the visually rich document.

18. The system according to claim 17, further comprising the computing system classifying each of the words of the identification text by the classifier.

19. The system according to claim 18, wherein each of the aforementioned words is classified as the key or value of a key-value pair.

20. The document model is a graph neural network, according to any one of claims 17 to 19.

21. The system according to any one of claims 17 to 20, wherein the language model is selected at least partially on the domain of the identified text.

22. The system according to any one of claims 17 to 21, wherein the domain includes the language or subject of the identification text.

23. The system according to any one of claims 16 to 22, further comprising the computing system classifying text in an incoming visually rich document by deploying the classifier.