System and method for field extraction from unlabeled data
The field extraction system addresses the challenge of extracting fields from unlabeled forms by using self-supervised training with pseudo-labels and a transformer-based structure, achieving efficient and accurate field extraction without the need for extensive annotations.
Patent Information
- Application Number
- JP2023571264
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-09-24
- Filing Date
- 2022-01-27
- Publication Date
- 2025-06-30
- Estimated Expiration
- 2042-01-27
AI Technical Summary
Existing methods for extracting fields from forms with unlabeled data require extensive human effort and large numbers of field-level annotations, which are costly and labor-intensive due to the need for confidential information handling.
A field extraction system that uses self-supervised training with pseudo-labels generated from label-free forms using simple rules, combined with a transformer-based structure to predict field tags, and a pseudo-label ensemble module to refine labels and reduce noise.
The system achieves efficient field extraction without requiring field-level annotations, improving processing efficiency and reducing human labor, while maintaining a balance between precision and recall through progressive label refinement.
Smart Images

Figure 0007700273000035 
Figure 0007700273000036 
Figure 0007700273000037
Abstract
Description
Technical Field
[0001] This application is a non-provisional application of U.S. Provisional Application No. 63 / 189,579, filed on May 17, 2021, and claims priority to U.S. Non-Provisional Applications No. 17 / 484,618 and 17 / 484,623, which claim priority under 35 U.S.C. § 119.
[0002] All of the above applications are hereby incorporated by reference in their entirety.
[0003] This embodiment generally relates to machine learning systems and computer vision, and more specifically, to mechanisms for extracting fields from forms having unlabeled data.
Background Art
[0004] Documents such as forms, such as invoices, payroll statements, medical information providing forms, etc., are commonly used in daily business workflows. Extracting fields from various forms can often be a difficult task. For example, document layouts and text representations can vary even for the same form type (registered trademark), and when forms are issued by different vendors, for example, invoices from different companies may have significantly different designs, and payroll statements from different systems (e.g., ADP and Workday) may have different text representations for similar information, etc. Conventionally, a great deal of human effort has been required to extract information from such form documents. For example, human workers are usually given a list of expected form fields, such as purchase_order, invoice_number, total_amount, etc., and based on that, extract the corresponding values based on an understanding of the form.
[0005] Therefore, an efficient system for extracting information from form documents is needed.
Brief Description of the Drawings
[0006]
Figure 1
[0007]
Figure 2
[0008]
Figure 3
[0009]
Figure 4
[0010]
Figure 5
[0011]
Figure 6
[0012]
Figure 7
[0013]
Figure 8A
Figure 8B
[0014]
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
[0015] In the figures, elements having the same reference numerals have the same or similar functions.
DETAILED DESCRIPTION OF THE INVENTION
[0016] Machine learning systems are widely used in computer vision, such as pattern recognition, object identification, etc. Some recent machine learning methods formulate form field extraction as field-value pairing or field tagging. For example, some existing systems take fields and value candidates as input, and adopt an expression learning method that utilizes metric learning techniques to apply high pairing scores to positive field-value pairs and low scores to negative pairs. Another system uses a pre-trained transformer that takes both text and its location as input. However, these existing methods generally require a large number of field-level annotations for training. Obtaining form field-level annotations is very costly, labor-intensive, and sometimes even impossible. This is because (1) forms usually contain confidential information, so the publicly available data that can be used for training purposes is limited, and (2) hiring external annotators is also infeasible due to the risk of exposing personal information.
[0017] Considering the need for an efficient system for extracting information from form documents, embodiments describe a field extraction system that does not require field-level annotations for training. Specifically, the training process is bootstrapped by mining pseudo-labels from label-free forms using simple rules. Then, a transformer-based structure is used to model the interactions between text tokens within the input form and accordingly predict the field tags for each token. The pseudo-labels are used to supervise the transformer training. Since the pseudo-labels are noisy, a refined module including a series of branches is used to refine the pseudo-labels. Each of the refined branches tags the fields and generates refined labels. At each stage, the branches are optimized by the labels ensembled from all previous branches to reduce label noise.
[0018] For example, a field extraction system is trained based on self-supervised pseudo-labels from unlabeled data. Specifically, the field extraction system detects a set of words within a form and their locations, and identifies field values based on geometric rules between the words. For example, fields and field values are typically horizontally aligned and separated by a colon. Then, the identified field values may be used as pseudo-labels to train a transformer network that encodes the detected words and locations for classification.
[0019] In one of several embodiments, a number of pseudo-label ensemble (PLE) branches may be used to refine the pseudo-labels for training. Specifically, the PLE branches are operated in parallel to generate a classification predicted from the encoded representations of the detected words and locations. In each branch, a loss component is computed by comparing the refined labels in this branch with the predicted labels generated by the "previous" PLE as pseudo-labels. Then, the loss components of the entire PLE branch are summed to update the PLE together.
[0020] As used herein, the term "network" may include any hardware or software-based framework that includes any artificial intelligence network or system, neural network or system, and / or any training or learning model implemented in or with it.
[0021] As used herein, the term "module" may include a hardware or software-based framework that executes one or more functions. In some embodiments, the module may be implemented on one or more neural networks.
[0022] FIG. 1 is a schematic diagram 100 illustrating an example of information extraction from a claim according to an embodiment described in this specification. Conventionally, in form processing, a human operator is usually given a list of expected form fields such as purchase_order, invoice_number, total_amount, etc., and the purpose is to extract corresponding values based on the understanding of the form. For example, keys such as INVOICE#, PO Number, and Total refer to the specific text expressions of the fields in the form and are important indicators for locating the values. Keys are generally the most important features for identifying values. Therefore, the field extraction system aims to automatically extract field values from irrelevant information in the form, which is essential for improving processing efficiency and reducing human labor.
[0023] As shown in FIG. 100, the form contains various phrases such as "invoice#", "1234", "PO Number", "000001", etc. The field extraction system may identify that "PO Number" 102 is the location key, and then determine whether any of the values "1234" 104, "00000001" 103, or "100.00" 105 matches the location key. Such a match may be determined based on the geometric relationship between the location key 102 and the values 103 - 105. For example, a rule-based algorithm may be applied to determine a match, such as that the value "0000001" 103 is likely to correspond to the location of the location key 102. This is because the value 103 has a position aligned vertically with the location of the location key 102.
[0024] Different from the conventional method of accessing large-scale labeled forms, a rule-based method may be used to generate noisy pseudo-labels (e.g., fields and values) from label-free data. The rule-based algorithm is constructed based on the following observations. (1) The field value (e.g., 103 in FIG. 1) is usually shown within a form together with some key (e.g., 102 in FIG. 1), and the key (e.g., 102 in FIG. 1) is the specific textual representation of the field. (2) The key and its corresponding value have a strong geometric relationship (as shown in FIG. 1, most keys are adjacent to their values either vertically or horizontally). (3) The layout of the form is very diverse, but usually there are some key texts that are frequently used in various form instances (e.g., the key texts for the field purchase_order are "PO Number", "PO#", etc.). (4) The field value is always associated with some date type (e.g., the data type of the value of "invoice_date" is date, and the data type of the value of "total_amount" is money amount or number).
[0025] Therefore, a rule - based method can be used to generate useful pseudo - labels for each field of interest from a large - scale form. As shown in FIG. 1, the key location identification 102 is first performed based on string matching between the text in the form and the possible key strings of the field. Then, based on the geometric relationship between the data type of the text and the location - identified key 102, the values 103 - 105 are estimated.
[0026] FIG. 2 is a schematic diagram illustrating an overall self - supervised training framework 200 of a field extraction system according to the embodiments described herein. The framework 200 includes an optical character recognition module 205, a transformer network 210, and a classifier 220. The form without labels 202, e.g., a check, an invoice, a pay stub, etc., may contain information about the fields within a predefined list {fd1, fd2, …, fd N}. When a form is given as input, the general OCR detection and recognition module 205 is applied to the form without labels 202 to obtain bounding boxes {b1, b2, …, bM Obtain a set of words {w1, w2, …, w M} having positions represented as such. Thus, the goal of the field extraction method is to match the target value v i to the field fd i when the information of the field exists in the input form, and automatically extract it from a large number of word candidates {w1, w2, …, w M}.
[0027] Next, the pair of word and bounding box position {w i , b i} may be input to the transformer encoder 210 for encoding into a feature representation. The pair {w i , b i} may also be sent to the pseudo-label inference module 215 configured to perform key position identification for identifying the position of the key corresponding to each predefined field and value estimation for determining the corresponding field value of the identified key.
[0028] For example, the key and value may include multiple words. When receiving the pair of word and bounding box position {w i , b i}, the pseudo-label inference module 215 uses the DBSCAN algorithm (Ester et al., 1996) to group nearby recognized words based on their positions and obtain phrase candidates
Number
Number
[0029] For each field fd i of interest, which is a list of frequently used keys
Number
Number
Number
Number
Number
[0030] Then, the key is located by finding the candidate with the highest key score, as follows.
Number
[0031] Next, the pseudo-label inference module 215 may determine the value (or, if applicable, one or more values) of the located key. Specifically, the values are estimated according to the following two criteria. First, the data type needs to be in line with the field. Second, their positions need to be in good harmony with the located key. For each field, a list of eligible data types may be determined in advance. For example, for the data field "invoice number", the data type may include a string or an integer. A pre-trained BERT-based model may be used to predict the data type of each phrase candidate, and candidates with the correct data type
Number
[0032] In one embodiment, the value score for each eligible candidate
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Equation
[0033] i , b i}, may extract features.
Equation
[0034] The classifier 220 for token classification may receive an input of the encoded feature representation from the transformer encoder 210, which generates a prediction field including the background for each token from the original label-free form 202. Specifically, the classifier 220 projects the features into the field space ({background, fd1, fd2, …, fd N}) to generate a field prediction score s k . The predicted field score from the classifier 220 and the pseudo label generated from the pseudo label inference 215 are then compared by the loss module 230 to generate a training objective. The training objective may be further utilized to update the transformer 210 and the classifier 220 via the backpropagation path (shown by the dashed line).
[0035] In one embodiment, multiple progressive pseudo label ensembles (PLE) may be employed for bootstrap training as further described in FIG. 3.
[0036] FIG. 3 is a block diagram illustrating an exemplary framework for refining the field extraction framework described in FIG. 2 with PLE according to the embodiments described herein. As described in FIG. 2, the transformer 210 receives an input 302 of words extracted from the label-free form 202 and the positions of the bounding boxes surrounding the words (w1, b1), (w2, b2), …, (w M , b M ), and based thereon, an initial word-level field label (also referred to as a bootstrap label)
Number
Number
[0037] However, if only noisy bootstrap labels are used as the training ground truth, the performance of the model may be compromised. After the Transformer 210, a sophisticated module 304 including multiple PLEs that each function as a classification branch is adopted. Specifically, in each branch j, the PLE independently performs field classification and, based on their predictions, the pseudo-label
Number
[0038] For example, in branch k, the refined label is generated according to the steps of (1) finding the predicted field label c for each word by argmax(sk
Number
Number
Number
Number
Number
Number
Number
Number
Number
[0039] Therefore, the final loss aggregates all the losses and
Number
[0040] In this way, the progressive refinement of the labels reduces label noise. However, if only the labels refined at each stage are used, the labels become more accurate after refinement, but some low-confidence values are filtered out and removed, resulting in a decrease in recall rate, so there is a limit to performance improvement. To mitigate this problem, each branch is enhanced with the ensemble of labels from all previous stages. The ensemble of labels not only maintains a better balance between precision and recall, but is also more diverse and can serve as regularization for model optimization. During inference, the average score predicted from all branches may be used. A similar procedure may be applied to obtain the final field value when generating refined labels.
[0041] Computer Environment FIG. 4 is a simplified diagram of a computing device 400 implementing a field extraction framework according to some embodiments described herein. As shown in FIG. 4, the computing device 400 includes a processor 410 coupled to a memory 420. The operation of the computing device 400 is controlled by the processor 410. Also, although the computing device 400 is shown as having only one processor 410, the processor 410 may represent one or more central processing units, multi-core processors, microprocessors, microcontrollers, digital signal processors, field programmable gate arrays (FPGAs), application specific integrated circuits, graphics processing units (GPUs), etc. within the computing device 400. It is understood that the computing device 400 may be implemented as a stand-alone subsystem, as a board added to a computing device, and / or as a virtual machine.
[0042] Memory 420 may be used to store software executed by computing device 400 and / or one or more data structures used during operation of computing device 400. Memory 420 may include one or more types of machine-readable media. Some common forms of machine-readable media include, for example, floppy disks, flexible disks, hard disks, magnetic tapes, any other magnetic media, CD-ROMs, any other optical media, punch cards, paper tapes, any other physical media with patterns of holes, RAM, PROM, EPROM, FLASH-EPROM, any other memory chip or cartridge, and / or any other media adapted to be read by a processor or computer.
[0043] Processor 410 and / or memory 420 may be arranged in any suitable physical configuration. In some embodiments, processor 410 and / or memory 420 may be implemented on the same board, in the same package (e.g., system-in-package), on the same chip (e.g., system-on-chip), etc. In some embodiments, processor 410 and / or memory 420 may include distributed, virtualized, and / or containerized computing resources. Matching such embodiments, processor 410 and / or memory 420 may be located in one or more data centers and / or cloud computing facilities.
[0044] In some examples, the memory 420 may include a non-transitory tangible machine-readable medium that includes executable code that, when operated on by one or more processors (e.g., processor 410), may cause the one or more processors to execute methods described in further detail herein. For example, as illustrated, the memory 420 may include instructions for extraction 430 that may be used to implement and / or emulate a system and model and / or implement any of the methods further described herein. In some examples, the field extraction module 430 may receive an input 440, such as an image instance of a form without labels, via the data interface 415. The data interface 415 may be either a user interface that receives an image instance of a form uploaded by a user or a communication interface that receives or retrieves an image instance of a form previously stored from a database. The field extraction module 430 may generate an output 450, such as an extracted field of the input 440.
[0045] In some embodiments, the field extraction module 430 may further include a pseudo-label estimation module 431 and a PLE module 432. The pseudo-label inference module 431 uses a rule-based method to mine noisy pseudo-labels from forms, as described, for example, in FIG. 2. The PLE module 432 (similar to the refinement module 304 of FIG. 3) may use an input of a set of tokens extracted from a form and an output of a predicted field that includes the background of each token to learn a data-driven model by using the estimated value of the field as a pseudo-label for the token classification task during training. Further details of the PLE module 432 are described below in connection with FIG. 3.
[0046] Field Extraction Workflow FIG. 5 is a schematic diagram of a method 500 for field extraction from a form having unlabeled data via a field extraction model, according to some embodiments. One or more of the processes of method 500 may be implemented in the form of executable code stored on a non-transitory tangible machine-readable medium that can cause one or more processors to execute one or more of the processes when executed by the one or more processors. In some embodiments, method 500 corresponds to the operation of a field extraction module 430 (FIG. 4) for performing a method of field extraction or training a field extraction model. As illustrated, method 500 includes a number of enumerated steps, although aspects of method 500 may include additional steps before, after, and in between the enumerated steps. In some aspects, one or more of the enumerated steps may be omitted or performed in a different order.
[0047] In step 502, an unlabeled form including a plurality of fields and a plurality of field values is received via a data interface (e.g., 415 of FIG. 4). For example, the unlabeled form can take a form similar to that shown in FIGS. 8A-8B.
[0048] In step 504, a set of words and a set of positions are detected within the unlabeled form with respect to the set of words. For example, the words and positions may be detected by the OCR module 205 of FIG. 2.
[0049] In step 506, the field value of a field is identified based at least in part on the geometric relationship between sets of words from a set of words and a set of positions. For example, the field value can be identified by applying a first rule that one or more words within the form of a key are related to the field name of the field. As another example, the field value may be identified by applying a second rule that a pair of horizontally or vertically aligned words are the key of the field and the field value. In another example, the field value may be identified by applying a third rule that a word from a set of words that matches pre-defined key text is the key of the field.
[0050] In one implementation, key identification corresponding to a field is determined. For example, a set of phrase candidates is determined from a set of words, and a corresponding set of phrase positions is determined from the set of positions by grouping nearby recognized words. For each phrase candidate, a key score indicating the likelihood that each phrase candidate is the key of the field is computed. The key score is computed based on the string distance between each phrase candidate and a pre-defined key (see, for example, Equation (1)). Then, the key is determined for the field based on the maximum key score among the set of phrase candidates (see, for example, Equation (2)).
[0051] Specifically, to compute the key score, a neural model may be used to predict each data type for each phrase candidate. Next, a subset of phrase candidates having data types that match the predefined data types of the fields is determined. For each phrase candidate within the subset, a value score indicating the likelihood that each phrase candidate is a field value of the field is computed. The value score is computed based on the key score of the position-specific key corresponding to the field and a geometric relationship criterion between each phrase candidate and the position-specific key (e.g., Equation (3)). The geometric relationship criterion is computed based on, for example, a string distance and an angle between each phrase candidate and the position-specific key (e.g., Equation (4)). Next, the field value is determined based on the maximum value score among the subset of phrase candidates.
[0052] In step 508, the encoder (e.g., the transformer encoder 210 of FIG. 2) may encode a pair of the first word and the first position corresponding to the field value into a first representation.
[0053] In step 510, the classifier (e.g., the classifier 220 of FIG. 2) may generate a field classification distribution from the first representation by the classifier.
[0054] In step 512, a first loss target is computed by comparing the field classification distribution with the field value as a pseudo label.
[0055] In step 514, the encoder is updated based on the first loss target via backpropagation.
[0056] FIG. 6 is a schematic diagram of a method 600 for label refinement in field extraction from a form having unlabeled data via a field extraction model, according to some embodiments. One or more of the processes of method 600 may be implemented in the form of executable code stored in a non-transitory tangible machine-readable medium that can cause one or more processors to execute one or more of the processes when executed, at least in part, by the one or more processors. In some embodiments, method 600 corresponds to the operation of a field extraction module 430 (FIG. 4) for performing a method of training a field extraction or field extraction model. As illustrated, method 600 includes a number of enumerated steps, although aspects of method 600 may include additional steps before, after, and in between the enumerated steps. In some aspects, one or more of the enumerated steps may be omitted or performed in a different order.
[0057] In step 602, an unlabeled form including a plurality of fields and a plurality of field values is received via a data interface (e.g., 415 in FIG. 4). For example, the unlabeled form can take a form similar to that shown in FIGS. 8A-8B.
[0058] In step 604, a first word and a first position of the first word are detected within the unlabeled form. For example, the word and position may be detected by the OCR module 205 of FIG. 2.
[0059] In step 606, an encoder (e.g., the transformer encoder 210 of FIG. 2) encodes the pair of the first word and the first position into a first representation (e.g., Equation (6)).
[0060] In step 608, a plurality of Progressive Label Ensemble (PLE) branches (see, e.g., 304a - n in FIG. 3) each generate a plurality of prediction labels in parallel based on a first representation. Each of the plurality of PLE branches includes a respective classifier that generates a respective prediction label based on the first representation. The prediction label in one PLE branch is generated by projecting the first representation into a set of field prediction scores via one or more fully connected layers and generating a prediction score based on the largest field prediction score among the set of words. When the largest field prediction score is greater than a predetermined threshold, for the field from the plurality of fields, the word corresponding to the largest field prediction score from the set of words is selected.
[0061] In step 610, a loss component in one PLE branch is computed by comparing the prediction label in one PLE branch with the prediction label from the previous PLE branch as a pseudo label.
[0062] In step 612, a loss target is computed as the sum of the loss components across the plurality of PLE branches (e.g., Equation (7)).
[0063] In step 614, the plurality of PLE branches are updated based on the first loss target via backpropagation. In one embodiment, the first PLE branch from the plurality of PLE branches uses the identified field value of the field as the first pseudo label. The combined loss target is computed by adding the first loss target computed in step 512 of FIG. 5 and the loss target. Then, the encoder and the plurality of PLE branches are jointly updated based on the combined loss target.
[0064] Exemplary Performance An exemplary training dataset may include actual claims collected from various vendors. For example, the training set includes 7,664 unlabeled claim forms of 2,711 templates. For example, the validation set includes 348 labeled claims of 222 templates. The test set includes 339 labeled claims of 222 templates. Each template has up to 5 images in each set. Seven frequently used fields are considered, including invoice_number, pur-chase_order, invoice_date, due_date, amount_due, total_amount, and total_tax.
[0065] For the Tobacco test set, 350 claims are collected from the Tobacco Collections of Industry Documents Library 2 for public release. The validation and test sets of the internal IN-Invoice dataset have similar statistical distributions of fields, but the publicly released Tobacco test set is different. For example, the claims in the Tobacco set (shown in Figure 8A) may have lower resolution and a more cluttered background compared to other claims in the training dataset (shown in Figure 8B).
[0066] The end-to-end macro-average F1 score on the fields is used as a criterion for evaluating the model. Specifically, exact string matching between the predicted value and the correct value is used to count true positives, false positives, and false negatives. For each field, precision, recall, and F1 score are obtained. The reported scores are averaged over 5 runs to reduce the influence of randomness.
[0067] Since there is no existing method to perform field extraction using only unlabeled data, the following baselines are constructed to verify this method. That is, Bootstrap Label (B-Label): Using the initial pseudo-labels inferred using the proposed simple rules, field extraction can be performed directly without training the data. Transformer training using B-Label: Since the Transformer is used as a backbone for extracting word features, the Transformer model is trained using B-Label as a baseline for evaluating (1) the data-driven model of the pipeline and (2) the performance gain from the refined module. Both the content of the text and its location are important for field prediction. An example of a Transformer backbone is LayoutLM, which takes both text and location as inputs. Additionally, two common Transformer models, BERT and RoBERTa, which take only text as inputs, are used.
[0068] Use the OCR engine to detect words and their positions, and rank the words in reading order. An exemplary key list and date type for each dataset are shown in Table 1 of FIG. 7. The key list and data types are quite broad. In Equation (4), α is set to 4.0. Further, to remove false positives, if the location-specific key is not within the adjacent zone, the value candidate is removed. Specifically, the adjacent zone around the value candidate extends to the left side of the image, with 4 candidate heights above it and 1 candidate height below it. In all experiments, the refined branch number k = 3. When the number of stages > 1, 1 hidden FC layer is added in units of 768 before classification. β in Equation (7) is set to 1.0 in all invoice experiments except that for the BERT-based refinement in Table 4 of FIG. 11, β = 5.0 due to its good performance in the validation set. For both the field extraction model and the baseline described herein, the model with the best F1 score is picked in the validation set. To prevent overfitting, a two-stage training strategy is adopted, using pseudo-labels to train the first branch of the model, and then the first branch is fixed together with the feature extractor during refinement. The batch size is set to 8, and the learning rate is 5e 5 and the Adam optimizer is used.
[0069] Next, since the proposed model contains large-scale unlabeled training data and a sufficient amount of validation / test data, it is verified using the IN-Invoice dataset, which fits better with the experimental setting. The proposed training method was first verified using LayoutLM as the backbone. The comparison results are shown in Table 2 of Figure 9 and Table 3 of Figure 10. The Bootstrap Labels (B-Labels) baseline achieved F1 scores of 43.8% and 44.1% on the validation set and test set, respectively, indicating that B-Labels have a reasonable accuracy but are still noisy. When training the LayoutLM transformer using B-labels, a significant performance improvement of about 15% on the validation set and about 17% on the test set is achieved. Adding the PLE refinement module significantly improves the model precision by about 6% on the validation set and about 7% on the test set, but slightly reduces the recall rate by about 2.5% on the validation set and about 3% on the test set. This indicates that the refined labels become increasingly reliable at later stages, leading to a higher model precision. However, during the refinement stage, less reliable false negatives are also removed, resulting in a decrease in the recall rate. Overall, the PLE refinement module further improves performance, resulting in a 3% gain in the F1 score.
[0070] Since both the text and its location are important for the task, LayoutLM is used as the default feature backbone. Additionally, to understand the impact of different transformer models as backbones, two additional models, BERT and RoBERTa, which use only text as input, were evaluated. The comparison results are shown in Table 4 of Figure 11 and Table 5 of Figure 12. When directly training BERT and RoBERTa using B-labels and the PLE improvement module, it has been observed that the baseline results of various transformer selections with different numbers of parameters (base or large) consistently improve. However, LayoutLM yields far higher results compared to the other two backbones, indicating that the location of the text is very important for achieving good performance on the task.
[0071] Next, the proposed model was tested in Table 6 of FIG. 13 using the introduced Tobacco test set. The simple rule-based method achieved an F1 score of 25.1%, which is reasonable but much lower compared to the results of the internal IN-Invoice dataset. The reason is that the Tobacco test set is visually noisy and results in more text recognition errors. The LayoutLM baseline is significantly improved when using the B-label. Also, the PLE refinement module further improves the F1 score by about 2%. As a result, it was suggested that the proposed method adapts well to various scenarios. FIGS. 8A-8B show that, despite the sample invoices being very diverse across different templates, having a cluttered background and low resolution, the proposed method achieves good performance.
[0072] An ablation study was further conducted on the invoice dataset with LayoutLM as the backbone. Influence of the number of stages: The proposed model was refined at k stages, and k = 3 was fixed in all experiments. The number of stages was changed for evaluation. FIG. 15 shows that as the number of stages k increases, the model generally performs better for both the validation set and the test set. The performance when using multiple stages is always higher than that of the single-stage model (Transformer baseline). The performance of the model reached its peak at k = 3. As shown in FIG. 16, during model refinement, the precision improves but the recall decreases. The best balance between precision and recall is achieved at k = 3. When k > 3, the recall decreases more than the precision improves, and a deterioration of the F1 score is observed.
[0073] Effect of refined labels (R-labels): To analyze the effect of this design, all refined labels were removed in the final loss, and three branches were independently trained using only B-labels and the predictions were ensembled during inference. As shown in Table 7 of Figure 14, removing the refined labels results in a 2.2% and 2.6% decrease in the F1 scores in the validation set and test set, respectively.
[0074] Effect of regularization using B-labels. At each stage, B-labels are used as a type of regularization to prevent the model from overfitting to the refined labels that are overfitted. Use of B-labels in the refinement stage by setting β = 0 in Equation (7). As shown in Table 7 of Figure 14, the model performance drops by approximately 2% in F1 score without this regularization.
[0075] Effect of two-stage training strategy: To avoid overfitting to noisy labels, a two-stage training strategy was adopted where the first branch was trained using B-labels and fixed during refinement. This effect was analyzed by training the model in a single step. As shown in Table 7 of Figure 14, single-step training leads to a 1.8% and 1.4% decrease in the F1 scores in the validation set and test set, respectively.
[0076] Some examples of computing devices, such as computing device 400, may include a non-transitory tangible machine-readable medium that includes executable code that, when operated on by one or more processors (e.g., processor 410), may cause the one or more processors to execute the processes of method 400. Some common forms of machine-readable media that may include the processes of method 400 are, for example, floppy disks, flexible disks, hard disks, magnetic tapes, any other magnetic media, CD-ROMs, any other optical media, punch cards, paper tapes, any other physical media with patterns of holes, RAM, PROM, EPROM, FLASH-EPROM, any other memory chip or cartridge, and / or any other media adapted to be read by a processor or computer.
[0077] This description and the accompanying drawings, which illustrate aspects, embodiments, implementations, or applications of the invention, should not be construed as limiting. Various mechanical, compositional, structural, electrical, and operational changes may be made without departing from the spirit and scope of this description and the claims. In some instances, well-known circuits, structures, or techniques have not been shown or described in detail in order not to obscure the embodiments of the present disclosure. Similar numerals in two or more figures represent the same or similar elements.
[0078] In this description, specific details of some embodiments consistent with the present disclosure are set forth. Numerous details are set forth in order to provide a thorough understanding of the embodiments. It will be apparent to those skilled in the art that some embodiments may be practiced without some or all of these specific details. The specific embodiments disclosed herein are illustrative but not limiting. Those skilled in the art may recognize other elements that are not specifically described herein but are within the scope and spirit of the present disclosure. Additionally, to avoid unnecessary repetition, one or more features shown and described in connection with one embodiment may be incorporated into other embodiments, unless specifically described otherwise in another way or unless one or more features render an embodiment non-functional.
[0079] This application is further described with respect to the attached document on page 9 of Appendix I entitled "Field Extraction from Forms with Unlabeled Data", which is considered part of this disclosure and is hereby incorporated by reference in its entirety.
[0080] Exemplary embodiments have been shown and described, but extensive modifications, changes, and substitutions are contemplated in the foregoing disclosure, and in some instances, some features of the embodiments may be employed without corresponding use of other features. Those skilled in the art will recognize many variations, alternatives, and modifications. Accordingly, the scope of the present invention should be limited only by the following claims, and it is appropriate that the claims be construed broadly in a manner matching the scope of the embodiments disclosed herein.
[0081] This application also discloses the following content. [Appendix 1] A method for field extraction from a form having unlabeled data via a field extraction model, comprising: Receiving, via a data interface, an unlabeled form including a plurality of fields and a plurality of field values; Detecting, by a processor, a set of words and a set of positions within the unlabeled form for the set of words; Identifying a field value of a field based at least in part on a geometric relationship between the sets of words from the set of words and the set of positions; Encoding, by an encoder, a first pair of a word and a first position corresponding to the field value into a first representation; Generating, by a classifier, a field classification distribution from the first representation; Computing a first loss target by comparing the field classification distribution with a field value as a pseudo-label; A method comprising updating the encoder based on the first loss target via backpropagation. [Appendix 2] The method according to Appendix 1, wherein identifying the field value of the field includes applying a first rule that one or more words in the form of a key are related to the field name of the field. [Appendix 3] The method according to Appendix 2, wherein identifying the field value of the field includes applying a second rule that a pair of horizontally or vertically aligned words are keys for the field and the field value. [Appendix 4] The method according to Appendix 3, wherein identifying the field value of the field includes applying a third rule that a word from a set of words that match a predefined key text is a key for the field. [Appendix 5] Determining a set of phrase candidates from the set of words and determining a corresponding set of phrase positions from the set of positions by grouping nearby recognized words; Computing, for each phrase candidate, a key score indicating the likelihood that the respective phrase candidate is a key for the field; The method according to Appendix 1, further comprising determining the key for the field based on the maximum key score in the set of phrase candidates. [Appendix 6] The method according to Appendix 5, wherein the key score is computed based on a string distance between the respective phrase candidate and a predefined key. [Appendix 7] Predicting, for each phrase candidate, a respective data type via a neural model; Determining a subset of phrase candidates having a data type that matches a predefined data type of the field; For each phrase candidate within the subset, computing a value score indicating the likelihood that each phrase candidate is the field value of the field; determining the field value based on a maximum key score among the subset of phrase candidates, the method according to appendix 5, further comprising. [Appendix 8] The value score is computed based on a key score of a location-specific key corresponding to the field and a geometric relationship criterion between each phrase candidate and the location-specific key, the method according to appendix 7. [Appendix 9] The geometric relationship criterion is computed based on a string distance and an angle between each phrase candidate and the location-specific key, the method according to appendix 8. [Appendix 10] generating, by a plurality of progressive label ensemble (PLE) branches, a plurality of prediction labels respectively based on the first representation; computing a loss component in one PLE branch by comparing a prediction label in the one PLE branch with a prediction label from a previous PLE branch as a pseudo label, further comprising. The first PLE branch from the plurality of PLE branches receives the identified field value of the field as a first pseudo label, the method according to appendix 1. [Appendix 11] A system for field extraction from a form having unlabeled data via a field extraction model, comprising: a data interface that receives an unlabeled form including a plurality of fields and a plurality of field values; a memory that stores a plurality of processor-executable instructions; a processor that executes the processor-executable instructions to perform operations, the operations including: detecting a set of words and a set of positions for the set of words within the unlabeled form; Identifying a field value of a field from the set of words and the set of positions, based at least in part on a geometric relationship between the sets of words; Encoding, by an encoder, a pair of a first word and a first position corresponding to the field value into a first representation; Generating, by a classifier, a field classification distribution from the first representation; Computing a first loss target by comparing the field classification distribution with a field value as a pseudo label; Updating the encoder based on the first loss target via backpropagation, a system comprising. [Appendix 12] The system according to Appendix 11, wherein identifying the field value of the field includes applying a first rule that one or more words within a key form are related to a field name of the field. [Appendix 13] The system according to Appendix 12, wherein identifying the field value of the field includes applying a second rule that a pair of horizontally or vertically aligned words are keys of the field and the field value. [Appendix 14] The system according to Appendix 13, wherein identifying the field value of the field includes applying a third rule that a word from the set of words matching a predefined key text is a key of the field. [Appendix 15] The operation is Determining a set of phrase candidates from the set of words and determining a corresponding set of phrase positions from the set of positions by grouping recognized neighboring words; Computing, for each phrase candidate, a key score indicating the likelihood that the respective phrase candidate is a key of the field; Determining the key of the field based on the maximum key score in the set of the phrase candidates, further comprising the system according to Supplementary Note 11. [Supplementary Note 16] The system according to Supplementary Note 15, wherein the key score is computed based on a string distance between each of the phrase candidates and a predefined key. [Supplementary Note 17] The operation Predicting, via a neural model, a respective data type for each phrase candidate, Determining a subset of phrase candidates having a data type matching the predefined data type of the field, For each phrase candidate in the subset, computing a value score indicating the likelihood that the respective phrase candidate is the field value of the field, Determining the field value based on the maximum key score in the subset of the phrase candidates, further comprising the system according to Supplementary Note 15. [Supplementary Note 18] The system according to Supplementary Note 17, wherein the value score is computed based on a key score of a position-specific key corresponding to the field and a geometric relationship criterion between each of the phrase candidates and the position-specific key. [Supplementary Note 19] The system according to Supplementary Note 18, wherein the geometric relationship criterion is computed based on a string distance and an angle between each of the phrase candidates and the position-specific key. [Supplementary Note 20] The operation Generating, in parallel, a plurality of prediction labels based on the first representation by a plurality of Progressive Label Ensemble (PLE) branches, Computing a loss component in one of the PLE branches by comparing a prediction label in one of the PLE branches with a prediction label from a previous PLE branch as a pseudo label, further comprising The method according to appendix 1, wherein a first PLE branch from the plurality of PLE branches receives the identified field value of the field as a first pseudo label. [Appendix 21] A method for field extraction from a form having label-free data via a field extraction model, comprising: Receiving, via a data interface, a label-free form including a plurality of fields and a plurality of field values; Detecting, by a processor, a first word and a first position of the first word in the label-free form; Encoding, by an encoder, a pair of the first word and the first position into a first representation; Generating, by a plurality of Progressive Label Ensemble (PLE) branches, a plurality of prediction labels in parallel based on the first representation, respectively; Further comprising computing a loss component in the one PLE branch by comparing a prediction label in the one PLE branch with a prediction label from a previous PLE branch as a pseudo label; Computing a loss target as a sum of loss components across the plurality of PLE branches; Updating the plurality of PLE branches via backpropagation based on the loss target. [Appendix 22] The method according to appendix 21, wherein each of the plurality of PLE branches includes a respective classifier that generates a respective prediction label based on the first representation. [Appendix 23] The prediction label in the one PLE branch is Projecting the first representation into a set of field prediction scores via one or more fully connected layers; Generating the prediction score based on the maximum field prediction score among a set of words, and is generated by the method according to appendix 21. [Appendix 24] For one of the plurality of fields, when the maximum field prediction score is greater than a predefined threshold, further comprising selecting, from the set of words, the word corresponding to the maximum field prediction score, the method according to appended note 23. [Appended note 25] By a processor, detecting, for the set of words, the set of words and the set of positions in a label-free form; Identifying a field value of a field, at least partially based on a geometric relationship between the sets of words, from the set of words and the set of positions; Generating, by a classifier, a field classification distribution from the first representation; Further comprising computing a first loss target by comparing the field classification distribution with the field value as a pseudo-label, the method according to appended note 21. [Appended note 26] The first PLE branch from the plurality of PLE branches uses the identified field value of the field as a first pseudo-label, the method according to appended note 25. [Appended note 27] Computing a combined loss target by adding the loss target and the first loss target; Further comprising updating the encoder and the plurality of PLE branches jointly via backpropagation based on the combined loss target, the method according to appended note 25. [Appended note 28] A method comprising updating the encoder based on the first loss target via backpropagation. [Appended note 29] After updating the encoder, further comprising updating the plurality of PLE branches via backpropagation based on the loss target while fixing the parameters of the encoder, the method according to appended note 28. [Appended note 30] A system for field extraction from a form having unlabeled data via a field extraction model, a data interface that receives an unlabeled form including a plurality of fields and a plurality of field values, a memory that stores a plurality of processor-executable instructions, a processor that executes the processor-executable instructions to perform operations, the operations including: detecting a first word and a first position of the first word within the unlabeled form; encoding, by an encoder, a pair of the first word and the first position into a first representation; generating, by a plurality of Progressive Label Ensemble (PLE) branches, a plurality of prediction labels in parallel based on the first representation; computing a loss component in one PLE branch by comparing a prediction label in the one PLE branch with a prediction label from a previous PLE branch as a pseudo label; computing a loss target as a sum of loss components across the plurality of PLE branches; updating the plurality of PLE branches via backpropagation based on the loss target. [Appendix 31] The system according to Appendix 30, wherein each of the plurality of PLE branches includes a respective classifier that generates a respective prediction label based on the first representation. [Appendix 32] The prediction label in the one PLE branch is projected, via one or more fully connected layers, the first representation into a set of field prediction scores; and generated based on the field prediction score having a maximum value among a set of words, the system according to Appendix 30. [Appendix 33] The operations further include The system according to Appendix 32, further comprising: for one field among the plurality of fields, when the maximum field prediction score is greater than a predefined threshold, selecting a word corresponding to the maximum field prediction score from the set of words. [Appendix 34] The operation by a processor, detecting, for the set of words, the set of words and the set of positions in a label-free form; identifying a field value of a field from the set of words and the set of positions, at least partially based on a geometric relationship between the sets of words; generating, by a classifier, a field classification distribution from the first representation; computing a first loss target by comparing the field classification distribution with the field value as a pseudo-label, the system according to Appendix 30, further comprising. [Appendix 35] The first PLE branch from the plurality of PLE branches uses the identified field value of the field as a first pseudo-label, the system according to Appendix 34. [Appendix 36] The operation computing a combined loss target by adding the loss target and the first loss target; updating the encoder and the plurality of PLE branches jointly via backpropagation based on the combined loss target, the system according to Appendix 34, further comprising. [Appendix 37] The operation updating the encoder based on the first loss target via backpropagation, a system comprising. [Appendix 38] The operation after updating the encoder, further comprising updating the plurality of PLE branches via backpropagation based on the loss target while fixing the parameters of the encoder, the system according to Appendix 37. [Appendix 39] A non-transitory storage processor-readable medium storing processor-executable instructions for field extraction from a form having unlabeled data via a field extraction model, the instructions being executed by a processor performing operations, the operations comprising: Receiving, via a data interface, an unlabeled form including a plurality of fields and a plurality of field values; Detecting, by a processor, a first word and a first position of the first word within the unlabeled form; Encoding, by an encoder, a pair of the first word and the first position into a first representation; Generating, in parallel, a plurality of predicted labels respectively based on the first representation by a plurality of Progressive Label Ensemble (PLE) branches; Computing a loss component in one PLE branch by comparing a predicted label in one PLE branch with a predicted label from a previous PLE branch as a pseudo label; Computing a loss target as a sum of loss components across the plurality of PLE branches; Updating the plurality of PLE branches via backpropagation based on the loss target. A non-transitory storage processor-readable medium comprising: [Appendix 40] Each of the plurality of PLE branches includes a respective classifier that generates a respective predicted label based on the first representation; The predicted label in one PLE branch is Projecting the first representation into a set of field prediction scores via one or more fully connected layers; Generating the predicted score based on the field prediction score having the maximum value among a set of words. The non-transitory storage processor-readable medium according to Appendix 39.
Claims
1. A system for field extraction from a form having unlabeled data via a field extraction model, comprising: a data interface that receives an unlabeled form including a plurality of fields and a plurality of field values; a memory storing a plurality of processor-executable instructions; a processor that executes the processor-executable instructions to perform operations, the operations including: detecting a set of words and a set of positions of the set of words within the unlabeled form for the set of words; identifying a field value of a field based at least in part on a geometric relationship between the sets of words from the set of words and the set of positions; encoding, by an encoder, a first word and a first position pair corresponding to the field value into a first representation; generating, by a classifier, a field classification distribution from the first representation; computing a first loss target by comparing the field classification distribution with the field value as a pseudo label; and updating the encoder based on the first loss target via backpropagation.
2. The system of claim 1, wherein identifying the field value of the field includes applying a first rule that one or more words within a key form are related to the field name of the field.
3. The system of claim 2, wherein identifying the field value of the field includes applying a second rule that a pair of horizontally or vertically aligned words are keys of the field and the field value.
4. The system of claim 3, wherein identifying the field value of the field includes applying a third rule that a word from the set of words that matches a predefined key text is a key of the field.
5. The operations further include: determining a set of phrase candidates from the set of words and determining a corresponding set of phrase positions from the set of positions by grouping recognized neighboring words; and computing, for each phrase candidate, a key score indicating the likelihood that the respective phrase candidate is a key of the field. The system according to claim 1, further comprising determining the key of the field based on the maximum key score among the set of phrase candidates.
6. The system according to claim 5, wherein the key score is computed based on a string distance between each of the phrase candidates and a predefined key.
7. The operation includes predicting, via a neural model, a respective data type for each phrase candidate; determining a subset of phrase candidates having a data type that matches the predefined data type of the field; computing, for each phrase candidate in the subset, a value score indicating the likelihood that the respective phrase candidate is the field value of the field; The system according to claim 5, further comprising determining the field value based on the maximum key score among the subset of phrase candidates.
8. The system according to claim 7, wherein the value score is computed based on a key score of a location-specific key corresponding to the field and a geometric relationship criterion between each of the phrase candidates and the location-specific key.
9. The system according to claim 8, wherein the geometric relationship criterion is computed based on a string distance and an angle between each of the phrase candidates and the location-specific key.
Citation Information
Patent Citations
Form recognition device and its program
JP2008204226A