Image analysis based document processing for inferring key-value pairs in non-fixed digital documents

The online system addresses the challenge of extracting key-value pairs from non-rigid digital documents by employing image analysis and machine learning to determine key and neighbor scores, enhancing the automation and accuracy of information extraction.

JP7767437B2Active Publication Date: 2025-11-11SALESFORCE INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2023540842
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-01-04
Filing Date
2021-10-25
Publication Date
2025-11-11
Estimated Expiration
2041-10-25

AI Technical Summary

Technical Problem

Existing technologies struggle to automatically extract information from non-rigid digital documents with varying layouts and arrangements of key-value pairs due to the lack of standardized formats, making it difficult to apply machine learning techniques effectively.

Method used

An online system that analyzes digital documents using image analysis and machine learning to identify key-value pairs by determining key scores, neighbor scores, and spatial relationships to accurately extract information from non-fixed form documents.

Benefits of technology

Enables efficient and accurate extraction of information from non-rigid forms without human intervention, improving processing efficiency by identifying key-value pairs through automated analysis of document images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007767437000003
    Figure 0007767437000003
  • Figure 0007767437000004
    Figure 0007767437000004
  • Figure 0007767437000005
    Figure 0007767437000005
Patent Text Reader

Abstract

The online system extracts information from unformed documents. The online system receives an image of the form document and obtains a set of phrases and the locations of the phrases on the form image. For at least one field, the online system determines a key score for the set of phrases. The online system identifies a set of candidate values ​​for the field from the identified set of phrases and identifies a set of neighbors for each candidate value from the identified set of phrases. The online system determines neighbor scores, where the neighbor scores for the candidate values ​​and their respective neighbors are determined based on the neighbors' key scores and the spatial relationship of the neighbors to the candidate values. The online system selects the candidate values ​​and their respective neighbors as values ​​and keys for the field based on the neighbor scores.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Utility Patent Application No. 17 / 140,987, filed January 4, 2021, the entire contents of which are incorporated herein by reference. [Background technology]

[0002] The present invention relates generally to processing digital documents, and more particularly to inferring key-value pairs in non-fixed digital documents using image analysis of the digital documents.

[0003] Entities such as organizations of different types process many digital documents that may contain information relevant to the entity's business processes. Information may be extracted from the documents to perform or support one or more tasks in the business processes. Forms, among other types of documents, may structure information into a set of fields, each having one or more key-value pairs. A field may characterize a different type of information to be extracted from the document. A field's key may refer to the label by which the respective field is called on the form document and may vary depending, for example, on the naming convention used by the entity responsible for the form.

[0004] Because entities frequently process a significant number of documents, it is advantageous to automatically extract information from key-value pairs on form documents without a human operator. Analyzing digital documents typically involves receiving an image representation of the document, performing image analysis, and using machine learning techniques, such as optical character recognition and deep learning-based neural networks, to generate digital documents. These techniques train a machine learning-based model, such as a convolutional neural network, and apply the trained neural network to images representing new digital documents. These techniques typically operate on a set of prefixed fields. However, while some types of form documents are standardized and have fixed locations for key-value pairs, many types of form documents are non-rigid, in that the type and format of information varies, for example, depending on the entity issuing the form. This variation for non-rigid forms makes it difficult to automatically extract information. [Brief explanation of the drawings]

[0005] [Figure 1] FIG. 1 is a block diagram of a system environment including an online system, according to one embodiment.

[0006] [Figure 2] FIG. 1 illustrates an example document and a set of phrases identified in the invoice document, according to one embodiment.

[0007] [Figure 3] 3 is an exemplary high-level process for determining key-value pairs in the exemplary invoice document of FIG. 2, according to one embodiment.

[0008] [Figure 4] FIG. 1 is a block diagram of an architecture of an online system, according to one embodiment.

[0009] [Figure 5]1 is a flowchart of a method for determining key-value pairs in a form document, according to one embodiment.

[0010] [Figure 6] 2 is a block diagram illustrating the architecture of a typical computer system for use in the environment of FIG. 1, according to one embodiment.

[0011] The drawings depict various embodiments of the present invention for purposes of illustration only. Those skilled in the art will readily recognize from the following description that alternative embodiments of the structures and methods illustrated herein may be employed without departing from the principles of the present invention as described herein.

[0012] The drawings use like reference numbers to identify like elements. Letters following a reference number, such as "110A," indicate that the text is specifically referring to the element with that particular reference number. A reference number in the text without a following letter, such as "110," refers to any or all of the elements in the figures with that reference number (e.g., "client device 110" in the text refers to reference numbers "client device 110A" and / or "client device 110B" in the figures). DETAILED DESCRIPTION OF THE INVENTION

[0013] overview An online system extracts information from digital documents. In one embodiment, a method employed by the online system enables information to be extracted from non-fixed form documents, which may have different layouts and arrangements of information. Specifically, the online system receives an image of the form document from a client device. The form document may include key-value pairs for a set of fields. In one embodiment, the online system also obtains a template indicating one or more fields to be extracted from the form image, where the fields may be associated with a set of candidate keys for that field. The online system obtains a set of phrases and their positions in the form image.

[0014] For at least one field, the online system determines a key score for a set of phrases, where the key score for a phrase indicates the likelihood that the phrase is a key for the field on the form image. The online system identifies a set of candidate values ​​for the field from the identified set of phrases and identifies a set of neighbors for each candidate value from the identified set of phrases. The online system determines neighbor scores, where the neighbor score for a candidate value and each neighbor is determined based on the neighbor's key score and the neighbor's spatial relationship to the candidate value. The online system selects a candidate value and each neighbor based on the neighbor scores, sets the selected candidate value as the value of the field, and sets the selected neighbor as the key for the field. The disclosed technology can automatically process non-rigid forms, i.e., forms with an unfixed structure.

[0015] System environment Figure 1 is a block diagram of a system environment 100 including an online system 130, according to one embodiment. The system environment 100 shown in Figure 1 includes the online system 130, client devices 110A, 110B, and a network 120. In alternative configurations, different and / or additional components may be included in the system environment 100.

[0016] Online system 130 is a system that receives requests to process digital documents and provides information extracted from the digital documents to the requesting user. Entities such as businesses and government organizations process many documents, which may contain information related to the entity's business processes, such as financial transactions, onboarding new employees to the company, or launching new products. Information may be extracted from the documents to perform or support one or more tasks in the business processes. For example, the document may be an invoice for services provided to the business organization, and the invoice may be processed by an employee of the organization so that payment can be made to a vendor. As another example, the document may be a mortgage application for a lending organization, and the mortgage application may be processed by an underwriter to determine whether the application can be approved.

[0017] Forms, among other types of documents, may structure information into a set of fields, each with one or more key-value pairs. A field may characterize a different type of information to be extracted from the document. For example, an invoice number field may characterize a unique identifier for each invoice. The key for a field may refer to the label by which the respective field is called on the form document and may vary depending, for example, on the naming convention used by the entity responsible for the form. For example, the key for one vendor's invoice number may be labeled "Invoice #," while the key for another vendor's invoice number may be labeled "Invoice No." The value of a field may refer to the data value of the field on the form document and may follow the format used by the entity responsible for the form. For example, the value of the invoice number may be "INV-023-US."

[0018] Typically, a human operator extracts information from a document and performs one or more tasks, such as entering it into a record or making a payment based on the extracted information. Because entities frequently process a significant number of documents, it is advantageous to automatically extract information from key-value pairs on form documents without a human operator, so that processing can be more efficient. Some types of form documents are standardized and have fixed locations for key-value pairs on the document, which may allow a computerized system to automatically extract information because the locations of the keys and values ​​are already known. For example, a driver's license for a certain state may contain the same set of fields (e.g., name, address, eye color) and a uniform layout of key-value pairs for these fields across individuals residing in that state.

[0019] However, many types of form documents are non-rigid in that the type and format of information varies depending, for example, on the entity issuing the form. For example, a mortgage application from one lender may contain different types of fields than a mortgage application from another lender because each lender considers different types of information from an applicant. As another example, an invoice from a plumbing vendor may arrange the key-value pairs for the invoice number and amount due fields differently from an invoice from another vendor, even though both must be processed by the same company. This variation for non-rigid forms makes it difficult to automatically extract information because the type and format of information on these forms is not standardized or fixed.

[0020] Thus, in one embodiment, online system 130 provides a method for extracting information from non-fixed digital documents. Online system 130 receives requests from client devices 110 to process digital documents and provides the information extracted from the digital documents to the requesting user. In one embodiment, online system 130 may be an internal system managed and owned by the same entity as requesting client device 110. In another embodiment, online system 130 may be a separate system managed by a different entity than requesting client device 110. For example, online system 130 may receive requests from client devices 110 from employees of a different company who need to process documents.

[0021] FIG. 2 illustrates an exemplary invoice document and a set of phrases identified within the invoice document, according to one embodiment. FIG. 3 illustrates an exemplary high-level process for determining key-value pairs in the exemplary invoice document of FIG. 2, according to one embodiment. Specifically, the online system 130 receives a request from the client device 110, the request including an image of the document. In one example, the requested document is a form document that includes information in the form of key-value pairs for a set of fields. The image of the document may be received in the form of a computerized image file, such as JPEG, GIF, PNG, EPIS, AI, PDF, RAW, TIFF, etc., although it is understood that the document may be received in other formats.

[0022] 2, this example shows an image 200 of a form for an invoice issued by a company named "ABC Services, Inc." Notably, image 200 may be a non-fixed form because the type and format of information for an invoice may vary, for example, depending on the company issuing the invoice. Among other things, image 200 includes key-value pairs for a set of fields to be extracted by online system 130. Specifically, image 200 includes the following key-value pairs that are as yet unknown and to be extracted by online system 130: {"Bill To:", "John Smith"} for the invoice recipient field; {"Invoice #", "US-001"} for the invoice number field; {"Invoice Date", "11 / 2 / 2019"} for the invoice date field; {"Qty", "2"} for the item quantity field; {"Description", "Front end brakes"} for the item description field; {"Unit Price", "100.00"} for the item unit price field; {"Amount", "200.00"} for the item price field; and {"Total Charges", "$200.00"} for the total invoice amount field.

[0023] The online system 130 obtains a set of phrases and their locations in the document. A phrase may include one or more words that are spatially proximate to one another on the document. In one example, the set and locations of phrases may be identified by having a bounding box around each phrase on the document and determining the location of the bounding box as the location of each phrase. For example, as shown in FIG. 2 , the online system 130 may identify a set of bounding boxes (dotted lines) for the set of phrases: “Bill To:,” “John Smith,” “Invoice #,” “US-001,” “Invoice Date,” “11 / 2 / 2019,” “Qty,” “Description,” “Unit Price,” “Amount,” “2,” “Front end brakes,” “100.00,” “200.00,” “Total Charges:,” and “$200.00.” Each bounding box surrounds a respective phrase, and the location of the bounding box (e.g., the location of the center of the box) may be determined as the location of each phrase on the document image 200.

[0024] In response to a request, the online system 130 may also retrieve a form template that includes a set of fields to extract from the form and one or more candidate or known keys for each field. Specifically, a candidate key for each field is a phrase that is a likely candidate for the field's key on a document when the label for the key is unknown. In the example of FIGS. 2 and 3, the online system 130 may retrieve a form template that specifies, for the invoice number field, one or more candidate keys {"Invoice No.", "Invoice #", "Invoice Number"} that are likely candidates for the field's key to be labeled in the form.

[0025] In one embodiment, the online system 130 may receive from the requesting user a customized form template that specifies at least a portion of the form template, e.g., candidate keys for fields. Because the requesting user is likely to be familiar with the document, receiving the customized form template along with the document allows the online system 130 to access phrases that are likely candidates for keys in the document. In another embodiment, the online system 130 may not receive a form template, or may receive a partially complete form template from the request and internally determine the set of fields and one or more candidate keys. For example, the online system 130 may determine this information based on the type of form by storing templates of previously processed documents based on categories, such as invoices, mortgage applications, documents from different government agencies, etc. The online system 130 may determine the appropriate category of the incoming form, retrieve a template from a previously processed document of the same category, and assign the retrieved template as the template for the incoming document.

[0026] For at least one field, the online system 130 determines a key score for a set of phrases, where the key score for a phrase indicates the likelihood that the phrase is a key for the field on the image. In one example, for a given field, the key score for a phrase is determined based on a match between the phrase and the candidate key for the field, such that the key score for the phrase is higher if the phrase is similar to the candidate key for the field. As shown in FIG. 3, for the invoice number field, the online system 130 determines a key score for the set of identified phrases based on the candidate key for that field. Specifically, the phrase "Invoice#" has the highest key score because it substantially or exactly matches the candidate key for the field.

[0027] The online system 130 also identifies a set of candidate values ​​for the field from the set of phrases and, for each candidate value, identifies a set of neighbors from the set of phrases. In one example, the candidate values ​​are determined by first identifying a data type for the field's value (e.g., cardinal, date, text string) and selecting only phrases that match that data type. As shown in Figure 3, the online system 130 identifies the data type of the invoice number field as a text string or cardinal and selects a subset of phrases containing numbers or text as candidate values ​​for the field, including "100.00," "US-001," "11 / 2 / 2019," and "2."

[0028] For each candidate value, the online system 130 determines neighbors that are spatially close to the candidate value on the document. For example, in FIG. 2, the set of neighbors for the candidate value "100.00" may include "Unit Price," "200.00," and "Front end brakes" because their spatial locations on the document are close to the location of the candidate value. The online system 130 determines neighbor scores, where the neighbor score for a candidate value and each neighbor is determined based on the neighbor's key score and the neighbor's spatial relationship to the candidate value. As shown in FIG. 3, the online system 130 identifies neighbor scores for candidate value-neighbor pairs, and among them, the candidate value "US-001" and its neighbor "Invoice #" have the highest neighbor score of 0.92 because the neighbor "Invoice #" is associated with a high key score and is spatially close to the candidate value "US-001."

[0029] The online system 130 selects candidate values ​​and their respective neighbors based on the neighbor scores, sets the selected candidate values ​​as the values ​​of the field, and sets the selected neighbors as the keys of the field. In the example of Figures 2 and 3, the online system 130 selects the key for the invoice number field as "Invoice #" and the value for the field as "US-001" because the neighbor score of that pair is the highest among the other pairs for that field. The online system 130 may repeat this process for other fields specified in the document template and extract the remaining information in the form of key-value pairs.

[0030] Returning to FIG. 1 , client device 110 is a computing device such as a smartphone, tablet computer, laptop computer, desktop computer, or any other type of network-enabled device with an operating system such as ANDROID® or APPLE® IOS®. A typical client device 110 includes hardware and software necessary to connect to network 122 (e.g., via WiFi and / or 4G, 5G, or other wireless communication standards). Client device 110 enables users to submit requests to online system 130 and extract information from documents. Client device 110 may include an operating system and various applications running on the operating system that enable users to submit requests. For example, client device 110 may include a browser application or a standalone application deployed by online system 130 that enables users of an organization to interact with online system 130 and submit requests.

[0031] The network 122 provides a communications infrastructure between the worker devices 110 and the process mining system 130. The network 122 is typically the Internet, but may be any network including, but not limited to, a local area network (LAN), a metropolitan area network (MAN), a wide area network (WAN), a mobile wired or wireless network, a private network, or a virtual private network.

[0032] Online System 4 is a high-level block diagram showing a detailed view of online system 130 according to one embodiment. Online system 130 is comprised of modules including request management module 410, template management module 415, recognition module 420, and key-value identifier module 425. Online system 130 also includes forms and templates data store 450 and key-value pair data store 455. Some embodiments of online system 130 have different modules than those described herein. Similarly, functionality may be distributed among modules in a manner different from that described herein.

[0033] The request management module 410 receives requests from the client device 110 to process digital documents and provides the extracted information to the user of the request. Specifically, the request management module 410 may receive requests that include images of documents and may apply preprocessing techniques to the images before information is extracted from the documents. For example, the request management module 410 may perform cropping, image enhancement techniques, scaling, translation, or rotation on the document and provide the preprocessed image to the recognition module 420. The request management module 410 may store the documents in the forms and templates data store 450.

[0034] Additionally, the request may also include a customized document template from the requesting user that specifies one or more fields to be extracted from the document and any candidate or known keys for the fields. The request management module 410 may receive the customized document template and forward the template to the template management module 415. The request management module 410 may also store any received customized templates in the forms and templates data store 450.

[0035] In response to receiving information extracted from the request document from a module of the online system 130, the request management module 410 provides this information to the user of the request. Specifically, the request management module 410 may receive the extracted information in the form of key-value pairs that can be provided to the user in an appropriate format. In one example, the request management module 410 provides the key-value pairs as text in a text file. In another example, the request management module 410 visually provides the key-value pairs by annotating the document with the locations of the identified key-value pairs. The annotations may be in the form of bounding boxes, which are rectangles that surround the key-value pairs, or segmentations that outline the actual text of the key-value pairs in the document.

[0036] The template management module 415 creates and manages templates for documents. In one embodiment, the template management module 415 receives a customized document template attached with a request and may flag any errors or incomplete information in the customized template. For example, a template may include a set of fields, but one or more of the template's fields may lack a candidate key for the one or more fields. In such cases, the template management module 415 may generate candidate keys for the fields, for example, based on previous instances of the document processed by the online system 130. For example, the template management module 415 may generate a candidate key for an invoice number field based on a key previously identified by the online system 130 for an invoice document.

[0037] In another embodiment, the template management module 415 may determine that a document for a request is not associated with any form template. In such a case, the template management module 415 may create a template for the document, for example, based on a previous instance of the document processed by the online system 130. For example, the template management module 415 may create form templates for documents based on templates generated or received for the same type of document (e.g., invoice, application, government form), documents from the same issuing entity, or documents with the same author. The template management module 415 stores the templates in the forms and templates data store 450, each associated with a respective document in the request.

[0038] The recognition module 420 receives the document included in the request and identifies a set of phrases and the locations of the phrases on the document. A phrase may be defined as a group of one or more words on the document that are spatially close to each other. In one embodiment, a group of one or more words is identified as a phrase on the document if the horizontal distance between the words is less than a predetermined threshold. The recognition module 420 may perform text recognition methods, such as optical character recognition (OCR) or applying machine learning models, to identify words and groups of words as phrases on the document. The recognition module 420 also associates each phrase with a location on the document. For example, the recognition module 420 may generate a bounding box around the phrase and determine the spatial coordinates of the bounding box as the location of the phrase. The spatial coordinates are {x min ,y min ,x max ,y ma x}, where x min is the leftmost horizontal coordinate, and y min is the bottom vertical coordinate, and x max is the rightmost horizontal coordinate, and y max is the top vertical coordinate of the bounding box.

[0039] The key-value identifier module 425 receives a document and a set of phrases and templates for the document and extracts information in the form of key-value pairs. For at least one field specified in the template, the key-value identifier module 425 determines a key score for the set of phrases for each document. In one embodiment, the key-value identifier module 425 determines a phrase key score for a given field based on string matching between the phrase and candidate keys for that field. In one example, the string matching is a fuzzy match between the phrase and the candidate keys for the field, and the key-value identifier module 425 generates a matching score between the phrase and the candidate keys for the field that indicates the similarity between the two pieces of text. The key score for a phrase is the maximum matching score between the candidate keys.

[0040] The key-value identifier module 425 also identifies a set of candidate values ​​for the field from a set of phrases in the document. In one embodiment, the key-value identifier module 425 determines the data type of the field's values ​​and selects only phrases that match that data type. In one example, the key-value identifier module 425 generates a set of categories and tags the field with one or more categories, including, but not limited to, person names, organizations, locations, cardinal numbers, medical codes, time expressions (e.g., dates or times), quantities, monetary values, percentages, etc. The key-value identifier module 425 may apply a named entity recognizer (NER) model to determine whether phrases belong to categories that match one or more of the field's categories and select only the matching phrases as candidate values ​​for the field.

[0041] For each candidate value from the set of phrases, the key-value identifier module 425 identifies a set of neighbors that are spatially close to the candidate value on the document. In one embodiment, the neighbor phrases for a candidate value are those that have spatial locations within a predetermined distance from the candidate value's location on the document. For example, the neighbor phrases for a candidate value may have bounding boxes that significantly overlap the bounding box area of ​​the respective candidate value, e.g., more than 90%, 80%, or 70% overlap.

[0042] The key-value identifier module 425 determines a neighborhood score for each candidate value and neighborhood pair, where the neighborhood score for the candidate value and each neighborhood is determined based on the neighborhood's key score and a spatial score indicating the neighborhood's spatial relationship to the candidate value. In one embodiment, the spatial score is given by a combination of a distance score and an angle score, for example, a weighted sum of the distance score and the angle score, and may be given as:

number

[0043] In particular, the key-value identifier module 425 calculates the distance between the candidate value and its neighboring locations on the document. ij The key-value identifier module 425 determines a distance score as a function of , where in one example the function is a Gaussian distribution (or any other probability distribution) that takes the distance as input and is centered around mean 0 and standard deviation z1. Similarly, the key-value identifier module 425 determines the angle angle between the location of the candidate value and its neighboring locations on the document. ij Determine the angle score as a function of , where in one example the function takes the angle as input and is a Gaussian distribution (or any other probability distribution) centered around mean 0 and standard deviation z2. ij is given as the minimum angular distance between candidate value i's location (e.g., the center of the bounding box) and neighbor j's location with respect to the set of anchor angles. For example, if key-value pairs are likely to be located horizontally from left to right on the document (e.g., anchor angle 0°) or vertically one above the other (e.g., anchor angle 90°), the set of anchor angles can be {0°, 90°}. Thus, the spatial score is increased for candidate value-neighbor pairs that have close distances from each other and are aligned with each other on the document.

[0044] In one example, the neighborhood score ns(·) for a candidate value and its respective neighbors is given by:

number

[0045] The key-value identifier module 425 selects candidate values ​​and their respective neighbors based on the ranking scores, sets the selected candidate values ​​as the values ​​of the field, and sets the selected neighbors as the keys of the field. For example, the key-value identifier module 425 may select the candidate value and neighbor pair with the highest ranking score as the final key-value pair for that field. The key-value identifier module 425 repeats this process for other fields specified in the document template, extracting the remaining information in the form of key-value pairs. The key-value identifier module 425 stores the extracted information in the key-value pair data store 455 and provides the extracted information to the request management module 410, which can provide this information to the user of the request.

[0046] How to determine key-value pairs from a digital document 5 illustrates a flowchart of a method for determining key-value pairs in a form document, according to one embodiment. In one embodiment, the process of FIG. 5 is performed by various modules of online system 130. In other embodiments, other entities may perform some or all of the steps of the process. Similarly, embodiments may include different and / or additional steps, or may perform steps in a different order.

[0047] The online system 130 receives a form image from a client device and obtains a template indicating one or more fields to extract from the form image (502). At least one field is associated with a set of candidate keys for that field. The online system 130 obtains a set of phrases from the form image and obtains the locations of the set of phrases on the document (504). For at least one field, the online system 130 determines a key score for the set of phrases (506). The key score of a phrase may indicate the likelihood that the phrase is a key for the field on the document. The online system 130 identifies a set of candidate values ​​for the field from the set of phrases and identifies a set of neighbors for each candidate value from the set of phrases (508).

[0048] The online system 130 determines (510) neighborhood scores for the set of candidate values ​​and the set of neighborhoods. The neighborhood scores for the candidate values ​​and their respective neighborhoods may be determined from the neighborhood's key scores and the neighborhood's spatial relationship to the candidate values. The online system 130 selects the candidate values ​​and their respective neighborhoods associated with neighborhood scores that exceed a threshold and sets (512) the selected candidate values ​​as values ​​of the field and the selected neighborhoods as keys of the field. The online system 130 may repeat steps 506 through 512 for the remaining fields specified in the document template.

[0049] Computer Architecture 6 is a block diagram illustrating the architecture of a typical computer system for use in the environment of FIG. 1, according to one embodiment. Shown is at least one processor 602 coupled to a chipset 604. Also coupled to chipset 604 are memory 606, a storage device 608, a keyboard 610, a graphics adapter 612, a pointing device 614, and a network adapter 616. A display 618 is coupled to graphics adapter 612. In one embodiment, the functionality of chipset 604 is provided by a memory controller hub 620 and an I / O controller hub 622. In another embodiment, memory 606 is coupled directly to processor 602 instead of chipset 604.

[0050] Storage device 608 is a non-transitory computer-readable storage medium such as a hard drive, compact disc read-only memory (CD-ROM), DVD, or solid-state memory device. Memory 606 holds instructions and data used by processor 602. Pointing device 614 may be a mouse, trackball, or other type of pointing device and is used in combination with keyboard 610 to input data into computer system 600. Graphics adapter 612 displays images and other information on display 618. Network adapter 616 couples computer system 600 to a network.

[0051] As is known in the art, computer 600 may have different and / or other components than those shown in Figure 6. Additionally, computer 600 may lack certain illustrated components. For example, a computer system 600 operating as online system 130 may lack keyboard 610 and pointing device 614. Furthermore, storage device 608 may be local and / or remote from computer 600 (such as may be embodied in a storage area network (SAN)).

[0052] The computer 600 is adapted to execute computer modules for providing the functionality described herein. As used herein, the term "module" refers to computer program instructions and other logic for providing a specified functionality. A module may be implemented in hardware, firmware, and / or software. A module may include one or more processes and / or may be provided by only a portion of a process. A module is typically stored in the storage device 608, loaded into the memory 606, and executed by the processor 602.

[0053] 1 may vary depending on the embodiment and the processing power used by the entity. For example, client device 110 may be a mobile phone with limited processing power, may have a small display 618, and may lack a pointing device 614. In contrast, online system 130 may include multiple blade servers that cooperate to provide the functionality described herein.

[0054] Additional Considerations The foregoing description of embodiments of the present invention has been presented for purposes of illustration and is not intended to be exhaustive or to limit the invention to the precise form disclosed. Those skilled in the art will recognize that many modifications and variations are possible in light of the above disclosure.

[0055] Some portions of this description describe embodiments of the invention in terms of algorithms and symbolic representations of operations on information. These algorithmic descriptions and representations are commonly used by those skilled in the data processing arts to effectively convey the substance of their work to others skilled in the art. While these operations are described functionally, computationally, or logically, they are understood to be implemented by computer programs or equivalent electrical circuits, microcode, or the like. Further, it has proven convenient at times to refer to arrangements of these operations as modules, without loss of generality. The described operations and their associated modules may be embodied in software, firmware, hardware, or any combination thereof.

[0056] Any of the steps, operations, or processes described herein may be performed or implemented in one or more hardware or software modules, alone or in combination with other devices. In one embodiment, the software modules are implemented in a computer program product that includes a computer-readable medium containing computer program code, which can be executed by a computer processor to perform any or all of the described steps, operations, or processes.

[0057] Embodiments of the present invention may also relate to apparatus for performing the operations herein. This apparatus may be specially constructed for the required purposes and / or may include a general-purpose computing device selectively activated or reconfigured by a computer program stored in the computer. Such a computer program may be stored on a non-transitory tangible computer-readable storage medium or any type of medium suitable for storing electronic instructions that may be coupled to a computer system bus. Furthermore, any computing system referred to herein may include a single processor or may be an architecture employing a multiple processor design for increased computing power.

[0058] Embodiments of the present invention may also relate to products produced by the computing processes described herein. Such products may include information obtained from the computing processes, where the information may be stored on a non-transitory, tangible, computer-readable storage medium, and may include any embodiment of the computer program product or other data combination described herein.

[0059] Finally, the language used in this specification has been chosen primarily for ease of reading and instructional purposes, and may not be chosen to define or limit the subject matter of the present invention. Accordingly, it is intended that the scope of the invention be limited not by this detailed description, but rather by the claims issued in an application based hereon. Accordingly, the disclosure of embodiments of the present invention is intended to be illustrative, but not limiting, of the scope of the invention, which is set forth in the following claims.

Claims

1. 1. A computer-implemented method comprising: receiving a form image from a client device; obtaining a template indicating one or more fields to be extracted from the form image, at least one field being associated with a set of candidate keys for that field; obtaining a set of phrases from the form image and obtaining positions of the phrases; For the at least one field: determining key scores for phrases from the set of phrases based on candidate keys for the at least one field, the key score for a phrase indicating the likelihood that the phrase is a key for the field on the form image; identifying a set of candidate values ​​for the at least one field from the set of phrases; determining a set of candidate value-neighborhood pairs by identifying a set of neighbors for candidate values ​​from the set of phrases based on the positions of the phrases; determining neighborhood scores for the set of candidate value-neighborhood pairs, each neighborhood score for each candidate value-neighborhood pair being determined from a combination of the key score of a neighborhood of the candidate value-neighborhood pair and a spatial relationship of the neighborhood to the candidate value of the candidate value-neighborhood pair; selecting a neighborhood value-candidate pair from the set of candidate value-neighborhood pairs that has a neighborhood score above a threshold; setting the candidate value of the selected neighborhood value-candidate pair as a value of the at least one field and setting the neighborhood of the selected neighborhood value-candidate pair as the key of the at least one field; 20. A computer-implemented method comprising:

2. The key score of the phrase is determined by performing string matching between the phrase and the candidate key for the at least one field to determine a similarity between the phrase and the candidate key. The computer-implemented method of claim 1 .

3. Identifying the set of neighbors for the candidate value from the set of phrases based on the position of the phrases includes identifying the set of neighbors for the candidate value from the set of phrases based on the position of the phrases such that the neighbors for a candidate value have positions within a threshold distance from the position of the candidate value in the form image. The computer-implemented method of claim 1 .

4. At least a portion of the information included in the template for the form image is received from the client device. The computer-implemented method of claim 1 .

5. The step of identifying a set of candidate values ​​comprises: determining a data type for the at least one field by assigning the at least one field to at least one of a set of predetermined categories; selecting a subset of phrases in the set of phrases that match the data type of the at least one field as the candidate values ​​for the at least one field; The computer-implemented method of claim 1 further comprising:

6. the spatial relationship of the neighbor to the candidate value is determined based on a combination of a distance between the position of the candidate value and the position of the neighbor on the form image and an angle between the position of the candidate value and the position of the neighbor on the form image. The computer-implemented method of claim 1 .

7. providing the key of the at least one field and the value of the at least one field to a user of the client device. The computer-implemented method of claim 1 .

8. A non-transitory computer-readable storage medium storing executable computer program instructions to perform operations, the operations comprising: receiving a form image from a client device; obtaining a template indicating one or more fields to be extracted from the form image, at least one field being associated with a set of candidate keys for that field; obtaining a set of phrases from the form image and obtaining positions of the phrases; For the at least one field: determining key scores for phrases from the set of phrases based on candidate keys for the at least one field, the key score for a phrase indicating the likelihood that the phrase is a key for the field on the form image; identifying a set of candidate values ​​for the at least one field from the set of phrases; determining a set of candidate value-neighborhood pairs by identifying a set of neighbors for candidate values ​​from the set of phrases based on the positions of the phrases; determining neighborhood scores for the set of candidate value-neighborhood pairs, each neighborhood score for each candidate value-neighborhood pair being determined from a combination of the key score of a neighborhood of the candidate value-neighborhood pair and a spatial relationship of the neighborhood to the candidate value of the candidate value-neighborhood pair; selecting a neighborhood value-candidate pair from the set of candidate value-neighborhood pairs that has a neighborhood score above a threshold; setting the candidate value of the selected neighborhood value-candidate pair as a value of the at least one field and setting the neighborhood of the selected neighborhood value-candidate pair as the key of the at least one field; 1. A non-transitory computer-readable storage medium comprising:

9. The method of claim 8, wherein the key score for the phrase is determined by performing string matching between the phrase and the candidate key for the at least one field to determine a similarity between the phrase and the candidate key. The non-transitory computer-readable storage medium of claim 8.

10. Identifying the set of neighbors for the candidate value from the set of phrases based on the position of the phrases includes identifying the set of neighbors for the candidate value from the set of phrases based on the position of the phrases such that the neighbors for a candidate value have positions within a threshold distance from the position of the candidate value in the form image. The non-transitory computer-readable storage medium of claim 8.

11. At least a portion of the information included in the template for the form image is received from the client device. The non-transitory computer-readable storage medium of claim 8.

12. The step of identifying a set of candidate values ​​comprises: determining a data type for the at least one field by assigning the at least one field to at least one of a set of predetermined categories; selecting a subset of phrases in the set of phrases that match the data type of the at least one field as the candidate values ​​for the at least one field; The non-transitory computer-readable storage medium of claim 8 , further comprising:

13. the spatial relationship of the neighbor to the candidate value is determined based on a combination of a distance between the position of the candidate value and the position of the neighbor on the form image and an angle between the position of the candidate value and the position of the neighbor on the form image. The non-transitory computer-readable storage medium of claim 8.

14. the operations further include providing the key of the at least one field and the value of the at least one field to a user of the client device. The non-transitory computer-readable storage medium of claim 8.

15. 1. A system comprising: a processor for executing computer program instructions; receiving a form image from a client device; obtaining a template indicating one or more fields to be extracted from the form image, at least one field being associated with a set of candidate keys for that field; obtaining a set of phrases from the form image and obtaining positions of the phrases; For the at least one field: determining key scores for phrases from the set of phrases based on candidate keys for the at least one field, the key score for a phrase indicating the likelihood that the phrase is a key for the field on the form image; identifying a set of candidate values ​​for the at least one field from the set of phrases; determining a set of candidate value-neighborhood pairs by identifying a set of neighbors for candidate values ​​from the set of phrases based on the positions of the phrases; determining neighborhood scores for the set of candidate value-neighborhood pairs, each neighborhood score for each candidate value-neighborhood pair being determined from a combination of the key score of a neighborhood of the candidate value-neighborhood pair and a spatial relationship of the neighborhood to the candidate value of the candidate value-neighborhood pair; selecting a neighborhood value-candidate pair from the set of candidate value-neighborhood pairs that has a neighborhood score above a threshold; setting the candidate value of the selected neighborhood value-candidate pair as a value of the at least one field and setting the neighborhood of the selected neighborhood value-candidate pair as the key of the at least one field; a non-transitory computer-readable storage medium storing computer program instructions executable to perform steps including: A system comprising:

16. The key score for the phrase is determined by performing string matching between the phrase and the candidate key for the at least one field to determine a similarity between the phrase and the candidate key.

16. The system of claim 15.

17. Identifying the set of neighbors for the candidate value from the set of phrases based on the position of the phrases includes identifying the set of neighbors for the candidate value from the set of phrases based on the position of the phrases such that the neighbors for a candidate value have positions within a threshold distance from the position of the candidate value in the form image.

16. The system of claim 15.

18. At least a portion of the information included in the template for the form image is received from the client device.

16. The system of claim 15.

19. The step of identifying a set of candidate values ​​comprises: determining a data type for the at least one field by assigning the at least one field to at least one of a set of predetermined categories; selecting a subset of phrases in the set of phrases that match the data type of the at least one field as the candidate values ​​for the at least one field; The system of claim 15 further comprising:

20. the spatial relationship of the neighbor to the candidate value is determined based on a combination of a distance between the position of the candidate value and the position of the neighbor on the form image and an angle between the position of the candidate value and the position of the neighbor on the form image.

16. The system of claim 15.

Citation Information

Patent Citations

  • Form recognition device and form recognition method

    JP2011248609A

  • Information processing device, information processing method, program, and document reading system

    JP2020016946A