Machine Learning System for Automatic Document Segmentation and Classification
A machine learning system automates document splitting and classification using neural networks and classification techniques, addressing the inefficiencies and errors of manual document processing by enhancing accuracy, speed, and adaptability.
Patent Information
- Application Number
- JP2024568107
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-04-20
- Filing Date
- 2024-04-16
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2044-04-16
AI Technical Summary
Manual document processing is time-consuming, labor-intensive, and prone to errors, especially when dealing with large volumes of diverse documents.
A machine learning system that automatically splits and classifies documents using a visual segmentation neural network, optical character recognition, title classification, document classification, and a grouper subsystem.
The system achieves improved accuracy, speed, and efficiency in document processing, reducing manual effort and enhancing data quality while being scalable and adaptable to different document formats.
Smart Images

Figure 2025516741000001_ABST
Abstract
Description
Background Art
[0001] The present disclosure generally relates to the field of document processing, and more specifically to a machine learning system for automatically splitting and classifying documents.
[0002] In today's digital age, a large number of documents are created and managed daily by organizations and individuals. These include, for example, contracts, invoices, legal documents, financial documents, and other important documents that are important for the functioning of an organization. Managing and processing these documents can be time-consuming and labor-intensive. In addition, manual processes can introduce risks of errors and inconsistencies that can affect the accuracy and reliability of the information being stored.
Summary of the Invention
[0003] Generally, one innovative aspect of the subject matter described herein can be embodied in a computer-implemented method for automatically splitting and classifying an input document into one or more sub-documents using a machine learning system. The input document includes a plurality of pages. The machine learning system includes a visual segmentation neural network, an optical character recognition subsystem, a title classifier, a document classifier, and a grouper subsystem. The method includes 〇receiving visual input representing a plurality of pages of the input document, and 〇using the visual segmentation neural network to classify each page of the input document into each of a plurality of templates, and 〇for each page of the input document, ·using the optical character recognition subsystem to generate a set of text lines from the text content of the page, ·using the title classifier to process each text line in the set of text lines to generate a corresponding confidence score representing the probability that the text line contains the title of the page, ·selecting the text line having the highest confidence score based on the confidence scores of the set of text lines, · Determine whether the highest confidence score exceeds a threshold, · In response to determining that the highest confidence score exceeds the threshold, use a document classifier to process the selected text lines to generate respective document scores for each of a plurality of document types, where each respective document score represents the probability that the page belongs to the document type, · Perform an operation that includes selecting the document type having the highest document score as the final document type to which the page belongs, 〇 Use a grouper subsystem to group a plurality of pages of an input document into one or more sub-documents based on (i) respective templates of each page and (ii) respective final document types to which each page belongs.
[0004] In some implementations, the machine learning system further includes an identification extraction neural network. In these implementations, the method includes, for each page of the input document, in response to determining that the highest confidence score does not exceed the threshold, using the identification extraction neural network to process each text line in a set of text lines to identify the identification number of the page. In these implementations, grouping a plurality of pages of the input document into one or more sub-documents using the grouper subsystem is further based on the identification number of the page.
[0005] Other embodiments of this aspect include corresponding systems, devices, and computer programs configured to perform the operations of the method, encoded on a computer storage device.
[0006] The subject matter described in this specification may be implemented in particular embodiments so as to realize one or more of the following technical advantages. ○ Improved accuracy: The described system takes into account both the visual and textual features of a document and thus achieves better performance in document processing tasks compared to existing systems. In particular, by using a specific combination of a visual segmentation neural network, a title classifier, a document classifier, and optionally an identification extraction neural network to process both the visual and textual features of a document, the machine learning system described herein can identify the template, document type, and identification number (if present) of each page. Using the identified information, the described system can split a document into one or more sub-documents with higher accuracy compared to existing systems. As a result, a more streamlined, efficient, and effective document processing system is obtained. ○ Scalability: The described machine learning system can handle a large number of documents in a short time and is suitable for use in organizations that must process a large number of documents daily. For example, the described system can process hundreds or thousands of documents in a few minutes. Such an operation cannot be performed manually by humans. ○ Improved speed: By using a visual segmentation neural network, a title classifier, a document classifier, and optionally an identification extraction neural network, the described system can process documents faster compared to existing document processing systems, thus saving the organization a significant amount of time and resources. ○ Robustness and customizability: The machine learning system is designed and trained to be robust to changes in document format, layout, and style. This means that it can still execute accurately even if the format of the document changes over time. Furthermore, by using multiple templates and document types, the machine learning system can learn and adapt to meet the specific needs of an organization. ○ Improved data quality: The described machine learning system can accurately classify each page of a document into the appropriate document type, which results in improved data quality and can reduce the manual effort required to review and correct errors in existing document processing systems. Additionally, the visual segmentation neural network, title classifier, document classifier, and identification extraction neural network enable improved data extraction from documents, allowing organizations to make more informed decisions based on the extracted data. ○ Improved document management: The described machine learning system's ability to automatically group sub-documents based on templates, document types, and identification numbers improves the overall management and organization of documents. This helps reduce the time and computational resources required to search for documents among large volumes of documents.
[0007] Details of one or more embodiments of the subject matter of this specification are set forth in the accompanying drawings and the following description. Other features, aspects, and advantages of the subject matter will become apparent from the description, drawings, and claims.
Brief Description of the Drawings
[0008]
Figure 1
[0009]
Figure 2
[0010]
Figure 3
[0011]
Figure 4
[0012]
Figure 5A
Figure 5B
[0013]
Figure 6
[0014] Like reference numerals and designations in the various drawings denote like elements.
DETAILED DESCRIPTION OF THE INVENTION
[0015] This specification describes a machine learning system implemented as a computer program on one or more computers in one or more locations configured to automatically split and classify an input document into one or more sub-documents. The machine learning system includes a visual segmentation neural network, an optical character recognition (OCR) subsystem, a title classifier, a document classifier, an identification extraction neural network, and a grouper subsystem.
[0016] Generally, an input document is a collection of multiple sub-documents that are combined with each other. Each of these sub-documents can belong to any document type in a set of document types.
[0017] For example, in some implementations, the set of document types includes one or more of invoices, delivery memos, purchase orders, or emails. For example, a 5-page input document may include a tax invoice on pages 1 and 2, a delivery memo on page 3, and the rest may be a purchase order. As another example, a 300-page input document may include an email on pages 1 to 57, a purchase order on pages 58 to 137, a delivery memo on pages 138 to 212, and the remaining pages may be invoices.
[0018] In some other implementations, the set of document types includes one or more financial document types. For example, the set of document types includes one or more of a balance sheet, an income statement, a cash flow statement, a statement of owner's equity, an annual report, a quarterly report, a memo to the financial statements, or a text return report.
[0019] In some other implementations, the set of document types includes one or more legal document types. For example, the set of document types includes one or more of an employment contract, an independent contractor contract, a non-disclosure agreement, a loan contract, a consulting contract, a partnership agreement, corporate law, a business contract, or a sales contract.
[0020] In some other implementations, the set of document types includes one or more medical document types. For example, the set of document types includes one or more of a doctor's memo, a prescription, an inoculation record, a medical record, a discharge summary, a medical test result, a surgical report, a consent form, or an email confirming an appointment.
[0021] The set of document types can be extended based on the header content and / or title of the document.
[0022] FIG. 1 shows an exemplary machine learning system 100 for automatically splitting and classifying documents. System 100 is an example of a system implemented as a computer program on one or more computers in one or more locations and can implement the systems, components, and techniques described below. Machine learning system 100 includes a visual segmentation neural network 104, an OCR subsystem 108, a title classifier 112, a document classifier 114, an identification extraction neural network 120, and a grouper subsystem 118.
[0023] To process the input document, system 100 first receives a visual input 102 representing multiple pages of the input document. For example, the visual input 102 can be a PDF file that is a combination of all scanned images of all pages of the input document.
[0024] The visual segmentation neural network 104 is configured to classify each page within the visual input 102 into respective templates 106 within a set of templates. Each template can be used to generate different document types. Different templates have different visual appearances. Specifically, each template has one or more corresponding visual features including, but not limited to, logos, padding, backgrounds, headers, footers, charts, and font styles.
[0025] The visual segmentation neural network 104 is a convolutional neural network (CNN) that includes one or more convolutional neural network layers. The CNN is trained on a document template recognition task. Generally, each of the one or more convolutional neural network layers includes a plurality of artificial neurons. The artificial neurons are arranged in a 2D or 3D grid called a filter. Each filter extracts different types of features from the input data. For example, from an image, one filter can extract edges, lines, circles, or more complex shapes.
[0026] To classify each page of the input document, the neural network 104 combines the page vertically with previous pages in the document to generate a combined two-page input and provides this combined two-page input to one or more trained CNN layers. The one or more trained CNN layers extract a set of visual features from each page of the combined two-page input and compare the two sets of visual features to determine whether the page and the previous page have the same template. One or more visual features can include logos, padding, backgrounds, headers, footers, charts, or font styles.
[0027] If the current page and the previous page have the same visual features, this means that they have the same template, and one or more trained convolutional neural network layers associate the current page with the known template of the previous page. If the current page and the previous page have different visual features, this means that they have different templates, and one or more trained convolutional neural network layers associate the current page with a new template. For example, as shown in FIG. 5A, one or more convolutional neural network layers 500 of the neural network 104 process the combined two-page input 502 and determine that the two pages in the input 502 belong to the same template. In contrast, FIG. 5B shows that one or more convolutional neural network layers 500 process the combined two-page input 504 and determine that the two pages in the input 504 belong to two different templates.
[0028] The results of the process of classifying each page of the input document are shown in FIG. 2, and the visual segmentation neural network 104 uses one or more trained convolutional neural network layers to classify pages 1-4 of the input document into template 1 and pages 5 and 6 of the input document into template 2. FIG. 2 shows a simplified implementation. In practice, the input document can have dozens, hundreds, or thousands of pages.
[0029] The following series of operations are performed on each page of the input document.
[0030] For each page, the OCR subsystem 108 is configured to generate a set of text lines 110 from the text content of the page. Next, the title classifier 112 processes each text line within the set of text lines to generate a corresponding confidence score representing the probability that the text line contains the title of the page. The title classifier 112 is configured to select the text line having the highest confidence score based on the confidence scores of the set of text lines. In some implementations, the title classifier 112 is a random forest classifier. The process for generating the confidence scores will be described in more detail below with reference to FIG. 3.
[0031] The system 100 determines whether the highest confidence score exceeds a threshold. If the highest confidence score exceeds the threshold, this means that the selected text line is very likely to contain the title, and the system 100 processes the selected text line to determine the document type 116 of the page. In particular, the system 100 uses the document classifier 114 to process the selected text line to generate a respective document score for each document type within the set of document types. Each document score represents the probability that the page belongs to that document type. The system 100 selects the document type having the highest document score as the final document type 116 to which the page belongs.
[0032] More specifically, as shown in FIG. 4, the system maintains (i) a positive set 402 of keywords representing document types that are likely to match the final document type of the page, and (ii) a negative set 404 of keywords representing document types that are unlikely to match the final document type of the page. For example, in the case of invoice classification, the positive set 402 of keywords can include "invoice", "tax invoice", "order invoice", and "commercial invoice", and the negative set 504 of keywords can include "delivery memo", "insurance", "purchase order", and "email". The document classifier 114 is configured to calculate a respective document score for each keyword within the positive set 402 of keywords and the negative set 404 of keywords. The document score of each keyword represents the probability that the page belongs to the document type represented by that keyword. The system 100 then selects the keyword having the highest document score. This keyword designates the final document type 116 to which the page belongs.
[0033] If the highest confidence score does not exceed the threshold, this means that the selected text line is unlikely to contain the title, and the system 100 uses the identification extraction neural network 120 to process each text line within the set of text lines to identify the identification number 122 of the page. The identification number can be, for example, an invoice number or a purchase number. In particular, the identification extraction neural network 120 extracts key information from each text line (by using, for example, the Spatial Dual-Modality Graph Reasoning (SDMGR) method or the graph method) and identifies the identification number 122 of the page from the extracted key information.
[0034] In some implementations, the system 100 can also determine the identification number even when the selected text line is likely to contain the title.
[0035] Once the system 100 determines the respective template for each page, the respective final document type for each page, and optionally the respective identification number for each page, the grouper subsystem 118 groups the pages of the input document represented by the visual input 102 into one or more sub-documents 124 based on the respective template for each page, the respective final document type for each page, and optionally the respective identification number for each page.
[0036] The components of the machine learning system 100 are designed and trained to be robust to changes in document format, layout, and style. This means that the system 100 can still execute accurately even if the document format changes over time. Additionally, the use of multiple templates and document types allows the machine learning system to learn and adapt to meet the specific needs of the organization.
[0037] FIG. 3 shows an exemplary process for processing page 302 to generate a confidence score for each text line among a plurality of text lines within page 302. First, the OCR subsystem 108 is configured to extract a plurality of text lines 304 from the text content of page 302. In particular, the OCR subsystem is a neural network trained to recognize text and extract text lines from the scanned image of the page. The OCR subsystem analyzes text across multiple levels (e.g., character level, word level, line level, etc.) and is trained to repeatedly process the image. This searches for different image attributes such as curves, lines, intersections, and loops, and combines the results of all these different levels of analysis to obtain a final result that enables the OCR subsystem to recognize text and extract text lines from page 302.
[0038] Next, the system 100 calculates (306) a set of features for each text line. For example, the set of features for each text line may include the x - coordinate, y - coordinate, width w, height h, and matching score of the text line. The system 100 calculates the matching score for each text line by using a fuzzy matching method. Specifically, for each text line and for each of the positive set of keywords and the negative set of keywords, the system 100 calculates a respective score for each keyword in the set, and each score represents the probability that the text line contains the keyword. The system 100 selects the highest score among the respective scores as the matching score of the text line.
[0039] The system 100 provides the calculated features of all text lines as input to a title classifier 112 (which is a random forest classifier 408 in this example). The random forest classifier 408 maintains a set of random trees (e.g., N random trees where N is a positive integer greater than 2) for classifying whether each line of text contains a title. Each random tree selects a respective random set of features from the set of features. Based on each selected random set of features, each random tree calculates a respective set of confidence scores including the corresponding confidence score for each text line, and each corresponding confidence score for each text line represents the probability that the text line contains a title. The random forest classifier calculates the final confidence score for each text line by taking the weighted average of the N corresponding confidence scores calculated by the N random trees for the text line. The final confidence score for each text line represents the final probability that the text line contains a title. The text line with the highest final confidence score is most likely to contain a title.
[0040] FIG. 6 is a flowchart of an exemplary process for processing an input document to split it into one or more sub - documents.
[0041] For the sake of simplicity, process 600 is described as being executed by a system of one or more computers located in one or more locations. For example, a machine learning system appropriately programmed according to this specification, such as machine learning system 100 of FIG. 1, can execute process 600.
[0042] The system receives visual input representing multiple pages of an input document (step 602). The visual input can be a PDF file that is a combination of all scanned images of all pages of the input document.
[0043] The system classifies each page of the input document into each of a plurality of templates (step 604). Each template can be used to generate different document types. Different templates have different visual appearances. Specifically, each template has one or more corresponding visual features including, but not limited to, logos, padding, backgrounds, headers, footers, charts, and font styles.
[0044] To classify each page of the input document, the system vertically combines the page with previous pages in the document to generate a combined two-page input, and provides this combined two-page input to one or more trained CNN layers. The one or more trained CNN layers extract one or more visual features from the combined two-page input and compare them to determine whether the page and the previous page have the same template. If the page and the previous page have the same template, the one or more trained convolutional neural network layers associate the page with the known template of the previous page. If the page and the previous page have different templates, the one or more trained convolutional neural network layers associate the current page with a new template.
[0045] For each page of the input document, the system performs steps 606 - 616 as follows.
[0046] The system uses optical character recognition technology to generate a set of text lines from the text content of the page (step 606).
[0047] The system processes each text line in the set of text lines to generate a corresponding confidence score representing the probability that the text line contains the title of the page (step 608). For example, the system extracts features of each text line and provides the extracted features as input to a random forest classifier. The system then uses the random forest classifier to process the extracted features and generate a corresponding confidence score for each text line using the random forest classification method.
[0048] The system selects the text line with the highest confidence score based on the confidence scores of the set of text lines (step 610).
[0049] The system determines whether the highest confidence score exceeds a threshold (step 612).
[0050] In response to determining that the highest confidence score exceeds the threshold, the system uses a document classifier to process the selected text line and generate a respective document score for each of a plurality of document types. Each document score represents the probability that the page belongs to that document type.
[0051] The system selects the document type with the highest document score as the final document type to which the page belongs (step 616).
[0052] In particular, in some implementations, the system maintains (i) a positive set of keywords representing document types that are likely to match the final document type of the page, and (ii) a negative set of keywords representing document types that are unlikely to match the final document type of the page. The system calculates a respective document score for each keyword in the positive set of keywords and the negative set of keywords. The system selects the keyword with the highest document score. The selected keyword designates the final document type to which the page belongs.
[0053] In response to determining that the highest confidence score does not exceed a threshold, the system processes each text line in the set of text lines to identify the identification number of the page. In particular, the system extracts key information for each text line and identifies the identification number of the page from the extracted key information. In some implementations, the system extracts key information from each text line by using a Spatial Dual Modality Graph Reasoning (SDMGR) method or a graph method.
[0054] The system uses a grouper subsystem to group a plurality of pages of an input document into one or more sub-documents (step 618) based on (i) respective templates of each page, and (ii) respective final document types of each page, and / or (ii) respective identification numbers of each page. By grouping a plurality of pages into one or more sub-documents, the system has successfully split the input document into one or more sub-documents.
[0055] The system may display one or more sub-documents on the user interface of the system. Alternatively, or in addition, the system can send one or more sub-documents to the computing device of the user of the system. Further, the system can store one or more sub-documents in an appropriate location within one or more data storages. The one or more data storages can be local or available on one or more cloud computing systems.
[0056] In this specification, the term "configured to" is used in relation to system and computer program components. For one or more computer systems to be configured to perform a particular operation or action means that the system has installed in it software, firmware, hardware, or a combination thereof that, during operation, causes the system to perform the operation or action. For one or more computer programs to be configured to perform a particular operation or action means that the one or more programs include instructions that, when executed by a data processing apparatus, cause the apparatus to perform the operation or action.
[0057] Embodiments of the subject matter and the functional operations described in this specification can be implemented in digital electronic circuitry, in tangibly embodied computer software or firmware, in computer hardware including the structures disclosed in this specification and their structural equivalents, or in one or more of them in combination. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., as one or more modules of computer program instructions encoded on a tangible non-transitory storage medium for execution by, or to control the operation of, a data processing apparatus. A computer storage medium can be, or can include, a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them. Alternatively, or in addition, the program instructions can be encoded on an artificially generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to a receiver device suitable for execution by a data processing apparatus.
[0058] The term "data processing apparatus" refers to data processing hardware and includes, by way of example, all kinds of apparatus, devices, and machines for processing data, including programmable processors, computers, or multiple processors or computers. The apparatus can also be, or further include, dedicated logic circuitry, such as a FPGA (Field Programmable Gate Array) or ASIC (Application Specific Integrated Circuit). Optionally, in addition to the hardware, the apparatus can include code for creating an execution environment for a computer program, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them.
[0059] A computer program, which may also be referred to as or described as a program, software, software application, app, module, software module, script, or code, can be written in any form of programming language, including compiled or interpreted languages, or declarative or procedural languages, and can be deployed in any form, such as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. The program can correspond to a file in a file system, but it does not have to. The program can be stored in a part of a file that holds other programs or data, such as one or more scripts stored in a markup language document, in a single file dedicated to the program in question, or in multiple cooperating files, such as files that store one or more modules, subprograms, or portions of code. The computer program can be deployed to be executed on one computer or on multiple computers located at one site or distributed across multiple sites and interconnected by a data communication network.
[0060] As used herein, the term "database" is used broadly to refer to any collection of data, which need not be structured in any particular way, or at all, and can be stored on a storage device in one or more locations. Thus, for example, an index database can include multiple collections of data, each of which can be differently organized and accessed.
[0061] Similarly, as used herein, the term "engine" is used broadly to refer to a software-based system, subsystem, or process programmed to perform one or more specific functions. In general, an engine is implemented as one or more software modules or components installed on one or more computers in one or more locations. In some cases, one or more computers are dedicated to a particular engine, and in other cases, multiple engines can be installed and executed on the same one or more computers.
[0062] The processes and logical flows described herein can be performed by one or more programmable computers executing one or more computer programs to operate on input data and generate output in order to perform functions. The processes and logical flows can also be performed by, for example, FPGAs or ASICs, which are dedicated logic circuits, or by a combination of dedicated logic circuits and one or more programmed computers.
[0063] A computer suitable for the execution of a computer program may be based on a general-purpose or special-purpose microprocessor or both, or any other kind of central processing unit. Generally, the central processing unit receives instructions and data from read-only memory or random access memory or both. Essential elements of a computer are a central processing unit for carrying out or executing instructions, and one or more memory devices for storing instructions and data. The central processing unit and the memory can be supplemented by, or incorporated in, dedicated logic circuitry. Generally, a computer also includes, or is operably coupled to receive data from, or transfer data to, or both, one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks. However, a computer need not have such devices. Further, a computer may be incorporated in another device, such as, by way of example only, a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device, such as a universal serial bus (USB) flash drive.
[0064] Computer-readable media suitable for storing computer program instructions and data include, by way of example, all forms of non-volatile memory, media, and memory devices, including semiconductor memory devices, such as EPROM, EEPROM, and flash memory devices, magnetic disks, such as internal hard disks or removable disks, magneto-optical disks, and CD ROM and DVD-ROM disks.
[0065] To provide interaction with a user, embodiments of the subject matter described herein may be implemented on a computer having a display device, such as a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, for displaying information to the user, and a keyboard and a pointing device, such as a mouse or trackball, by which the user can provide input to the computer. Other types of devices may be used to provide interaction with the user, for example, feedback provided to the user may be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback, and input received from the user may be in any form including acoustic, voice, or tactile input. Additionally, the computer can interact with the user by sending documents to and receiving documents from the devices used by the user, for example, by sending a web page to a web browser on the user's device in response to a request received from the web browser, or by sending a text message or other form of message to a personal device, such as a smartphone running a messaging application, and receiving a response message from the user as a reply.
[0066] Embodiments of the subject matter described in this specification can be implemented in a computing system that includes, for example, backend components as a data server, or includes middleware components, such as an application server, or includes frontend components, such as a graphical user interface, a web browser, or a client computer having an app with which a user can interact with an implementation of the subject matter described in this specification, or any combination of one or more such backend, middleware, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication, such as a communication network. Examples of communication networks include local area networks (LANs) and wide area networks (WANs), such as the Internet.
[0067] A computing system can include clients and servers. Clients and servers are generally remote from each other and typically interact over a communication network. The relationship of client and server arises by virtue of computer programs running on respective computers and having a client-server relationship to each other. In some embodiments, a server can send data, such as an HTML page, to a user device for the purpose of displaying data to a user interacting with a device acting as a client and receiving user input from the user. Data generated at the user device, such as the results of user interaction, can be received at the server from the device.
[0068] This specification includes details of many specific implementations, which should not be construed as limitations on the scope of any invention or what may be claimed, but rather as descriptions of features that may be specific to particular embodiments of a particular invention. The specific features described herein in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, the various features described in the context of a single embodiment may also be implemented separately, or in any suitable sub-combination, in multiple embodiments. Furthermore, features are described above as acting in some combinations and may even be initially claimed as such, but one or more features from a claimed combination may, in some cases, be deleted from that combination, and the claimed combination may be directed to a sub-combination or a variant of a sub-combination.
[0069] Similarly, operations are shown in the drawings and described in the claims in a particular order, but this should not be understood as requiring that such operations be performed in the particular order shown, or in a sequential order, or that all of the shown operations be performed, in order to achieve a desired result. Multitasking and parallel processing may be advantageous in certain circumstances. Further, the separation of the various system modules and components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems may generally be integrated together in a single software product or packaged into multiple software products.
[0070] Specific embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. For example, the actions recited in the claims can be performed in a different order and still achieve desirable results. As one example, the processes illustrated in the accompanying figures do not necessarily require the particular order shown, or sequential order, to achieve the desired results. In some cases, multitasking and parallel processing may be advantageous.
Claims
1. 1. A computer-implemented method for automatically segmenting and classifying an input document into one or more sub-documents using a machine learning system, the input document including a plurality of pages, the machine learning system including a visual segmentation neural network, an optical character recognition subsystem, a title classifier, a document classifier, and a grouper subsystem, the method comprising: receiving a visual input representing the plurality of pages of the input document; classifying each page of the input document into a respective one of a plurality of templates using the visual segmentation neural network; For each page of the input document: generating a set of text lines from the text content of the page using the optical character recognition subsystem; using the title classifier to process each line of text in the set of lines of text to generate a corresponding confidence score representing the probability that the line of text contains a title of the page; selecting a text line having a highest confidence score based on the confidence scores of the set of text lines; determining whether the highest confidence score exceeds a threshold; in response to determining that the highest confidence score exceeds the threshold, processing the selected lines of text using the document classifier to generate a respective document score for each of a plurality of document types, each respective document score representing a probability that the page belongs to that document type; performing operations including selecting the document type having the highest document score as the final document type to which the page belongs; using a grouping subsystem to group the pages of the input document into one or more sub-documents based on (i) the respective template of each page and (ii) the respective final document type to which each page belongs.
2. The machine learning system further includes a discrimination extraction neural network; responsive to determining that, for each page of the input document, the highest confidence score does not exceed the threshold, processing each line of text in the set of lines of text using the discriminative extraction neural network to identify an identification number for the page; The method of claim 1 , wherein using the grouper subsystem to group the multiple pages of the input document into the one or more sub-documents is further based on the identification numbers of the pages.
3. responsive to determining that for each page of the input document, the highest confidence score does not exceed the threshold, processing each line of text in the set of lines of text using the discriminative extraction neural network to identify the identification number of the page; extracting key information within each line of text using said discriminative extraction neural network; and identifying the identification number of the page from the extracted key information using the identification extraction neural network.
4. Extracting key information from each line of text The method of claim 3 , comprising extracting key information using spatial dual-modality graph reasoning (SDMGR) or graph methods.
5. The method of claim 1 , wherein the visual input is a PDF file.
6. The method of claim 5 , wherein each page of the PDF file is a scanned image.
7. The method of claim 1 , wherein different templates of the plurality of templates have different visual appearances.
8. 2. The method of claim 1 , wherein each of the plurality of templates is used to generate a different document type, and each template has one or more visual features, the one or more visual features including a logo, padding, background, header, footer, chart, or font style.
9. The method of claim 1 , wherein the plurality of document types includes an invoice, a delivery note, a purchase order, an insurance policy, and an email.
10. classifying each page of the input document into a respective one of the plurality of templates using the visual segmentation neural network, vertically combining the page with a previous page in the document to generate a combined two-page input for the visual partitioning neural network, the visual partitioning neural network including one or more trained convolutional neural network layers; processing the combined two-page input using the one or more trained convolutional neural network layers to determine whether the page and the previous page have the same template; 2. The method of claim 1, comprising: in response to determining that the page and the previous page have the same template, associating the page with a known template of the previous page; or in response to determining that the page and the previous page have different templates, associating the current page with a new template.
11. processing the combined two-page input using the visual segmentation neural network to determine whether the page and the previous page have the same template; 11. The method of claim 10, comprising using the one or more trained convolutional neural network layers to extract one or more visual features from the combined two-page input, the one or more visual features comprising a logo, a padding, a background, or a font style.
12. The method of claim 1 , wherein the title classifier is a random forest classifier.
13. processing each line of text in the set of lines of text using the title classifier to generate the corresponding confidence score; extracting features for each line of text and providing the extracted features as inputs to the random forest classifier; and processing the extracted features using the random forest classifier to generate a corresponding confidence score for each line of text using a random forest classification method.
14. The method of claim 13 , wherein the features of each text line include an x-coordinate, a y-coordinate, a width w, a height h, and a matching score of the text line.
15. responsive to determining that the highest confidence score exceeds the threshold for each page of the input document, processing the selected lines of text using the document classifier to generate a respective document score for each of a plurality of document types; (i) maintaining a set of positive keywords representing document types that are likely to match the final document type of the page, and (ii) a set of negative keywords representing document types that are unlikely to match the final document type of the page; calculating a respective document score for each keyword in the set of positive keywords and the set of negative keywords; 2. The method of claim 1, wherein selecting the document type having the highest document score as the final document type to which the page belongs comprises selecting the keyword having the highest document score, the keyword specifying the final document type to which the page belongs.
16. 1. A system comprising: one or more computers; and one or more storage devices storing instructions that, when executed by the one or more computers, cause the one or more computers to perform operations for automatically segmenting and classifying an input document into one or more sub-documents using a machine learning system, the input document including a plurality of pages, the machine learning system including a visual segmentation neural network, an optical character recognition subsystem, a title classifier, a document classifier, and a grouper subsystem, the operations comprising: receiving a visual input representing the plurality of pages of the input document; classifying each page of the input document into a respective one of a plurality of templates using the visual segmentation neural network; For each page of the input document: generating a set of text lines from the text content of the page using the optical character recognition subsystem; using the title classifier to process each line of text in the set of lines of text to generate a corresponding confidence score representing the probability that the line of text contains a title of the page; selecting a text line having a highest confidence score based on the confidence scores of the set of text lines; determining whether the highest confidence score exceeds a threshold; in response to determining that the highest confidence score exceeds the threshold, processing the selected lines of text using the document classifier to generate a respective document score for each of a plurality of document types, each respective document score representing a probability that the page belongs to that document type; performing operations including selecting the document type having the highest document score as the final document type to which the page belongs; using a grouper subsystem to group the multiple pages of the input document into one or more sub-documents based on (i) the respective template of each page and (ii) the respective final document type to which each page belongs.
17. The operation includes:
17. The system of claim 16, further comprising: in response to determining that for each page of the input document, the highest confidence score does not exceed the threshold, processing each line of text in the set of lines of text to identify an identification number of the page; and wherein the operation of grouping the plurality of pages of the input document into the one or more sub-documents using the grouper subsystem is further based on the identification number of the page.
18. 17. The system of claim 16, wherein the visual input is a PDF file and each page of the PDF file is a scanned image.
19. One or more non-transitory computer storage media storing instructions that, when executed by one or more computers, cause the one or more computers to perform operations for automatically segmenting and classifying an input document into one or more sub-documents using a machine learning system, the input document including a plurality of pages, the machine learning system including a visual segmentation neural network, an optical character recognition subsystem, a title classifier, a document classifier, and a grouper subsystem, the operations including: receiving a visual input representing the plurality of pages of the input document; classifying each page of the input document into a respective one of a plurality of templates using the visual segmentation neural network; For each page of the input document: generating a set of text lines from the text content of the page using the optical character recognition subsystem; using the title classifier to process each line of text in the set of lines of text to generate a corresponding confidence score representing the probability that the line of text contains a title of the page; selecting a text line having a highest confidence score based on the confidence scores of the set of text lines; determining whether the highest confidence score exceeds a threshold; in response to determining that the highest confidence score exceeds the threshold, processing the selected lines of text using the document classifier to generate a respective document score for each of a plurality of document types, each respective document score representing a probability that the page belongs to that document type; performing operations including selecting the document type having the highest document score as the final document type to which the page belongs; using a grouper subsystem to group the multiple pages of the input document into one or more sub-documents based on (i) the respective template of each page and (ii) the respective final document type to which each page belongs.
20. The operation includes:
20. The one or more non-transitory computer storage media of claim 19, further comprising: in response to determining that, for each page of the input document, the highest confidence score does not exceed the threshold, processing each line of text in the set of lines of text to identify an identification number of the page; and wherein the operation of grouping the plurality of pages of the input document into the one or more sub-documents using the grouper subsystem is further based on the identification number of the page.
Citation Information
Patent Citations
Template-based document extraction
EP3955130A1
Document automated dividing device
JP2002312385A
Document processor and document processing method
JP2009145963A
System and method for automated file reporting
US20220237230A1