Information processing system, manuscript type identification method, model generation method, and program
The system uses character recognition and machine learning to identify document types by analyzing frequently occurring word sequences and their positions, effectively classifying documents with varied layouts.
Patent Information
- Application Number
- JP2023553861
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-10-14
- Publication Date
- 2025-07-09
- Estimated Expiration
- 2041-10-14
AI Technical Summary
Conventional methods struggle to accurately identify the type of a document with undetermined layouts, such as semi-standard forms, due to variations in word positions and layouts even for the same document type.
A system that utilizes character recognition, machine learning, and feature generation to identify document types by detecting frequently occurring word sequences and their positional relationships, generating feature amounts, and using a learned model to determine the document type.
Enables accurate identification of document types with diverse layouts by leveraging positional relationship feature amounts, improving identification accuracy and enabling automatic, complex, and precise classification.
Smart Images

Figure 0007705468000001 
Figure 0007705468000002 
Figure 0007705468000003
Abstract
Description
Technical Field
[0001] The present disclosure relates to a technique for identifying the type of a document.
Background Art
[0002] Conventionally, there has been proposed an apparatus that includes a scanner for reading a document image, and a document type registration and document type determination circuit that classifies color information such as RGB signals of the read document for each pre-divided color space to extract feature amounts of the image, and compares the extracted feature amounts with the feature amounts stored in advance to determine the type of the read document, and switches the image processing content based on the determination result of the document type registration and document type determination circuit (see Patent Document 1).
[0003] Also, there has been proposed an image reading apparatus that acquires image information of an image formed on a document, executes a first recognition process for classification from the feature amounts of the image, executes a second recognition process for classification from the character information of the image, and classifies the image using either one of the recognition processes or both recognition processes based on the processing result of either one of the recognition processes (see Patent Document 2).
[0004] Also, there has been proposed a document classification apparatus that is a model for classifying documents, and generates a document classification model by machine learning that outputs identification information for identifying the classification result based on the input document, acquires learning data including the document and the identification information associated with the document, extracts as feature amounts a word included in the document and character information that is a string composed of one character or a plurality of consecutive characters among the characters constituting the word and that is information that can be extracted one or more times from the word, performs machine learning based on the feature amounts extracted from the document and the identification information associated with the document, and generates a document classification model (see Patent Document 3).
[0005] Furthermore, there has been proposed a document classification apparatus that acquires image data representing an image of a document, obtains layout information representing the layout of components constituting each page of the document by analyzing the image represented by the image data, extracts text areas where the text is spatially continuous within a page, recognizes the character strings included in the text areas, extracts the character strings visually emphasized from the recognized character strings, uses the extracted character strings as keywords, generates structure data representing the hierarchical structure in the layout of the text areas for each page, extracts the logical structure of the document using the structure data and the keywords, and classifies and stores the document using the extracted logical structure (see Patent Document 4).
Prior Art Documents
Patent Documents
[0006]
Patent Document 1
Patent Document 2
Patent Document 3
Patent Document 4
Summary of the Invention
Problems to be Solved by the Invention
[0007] Conventionally, as techniques for identifying the type of manuscript, various techniques have been proposed, such as a method using ruled line information and a method of identifying a specific manuscript type based on the presence or absence of specific words described only in a specific manuscript type and their positions.
[0008] However, in the case of documents with various layouts (formats) even for the same type of document such as semi-standard forms, depending on the manuscript, the words written, the ruled lines, the positions of the words, etc. are different. Therefore, with the conventional methods described above, it is difficult to identify the type of manuscript for a manuscript of a document with an undetermined layout.
[0009] In view of the above problems, an object of the present disclosure is to appropriately identify the type of a manuscript even if the layout of the manuscript is not determined.
Means for Solving the Problems
[0010] An example of the present disclosure includes a recognition result acquisition means for acquiring a character recognition result for an identification target image that is an image of a manuscript to be identified, a frequently occurring word storage means for storing a frequently occurring word sequence of a predetermined manuscript type, a detection means for detecting the frequently occurring word sequence from the character recognition result of the identification target image to acquire information regarding the position of the frequently occurring word sequence in the manuscript to be identified, a feature generation means for generating a feature amount related to the manuscript to be identified including a positional relationship feature amount regarding the positional relationship between the frequently occurring word sequence and other word sequences in the manuscript to be identified using the information regarding the position, a model storage means for storing a learned model for identifying the predetermined manuscript type, which is generated by machine learning so that information indicating the validity that the manuscript is a manuscript of the predetermined manuscript type is output when a feature amount related to the manuscript including a positional relationship feature amount regarding the positional relationship between the frequently occurring word sequence and other word sequences in the manuscript is input, and an identification means for identifying whether or not the manuscript to be identified is a manuscript of the predetermined manuscript type by inputting the feature amount related to the manuscript to be identified into the learned model.
[0011] The present disclosure can be understood as an information processing apparatus, a system, a method executed by a computer, or a program to be executed by a computer. Further, the present disclosure can also be understood as a recording medium in which such a program is recorded and can be read by a computer or other devices, machines, etc. Here, a recording medium readable by a computer or the like refers to a recording medium that accumulates information such as data and programs by an electrical, magnetic, optical, mechanical, or chemical action and can be read by a computer or the like.
Advantages of the Invention
[0012] According to the present disclosure, even for a manuscript of a document with an undetermined layout, it is possible to appropriately identify the type of the manuscript.
Brief Description of Drawings
[0013]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21
Mode for Carrying Out the Invention
[0014] Hereinafter, embodiments of an information processing system, method, and program according to the present disclosure will be described with reference to the drawings. However, the embodiments described below are merely illustrative of the embodiments, and do not limit the information processing system, method, and program according to the present disclosure to the specific configurations described below. In practice, specific configurations according to the implementation mode may be appropriately adopted, and various improvements and modifications may be made.
[0015] In the present embodiment, embodiments in the case where the information processing system, method, and program according to the present disclosure are implemented in a system for identifying INVOICE (INVOICE manuscript) will be described. However, the information processing system, method, and program according to the present disclosure can be widely used for technologies for identifying any manuscript type (manuscript type), and the application target of the present disclosure is not limited to the examples shown in the embodiments.
[0016] [First Embodiment] [Configuration of the System] FIG. 1 is a schematic diagram showing the configuration of an information processing system 9 according to the present embodiment. The information processing system 9 according to the present embodiment includes one or a plurality of information processing devices 1, a learning device 2, and document reading devices 3 (3A, 3B) that can communicate with each other by being connected to a network. In the learning device 2, learning processing for identifying a predetermined document type (hereinafter, the document type is referred to as a "document kind") is performed, and a learned model for identifying a predetermined document kind is generated. In the information processing device 1, the learned model generated in the learning device 2 is used to identify the document kind of the document to be identified (whether the document to be identified is a document of a predetermined document kind).
[0017] In the present embodiment, "INVOICE" is exemplified as a predetermined document kind, and learning processing and identification processing for identifying INVOICE (INVOICE document) are exemplified. However, the document kind to be identified (predetermined document kind) may be any document kind other than INVOICE, and may be, for example, an invoice, an irregular receipt, a notice, a certificate, or the like. Further, in the present embodiment, the document includes not only a paper medium document but also an electronic document (image).
[0018] The information processing device 1 is a computer including a CPU (Central Processing Unit) 11, a ROM (Read Only Memory) 12, a RAM (Random Access Memory) 13, a storage device 14 such as an EEPROM (Electrically Erasable and Programmable Read Only Memory) or an HDD (Hard Disk Drive), a communication unit 15 such as a NIC (Network Interface Card), an input device 16 such as a keyboard or a touch panel, and an output device 17 such as a display. However, regarding the specific hardware configuration of the information processing device 1, appropriate omission, replacement, or addition can be made according to the implementation mode. Further, the information processing device 1 is not limited to a device consisting of a single housing. The information processing device 1 may be realized by a plurality of devices using so-called cloud or distributed computing technology or the like.
[0019] The information processing apparatus 1 acquires and stores the learned model and the high-frequency word list generated by the learning apparatus 2 from the learning apparatus 2. Further, the information processing apparatus 1 acquires a document image (image to be identified), which is an image of the manuscript to be identified, from the document reading apparatus 3A. Then, the information processing apparatus 1 identifies the type of the manuscript (the manuscript shown in the image to be identified) by using the learned model and the high-frequency word list.
[0020] Note that the document image is not limited to electronic data (image data) such as TIFF (Tagged Image File Format), JPEG (Joint Photographic Experts Group), and PNG (Portable Network Graphics), and may be electronic data in PDF (Portable Document Format). Therefore, the document image may be electronic data (PDF file) obtained by scanning a manuscript and converting it into PDF, or electronic data (electronic manuscript) created as a PDF file from the beginning.
[0021] Note that the method of acquiring the image to be identified is not limited to the above-described example, and any method may be used, such as a method of acquiring through another device, or a method of reading from an external recording medium or storage device 14 such as a USB (Universal Serial Bus) memory, an SD memory card (Secure Digital memory card), and an optical disk. When the image to be identified is not acquired from the document reading apparatus 3A, the information processing system 9 may not include the document reading apparatus 3A. Similarly, the method of acquiring the learned model and the high-frequency word list is not limited to the above-described example, and any method may be used.
[0022] The learning device 2 is a computer including a CPU 21, a ROM 22, a RAM 23, a storage device 24, a communication unit 25, and the like. However, with respect to the specific hardware configuration of the learning device 2, appropriate omission, replacement, or addition can be made according to the implementation mode. Further, the learning device 2 is not limited to a device consisting of a single housing. The learning device 2 may be realized by a plurality of devices using so-called cloud or distributed computing technologies and the like.
[0023] The learning device 2 acquires a document image (learning image) from the document reading device 3B. Then, the learning device 2 performs a learning process using the learning image to generate a learned model and a high-frequency word list for identifying a predetermined document type (a document of a predetermined document type).
[0024] Note that the method of acquiring the learning image is not limited to the example described above, and any method such as a method of acquiring through another device, a method of acquiring by reading from an external recording medium or the storage device 24, etc. may be used. Note that when the learning image is not acquired from the document reading device 3B, the information processing system 9 does not have to include the document reading device 3B. Further, in this embodiment, the information processing device 1 and the learning device 2 which are separate devices (separate housings) are exemplified, but it is not limited to this example, and the information processing device 9 may be configured to include a single device (housing) that performs both the learning process and the document type identification process.
[0025] The document reading device 3 (3A, 3B) is a device that receives a scan instruction or the like from a user and optically reads a paper document (manuscript) to obtain a document image (manuscript image), and is exemplified by a scanner, a multifunction peripheral, or the like. The document reading device 3A obtains an identification target image by reading a manuscript whose manuscript type the user wants to identify. The document reading device 3B obtains a plurality of learning images by reading manuscripts of a plurality of manuscript types including a predetermined manuscript type (for example, INVOICE). Note that the document reading device 3A and the document reading device 3B may be the same device (housing). Also, the document reading device 3 is not limited to a device having a function of transmitting an image to another device, and may be an imaging device such as a digital camera or a smartphone. Also, the document reading device 3 does not necessarily have to have an optical character recognition (OCR) function.
[0026] <Functional Configuration> FIG. 2 is a diagram showing an outline of the functional configuration of the learning device according to the present embodiment. In the learning device 2, the program recorded in the storage device 24 is read out to the RAM 23 and executed by the CPU 21, and each hardware provided in the learning device 2 is controlled, so that an image acquisition unit 51, a recognition result acquisition unit 52, a correct answer definition acquisition unit 53, a frequently used word acquisition unit 54, a detection unit 55, a feature generation unit 56, a model generation unit 57, and a storage unit 58 are provided. In the present embodiment and other embodiments described later, each function provided in the learning device 2 is executed by the CPU 21 which is a general-purpose processor, but a part or all of these functions may be executed by one or a plurality of dedicated processors. Also, each functional unit provided in the learning device 2 is not limited to being mounted on a device (one device) composed of a single housing, and may be mounted remotely and / or distributedly (for example, on the cloud).
[0027] The image acquisition unit 51 acquires a plurality of document images (learning images) used in the learning process. In the present embodiment, the image acquisition unit 51 acquires scan images of documents of a plurality of document types including a predetermined document type (INVOICE) as learning images. Note that the image acquisition unit 51 acquires images of documents of a predetermined document type (hereinafter referred to as "predetermined document type images") that are images of documents of a predetermined document type with different layouts from each other (a plurality of documents). For example, when a document of a plurality of document types including a predetermined document type is read by the document reading device 3B according to a user's scan instruction, the image acquisition unit 51 acquires the scan image that is the reading result as a learning image.
[0028] Note that the information in the document is included in the document image as an image. Also, the learning image and the identification target image described later are images that have been pre-processed (such as trimming processing to match the size of the document) so as to match the target document (the document shown in the image). Thus, the position within the document can be treated as equivalent to the position within the image. In the present embodiment, document images of document types other than the predetermined document type are used as incorrect learning data during learning, but the number of learning images for each of the predetermined document type and other document types is arbitrary.
[0029] The recognition result acquisition unit 52 acquires a character recognition result (string data) for each learning image. The recognition result acquisition unit 52 reads the entire learning image (entire area) using OCR to acquire a character recognition result (full text OCR result) for the learning image. Note that the data structure of the character recognition result may be arbitrary as long as it includes the character recognition result for each character string (character string image) in the learning image. Note that the method for acquiring the character recognition result is not limited to the example described above, and any method may be used, such as a method of acquiring through another device such as a character recognition device that performs OCR processing, or a method of acquiring by reading from an external recording medium or the storage device 24. In the present embodiment, a character string is a sequence (a continuous series of characters) consisting of one or more characters, and the characters include hiragana, katakana, kanji, alphabet, numbers, symbols, and the like.
[0030] The correct definition acquisition unit 53 acquires a correct definition (correct definition table) in which a learning image (identification information of the learning image) and information indicating whether the document shown in the learning image is a document of a predetermined document type are associated with each other for each learning image. For example, in the correct definition, for a learning image that is an image of a predetermined document type (INVOICE), as information indicating that it is a predetermined document type, a document type name (INVOICE), a label "1", etc. are stored. Also, for a learning image used as incorrect data, as information indicating that it is not a predetermined document type, the document type name of the learning image, a label "0", etc. are stored. Note that the identification information of the learning image may be arbitrary as long as it is information indicating the learning image, such as a file name, number, symbol, etc. In the present embodiment, the correct definition acquisition unit 53 acquires the correct definition by inputting the correct definition generated (defined) by the user into the learning device 2.
[0031] Note that the data structure for storing information indicating whether a document is a document of a predetermined document type is not limited to a table format such as the CSV (comma - separated values) format, and may be any format. Also, the method of acquiring the correct definition is not limited to the example described above, and any method may be used, such as a method of acquiring through another device, a method of acquiring by reading from an external recording medium or the storage device 24.
[0032] The frequently - occurring word acquisition unit 54 acquires (extracts) one or more frequently - occurring word sequences (frequently - occurring word sequences of a predetermined document type) that are word sequences frequently occurring in a document (image) of a predetermined document type. In the present embodiment, in a plurality of learning images that are images of a predetermined document type, a word sequence that more commonly appears is extracted as the frequently - occurring word sequence. Thus, a word sequence that is a feature of a predetermined document type can be obtained. Note that a word sequence means a sequence (arrangement of words) consisting of one or more words, and includes word sequences consisting of multiple words and single words. Hereinafter, an image (learning image) of a document of a predetermined document type will be referred to as a "predetermined document type image". Hereinafter, a more specific method of extracting the frequently - occurring word sequence will be described.
[0033] The frequent word acquisition unit 54 extracts a frequently occurring word sequence (frequent word sequence) in a document (image) of a predetermined document type by performing frequency analysis on a plurality of predetermined document type images. In the present embodiment, frequency analysis is performed on each word sequence consisting of two consecutive words and each word included in the character recognition result of each predetermined document type image, and a predetermined number (N (N≧1)) of word sequences are extracted as frequent word sequences in descending order of frequency. The frequent word acquisition unit 54 generates a high-frequency word list storing the extracted frequent word sequences.
[0034] FIG. 3 is a diagram showing an example of the high-frequency word list according to the present embodiment. As shown in FIG. 3, in the high-frequency word list for a predetermined document type, frequent word sequences (word sequences 1 to M (M frequent word sequences)) of the predetermined document type and identification information of a learned model for identifying the predetermined document type are stored. The identification information of the learned model can be arbitrarily set such as a model name (Model1 etc.), number, symbol, etc., as long as it is information indicating the learned model. In this way, by storing the frequent word sequences of the document to be identified and the identification information of the corresponding learned model in the high-frequency word list, the frequent word sequences and the learned model may be associated with each other. In the present embodiment, since the case where there is one predetermined document type is exemplified, the identification information of the learned model may not be stored.
[0035] The high-frequency word list generated in this way is stored by the storage unit 58. In the frequency analysis, the degree of appearance (number of appearances etc.) of each word sequence included in each predetermined document type image may be acquired, or a word sequence with a high appearance frequency in a plurality of predetermined document type images may be acquired. Further, the method of extracting the frequent word sequences is not limited to the above-described example, and a predetermined threshold value for the frequency (number of appearances) may be set, and a word sequence whose frequency exceeds the threshold value may be extracted as a frequent word sequence. Further, as a method for acquiring the frequent word sequences (high-frequency word list), any method may be used, such as a method of acquiring via another device or a method of acquiring by reading from an external recording medium or the storage device 24, other than the above-described example.
[0036] The detection unit 55 performs detection processing on the frequently occurring word sequences (the frequently occurring word sequences stored in the high-frequency word list) extracted by the frequently occurring word acquisition unit 54 in each learning image. In the detection processing, the detection unit 55 acquires information regarding the position of the frequently occurring word sequence within the manuscript (learning image) (position information related to the frequently occurring word sequence) for each learning image. For example, the detection unit 55 detects the frequently occurring word sequences included in the character recognition result of the learning image among the frequently occurring word sequences stored in the high-frequency word list. Then, the detection unit 55 acquires information regarding the position of the detected frequently occurring word sequence within the learning image (manuscript) (position information related to the frequently occurring word sequence) from, for example, the character recognition result of the learning image. By executing these processes for each learning image, the detection unit 55 acquires information regarding the position of the frequently occurring word sequence within each manuscript (learning image).
[0037] The position information related to the frequently occurring word sequence is the position information of the frequently occurring word sequence and / or the position information of the line including the frequently occurring word sequence. In this embodiment, both pieces of position information will be used. Also, in this embodiment, position coordinates are used as the position information. Therefore, in this embodiment, as the position information related to the frequently occurring word sequence, the position coordinates of the frequently occurring word sequence and the position coordinates (line coordinates) of the line including the frequently occurring word sequence are used.
[0038] The position coordinates of the frequently occurring word sequence are, for example, the coordinates indicating the position of the circumscribed rectangle of the frequently occurring word sequence in the manuscript (learning image) (coordinates of each vertex of the circumscribed rectangle, etc.). Also, for example, the line coordinates are the coordinates indicating the position of the circumscribed rectangle of the line including the frequently occurring word sequence (the circumscribed rectangle surrounding all the characters included in the line) (coordinates of each vertex of the circumscribed rectangle, etc.). Note that the position information related to the frequently occurring word sequence is not limited to the above example, and any position information may be used as long as it can generate (calculate) the feature amounts described later. For example, the position information is not limited to the position coordinates, and may be, for example, a combination of the coordinates of one point of the circumscribed rectangle and the information indicating the size of the circumscribed rectangle. Also, the position coordinates are not limited to the coordinates of each vertex of the circumscribed rectangle, and may be the coordinates of two vertices located on the diagonal of the circumscribed rectangle, etc.
[0039] The feature generation unit 56 generates feature amounts related to the manuscript shown in each learning image. The feature generation unit 56 generates the feature amounts related to the manuscript shown in the learning image by using the position information related to the frequent word sequence acquired by the detection unit 55. Then, the feature generation unit 56 generates a feature array in which the feature amounts related to the manuscript shown in each learning image are aggregated in the form of an array. In the learning process described later, the feature amounts (feature array) related to the manuscript shown in each learning image are used as the feature amounts (inputs of the learned model) for identifying the manuscript type.
[0040] In this embodiment, the feature generation unit 56 calculates the feature amounts related to the manuscript shown in the learning image based on the information regarding the frequent word sequence. That is, as the feature amounts related to the manuscript shown in the learning image, the feature amounts related to the frequent word sequence are calculated. In this embodiment, by using four pieces of information (the position of the frequent word sequence, the distance between the frequent word sequences, the size of the frequent word sequence, and the size of the line including the frequent word sequence) as the information regarding the frequent word sequence, the feature amounts related to the manuscript shown in the learning image are generated. More specifically, the feature amounts related to the manuscript shown in the learning image are generated as feature amounts including a feature amount indicating the position of the frequent word sequence (hereinafter referred to as "position feature amount"), a feature amount indicating the distance between the frequent word sequences (hereinafter referred to as "distance feature amount"), a feature amount indicating the size of the frequent word sequence (hereinafter referred to as "size feature amount"), and a feature amount indicating the size of the line including the frequent word sequence (hereinafter referred to as "line feature amount").
[0041] Note that the position feature amount and the size feature amount are each an example of a feature amount indicating the attribute of the frequent word sequence (itself). Also, the distance feature amount and the line feature amount are each an example of a feature amount (hereinafter referred to as "positional relationship feature amount") regarding the positional relationship between the frequent word sequence and other word sequences in the manuscript (learning image). The feature amount indicating the size of the line including the frequent word sequence (line feature amount) is, in other words, a feature amount indicating the possibility that other word sequences are included in the same line as the frequent word sequence, and thus corresponds to a feature amount regarding the positional relationship between the frequent word sequence and other word sequences.
[0042] In addition, in the present embodiment, a case where the feature amount of the manuscript includes the four feature amounts described above is exemplified. However, the present invention is not limited to the above example, and it may include only one of the four feature amounts, or may include a combination of two or three feature amounts. Hereinafter, the above four pieces of information will be described.
[0043] <Position of frequently occurring word sequence> In manuscripts of the same manuscript type, frequently occurring word sequences (frequently occurring word sequences) are often described at similar positions, even if they are not described at exactly the same positions among the manuscripts of that manuscript type.
[0044] FIG. 4 is a diagram showing an example of an INVOICE manuscript according to the present embodiment. As shown in FIG. 4, in the case of an INVOICE manuscript, for example, "Invoice" representing the manuscript type tends to be described at the upper part of the manuscript, and "Amount" representing the amount tends to be described at the right part of the manuscript. That is, it can be said that each manuscript type has a tendency to be at the position where the frequently occurring word sequence of that manuscript type is described. Therefore, in the present embodiment, as a feature amount for identifying the manuscript type, a feature amount (position feature amount) indicating the position of the frequently occurring word sequence is used.
[0045] <Distance between frequently occurring word sequences> In manuscripts of the same manuscript type, the positions where frequently occurring word sequences (frequently occurring word sequences) appear may vary among manuscripts of that manuscript type, but the distances between frequently occurring word sequences are often approximately the same among manuscripts. For example, in the case of INVOICE manuscripts, the positions of "VAT." representing tax and "Total" representing the total amount may vary depending on the manuscript, but as shown in Figure 4, there is a tendency for them to be arranged side by side vertically. That is, it can be said that each manuscript type has a tendency in the distances between frequently occurring word sequences of that manuscript type. Therefore, in this embodiment, as a feature quantity for identifying the manuscript type, a feature quantity (distance feature quantity) indicating the distance between frequently occurring word sequences is used. In this way, even when the positions where frequently occurring word sequences are written vary depending on the manuscript, or when the frequently occurring word sequences of a given manuscript type are word sequences that are also used in manuscripts of manuscript types other than the given manuscript type, it becomes possible to identify the manuscript type by using the distance feature quantity. When using the distance feature quantity as a feature quantity related to the manuscript shown in the learning image, a plurality of frequently occurring word sequences of a given manuscript type are required.
[0046] <Size of frequently occurring word sequence> Among the word sequences written in manuscripts of each manuscript type, there are word sequences that are likely to be written in large characters like the title part and word sequences that are likely to be written in small characters like annotations. For example, in the case of INVOICE manuscripts, as shown in Figure 4, the word "Invoice" representing the manuscript type is likely to be written in large characters, and words such as "e-mail" and "Tel" are likely to be written in small characters. That is, in each manuscript type, it can be said that there is a tendency in the size of the frequently occurring word sequences of that manuscript type. Therefore, in this embodiment, as a feature quantity for identifying the manuscript type, a feature quantity (size feature quantity) indicating the size of the frequently occurring word sequences is used.
[0047] <Size of the line containing the frequently occurring word sequence> Among the word sequences described in the manuscripts of each manuscript type, there are word sequences that are likely to exist in short sentences. For example, in the case of INVOICE manuscripts, as shown in FIG. 4, the word "Invoice" tends to exist in short sentences such as "Invoice", "Invoice Date", "Invoice NO", etc., while it tends to be less likely to exist in long sentences. On the other hand, in manuscripts of manuscript types other than INVOICE, the word "Invoice" is contained in quite a few long sentences. Thus, there are differences in the usage methods of word sequences between the target manuscript type and other manuscript types. That is, it can be said that each manuscript type has a tendency regarding whether the frequently occurring word sequences of that manuscript type are included in short sentences. Therefore, in the present embodiment, as a feature quantity for identifying the manuscript type, a feature quantity (line feature quantity) indicating the size of the line containing the frequently occurring word sequence, which is a feature quantity regarding the possibility that the frequently occurring word sequence is included in a short (long) sentence, is used.
[0048] The feature generation unit 56 generates the four feature quantities described above for each learning image, and generates a feature array that aggregates (stores) the four feature quantities for all learning images. In the present embodiment, arrays in which the position feature quantity, distance feature quantity, size feature quantity, and line feature quantity are respectively stored are referred to as "information arrays". In the present embodiment, the feature array is formed in a form in which four information arrays are aggregated. Hereinafter, each information array and each feature quantity stored in the feature array will be described.
[0049] <Array A: Coordinate information array (position feature quantity)> FIG. 5 is a diagram for explaining the position feature quantity according to the present embodiment. FIG. 6 is a diagram showing an example of the coordinate information array according to the present embodiment. In FIG. 6, an information array (coordinate information array) storing a feature quantity (position feature quantity) indicating the position of the frequently occurring word sequence in the manuscript (learning image) shown in FIG. 5 is exemplified. The position feature quantity is calculated (generated) using the position coordinates (lower left coordinates of the frequently occurring word sequence) of the frequently occurring word sequence acquired by the detection unit 55.
[0050] As shown in FIG. 6, the coordinate information array (array A) stores position feature amounts for all frequently occurring word sequences (such as "invoice", "total", "amount", "payment", etc.). In the present embodiment, the coordinates of the frequently occurring word sequences on the manuscript (x coordinates, y coordinates) normalized to values from 0 to 1 obtained by dividing the coordinates of the frequently occurring word sequences on the manuscript by the size of the manuscript are calculated as the position feature amounts. For example, the normalized coordinate obtained by dividing the x coordinate of the frequently occurring word sequence by the length of the manuscript in the x-axis direction is acquired as the position feature amount in the x-axis direction. In the present embodiment, the lower left coordinates of the frequently occurring word sequence (the coordinates of the lower left vertex of the circumscribed rectangle of the frequently occurring word sequence (the dotted rectangle in FIG. 5) (the coordinates of the circled point in FIG. 5)) are used as the coordinates of the frequently occurring word sequence, but the present invention is not limited to this example, and any coordinates such as the upper, lower, left, or right coordinates of the frequently occurring word sequence or the center of gravity coordinates may be used.
[0051] Note that the frequently occurring word sequence "amount" in the coordinate information array in FIG. 6 is a word sequence not included in the INVOICE manuscript (learning image) shown in FIG. 5. For example, it is a word sequence determined as a frequently occurring word sequence because it appears frequently in other INVOICE manuscripts (learning images). Thus, the position feature amount of the frequently occurring word sequence not included in the target manuscript (learning image) is set in advance as a value (for example, 0) when the frequently occurring word sequence does not exist in the manuscript (see FIG. 6).
[0052] Note that the position feature amount is not limited to the above-described normalized coordinates, and may be the coordinates of the frequently occurring word sequence on the manuscript itself. Also, in the example of FIG. 6, the coordinates of the frequently occurring word sequence are acquired with the upper left vertex of the manuscript as the origin, but the present invention is not limited to this example, and any position such as the upper right vertex, lower right vertex, or lower left vertex of the manuscript may be used as the origin for acquisition.
[0053] <Array B: Word sequence distance information array (distance feature amount)> FIG. 7 is a diagram for explaining the distance feature amount according to the present embodiment. FIG. 8 is a diagram showing an example of a word sequence distance information array according to the present embodiment. In FIG. 8, an information array (word sequence distance information array) storing a feature amount (distance feature amount) indicating the distance between frequently occurring word sequences in the manuscript (learning image) shown in FIG. 7 is illustrated. Note that the distance feature amount is calculated (generated) using the position coordinates (lower left coordinates of the frequently occurring word sequence) of the frequently occurring word sequence acquired by the detection unit 55.
[0054] As shown in FIG. 8, the word sequence distance information array (array B) stores the distance feature amounts for all combinations (combinations of two word sequences) of frequently occurring word sequences (such as "invoice", "total", "amount", "payment", etc.). In the present embodiment, the distance between frequently occurring word sequences on the manuscript (in the x-axis direction and y-axis direction) divided by the size of the manuscript is calculated as the distance feature amount, which is a value normalized to a value from 0 to 1. For example, the normalized distance obtained by dividing the x-axis component (distance) of the distance between frequently occurring word sequences (the distance between the coordinates of the frequently occurring word sequences (the length of the double arrow in FIG. 7)) by the length of the manuscript in the x-axis direction is acquired as the distance between frequently occurring word sequences in the x-axis direction.
[0055] Note that the INVOICE manuscript (learning image) shown in FIG. 7 does not include the frequently occurring word sequence "amount". Thus, the feature amount (distance feature amount) indicating the distance to a frequently occurring word sequence not included in the manuscript (learning image) is set in advance as a preset value (for example, 1) when the frequently occurring word sequence does not exist in the manuscript (see FIG. 8). Also, the distance feature amount is not limited to the above-described normalized distance, and may be the distance itself between frequently occurring word sequences on the manuscript.
[0056] <Array C: Size information array (size feature amount)> FIG. 9 is a diagram for explaining the size feature amount according to the present embodiment. FIG. 10 is a diagram showing an example of the size information array according to the present embodiment. In FIG. 10, an information array (size information array) storing a feature amount (size feature amount) indicating the size of frequently appearing word sequences in the manuscript (learning image) shown in FIG. 9 is illustrated. Note that the size feature amount is calculated (generated) using the position coordinates (coordinates such as the top, bottom, left, and right of the frequently appearing word sequence) of the frequently appearing word sequence acquired by the detection unit 55.
[0057] As shown in FIG. 10, the size information array (array C) stores size feature amounts for all frequently appearing word sequences ("invoice", "total", "amount", "payment", etc.). In the present embodiment, the area of the circumscribed rectangle of the frequently appearing word sequence on the manuscript (the area of the shaded part in FIG. 9) is calculated as the size feature amount. Note that in the present embodiment, the area of the circumscribed rectangle is expressed in square millimeters, but the unit of the area of the circumscribed rectangle is not limited to this example.
[0058] Also, the frequently appearing word sequence "amount" in the size information array of FIG. 10 is a word sequence not included in the INVOICE manuscript (learning image) shown in FIG. 9. Thus, the size feature amount of a frequently appearing word sequence not included in the manuscript (learning image) is set in advance as a value (for example, 0) when the frequently appearing word sequence does not exist in the manuscript (see FIG. 10).
[0059] Note that the size feature amount is not limited to the area of the circumscribed rectangle of the frequently appearing word sequence on the manuscript as described above, and may be, for example, the size of the frequently appearing word sequence on the manuscript (the area of the circumscribed rectangle) normalized to a value from 0 to 1 obtained by dividing the size of the frequently appearing word sequence on the manuscript by the size of the manuscript.
[0060] <Array D: Row information array (row feature amount)> FIG. 11 is a diagram for explaining the line feature amount according to the present embodiment. FIG. 12 is a diagram showing an example of the market information array according to the present embodiment. In FIG. 12, an information array (market information array) storing a feature amount (line feature amount) indicating the size of a line including a frequently appearing word sequence in the manuscript (learning image) shown in FIG. 11 is exemplified. Note that the line feature amount is calculated (generated) using the position coordinates (line coordinates) of the line including the frequently appearing word sequence acquired by the detection unit 55.
[0061] As shown in FIG. 12, the market information array (array D) stores the line feature amounts for all frequently appearing word sequences (“invoice”, “total”, “amount”, “payment”, etc.). In the present embodiment, the length of the line including the frequently appearing word sequence on the manuscript (the length of the double arrow in FIG. 11) is normalized to a value from 0 to 1 obtained by dividing by the length of the manuscript in the same direction as the length direction of the line, and this normalized line length is calculated as the line feature amount.
[0062] Also, the frequently appearing word sequence “amount” in the market information array of FIG. 12 is a word sequence not included in the INVOICE manuscript (learning image) shown in FIG. 11. Thus, the line feature amount for a frequently appearing word sequence not included in the manuscript (learning image) is set in advance as a value (for example, 0) when the frequently appearing word sequence does not exist in the manuscript (see FIG. 12). Also, the line feature amount is not limited to the above-described normalized line length, and may be, for example, the length of the line including the frequently appearing word sequence on the manuscript itself, a value obtained by dividing the length of the line including the frequently appearing word sequence on the manuscript by the length of the frequently appearing word sequence (the magnification of the length with respect to the frequently appearing word sequence), the area of the line including the frequently appearing word sequence on the manuscript (the area of the circumscribed rectangle of the line), a value obtained by dividing the area of the line by the area of the manuscript (the magnification with respect to the size of the manuscript), etc.
[0063] <Feature array> FIG. 13 is a diagram showing an example of the feature array according to the present embodiment. As shown in FIG. 13, the feature array is formed in a form in which the above-described respective information arrays (array A, array B, array C, array D) are aggregated. The feature array stores the respective information arrays (array A, array B, array C, array D) generated for each manuscript (each learning image).
[0064] In addition, when a plurality of identical word sequences appear in a single manuscript (image), it may be possible to select which word sequence among the plurality of identical word sequences to use as a feature quantity. Any method may be used as a method for determining which word sequence to use. In the case of Array A, for example, among the plurality of identical word sequences, only one of the word sequence with the maximum y - coordinate and the word sequence with the minimum y - coordinate may be used, or both may be used. In the case of Array B, for example, the word sequence with the smallest distance between frequently - occurring word sequences may be used. In the case of Array C, for example, only one of the word sequence with the maximum size of frequently - occurring word sequences and the word sequence with the minimum size may be used, or both may be used. In the case of Array D, for example, the word sequence used in Array A may be used, or only one of the word sequence with the maximum row size and the word sequence with the minimum row size may be used.
[0065] The model generation unit 57 generates a learned model for identifying a predetermined manuscript type by performing machine learning (supervised learning). For machine learning, a feature quantity (feature array) related to the manuscript shown in the learning image and information (correct label) indicating whether the manuscript shown in the learning image is a manuscript of a predetermined manuscript type are used as learning data (a dataset (teacher data) of the feature quantity and information on whether it is a predetermined manuscript type) associated with each learning image. The information indicating whether the manuscript shown in the learning image, which is the correct label, is a manuscript of a predetermined manuscript type is information based on the correct definition acquired by the correct definition acquisition unit 53. By performing machine learning using this learning data, it becomes possible to learn the feature quantities of a predetermined manuscript type.
[0066] Accordingly, by inputting feature quantities related to the target manuscript (including at least positional relationship feature quantities indicating the positional relationship within the manuscript between frequently occurring word sequences and other word sequences), it is possible to generate a discriminator capable of determining whether the target manuscript is a manuscript of a predetermined manuscript type. More specifically, by inputting feature quantities related to the manuscript, it is possible to generate a discriminator (trained model) capable of outputting information indicating the validity that the manuscript is a manuscript of a predetermined manuscript type. Note that the information indicating the validity that the manuscript is a manuscript of a predetermined manuscript type is information (such as a label) indicating whether the manuscript is a manuscript of a predetermined manuscript type and / or information (such as a confidence level or probability) indicating the likelihood that the manuscript is a manuscript of a predetermined manuscript type. The generated trained model is stored by the storage unit 58.
[0067] Note that the machine learning method is arbitrary, and any method among decision trees, random forests, gradient boosting, linear regression, support vector machines (SVMs), neural networks, etc. may be used.
[0068] The storage unit 58 stores frequently occurring word sequences (high-frequency word lists) for a predetermined manuscript type extracted by the frequently occurring word acquisition unit 54 and trained models for a predetermined manuscript type generated by the model generation unit 57. The storage unit 58 may store the high-frequency word list (frequently occurring word sequence) and the trained model in association with each other.
[0069] FIG. 14 is a diagram showing an outline of the functional configuration of the information processing apparatus according to the present embodiment. In the information processing apparatus 1, a program recorded in the storage device 14 is read into the RAM 13 and executed by the CPU 11, and each hardware provided in the information processing apparatus 1 is controlled, whereby the information processing apparatus 1 functions as a device including an image acquisition unit 41, a recognition result acquisition unit 42, a frequent word storage unit 43, a model storage unit 44, a detection unit 45, a feature generation unit 46, and an identification unit 47. In the present embodiment and other embodiments described later, each function provided in the information processing apparatus 1 is executed by the CPU 11 which is a general-purpose processor, but a part or all of these functions may be executed by one or more dedicated processors. Further, each functional unit provided in the information processing apparatus 1 is not limited to being mounted on a device (device 1) composed of a single housing, and may be mounted remotely and / or dispersedly (for example, on the cloud).
[0070] The image acquisition unit 41 acquires a document image to be identified (an image of a document to be identified (hereinafter referred to as an "identification target image")) in the identification process of the document type. In the present embodiment, for example, when a document (document) to be identified is read by the document reading device 3A according to a user's scan instruction, the image acquisition unit 41 acquires the scan image which is the read result as the identification target image.
[0071] The recognition result acquisition unit 42 acquires a character recognition result (full text OCR result) for the identification target image. Since the processing in the recognition result acquisition unit 42 is substantially the same as the description of the processing in the recognition result acquisition unit 52, a detailed description thereof is omitted.
[0072] The frequent word storage unit 43 stores a high-frequency word list for identifying a predetermined document type, which is generated in the learning device 2. Since the details of the high-frequency word list have been described in the description of the functional configuration (frequent word extraction unit 54) of the learning device 2, the description thereof is omitted.
[0073] The model memory unit 44 stores a learned model for identifying a predetermined manuscript type, which is generated in the learning device 2. Since the details of the learned model have been described in the explanation of the functional configuration (model generation unit 57) of the learning device 2, the explanation is omitted.
[0074] The detection unit 45 performs a detection process of frequent word sequences (frequent word sequences stored in the high-frequency word list stored in the frequent word memory unit 43) in the identification target image. In the detection process, the detection unit 45 acquires information regarding the position of the frequent word sequence in the manuscript (manuscript to be identified) shown in the identification target image (position information related to the frequent word sequence). Since the process in the detection unit 45 is substantially the same as the explanation of the process in the detection unit 55, the detailed explanation is omitted.
[0075] The feature generation unit 46 generates a feature amount related to the manuscript (manuscript to be identified) shown in the identification target image. The feature generation unit 46 uses the position information related to the frequent word sequence acquired by the detection unit 45 to generate a feature amount related to the manuscript to be identified. Then, the feature generation unit 46 generates a feature array in which the feature amount related to the manuscript to be identified is formed in the form of an array. In the identification process described later, the feature amount (feature array) related to the manuscript to be identified is used as the feature amount (input of the learned model) for identifying the manuscript type. Similar to the feature amount related to the manuscript shown in the learning image described above, the feature amount related to the manuscript to be identified is generated as a feature amount including a position feature amount, a distance feature amount, a size feature amount, and a line feature amount.
[0076] Since the feature amount (feature array) related to the manuscript to be identified and its generation method are substantially the same as the feature amount (feature array) related to the manuscript shown in the learning image and its generation method described above, the detailed explanation is omitted. The order of each feature amount in the feature array related to the identification target image (the position of each feature amount in the array) is the same as the order of each feature amount in the feature array related to the learning image.
[0077] The identification unit 47 identifies whether the document to be identified is a document of a predetermined document type by inputting the feature amount (feature array) related to the document to be identified into the learned model. Specifically, the identification unit 47 receives the learned model for identifying a predetermined document type stored by the model storage unit 44, and inputs the feature amount (feature array) related to the document to be identified generated by the feature generation unit 46 into the learned model, thereby identifying whether the document is a document of a predetermined document type. The identification unit 47 outputs the identified result.
[0078] As described above, when the feature amount related to the document is input into the learned model, information (label and / or probability) indicating the validity that the document is a document of a predetermined document type is output from the learned model. In the present embodiment, the identification unit 47 inputs the feature amount related to the document to be identified into the learned model, thereby obtaining information (label (for example, label "1" in the case of a predetermined document type, and label "0" otherwise)) indicating whether the document image to be identified is a document of a predetermined document type and information (confidence level, probability, etc.) indicating the likelihood that the document to be identified is a document of a predetermined document type.
[0079] Note that, for example, when the probability that the document is a document of a predetermined document type exceeds the probability that it is not a document of a predetermined document type or exceeds a predetermined threshold value, etc., it can be determined that the document is a document of a predetermined document type. Therefore, the identification unit 47 may obtain only the probability that the document is a document of a predetermined document type from the learned model, and determine whether the document is a document of a predetermined document type based on the obtained probability.
[0080] <Flow of processing> Next, the flow of the learning process executed by the learning device 2 according to the present embodiment will be described. Note that the specific content and processing order of the processing described below are an example for implementing the present disclosure. The specific processing content and processing order may be appropriately selected according to the implementation mode of the present disclosure.
[0081] FIG. 15 is a flowchart showing an outline of the flow of the learning process according to the present embodiment. The processes shown in this flowchart are executed in the learning device 2 triggered by, for example, receiving a scan instruction for a document. Note that this flowchart may also be executed triggered by, for example, receiving an instruction from the user to acquire a document image stored in the storage device 24. In this flowchart, the processing in the case where the document type to be identified (predetermined document type) is "INVOICE" will be exemplified.
[0082] In step S101, a plurality of document images (learning images) are acquired. The image acquisition unit 51 acquires a learning image (scan image) including a plurality of predetermined document type images that are images of documents of a predetermined document type (INVOICE) with different layouts from each other. Then, the process proceeds to step S102.
[0083] In step S102, a correct answer definition is acquired. The correct answer definition acquisition unit 53 acquires a correct answer definition in which a learning image (identification information of the learning image) and information indicating whether the document shown in the learning image is a document of a predetermined document type (INVOICE) are associated with each other for each learning image. Then, the process proceeds to step S103.
[0084] In step S103, a character recognition result (full text OCR result) is acquired. The recognition result acquisition unit 52 performs character recognition on each learning image acquired in step S101 to acquire a character recognition result for each learning image. Note that steps S102 and S103 may be in any order. Also, steps S101 and S102 may be in any order. Then, the process proceeds to step S104.
[0085] In step S104, an extraction process of frequently occurring word sequences is performed. In the extraction process of frequently occurring word sequences, frequently occurring word sequences of a predetermined document type are extracted using the character recognition results of a plurality of learning images (predetermined document type images) that are images of a predetermined document type (INVOICE). Details of the extraction process of frequently occurring word sequences will be described later with reference to FIG. 16. Then, the process proceeds to step S105.
[0086] In step S105, a frequent word sequence detection process is performed. In the frequent word sequence detection process, in the learning image acquired in step S101, a detection process of the frequent word sequence extracted in step S104 is performed. In the frequent word sequence detection process, position information related to the frequent word sequence (position information of the frequent word sequence in the manuscript (learning image) and position information of the line including the frequent word sequence in the manuscript (learning image)) is acquired. Details of the frequent word sequence detection process will be described later with reference to FIG. 17. Thereafter, the process proceeds to step S106.
[0087] In step S106, a feature quantity generation process is performed. In the feature quantity generation process, based on the position information acquired in step S105, a feature quantity (feature array) related to the manuscript shown in the learning image acquired in step S101 is generated. Details of the feature quantity generation process will be described later with reference to FIG. 18. Thereafter, the process proceeds to step S107.
[0088] In step S107, it is determined whether or not feature quantities have been generated for all the learning images (whether the processes of step S105 and step S106 have been executed). The CPU 21 determines whether or not a feature quantity related to the manuscript shown in the learning image has been generated for each of all the learning images. If the processing has not been completed for all the learning images (NO in step S107), the process returns to step S105, and the processing for the learning images for which the processing has not been completed is executed. On the other hand, if the processing has been completed for all the learning images (YES in step S107), the process proceeds to step S108.
[0089] In step S108, a learned model for identifying a predetermined document type is generated. The model generation unit 57 performs machine learning using the learning data in which the feature amounts (feature arrays) related to the documents shown in each learning image generated in step S107 are associated with the information indicating whether the document shown in each learning image is a document of a predetermined document type (INVOICE) (information based on the correct definition acquired in step S102), thereby generating a learned model for identifying a predetermined document type (INVOICE). Then, the processing shown in this flowchart ends.
[0090] FIG. 16 is a flowchart showing an outline of the flow of the frequently-occurring word sequence extraction process according to the present embodiment. The processing shown in this flowchart is executed starting from the completion of the processing in step S103 in FIG. 15. Note that, also in this flowchart, the processing in the case where the predetermined document type is "INVOICE" is exemplified.
[0091] In step S1041, frequency analysis of words (individual words) in a plurality of predetermined document type images is performed. For example, the frequently-occurring word acquisition unit 54 uses the character recognition results of the plurality of predetermined document type images acquired in step S103 to acquire (totalize) the number of occurrences of each word included in each of the plurality of predetermined document type images in the plurality of predetermined document type images. Then, the processing proceeds to step S1042.
[0092] In step S1042, frequency analysis of word sequences consisting of two consecutive words in a plurality of predetermined document type images is performed. The frequently-occurring word acquisition unit 54 uses the character recognition results of the plurality of predetermined document type images acquired in step S103 to acquire (totalize) the number of occurrences of each word sequence (a word sequence consisting of two consecutive words) included in each of the plurality of predetermined document type images in the plurality of predetermined document type images. Then, the processing proceeds to step S1043.
[0093] In step S1043, a predetermined number (N) of word sequences are extracted as frequently occurring word sequences in descending order of frequency (number of occurrences). The frequently occurring word acquisition unit 54 extracts, as frequently occurring word sequences for a predetermined manuscript type (INVOICE), a predetermined number (N) of word sequences from among the word sequences (including words) included in each predetermined manuscript type image, in descending order of the number of occurrences, based on the results of the frequency analysis in steps S1041 and S1042. Thereafter, the process proceeds to step S1044.
[0094] In step S1044, a high-frequency word list is generated. The frequently occurring word acquisition unit 54 generates a high-frequency word list storing the frequently occurring word sequences extracted in step S1043. Then, the storage unit 58 stores the generated high-frequency word list. Thereafter, the process shown in this flowchart ends.
[0095] FIG. 17 is a flowchart showing an outline of the flow of the frequently occurring word sequence detection process according to the present embodiment. The process shown in this flowchart is executed when the process of step S104 in FIG. 15 ends.
[0096] In step S1051, a high-frequency word list is acquired. The detection unit 55 acquires the high-frequency word list stored in step S1044. Thereafter, the process proceeds to step S1052.
[0097] In step S1052, position information of the frequently occurring word sequences is acquired. The detection unit 55 detects, from among the frequently occurring word sequences stored in the high-frequency word list acquired in step S1051, the frequently occurring word sequences included in the character recognition result of the learning image, and for each detected frequently occurring word sequence, acquires information (coordinate information) on the position of the frequently occurring word sequence in the manuscript shown in the learning image. Thereafter, the process proceeds to step S1053.
[0098] In step S1053, position information of a line including a frequent word string is obtained. The detection unit 55 detects frequent word strings included in the character recognition result of the learning image from among the frequent word strings stored in the high-frequency word list obtained in step S1051, and obtains position information (coordinate information) of the line including the frequent word string in the document shown in the learning image for each detected frequent word string. After that, the process shown in this flowchart ends. Note that steps S1052 and S1053 can be performed in any order.
[0099] 18 is a flowchart showing an outline of the flow of feature generation processing according to this embodiment. The processing shown in this flowchart is executed when the processing in step S105 in FIG.
[0100] In step S1061, a feature indicating the position of a frequent word string is generated. The feature generation unit 56 generates a feature indicating the position of a frequent word string (the feature stored in array A in FIG. 6) using the position information acquired in step S1052. Then, the process proceeds to step S1062.
[0101] In step S1062, a feature indicating the distance between frequent word strings is generated. The feature generation unit 56 generates a feature indicating the distance between frequent word strings (the feature stored in array B in FIG. 8) using the position information acquired in step S1052. Then, the process proceeds to step S1063.
[0102] In step S1063, a feature indicating the size of the frequent word string is generated. The feature generation unit 56 generates a feature indicating the size of the frequent word string (the feature stored in array C in FIG. 10) using the position information acquired in step S1052. Then, the process proceeds to step S1064.
[0103] In step S1064, a feature amount indicating the size of a line including a frequently occurring word sequence is generated. The feature generation unit 56 generates a feature amount (the feature amount stored in the array D in FIG. 12) indicating the size of a line including a frequently occurring word sequence by using the position information acquired in step S1053. Note that steps S1061 to S1064 may be in any order. Thereafter, the process proceeds to step S1065.
[0104] In step S1065, the feature amounts are shaped into an array. The feature generation unit 56 generates a feature array (each row in FIG. 13) that aggregates the respective feature amounts generated in steps S1061 to S1064. Thereafter, the process shown in this flowchart ends. Note that by executing the process of step S106 for each learning image, the feature amounts related to each learning image (the feature amounts of the manuscript shown in the learning image) are stored in the feature array, and a feature array as shown in FIG. 13 is generated.
[0105] FIG. 19 is a flowchart showing an outline of the flow of the identification process according to the present embodiment. The process shown in this flowchart is executed in the information processing apparatus 1 triggered by, for example, receiving a scan instruction for a manuscript. Note that this flowchart may be executed triggered by, for example, receiving an instruction from a user to acquire a manuscript image stored in the storage device 14. In this flowchart as well, the process in the case where the manuscript type to be identified is "INVOICE" is exemplified.
[0106] In step S201, a document image (image to be identified) is acquired. The image acquisition unit 41 acquires a scan image of the manuscript to be identified. Thereafter, the process proceeds to step S202.
[0107] In step S202, a character recognition result (full text OCR result) is acquired. The recognition result acquisition unit 42 performs character recognition on the image to be identified acquired in step S201 to acquire a character recognition result for the image to be identified. Thereafter, the process proceeds to step S203.
[0108] In step S203, a frequent word sequence detection process is performed. In the frequent word sequence detection process, in the identification target image acquired in step S201, a detection process of the frequent word sequences stored in the frequent word storage unit 43 is performed. In the frequent word sequence detection process, position information related to the frequent word sequences (information on the position of the frequent word sequences in the original document to be identified and information on the position of the line containing the frequent word sequences in the original document to be identified) is acquired. Since the frequent word sequence detection process is substantially the same as the process shown in FIG. 17, a detailed description thereof is omitted. Thereafter, the process proceeds to step S204.
[0109] In step S204, a feature quantity generation process is performed. In the feature quantity generation process, based on the position information acquired in step S203, a feature quantity (feature array) related to the original document (the original document to be identified) shown in the identification target image acquired in step S201 is generated. Since the details of the feature quantity generation process are substantially the same as the process shown in FIG. 18, a detailed description thereof is omitted. Thereafter, the process proceeds to step S205.
[0110] In step S205, the document type of the original document to be identified is identified. The identification unit 47 receives a learned model for identifying a predetermined document type (INVOICE) stored in the model storage unit 44, and the identification unit 47 inputs the feature quantity (feature array) related to the original document to be identified generated in step S204 into the received learned model, thereby identifying whether the original document to be identified is a document of a predetermined document type (INVOICE). The identification unit 47 outputs the identified result. Thereafter, the process shown in this flowchart ends.
[0111] As described above, according to the present embodiment, the learning device 2 can generate a learned model capable of identifying whether a document is a document of a predetermined document type from the feature amounts related to the document (including the positional relationship feature amounts related to the positional relationship within the document between the frequently occurring word sequences of a predetermined document type and other word sequences). Therefore, it is possible to generate a model (identifier) that can appropriately identify the type of a document (semi-standard form, etc.) whose layout is not determined (whose layout is diverse). Further, according to the present embodiment, the information processing device 1 can identify whether the document to be identified is a document of a predetermined document type by using a learned model capable of identifying whether a document is a document of a predetermined document type from the feature amounts related to the document. Therefore, it is possible to appropriately identify the type of a document whose layout is not determined. That is, even for documents with different layouts, it is possible to identify them as documents of the same document type.
[0112] Also, in the case of a document whose layout is not determined, the positions of the frequently occurring word sequences differ depending on the document. However, according to the present embodiment, as the feature amounts related to the document, positional relationship feature amounts (distance feature amounts and line feature amounts) related to the positional relationship within the document between the frequently occurring word sequences of a predetermined document type and other word sequences are used. Therefore, it is possible to improve the identification accuracy as compared with the case of using only the feature amounts indicating the positions of the frequently occurring word sequences.
[0113] Conventionally, although there has been a demand to identify INVOICE documents, there are various layouts for INVOICE documents, there are no specific words that are always described only in INVOICE documents, and the positions of frequently occurring words are not determined (vary depending on the document). Therefore, there is a problem that it is difficult to identify INVOICE documents with simple rules. Conventionally, document types such as receipts and business cards have been identified based on the size of the document. However, most INVOICE documents are basically of A4 size and do not have characteristics in terms of document size. Therefore, it is difficult to identify INVOICE documents by this method.
[0114] Conventionally, there is also a method of identifying a specific manuscript type based on the presence or absence of specific words described only in a specific manuscript type and their positions. However, there is no word that is always described only in INVOICE manuscripts, and words that frequently appear in INVOICE manuscripts also exist (appear) in other manuscript types. Even for the same item (information), different words may be used for description. Therefore, it is difficult to establish rules based on the presence or absence of specific words.
[0115] In addition, there is a method of identifying forms using ruled line information. However, regarding INVOICE manuscripts, since there are various layouts, the ruled lines also vary depending on the manuscript. Therefore, it is difficult to identify INVOICE manuscripts using this method.
[0116] However, according to the present embodiment, as a feature amount related to a manuscript, by using a positional relationship feature amount (distance feature amount or line feature amount) regarding the positional relationship within the manuscript between the frequently appearing word sequence of a predetermined manuscript type and other word sequences, it becomes possible to identify an INVOICE manuscript whose layout is not determined.
[0117] In addition, according to the present embodiment, in the learning device 2, since learning is performed by machine learning, it becomes possible to automatically generate an identifier (trained model). Also, by performing learning by machine learning, more complex and accurate identification becomes possible.
[0118] [Second Embodiment] In the first embodiment described above, the implementation mode in the case where there is one predetermined manuscript type (the manuscript type to be identified) (when identifying only one manuscript type) was explained. In this embodiment, the implementation mode in the case where there are a plurality of predetermined manuscript types (when identifying a plurality of manuscript types) will be explained. In this embodiment, an implementation mode of identifying a plurality of manuscript types by using a plurality of trained models for identifying only one manuscript type will be explained.
[0119] The configuration of the system according to this embodiment is substantially the same as that described in the first embodiment with reference to FIG. 1, so the description is omitted. Also, the functional configuration of the learning device according to this embodiment is substantially the same as that described in the first embodiment with reference to FIG. 2, so the description is omitted. However, different from the first embodiment, in the learning device 2, for each of a plurality of predetermined manuscript types, the above-described learning process (see FIG. 15) is performed, and a high-frequency word list and a learned model are generated for each of the plurality of predetermined manuscript types. Note that the high-frequency word list may be generated for each manuscript type, or may be a list in which the frequently occurring word sequences of each manuscript type are stored.
[0120] FIG. 20 is a diagram showing an example of the high-frequency word list according to this embodiment. As shown in FIG. 20, in the high-frequency word list, identification information of a predetermined manuscript type, frequently occurring word sequences (word sequences 1 to M (M frequently occurring word sequences)) of the predetermined manuscript type, and identification information (model name, etc.) of the learned model for identifying the predetermined manuscript type are stored in association with each other. Note that the identification information of the manuscript type may be arbitrarily selected, such as a manuscript type name (manuscript type 1, manuscript type 2, etc.), a number, a symbol, etc., as long as it indicates the manuscript type. In this way, the high-frequency word list may be a list in which the frequently occurring word sequences of each of a plurality of predetermined manuscript types are stored. Note that the number of frequently occurring word sequences does not have to be common (the same number) for all manuscript types.
[0121] Also, the functional configuration of the information processing device according to this embodiment is substantially the same as that described in the first embodiment with reference to FIG. 14, so the description is omitted. However, in this embodiment, different from the first embodiment, in the information processing device 1, for each of a plurality of predetermined manuscript types, it is identified whether the image to be identified is an image of a predetermined manuscript type. Therefore, each functional unit other than the image acquisition unit 41 performs processing for each of the plurality of predetermined manuscript types. Note that the identification unit 47 identifies the manuscript type of the manuscript to be identified based on the result of identifying whether the manuscript to be identified (the manuscript targeted by the image to be identified) corresponds to each of the plurality of predetermined manuscript types. Specifically, the manuscript type of the manuscript to be identified is identified by adopting one result from a plurality of identification results.
[0122] As a result of determining whether or not the document to be identified corresponds to each of a plurality of predetermined document types, if only one document type is determined to correspond, the identification unit 47 identifies (determines) that document type as the document type of the document to be identified. On the other hand, if there are a plurality of document types determined to correspond, the identification unit 47 selects one document type from these plurality of document types by the following method or the like, and identifies (determines) the selected document type as the document type of the document to be identified.
[0123] (Selection based on the output (probability, etc.) of the learned model) One document type may be selected based on the likelihood (probability, reliability, etc.) that the document to be identified, output by the learned model, is a document of a predetermined document type. For example, the document type with the highest such likelihood is determined (estimated) as the document type of the document to be identified.
[0124] (Selection based on the past identification degree) It may be selected based on the identification result (identification degree) for past images to be identified. For example, one document type may be selected based on the frequency (number of times) that past documents to be identified were identified as corresponding to documents of a predetermined document type. Specifically, the document type with the highest number of times of being identified (determined) as corresponding to a predetermined document type in past documents to be identified is determined (estimated) as the document type of the document to be identified. When determining the document type using this method, the information processing apparatus 1 is provided with a history information storage unit (not shown) to store past identification results.
[0125] (Selection based on the past identification time) It may be selected based on the identification time (the time when it was identified) for the past target image to be identified. For example, based on the time when a past target manuscript was identified as corresponding to a manuscript of a predetermined manuscript type, one manuscript type may be selected. Specifically, the manuscript type that is closest (the most recent) to the time when the past target manuscript was identified (determined) as corresponding to a predetermined manuscript type is determined (estimated) as the manuscript type of the target manuscript to be identified. When determining the manuscript type using this method, the information processing apparatus 1 is provided with a history information storage unit (not shown) to store the past identification time.
[0126] (Selection by user's selection) A plurality of manuscript types determined to be applicable are displayed, and one manuscript type may be selected by the user selecting one manuscript type from among the plurality of displayed manuscript types. When determining the manuscript type using this method, the information processing apparatus 1 is provided with a display unit (not shown) to display the determined applicable manuscript types, and is provided with an instruction reception unit (not shown) to receive a selection instruction from the user.
[0127] FIG. 21 is a flowchart showing an outline of the flow of the identification process according to the present embodiment. The process shown in this flowchart is executed in the information processing apparatus 1 triggered by, for example, receiving a scan instruction for a manuscript (document). Note that this flowchart may also be executed triggered by receiving an instruction from the user to acquire a form image stored in the storage device 14. In this flowchart, the case where there are two types of manuscript types to be identified (manuscript type 1 and manuscript type 2) is exemplified, but even when there are three or more types of manuscript types to be identified, the same process as this flowchart is performed to enable identification of the manuscript type.
[0128] In step S301, a document image (image to be identified) is acquired. The image acquisition unit 41 acquires a scanned image of the original document to be identified. Thereafter, the process proceeds to steps S302 and S306. Thereafter, the processes of steps S302 to S305 (identification process of whether the original document to be identified corresponds to document type 1) and the processes of steps S306 to S309 (identification process of whether the original document to be identified corresponds to document type 2) are executed in parallel.
[0129] In step S302, a character recognition result (full-text OCR result) is acquired. Since the process of step S302 is substantially the same as the process of step S202 in FIG. 19, a detailed description thereof is omitted. Thereafter, the process proceeds to step S303.
[0130] In step S303, a frequent word sequence detection process is performed. The detection unit 45 receives the high-frequency word list for document type 1 stored in the frequent word storage unit 43, and performs a detection process on the frequent word sequences of document type 1 stored in the high-frequency word list. Since the process of step S303 is substantially the same as the process of step S203 in FIG. 19, a detailed description thereof is omitted. Thereafter, the process proceeds to step S304.
[0131] In step S304, a feature quantity generation process is performed. The feature generation unit 46 generates a feature quantity (feature array) related to the original document shown in the image to be identified acquired in step S301 based on the position information acquired in step S303. Since the process of step S304 is substantially the same as the process of step S204 in FIG. 19, a detailed description thereof is omitted. Thereafter, the process proceeds to step S305.
[0132] In step S305, it is identified whether the document to be identified is a document of a predetermined document type (document type 1). The identification unit 47 receives the learned model for document type 1 stored in the model storage unit 44, and inputs the feature amount generated in step S304 to the learned model, thereby identifying whether the document to be identified is a document of document type 1. Since the process of step S305 is substantially the same as the process of step S205 in FIG. 19, detailed description thereof is omitted. Thereafter, the process proceeds to step S310.
[0133] Note that the identification process for document type 2 (steps S306 to S309) is substantially the same as the identification process for document type 1 (steps S302 to S305) described above, except that the target document type is different, and thus the description thereof is omitted.
[0134] In step S310, by aggregating the identification results, the document type of the document to be identified is identified, and the identified result is output. The identification unit 47 identifies the document type of the document to be identified based on the identification result as to whether the document to be identified corresponds to document type 1 and the identification result as to whether the document to be identified corresponds to document type 2. For example, if the identification result in step S305 is "corresponds to document type 1" and the identification result in step S309 is "does not correspond to document type 2", the document to be identified is identified (judged) as corresponding to document type 1 (being a document of document type 1), and the result is output. Thereafter, the process shown in this flowchart ends.
[0135] Note that in the above example, the identification processes for document type 1 and document type 2 are executed in parallel, but the present invention is not limited to this example, and the identification process for document type 2 may be executed after the identification process for document type 1 is completed. Further, the process of obtaining the character recognition result may not be performed for each document type as in the example shown in FIG. 21, but the character recognition result may be obtained only once for the image to be identified, and the obtained result may be used for all document types.
[0136] [Third Embodiment] In the second embodiment described above, an embodiment of identifying a plurality of manuscript types by using a plurality of learned models for identifying only one manuscript type has been described. However, in this embodiment, an embodiment of identifying a plurality of manuscript types by using one learned model capable of identifying a plurality of manuscript types will be described.
[0137] The configuration of the system according to this embodiment is substantially the same as that described in the first embodiment with reference to FIG. 1, and thus the description thereof will be omitted. Also, the functional configuration of the learning device according to this embodiment is substantially the same as that described in the first embodiment with reference to FIG. 2, and thus a detailed description thereof will be omitted. Also, the flow of the learning process in this embodiment is substantially the same as that described in the first embodiment with reference to FIG. 15, and thus the description thereof will be omitted. However, different from the first embodiment, in the learning device 2, one learned model capable of identifying a plurality of predetermined manuscript types is generated by the learning process. Therefore, the correct definition acquired by the correct definition acquisition unit 53, the high-frequency word list generated by the frequent occurrence unit acquisition unit 54, the feature amount (feature array) generated by the feature generation unit 56, etc. are different from those in the first embodiment.
[0138] Specifically, the correct definition acquisition unit 53 acquires a correct definition in which a learning image (identification information of the learning image) and information (such as a label) indicating whether the manuscript shown in the learning image is a manuscript of which manuscript type among a plurality of predetermined manuscript types are associated for each learning image. For example, when the manuscript types to be identified (predetermined manuscript types) are manuscript type 1 (INVOICE) and manuscript type 2 (bill), in the correct definition, the label "1" is associated with each learning image when it is manuscript type 1, the label "2" is associated with each learning image when it is manuscript type 2, and the label "0" is associated with each learning image when it does not correspond to either manuscript type. Note that it is optional whether to use an image of a manuscript that does not correspond to any manuscript type for the learning process.
[0139] The frequent word acquisition unit 54 acquires (extracts) the frequent word sequences of each of a plurality of predetermined manuscript types, and generates a high-frequency word list storing the frequent word sequences of each of the plurality of acquired predetermined manuscript types. Specifically, the frequent word acquisition unit 54 groups the learning images for each manuscript type (predetermined manuscript type), and extracts the frequent word sequence for each group (manuscript type). For example, in a plurality of learning images (INVOICE images) corresponding to manuscript type 1 (INVOICE), by executing the processes shown in steps S1041 to S1044, the frequent word sequence of manuscript type 1 is extracted, and a high-frequency word list for manuscript type 1 storing the frequent word sequence is generated. By performing the same process for other manuscript types, the frequent word sequences for each manuscript type are extracted (high-frequency word lists are generated). Note that, as described above, the high-frequency word list may be one list including the frequent word sequences of each manuscript type, rather than being generated for each manuscript type. Also, in the present embodiment, since a learned model is not generated for each manuscript type, identification information (model name) of the learned model does not need to be stored as in the high-frequency word list shown in FIG. 20.
[0140] The detection unit 55 acquires the position information related to the frequent word sequences of each of a plurality of predetermined manuscript types for each learning image (manuscript). The detection unit 55 acquires the position within the manuscript of each frequent word sequence stored in the high-frequency word list (if a list is generated for each manuscript type, all high-frequency word lists). That is, for each learning image (manuscript), the position information related to each frequent word sequence of each manuscript type (when the identified manuscript types (predetermined manuscript types) are manuscript type 1 and manuscript type 2, each frequent word sequence of manuscript type 1 and each frequent word sequence of manuscript type 2) is acquired.
[0141] Based on the position information acquired by the detection unit 55, the feature generation unit 56 generates feature amounts (feature arrays) related to the manuscript shown in the learning image. In this embodiment, the feature array stores feature amounts (position feature amounts, distance feature amounts, size feature amounts, line feature amounts) related to the frequently occurring word sequences of each of a plurality of predetermined manuscript types (all manuscript types to be identified). For example, when the manuscript types to be identified (predetermined manuscript types) are manuscript type 1 and manuscript type 2, the feature amounts related to each frequently occurring word sequence of manuscript type 1 and the frequently occurring word sequence of manuscript type 2 are stored. However, regarding the distance feature amount, only those calculated between the frequently occurring word sequences of the same manuscript type are stored.
[0142] The model generation unit 57 performs machine learning using the learning data in which the feature amounts (feature arrays) related to the manuscript shown in the learning image generated by the feature generation unit 56 are associated with information indicating which manuscript type among a plurality of predetermined manuscript types the manuscript shown in the learning image is (information based on the correct definition), thereby generating a learned model for identifying a plurality of predetermined manuscript types. That is, an identifier (learned model) is generated in which, when the feature amounts related to the manuscript including the positional relationship feature amounts regarding the positional relationship within the manuscript between the frequently occurring word sequences of each of the plurality of predetermined manuscript types and other word sequences are input, information indicating the validity that the manuscript is each of the plurality of predetermined manuscript types is output.
[0143] In this embodiment, in order to enable identification of a plurality of predetermined manuscript types, feature amounts related to the frequently occurring word sequences of each of the plurality of predetermined manuscript types are generated (stored in one feature array). Therefore, it is considered that the generated feature amounts (the feature amounts stored in the feature array) will become enormous. Thus, as a measure to reduce the feature amounts (the feature amounts stored in the feature array (the position feature amounts, distance feature amounts, size feature amounts, line feature amounts of each frequently occurring word sequence)), it is possible to use the method shown below.
[0144] (Removal of frequently occurring word sequences that overlap between manuscript types) If there are frequently occurring word sequences that overlap among a plurality of (two or more) manuscript types, the overlapping frequently occurring word sequences may be excluded from the frequently occurring word sequences used when generating the feature amounts.
[0145] (Only use frequent word sequence pairs where the average distance between word sequences is below a threshold) Among combinations (pairs) consisting of two frequent word sequences of a predetermined manuscript type (e.g., INVOICE), the distance between frequent word sequences of only combinations that meet a predetermined condition may be used for calculating feature quantities. A combination that meets the predetermined condition is a combination where the representative value (average value) of the distance between frequent word sequences in a plurality of learning images that are images of the predetermined manuscript type (INVOICE) is below a predetermined value. For example, when frequent word sequences are extracted from a plurality of learning images (e.g., 100 images) that are images of manuscript type 1 (INVOICE), the distance between word sequences for all combinations (pairs) of frequent word sequences is calculated for each learning image (in each of the 100 images). Then, only pairs of frequent word sequences where the average value of the distance between frequent word sequences in the 100 learning images is below a predetermined threshold may be determined as the word sequence pairs used for the distance feature quantity.
[0146] (Only use feature quantities with high usage frequencies) As a result of performing manuscript type identification processing using the generated learned model, it is possible to obtain, using the learned model, which feature quantities were the feature quantities used for identification. Therefore, the feature array may be changed so as to only use feature quantities that are frequently used in actual identification processing (feature quantities with high usage frequencies).
[0147] (Removing highly correlated feature quantities) When there are feature quantities with high correlation among feature quantities, only one of the feature quantities with high correlation may be used as the feature quantity related to the manuscript shown in the learning image, and the other feature quantities may be excluded from the feature quantities related to the manuscript shown in the learning image.
[0148] (Dimensionality reduction by principal component analysis) The dimensionality of feature quantities may be reduced by using principal component analysis (PCA).
[0149] The functional configuration of the information processing apparatus according to the present embodiment is substantially the same as that described in the first embodiment with reference to FIG. 14, and thus the description thereof is omitted. Further, the flow of the identification process in the present embodiment is substantially the same as that described in the first embodiment with reference to FIG. 19, and thus the description thereof is omitted.
[0150] However, in the present embodiment, the frequently-occurring word storage unit 43 stores the frequently-occurring word sequences (high-frequency word lists storing the frequently-occurring word sequences of each of a plurality of predetermined manuscript types generated by the frequently-occurring word acquisition unit 54) of each of the plurality of predetermined manuscript types described above. Further, the model storage unit 44 stores the learned models for identifying a plurality of predetermined manuscript types generated by the model generation unit 57 described above. Further, the detection unit 45 acquires information regarding the positions of the frequently-occurring word sequences of each of the plurality of predetermined manuscript types within the manuscript to be identified. The feature generation unit 46 generates feature quantities (feature quantities regarding the frequently-occurring word sequences of each of the plurality of predetermined manuscript types) regarding the manuscript to be identified using the information acquired by the detection unit 45. Note that the details of the feature quantities regarding the frequently-occurring word sequences are the same as those in the first embodiment.
[0151] The identification unit 47 inputs the feature quantities regarding the manuscript to be identified into the learned models for identifying a plurality of predetermined manuscript types, thereby acquiring information indicating the validity that the manuscript to be identified is a manuscript of each of the plurality of predetermined manuscript types (for example, when the manuscript types to be identified are manuscript type 1 and manuscript type 2, information indicating the validity of being manuscript type 1 and information indicating the validity of being manuscript type 2). From this, the identification unit 47 identifies whether the manuscript to be identified is a manuscript of which manuscript type among the plurality of predetermined manuscript types based on the acquired information indicating the validity. For example, from the probabilities (confidence levels, etc.) that the manuscript is a manuscript of each manuscript type output from the learned model, the manuscript type with the highest probability can be determined (identified) as the manuscript type of the manuscript to be identified.
Description of Reference Numerals
[0152] 1 Information processing apparatus 2 Learning apparatus 3 Document reading apparatus 9 Information processing system
Claims
1. Recognition result acquisition means for acquiring a character recognition result for an identification target image which is an image of a manuscript to be identified; Frequent word storage means for storing a frequent word sequence of a predetermined manuscript type; Detection means for detecting the frequent word sequence from the character recognition result of the identification target image, thereby acquiring information regarding the position of the frequent word sequence within the manuscript to be identified; Feature generation means for generating a feature amount related to the manuscript to be identified, including a positional relationship feature amount regarding the positional relationship within the manuscript to be identified between the frequent word sequence and another word sequence, using the information regarding the position; Model storage means for storing a learned model for identifying the predetermined manuscript type, which is generated by machine learning such that information indicating the validity that the manuscript is a manuscript of the predetermined manuscript type is output when a feature amount related to the manuscript, including a positional relationship feature amount regarding the positional relationship within the manuscript between the frequent word sequence and another word sequence, is input; Identification means for identifying whether or not the manuscript to be identified is a manuscript of the predetermined manuscript type by inputting the feature amount related to the manuscript to be identified into the learned model; An information processing system comprising the above.
2. The learned model is a model generated by machine learning using learning data in which, for each of a plurality of learning images including a plurality of predetermined manuscript type images which are images of manuscripts of the predetermined manuscript type having different layouts from each other, a feature amount related to the manuscript shown in the learning image, including a positional relationship feature amount regarding the positional relationship within the manuscript shown in the learning image between the frequent word sequence and another word sequence, is associated with information indicating whether or not the manuscript shown in the learning image is a manuscript of the predetermined manuscript type. The information processing system according to Claim 1.
3. The frequent word sequence is one of a plurality of frequent word sequences, The positional relationship feature amount includes a feature amount indicating the distance within the target manuscript between the frequent word sequence and another frequent word sequence. The information processing system according to Claim 1 or 2.
4. The positional relationship feature amount includes a feature amount indicating the size of the line including the frequent word sequence. The information processing system according to any one of Claims 1 to 3.
5. The feature amount related to the manuscript to be identified includes, in addition to the positional relationship feature amount, a feature amount indicating an attribute of the frequent word sequence. The information processing system according to any one of Claims 1 to 4.
6. The feature amount indicating the attribute of the frequent word sequence includes at least one of the feature amount indicating the position of the frequent word sequence and the feature amount indicating the size of the frequent word sequence. The information processing system according to claim 5.
7. The model storage means stores the learned model generated by learning data in which a feature array in which the feature amounts related to the manuscripts shown in each learning image are aggregated in an array form is associated with information indicating whether the manuscripts shown in each learning image are manuscripts of the predetermined manuscript type. The feature generation means shapes the feature amount related to the manuscript to be identified into an array in the same order as the feature array. The identification means inputs the feature amount related to the manuscript to be identified shaped into the array into the learned model, and identifies whether the manuscript to be identified is a manuscript of the predetermined manuscript type. The information processing system according to claim 2.
8. The predetermined manuscript type is one of a plurality of predetermined manuscript types. The model storage means stores a learned model for identifying a predetermined manuscript type for each of the plurality of predetermined manuscript types. The identification means uses a learned model for identifying a predetermined manuscript type for each of the plurality of predetermined manuscript types to identify whether the image to be identified corresponds to the predetermined manuscript type, and based on the results of identification for each of the plurality of predetermined manuscript types, identifies which manuscript type among the plurality of predetermined manuscript types the manuscript to be identified is. The information processing system according to any one of claims 1 to 7.
9. When, as a result of identification for each of the plurality of predetermined manuscript types, the manuscript to be identified is identified as a manuscript of a predetermined manuscript type in two or more predetermined manuscript types, the identification means selects one manuscript type from the two or more predetermined manuscript types, and determines the selected manuscript type as the manuscript type of the manuscript to be identified. The information processing system according to claim 8.
10. The identification means selects one manuscript type from the two or more predetermined manuscript types based on the likelihood that the manuscript to be identified is a manuscript of the predetermined manuscript type. The information processing system according to claim 9.
11. The identification means selects one manuscript type from the two or more predetermined manuscript types based on the number of times each of the two or more predetermined manuscript types has been identified as the manuscript type of the manuscript to be identified by the learned model in the past. The information processing system according to claim 9.
12. The identification means selects one manuscript type from the two or more predetermined manuscript types based on the time when each of the two or more predetermined manuscript types was identified as the manuscript type of the manuscript to be identified by the learned model in the past. The information processing system according to claim 9.
13. The predetermined manuscript type is one of a plurality of predetermined manuscript types. The frequent word storage means stores the frequent word sequences of each of the plurality of predetermined manuscript types. The detection means acquires information regarding the positions of the frequent word sequences of each of the plurality of predetermined manuscript types within the manuscript to be identified. The feature generation means generates a feature quantity related to the manuscript to be identified, including a positional relationship feature quantity regarding the positional relationship between the frequent word sequence of each of the plurality of predetermined manuscript types and another word sequence within the manuscript to be identified, using the information regarding the positions. The model storage means stores a learned model for identifying the plurality of predetermined manuscript types, which is generated by machine learning such that when a feature quantity related to a manuscript, including a positional relationship feature quantity regarding the positional relationship between the frequent word sequence of each of the plurality of predetermined manuscript types and another word sequence within the manuscript, is input, information indicating the validity that the manuscript is each of the plurality of predetermined manuscript types is output. The identification means inputs the feature quantity related to the manuscript to be identified into the learned model for identifying the plurality of predetermined manuscript types, and thereby identifies whether the manuscript to be identified is a manuscript of any of the plurality of predetermined manuscript types. The information processing system according to any one of claims 1 to 7.
14. When there is a frequent word sequence that overlaps among the plurality of predetermined manuscript types, the positional relationship feature quantity is a positional relationship feature quantity regarding the positional relationship between the frequent word sequence of each of the plurality of predetermined manuscript types that does not correspond to the overlapping frequent word sequence and another word sequence. The information processing system according to claim 13.
15. The positional relationship feature quantity includes a feature quantity indicating the distance between frequent word sequences related to a combination that satisfies a predetermined condition among combinations consisting of two frequent word sequences of the predetermined manuscript type. The combination that satisfies the predetermined condition is a combination in which a representative value of the distance between the frequent word sequences in a plurality of learning images that are images of the predetermined manuscript type is equal to or less than a predetermined value. The information processing system according to claim 13.
16. Recognition result acquisition means for acquiring, for each of a plurality of learning images including a plurality of predetermined manuscript type images that are images of manuscripts of a predetermined manuscript type with different layouts from each other, the character recognition result; Frequent word acquisition means for acquiring the frequent word sequence of the predetermined manuscript type; For each learning image, detection means for acquiring information regarding the position in the manuscript shown in the learning image of the frequent word sequence by detecting the frequent word sequence from the character recognition result of the learning image; For each learning image, feature generation means for generating a feature amount related to the manuscript shown in the learning image, including a positional relationship feature amount regarding the positional relationship in the manuscript shown in the learning image between the frequent word sequence and other word sequences, using the information regarding the position in the manuscript shown in the learning image of the frequent word sequence; Model generation means for generating a learned model for identifying the predetermined manuscript type by performing machine learning using learning data in which the feature amount related to the manuscript shown in the learning image and information indicating whether the manuscript shown in the learning image is a manuscript of the predetermined manuscript type are associated with each other for each learning image; An information processing system comprising the above.
17. The frequent word acquisition means extracts a word sequence that frequently appears in the manuscripts shown in the plurality of predetermined manuscript type images based on the character recognition results of the plurality of predetermined manuscript type images, and acquires the extracted word sequence as the frequent word sequence of the predetermined manuscript type. The information processing system according to claim 16.
18. The information processing system further comprises correct definition acquisition means for acquiring a correct definition in which the identification information of the learning image and information indicating whether the manuscript shown in the learning image is a manuscript of the predetermined manuscript type are associated with each other for each learning image. The model generation means acquires information indicating whether the manuscript shown in the learning image is a manuscript of the predetermined manuscript type based on the correct definition. The information processing system according to claim 16 or 17.
19. A computer A recognition result acquisition step of acquiring a character recognition result for an identification target image that is an image of a manuscript to be identified; A frequent word storage step of storing a frequent word sequence of a predetermined manuscript type; A detection step of acquiring information regarding the position in the manuscript to be identified of the frequent word sequence by detecting the frequent word sequence from the character recognition result of the identification target image; A feature generation step of generating a feature amount related to the manuscript to be identified, including a positional relationship feature amount regarding the positional relationship in the manuscript between the frequently occurring word sequence and other word sequences, using the information regarding the position; A model storage step of storing a learned model for identifying the predetermined manuscript type, which is generated by machine learning so that information indicating the validity that the manuscript is a manuscript of the predetermined manuscript type is output when a feature amount related to the manuscript including a positional relationship feature amount regarding the positional relationship in the manuscript between the frequently occurring word sequence and other word sequences is input; An identification step of identifying whether or not the manuscript to be identified is a manuscript of the predetermined manuscript type by inputting the feature amount related to the manuscript to be identified into the learned model; A manuscript type identification method for executing the above.
20. A computer, A recognition result acquisition means for acquiring a character recognition result for an identification target image which is an image of a manuscript to be identified; A frequently occurring word storage means for storing a frequently occurring word sequence of a predetermined manuscript type; A detection means for acquiring information regarding the position in the manuscript to be identified of the frequently occurring word sequence by detecting the frequently occurring word sequence from the character recognition result of the identification target image; A feature generation means for generating a feature amount related to the manuscript to be identified, including a positional relationship feature amount regarding the positional relationship in the manuscript between the frequently occurring word sequence and other word sequences, using the information regarding the position; A model storage means for storing a learned model for identifying the predetermined manuscript type, which is generated by machine learning so that information indicating the validity that the manuscript is a manuscript of the predetermined manuscript type is output when a feature amount related to the manuscript including a positional relationship feature amount regarding the positional relationship in the manuscript between the frequently occurring word sequence and other word sequences is input; An identification means for identifying whether or not the manuscript to be identified is a manuscript of the predetermined manuscript type by inputting the feature amount related to the manuscript to be identified into the learned model; A program for functioning as the above.
21. A computer performs A recognition result acquisition step of acquiring a character recognition result for each of a plurality of learning images including a plurality of predetermined manuscript type images which are images of manuscripts of a predetermined manuscript type having different layouts from each other; A frequently occurring word acquisition step of acquiring a frequently occurring word sequence of the predetermined manuscript type; For each learning image, a detection step of obtaining information regarding the position in the document shown in the learning image of the frequent word sequence by detecting the frequent word sequence from the character recognition result of the learning image; For each learning image, a feature generation step of generating a feature amount related to the document shown in the learning image, including a positional relationship feature amount regarding the positional relationship in the document shown in the learning image between the frequent word sequence and other word sequences, using the information regarding the position in the document shown in the learning image of the frequent word sequence; A model generation step of generating a learned model for identifying the predetermined document type by performing machine learning using the learning data in which the feature amount related to the document shown in the learning image and the information indicating whether the document shown in the learning image is a document of the predetermined document type are associated with each other for each learning image; A model generation method for executing the above.
22. A computer, A recognition result acquisition means for acquiring a character recognition result for each of a plurality of learning images including a plurality of predetermined document type images that are images of documents of a predetermined document type having different layouts from each other; A frequent word acquisition means for acquiring a frequent word sequence of the predetermined document type; A detection means for obtaining information regarding the position in the document shown in the learning image of the frequent word sequence by detecting the frequent word sequence from the character recognition result of the learning image for each learning image; A feature generation means for generating a feature amount related to the document shown in the learning image, including a positional relationship feature amount regarding the positional relationship in the document shown in the learning image between the frequent word sequence and other word sequences, using the information regarding the position in the document shown in the learning image of the frequent word sequence for each learning image; A model generation means for generating a learned model for identifying the predetermined document type by performing machine learning using the learning data in which the feature amount related to the document shown in the learning image and the information indicating whether the document shown in the learning image is a document of the predetermined document type are associated with each other for each learning image; A program for functioning as the above.
Citation Information
Patent Citations
Image processor
JP1999146220A
Document classification device, program and document classification method
JP2005122550A
Image processing device and program
JP2017090974A
Document classification device and trained model
WO2020021845A1