Document reading system, and document reading method
The document reading system employs character recognition and natural language processing to reduce the reading load by generating accurate target character strings from document images, addressing the high inspection demands in existing systems.
Patent Information
- Application Number
- JP2023190478
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-08
- Publication Date
- 2025-05-20
AI Technical Summary
Existing document reading systems face a high load due to the need to inspect numerous items in documents such as contracts and account books, necessitating a reduction in reading load.
A document reading system and method that utilizes a text conversion unit for character recognition and a natural language processing unit to generate target character strings from document image data, incorporating morphological, syntactic, and semantic analyses, along with machine learning for inference.
Reduces the reading load on users by accurately generating target character strings, increasing accuracy and certainty in identifying specified entry items within documents.
Smart Images

Figure 2025078130000001_ABST
Abstract
Description
[Technical field]
[0001] The present disclosure relates to a document reading system and a document reading method for reading the contents of a document. [Background technology]
[0002] Character recognition technology is used to read various documents such as contract documents and account books. One example of a document reading system includes multiple character recognition engines, and each character recognition engine is associated with a document type. The document reading system then performs character recognition processing on the document to be read using the character recognition engine associated with the document type to be read (see, for example, Patent Document 1). [Prior art documents] [Patent documents]
[0003] [Patent Document 1] JP 2020-181369 A Summary of the Invention [Problem to be solved by the invention]
[0004] Since numerous items must be inspected when issuing and using the above documents, there is a strong demand for reducing the document reading load on the document reading system described above. [Means for solving the problem]
[0005] A document reading system for solving the above problem includes a target string, which is a string indicating the content of an item described in a document, image data of the document is document image data, and an electronic file including the document image data is an electronic document file, the system including a text conversion unit configured to convert a string in the document image data included in the electronic document file into text data using character recognition processing, and a natural language processing unit configured to generate the target string from the text data converted by the text conversion unit using a natural language processing model.
[0006] The document reading system may further include a text data processing unit configured to process the text data into a format suitable for generating the target character string by the natural language processing model, and the text data processed by the text data processing unit may be applied to the natural language processing model.
[0007] In the document reading system, the text data processing unit may process the text data differently for each type of document. A document reading method for solving the above problem includes: a character string indicating the content of an item described in a document is a target character string; image data of the document is document image data; an electronic file including the document image data is an electronic document file; a text conversion unit uses character recognition processing to convert a character string in the document image data included in the electronic document file into text data; and a natural language processing unit uses a natural language processing model to generate the target character string from the text data converted by the text conversion unit.
[0008] A document reading method for solving the above problem includes the steps of: a target string is a string indicating the content of an item described in a document; image data of the document is document image data; and an electronic file including the document image data is an electronic document file; causing a control unit to acquire the text data from a text conversion unit configured to convert a string in the document image data included in the electronic document file into text data using character recognition processing; and causing a natural language processing unit configured to generate the target string from the text data using a natural language processing model to generate the target string from the acquired text data. Effect of the Invention
[0009] According to the document reading system and document reading method of the present disclosure, the load of reading documents is reduced. [Brief description of the drawings]
[0010] [Figure 1] FIG. 1 is a block diagram showing a system configuration of a document reading system. [Diagram 2] FIG. 2 is a flow chart showing a process flow in the document reading method. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0011] An embodiment of a document reading system and a document reading method will be described below. [System Overview] The document reading system performs character recognition processing to convert character strings in document image data into text data. The document reading system performs natural language processing to generate target character strings indicating the contents of specified description items from the text data. Natural language processing includes morphological analysis, syntactic analysis, semantic analysis, context analysis, etc., and includes inference by machine learning.
[0012] The designated description item is a designated description item among the description items described in the document. The designated description item may be one description item in the document, or may be two or more description items in the document. The designated description item is input to the document reading system by a user of the document reading system.
[0013] The target character string is a character string that is read as the content of an item described in a document. The target character string is an object that is output to a document reading system using document image data as input. The target character string may be a character string that indicates the content of one item described in a document, or may be a character string that indicates the content of two or more items described in a document separately.
[0014] The document image data is data for reproducing the contents of the written items as an image. The document image data is described by a set of drawing commands. The document image data may be data described in a standard drawing language, or may be an image file in a file format such as JPEG, GIF, PNG, BMP, EPS, TIFF, or PSD.
[0015] The electronic document file F11 (see FIG. 2) is an electronic file including document image data. The electronic document file F11 causes an application to execute a drawing command for the document image data. The electronic document file F11 reproduces a target character string as an image by executing the drawing command. The application that executes the drawing command is a viewer, a browser, an editor, a converter, an analyzer, or the like. If the document image data is an image file, the electronic document file F11 may be the document image data. In this case, the electronic document file F11 may be an image file described in a file format such as JPEG, GIF, PNG, BMP, EPS, TIFF, or PSD. The electronic document file F11 may be an application file that incorporates a document image file, such as PDF, JSON, XML, DOC, or XLS, or may be an application file configured to be able to read the document image file in a document reading system.
[0016] The electronic document file F11 may include text data that can be read by a document reading system in addition to the document image data. The text data may be composed of a character string other than the target string, or may be composed of a character string including the target string. The electronic document file F11 may incorporate audio data. The audio data may be configured to reproduce a character string other than the target string as audio, or to reproduce a character string including the target string as audio.
[0017] The documents handled by the document reading system are documents handled by companies such as joint stock companies and private companies. The documents handled by companies may be contract documents such as sales contracts, lease agreements, labor contracts, certified copies of real estate registers, and property overviews. The documents handled by companies may be account books such as pay slips, delivery notes, invoices, expense reports, and receipts. The documents handled by companies may be labor-related documents such as work regulations, employment contracts, time cards, resumes, and CVs. The documents handled by companies may be insurance-related documents such as health insurance cards, employee pension insurance cards, employment insurance certificates, and workers' accident insurance certificates. The documents handled by companies may be questionnaires, manufacturing drawings, architectural drawings, catalogs, and the like. The documents handled by companies may be documents with different samples for each company depending on the type and size of the company, or documents with samples common to all companies.
[0018] The document handled by the document reading system includes the contents of the items that are predetermined for the document. For example, if the document is a contract, the items include the parties, the subject matter of the contract, the purpose of the contract, the due date of the contract, the conditions for canceling the contract, the method of resolving disputes in the contract, and the terms and conditions of the contract. For example, if the document is an invoice, the items include the invoice number, the date of issue of the invoice, the recipient of the invoice, the billing destination, the product name, the service name, the quantity, the unit price, the total amount, the amount of consumption tax, and the due date of payment. For example, if the document is a work rule, the items include the working hours, break time, holidays, wages, salary increases, dismissal, and safety and health. If the document is a health insurance card, the items include the name, date of birth, sex, address, insurer number, date of qualification acquisition, dependent classification, expiration date, and type of health insurance card. For example, if the document is a certified copy of the registry, the items include the location of the real estate, floor area, owner, type of ownership, amount of mortgage, and start date of registration record.
[0019] [System Configuration] As shown in Fig. 1, the document reading system includes a control unit 10, a text conversion unit 20, and a target character string generation unit 30. The control unit 10 and the target character string generation unit 30 configure a natural language processing unit. The control unit 10 includes an electronic file acquisition unit 11, a text data processing unit 12, and a target character string output unit 13. The text conversion unit 20 includes a character recognition processing unit 21. The target character string generation unit 30 includes a natural language processing model 31.
[0020] The control unit 10 is connected to the text conversion unit 20 and the target character string generation unit 30 via a network. The control unit 10 accepts input of designated entry items. The control unit 10 controls the conversion process by the text conversion unit 20 and the generation process by the target character string generation unit 30.
[0021] The electronic file acquisition unit 11 is configured to acquire an electronic document file F11 from an external device. The external device may be an external server connected to the control unit 10 via a network. The external device may be an input device such as a scanner or a camera connected to the control unit 10. The control unit 10 is configured to transmit the electronic document file F11 acquired by the electronic file acquisition unit 11 to the text conversion unit 20.
[0022] [Text conversion section 20] The text conversion unit 20 is configured to receive the electronic document file F11 from the control unit 10. The text conversion unit 20 performs image processing on the document image data included in the electronic document file F11. The text conversion unit 20 uses the results of the image processing to cause the character recognition processing unit 21 to convert character strings in the document image data into text data. The text conversion unit 20 is configured to transmit the converted text data to the control unit 10.
[0023] As image processing of the document image data, the text conversion unit 20 may generate, from the document image data, character shapes for identifying character types, character sizes for identifying character font sizes, character positions for identifying character string lengths and separators, etc. The character recognition processing unit 21 may output character strings corresponding to the results of the image processing as text data.
[0024] The text conversion unit 20 may generate character combinations, character thicknesses, and the like in addition to character shapes, character sizes, and character positions as image processing of the document image data. The character recognition processing unit 21 may include a text inference unit, and may input the results of the image processing to the text inference unit. The character recognition processing unit 21 may output the inference results by the text inference unit as text data. The text inference unit is a neural network that is trained to infer character strings contained in the document image data from the results of image processing using the learning document image data.
[0025] For example, if the document image data is image data of a contract, the text conversion unit 20 outputs, as text data, character strings indicating the contents of the parties, the subject matter of the contract, the purpose of the contract, the due date of the contract, the conditions for canceling the contract, the method of resolving disputes in the contract, and the terms and conditions of the contract. The text conversion unit 20 transmits character strings indicating the contents of each description item to the control unit 10. The control unit 10 is configured to receive, from the text conversion unit 20, the text data converted by the text conversion unit 20.
[0026] [Text data processing section 12] The text data processing unit 12 is configured to process the text data received from the text conversion unit 20. The text data processing unit 12 includes a processing inference model and a neural network that processes the text data using the processing inference model. The processing inference model includes at least one of a line break normalization model, a colloquial expression normalization model, and a grammatical error correction model. The control unit 10 is configured to include a query for outputting the contents of the designated description item in the processed text data. The control unit 10 is configured to transmit the processed text data including the query to the target character string generation unit 30.
[0027] The line break normalization model is a set of data for normalizing line breaks according to a line break normalization algorithm. The data for normalizing line breaks includes line break rules suitable for inference by the target string generation unit 30, rules for combining broken text, etc. The text data processing unit 12 corrects line breaks in the text data using the line break normalization model.
[0028] The colloquial expression normalization model is a collection of data for converting colloquial expressions to sentence expressions according to a colloquial expression normalization algorithm. The data for converting colloquial expressions to sentence expressions includes grammar rules, vocabulary knowledge, context knowledge, etc. The text data processing unit 12 corrects the colloquial expressions of the text data to sentence expressions using the colloquial expression normalization model.
[0029] The grammatical error correction model is a set of data for correcting grammatical errors according to a grammatical error correction algorithm. The data for correcting grammatical errors includes grammatical rules, dictionary data, etc. The text data processing unit 12 corrects grammatical errors in the text data using the grammatical error correction model.
[0030] The text data processing unit 12 may be configured to generate a document type from the text data received from the text conversion unit 20. The document type may be text data indicating a group to which the document belongs, such as a contract document or a book, or may be text data indicating the name of the document, such as a sales contract or a lease agreement. The document type may be a named entity used specifically for the document. In either case, the text data processing unit 12 may include a type inference model and a neural network that classifies the text data using the type inference model. The type inference model may include at least one of a named entity extraction model and a document determination model. The control unit 10 may be configured to include an inference result using the type inference model in the processed text data and transmit it to the target character string generation unit 30.
[0031] The named entity extraction model is a collection of data for extracting named entities according to a named entity extraction algorithm. The data for extracting named entities includes grammar rules, dictionary data, and knowledge data of named entities. The text data processing unit 12 extracts named entities from the text data using the named entity extraction model. For example, when the text data is converted from document image data of a contract, the named entities are the due date of performance, the termination conditions, and the terms of agreement. When the text data is converted from document image data of an invoice, the named entities are the due date of payment, the amount of consumption tax, and the like. When the text data is converted from document image data of a health insurance card, the named entities are the dependent category, the type of insurance card, and the like.
[0032] The document determination model is a collection of data for determining the type of document according to a document determination algorithm. The data for determining the type of document is dictionary data, knowledge data related to document types and named entities, etc. The text data processing unit 12 identifies the type of document from the text data using the document determination model. If the type of document is other than a named entity, the text data processing unit 12 may identify the type of document from the text data using a named entity extracted using a named entity extraction model and the document determination model.
[0033] The text data processing unit 12 may be configured to divide (chunk) and embed (embedding) the text data. In this case, the text data processing unit 12 divides the text data received from the text conversion unit 20 into data elements of a predetermined unit. The data elements of a predetermined unit may be word groups consisting of multiple words, or may be sentences which are a collection of word groups. The text data processing unit 12 may be configured to understand the meaning of each data element of the text data, and embed the meaning of each data element in the text data.
[0034] The text data processing unit 12 may determine whether to perform division and embedding based on the type of document. Contract documents and labor-related documents use more written expressions than account books and insurance-related documents. Ledgers and insurance-related documents use more numerical expressions than contract documents and labor-related documents. The text data processing unit 12 may be configured not to perform division and embedding when the type of document is account books or insurance-related documents, and to perform division and embedding when the type of document is contract documents or labor-related documents.
[0035] For example, when image data of a contract is converted into text data, the text data processing unit 12 divides the text data into sentences of clause numbers. The text data processing unit 12 embeds the meaning of each sentence, such as contract cancellation conditions and prohibited conditions, into the text data.
[0036] The text data processing unit 12 may be configured to search the text data for data elements that are likely to contain the target string, using the meaning of each data element in the text data. In this case, the text data processing unit 12 may include a target search model and a neural network that searches for data elements using the target search model. The control unit 10 may be configured to transmit the data elements searched for using the target search model to the target string generation unit 30 as text data.
[0037] The target search model is a collection of data for searching for data elements that are likely to contain a target string according to a target search algorithm. Data for extracting data elements includes grammar rules, dictionary data, knowledge data of description items, etc. The neural network provided in the text data processing unit 12 learns to infer whether or not a target string of a specified description item is included in a data element for learning.
[0038] For example, when text data is converted from image data of a contract and a termination condition is specified as a specified description item, the text data processing unit 12 searches the text data for data elements that are likely to include the termination condition. The control unit 10 transmits the data elements searched for using the target search model to the target string generation unit 30 as text data.
[0039] The target string output unit 13 is configured to receive a target string from the target string generation unit 30. The target string output unit 13 may output the target string in a conversational format, such as to respond to an inquiry about a specified entry item. The target string output unit 13 may output the target string in a data input format, such that the target string is input into an input field for subsequent processing.
[0040] [Target string generation part 30] The target string generation unit 30 is configured to receive processed text data from the control unit 10. The target string generation unit 30 includes a language inference unit, and inputs text data to the language inference unit. The language inference unit automatically generates natural language based on the text data using a natural language processing model 31. The language inference unit is a neural network that has been trained to infer the contents of designated description items in the learning text data based on the text data. The natural language processing model 31 is a neural network language model that has learned the characteristics of the text data by a neural network. The natural language processing model 31 is a language model that is larger in scale than the model included in the control unit 10. The target string generation unit 30 is configured to transmit the inference result by the language inference unit to the control unit 10 as a target string.
[0041] For example, if the text data is data converted from image data of a contract and the contract execution date is a specified item, the target character string generation unit 30 infers the content of the execution date in the text data. The target character string generation unit 30 transmits the year, month, and date of the execution date, which is the inference result, to the control unit 10.
[0042] For example, if the text data is data converted from image data of an invoice and the recipient and the payment due date are the specified items, the target character string generation unit 30 infers the contents of the recipient and the payment due date in the text data. The target character string generation unit 30 transmits the inference results, that is, the recipient's name and the date of the payment due date, to the control unit 10.
[0043] For example, if the text data is data converted from image data of work regulations and the working hours are the designated description items, the target character string generation unit 30 infers the contents of the working hours in the text data. The target character string generation unit 30 transmits the inference results, that is, the starting time, the finishing time, and the statutory working hours, to the control unit 10.
[0044] The control unit 10, the text conversion unit 20, and the target character string generation unit 30 each include an electronic circuit that is a computing device such as a CPU, MPU, or GPU, a memory such as a ROM, RAM, registered memory, or unbuffered memory, and a storage such as an SSD or HDD. The computing device loads an operating system or various programs from the storage into the memory and executes instructions retrieved from the memory. The control unit 10, the text conversion unit 20, and the target character string generation unit 30 are not limited to those that execute various processes all by software. The control unit 10, the text conversion unit 20, and the target character string generation unit 30 may each include an integrated circuit for a specific application that executes at least a part of the various processes. The control unit 10, the text conversion unit 20, and the target character string generation unit 30 may each be configured as a circuit that includes one or more dedicated hardware circuits such as an FPGA or an ASIC, one or more processors that operate according to a computer program, or a combination of these.
[0045] The text conversion unit 20 and the target string generation unit 30 may be configured as a single server equipped with an electronic circuit, memory, and storage, and the server may execute various processes to function as the control unit 10, the text conversion unit 20, and the target string generation unit 30. The document reading system may be configured as a single server equipped with an electronic circuit, memory, and storage, and the server may execute various processes to function as the control unit 10, the text conversion unit 20, and the target string generation unit 30. The document reading system may be configured as a single server equipped with an electronic circuit, memory, and storage, and may include the text conversion unit 20 in a front-end module and the target string generation unit 30 in a back-end module.
[0046] [How to read documents] An example of a document reading method executed by the document reading system will be described with reference to FIG. As shown in Fig. 2, the document reading system causes the electronic file acquisition unit 11 to acquire an electronic document file F11 from an external device. The document reading system accepts settings of designated description items. The document reading system causes the text conversion unit 20 to convert the document image data of the electronic document file F11 into text data. The control unit 10 acquires the text data from the text conversion unit 20 (step S11).
[0047] When the control unit 10 acquires the text data, the control unit 10 performs formatting processing such as line break normalization, colloquial expression normalization, and grammatical error correction on the text data (step S12). The control unit 10 performs expression extraction to extract named entities from the formatted text data (step S13). The control unit 10 uses the extracted named entities and the like to determine the type of document, and determines whether or not to divide the processed text data (step S14).
[0048] When the control unit 10 determines that the text data should be divided, it performs division and embedding on the formatted text data (step S15) and searches for data elements that are likely to contain the contents of the designated description items (step S16). The control unit 10 generates text data inquiring about the contents of the designated description items from the searched data elements, and causes the target character string generation unit 30 to infer the contents of the pointed description items in the text data (step S17).
[0049] When the control unit 10 determines not to divide the text data, it generates text data inquiring about the contents of the specified description items from the format-processed text data, and causes the target string generation unit 30 to infer the contents of the pointed-out description items in the text data (step S18).
[0050] Then, the control unit 10 outputs the inference result by the target string generation unit 30 as a target string (step S19).
[0051] As described above, the following effects can be obtained. (1) The document reading system causes the target character string generation unit 30 to generate a target character string from the text data converted by the text conversion unit 20 using the natural language processing model 31. This allows the document reading system to reduce the reading load on the user, such as searching for the location of a specified entry item in a document.
[0052] (2) The document reading system causes the text data processing unit 12 to process the text data into a format suitable for generating the target character string. As a result, the document reading system increases the accuracy that the target character string is the content of the designated entry item.
[0053] (3) The document reading system varies the processing by the text data processing unit 12 for each type of document. As a result, the document reading system increases the accuracy that the target character string is the content of the designated entry item while suppressing an increase in the processing load required for generating the target character string.
[0054] (4) The document reading system causes the text data processing unit 12 to divide and embed the text data, and then causes the target string generation unit 30 to generate a target string from data elements determined to be likely to contain the target string. Therefore, when the text data contains many character strings different from the target string, the document reading system increases the certainty that the target string is the content of the designated entry item.
[0055] (5) On condition that the document type is a predetermined type, the document reading system causes the text data processing unit 12 to divide and embed the text data. Therefore, the document reading system prevents an excessive increase in the processing load required to increase the certainty that the target character string is the content of the specified entry item.
[0056] The above-described embodiment can be modified as follows. When text data is included in the electronic document file F11, the document reading system may be configured to omit the process of converting character strings in the document image data into text data. In this case, the document reading system may be configured to extract the text data included in the electronic document file F11 from the electronic document file F11 and cause the target character string generation unit 30 to generate a target character string from the text data.
[0057] When text data is included in the document electronic file F11, the document reading system may be configured to determine whether or not the text data included in the document electronic file F11 includes a designated description item before converting a character string in the document image data into text data. If the document reading system determines that the text data included in the document electronic file F11 includes a designated description item, the document reading system may be configured to extract the text data included in the document electronic file F11 and cause the target string generating unit 30 to generate a target string from the text data.
[0058] When voice data is included in the electronic document file F11, the document reading system may be configured to convert the voice reproduced by the voice data into text data. In this case, the document reading system may be configured to cause the target string generating unit 30 to generate a target string from the text data converted from the voice data.
[0059] When voice data is included in the electronic document file F11, the document reading system may be configured to determine whether or not the text data converted from the voice data includes a designated entry item before transmitting the text data to the target string generation unit 30. Then, when the document reading system determines that the text data converted from the voice data includes a designated entry item, the document reading system may be configured to cause the target string generation unit 30 to generate a target string from the text data converted from the voice data. [Explanation of symbols]
[0060] F11...Electronic document file 10...Control section 11…Electronic File Acquisition Section 12…Text data processing section 13...Character string output section 20…Text conversion section 21...Character recognition processing section 30...Target string generation section 31…Natural language processing model
Claims
1. A character string indicating the content of an item to be described in a document is a target character string, the image data of the document is document image data, the electronic file including the document image data is an electronic document file, a text conversion unit configured to convert a character string in the document image data included in the electronic document file into text data using a character recognition process; A natural language processing unit configured to generate the target character string from the text data converted by the text conversion unit using a natural language processing model. A document reading system comprising:
2. a text data processing unit configured to process the text data into a format suitable for generating the target character string by the natural language processing model, and applying the text data processed by the text data processing unit to the natural language processing model.
2. The document reading system according to claim 1.
3. The text data processing unit processes the text data differently for each type of document.
3. The document reading system according to claim 2.
4. A character string indicating the content of an item to be described in a document is a target character string, the image data of the document is document image data, the electronic file including the document image data is an electronic document file, a text conversion unit converting a character string in the document image data included in the electronic document file into text data using a character recognition process; A natural language processing unit generates the target character string from the text data converted by the text conversion unit using a natural language processing model. A document reading method comprising:
5. A character string indicating the content of an item to be described in a document is a target character string, the image data of the document is document image data, the electronic file including the document image data is an electronic document file, The control unit obtaining the text data from a text conversion unit configured to convert a character string in the document image data included in the electronic document file into text data using a character recognition process; causing a natural language processing unit configured to generate the target character string from the text data using a natural language processing model to generate the target character string from the acquired text data; A document reading method comprising:
Citation Information
Patent Citations
Document reading system
JP2020181369A