Information processing device, information processing method, and program
The image processing device uses a neural network model to generate and adjust tokens from text data of consecutive pages, enhancing document segmentation accuracy by considering semantic relationships and continuity, addressing the limitations of existing methods.
Patent Information
- Application Number
- JP2021178618
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-11-01
- Publication Date
- 2026-03-03
- Estimated Expiration
- 2041-11-01
AI Technical Summary
Existing methods for segmenting scanned documents based on word frequency fail to consider the order of words, interdependencies, and contextual meaning, leading to inaccurate document segmentation, particularly when documents share a common template with varying names.
An image processing device employs a neural network model to determine document break positions by generating tokens from text data of consecutive page pairs, adjusting them to fit the model, and using a delimiter to accurately divide scanned images into document units, considering semantic relationships and continuity.
The method accurately segments scanned images into individual documents, improving segmentation accuracy by accounting for contextual information and reducing errors in documents with similar templates.
Smart Images

Figure 0007822746000001 
Figure 0007822746000002 
Figure 0007822746000003
Abstract
Description
[Technical Field]
[0001] The technology disclosed herein relates to a technology for dividing scanned images obtained by collectively scanning a plurality of documents into document units. [Background technology]
[0002] In recent years, digitalization has been promoted in various business fields, and one example of this is the paperless (digitalization) of various documents used in business, such as approval documents. When digitizing documents, a process of reading printed documents using a scanner or the like and converting them into image data is required. Since scanning a large number of forms one by one is inefficient during this process, a process of reading all pages across multiple documents at once and automatically dividing the obtained scanned image data into image data for each document is performed. Patent Document 1 discloses a method for automatically dividing a document by calculating the frequency of occurrence of each word from the text information of two consecutive pages, calculating the text similarity between the two pages, and dividing the document between pages with low similarity. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2002-312385 [Non-patent literature]
[0004] [Non-Patent Document 1] “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding” by Jacob Devlin, Ming-Wei Chang, Kenton Lee, Kristina Toutanova 2018 Summary of the Invention [Problem to be solved by the invention]
[0005] The above-mentioned method that uses the frequency of occurrence of words is known to have limitations in segmentation accuracy because it does not take into account the order in which words appear, the interdependencies between words, or the meaning and content of the context. For example, in the case of the method of Patent Document 1, if two documents created using the same template differ only in the names of people or organizations in the documents, the similarity between the pages of the two documents will be high, which may result in incorrect segmentation positions. [Means for solving the problem]
[0006] The image processing device according to the present disclosure is an image processing device that divides read image data consisting of a plurality of page images obtained by collectively reading a plurality of documents page by page into image data for each document, and includes a generating means that performs character recognition processing on the page images to generate text data, a determining means that sequentially acquires pairs of consecutive page images from the plurality of page images and determines document break positions based on the text data of the two page images that make up the pair, and a dividing means that divides the read image data at the break positions determined by the determining means, wherein the determining means has a neural network model and generates tokens obtained by breaking down the text of each of the two page images that make up the pair. The tokens were generated by converting the tokens obtained by performing an adjustment process to match the specifications of the neural network model. The vector is input to the neural network model, and the delimiter position is determined using a score output from the neural network that numerically represents the likelihood that the two page images constituting the pair belong to different documents. the adjustment process includes a reduction process for reducing the number of tokens so that they can be input to the neural network model, and the reduction process is a process for discarding some of the tokens obtained by decomposing the text of each page, and for the two page images constituting the pair, extracts only tokens corresponding to the text in the upper and lower regions of the page image for the previous page, and extracts only tokens corresponding to the text in the upper region of the page image for the next page. It is characterized by: [Effects of the Invention]
[0007] According to the technology of the present disclosure, scanned image data obtained by collectively scanning a plurality of documents can be accurately divided into image data for each document. [Brief explanation of the drawings]
[0008] [Figure 1] FIG. 1 is a block diagram showing the hardware configuration of an MFP. [Figure 2] FIG. 3 is a functional block diagram showing details of an image processing unit. [Figure 3] 10 is a flowchart showing the flow of image division processing. [Figure 4] A diagram showing an overview of a neural network model. [Figure 5] Conceptual diagram of segmentation detection using BERT. [Figure 6] 10 is a flowchart showing details of a token adjustment process. [Figure 7] 10A and 10B are diagrams illustrating a token extraction range. DETAILED DESCRIPTION OF THE INVENTION
[0009] The following describes embodiments of the present invention with reference to the drawings. Note that the following embodiments do not limit the scope of the invention as claimed, and not all of the combinations of features described in the embodiments are necessarily essential to the solution of the invention.
[0010] [Embodiment 1] <Hardware configuration of information processing device> 1 is a block diagram showing the hardware configuration of an MFP as an information processing device having a function for automatically dividing multiple documents according to this embodiment. The MFP 100 includes a CPU 101, a ROM 102, a RAM 103, a mass storage device 104, a UI unit 105, an image processing unit 106, an engine interface (I / F) 107, a network I / F 108, and a scanner I / F 109. These units are connected to each other via a system bus 110. The MFP 100 also includes a printer engine 111 and a scanner unit 112. The printer engine 111 and the scanner unit 112 are connected to the system bus 110 via the engine I / F 107 and the scanner I / F 109, respectively. The image processing unit 106 may be configured as an image processing device (image processing controller) independent of the MFP 100.
[0011] The CPU 101 controls the overall operation of the MFP 100. The CPU 101 executes various processes, described below, by loading programs stored in the ROM 102 into the RAM 103 and executing them. The ROM 102 is a read-only memory that stores a system startup program, a program for controlling the printer engine, character data, character code information, and the like. The RAM 103 is a volatile random access memory that is used as a work area for the CPU 101 and a temporary storage area for various data. For example, the RAM 103 is used as a storage area for storing font data additionally registered by downloading, image files received from an external device, and the like. The mass storage device 104 is, for example, an HDD or SSD, and is used to spool various data and store programs, various tables, information files, image data, and the like, as well as a work area.
[0012] The UI (user interface) unit 105 is configured, for example, with a liquid crystal display (LCD) equipped with a touch panel function, and displays the setting status of the MFP 100, the status of ongoing processing, error status, etc. It is also used to display the results of the process of dividing scanned image data of multiple documents into image data for each document, according to the present disclosure. The UI unit 105 also accepts various user instructions, such as input of values for various settings of the MFP 100 and selection of various buttons. For example, an instruction to execute a scan to read multiple documents collectively is also given via the UI unit 105. In this embodiment, the multiple documents to be scanned are assumed to be a collection (document group) of multiple types of documents with different formats, such as a stack of documents bound in a binder. The user issues an instruction to execute a scan after placing the stack of documents in an ADF (auto document feeder), not shown. The UI unit 105 may also be provided with a separate input device, such as a hard key.
[0013] When multiple documents are scanned collectively by the scanner unit 112, the image processing unit 106 performs a process of dividing the scanned image data into image data for each document. The image processing unit 106 also performs other image processing, such as generating print image data for the printer engine 111 from PDL data input from an external device. Details of the image processing unit 106 will be described later with reference to FIG. 2. Although the CPU 101 and the image processing unit 106 are shown separately in FIG. 1, in this embodiment, the CPU 101 functions as the image processing unit 106 by executing a program. However, this is not limiting, and the image processing unit 106 may be realized by a dedicated circuit such as a GPU or ASIC.
[0014] The engine I / F 107 functions as an interface for controlling the printer engine 111 in response to instructions from the CPU 101 when printing is performed. Engine control commands and the like are transmitted and received between the CPU 101 and the printer engine 111 via the engine I / F 107. The network I / F 108 functions as an interface for connecting the MFP 100 to a network 113. The network 108 may be, for example, a LAN or a public switched telephone network (PSTN). The printer engine 111 forms a multicolor image on a recording medium such as paper using color materials (toners in this case) of multiple colors (four colors: CMYK in this case) based on print image data received from the system bus 110. The scanner I / F 109 functions as an interface for controlling the scanner unit 112 in response to instructions from the CPU 101 when the scanner unit 112 reads a document. Scanner unit control commands and the like are transmitted and received between the CPU 101 and the scanner unit 112 via the scanner I / F 109. The scanner unit 112 optically reads a document stack set in the ADF page by page and generates read image data under the control of the CPU 101. The generated read image data (scanned image data) is transmitted to the mass storage device 104 via the scanner I / F 109.
[0015] <Details of the image processing unit> 2 is a functional block diagram showing the details of the image processing unit 106, and in particular shows only the functions related to the process of dividing the scanned image data of a document batch into image data for each document (hereinafter referred to as "image division process"). As shown in FIG. 2, the image processing unit 106 has a text conversion unit 210, a segmentation determination unit 220, and an image division unit 230.
[0016] The text conversion unit 210 converts the page image data of the acquired page pair into text data that can be processed using language. The text conversion unit 210 is composed of an OCR processing unit 211 that performs OCR (Optical Character Recognition) processing on character areas in the page image, and a text data generation unit 212 that generates text data representing character strings and sentences in the page based on the results of the OCR processing.
[0017] The delimiter determination unit 220 determines document delimiter positions (boundaries between documents) in the scanned image data of multiple documents based on the text data generated by the text conversion unit 210 in page units. The delimiter determination unit 220 is composed of four processing units: a text decomposition unit 221, a token adjustment unit 222, a vectorization unit 223, and a delimiter position determination unit 224. The text decomposition unit 221, commonly called a tokenizer, performs a process of decomposing text into tokens. Here, a token is the smallest unit of linguistic information to be input to a neural network model. The token adjustment unit 222 performs a process of adjusting the decomposed tokens so that they are suitable for the neural network model. The vectorization unit 223 performs a process of converting the adjusted tokens into vectors in a format that can be input to the neural network model. The delimiter position determination unit 224 inputs the vectors generated by the vectorization unit 223 into the neural network model to determine document delimiter positions in the scanned image data of multiple documents.
[0018] Based on the determination result of the division determination unit 220, the image division unit 230 divides the input scanned image data of multiple documents into image data for each document.
[0019] 2 are realized by the CPU 101 executing a program stored in the ROM 103 or a program such as an application loaded from the mass storage device 104 to the RAM 102.
[0020] <Image division processing> Fig. 3 is a flowchart showing the flow of image segmentation processing by the image processing unit 106 according to this embodiment. The image segmentation processing will be described in detail below with reference to the flowchart in Fig. 3. In the following description, the symbol "S" means step.
[0021] In S301, scanned image data obtained by scanning a document stack page by page with the scanner unit 112 is read from the mass storage device 104 and acquired as image data to be processed. Note that the scanned image data obtained by the scanner unit 112 may also be acquired directly via the system bus 111.
[0022] In S302, a combination of a page image of interest and its next page image, i.e., a pair of two consecutive page images (hereinafter referred to as a "page pair"), is obtained from the input scanned image data. All page images in the input scanned image data are selected as pages of interest, starting from the first page, such as a pair of pages 1 and 2, a pair of pages 2 and 3, and so on, and page pairs are obtained in order. Here, since the scanned image data is obtained by scanning a batch of documents, there are two patterns: one where the two page images constituting a page pair belong to the same document, and one where they belong to different documents. If both page images belong to the same document, the page pair is not a boundary between documents, so the previous page and the next page in the page pair are not divided. On the other hand, if the two page images belong to different documents, the page pair is a boundary between documents, so the previous page and the next page in the page pair are divided.
[0023] In S303, the text conversion unit 210 converts the image data of the two pages that make up the page pair acquired in S302 into text data. The specific procedure is as follows.
[0024] <OCR processing> First, the OCR processor 211 performs OCR processing on the two page image data. During OCR processing, a so-called block selection process is performed on each of the two page images to extract character blocks. Then, character recognition processing is performed on each extracted character block, and a specific character code is assigned to each character in the character block.
[0025] <<Generating text data>> Next, the text data generation unit 212 generates text data for each of the pages constituting the page pair based on the results of the OCR processing. The text data is sentence data that compiles characters present on the page into one, and is obtained by combining characters recognized from the same page image. When combining characters, they may be combined in the order in which they were recognized during the OCR processing, or they may be combined in the order in which humans read characters from left to right for each line. Furthermore, if there is a space between characters, they may be combined by inserting a space character or a symbol indicating a space. Furthermore, when characters are combined after a line break in the document, they may be combined by inserting a space character or a symbol indicating a line break. Furthermore, characters that are clearly misrecognized due to blurring or the like that occurred during reading may be removed.
[0026] Returning to the explanation of the flowchart in FIG.
[0027] Steps S304 to S310 are processes performed by the break determination unit 220 to determine whether it is appropriate to divide the scanned image data between both pages of a page pair (i.e., between the previous and next pages). A neural network model obtained by learning is used for this break determination. More specifically, a neural network model is used in which a unique discrimination layer (unique layer), such as a fully connected layer, is added to a general-purpose natural language processing model that has been pre-trained using a large amount of text data. FIG. 4 is a diagram showing an overview of the neural network model according to this embodiment. In FIG. 4, reference numeral 401 denotes input text data, and reference numeral 402 denotes the pre-trained natural language processing model. The natural language processing model 402 according to this embodiment has a structure in which neurons are bidirectionally connected. However, a natural language processing model in which neurons are unidirectionally connected, such as GPT, may also be used. Reference numeral 403 denotes a vector output from the natural language processing model 402. Reference numeral 404 denotes a unique layer connected to the vector output from the natural language processing model 402. Examples of pre-trained natural language processing models include BERT (Bidirectional Encoder Representations from Transformers) and XLNet. Note that instead of using publicly available models such as those described above, a model that has been pre-trained from scratch may also be used. Furthermore, the model does not necessarily have to have a transformer-based structure, as long as it is a pre-trained, highly accurate natural language processing model. For example, a model with a uniquely designed structure or a model with a structure automatically designed using AUTOML or the like may also be used. The following description will be given taking as an example a case where BERT is used as a pre-trained natural language processing model.
[0028] FIG. 5 is a conceptual diagram of segmentation detection using BERT according to this embodiment. In FIG. 5, BERT 500 receives text 501 and 502 from two page images of a page pair as input and outputs a 768-dimensional vector 503 representing the similarity, semantic relationship, and continuity between the two texts. In this embodiment, the 768-dimensional vector 503 output from BERT 500 is input to a unique layer 504 to obtain a solution to the binary classification problem, i.e., a determination result as to whether the target page pair corresponds to a "document segment." As a pre-processing step, a unique layer 504 is added to the output layer of BERT 500, and fine-tuning is performed to adjust parameters for a segmentation detection model that uses the two text data of the page pair as input. Fine-tuning is a learning technique that uses neural network parameters obtained in pre-training as initial values and fine-tunes them to parameters specific to the problem being solved. This fine-tuning is known to be a highly effective technique for improving the accuracy of language processing. In this embodiment, the text data of each of the two page images of a page pair is input to a model (trained model) obtained by performing fine tuning on BERT 500 for the purpose of determining breaks, and document breaks are determined. This makes it possible to determine breaks that take into consideration not only the similarity of the text in both page images, but also the semantic relationship and continuity between the text. Specific procedures in the break determination unit 220 are described below.
[0029] <<Text Decomposition>> First, in S304, the text decomposition unit 221 decomposes the text of each of the two page images of the page pair into tokens. If the text is in Japanese, it will be decomposed into words (character strings) consisting of one or more characters, as indicated by reference numeral 505 in FIG. 5. If the text is in a European language such as English in which words are separated by spaces, it will be decomposed into words, their prefixes, suffixes, etc. When decomposing the text into tokens, it is desirable to use the same tokenizer as that used in the pre-trained natural language processing model (here, BERT).
[0030] In the next step S305, the token adjustment unit 222 adjusts the tokens for both pages obtained in S304 so that they conform to the specifications of the neural network model to be used. FIG. 6 is a flowchart showing details of the token adjustment process according to this embodiment. In the case of BERT, the input vectors required are a vector obtained by converting the tokens into numerical values, a vector indicating the start position of the token on the subsequent page, and a vector indicating the start position of the padding token. The upper limit of the number of tokens (number of words) that can be input is 512, so it is not always possible to input tokens corresponding to all the text present on a page. In addition, it is necessary to input a continuous vector that combines the vector of the token corresponding to the text on the previous page and the vector of the token corresponding to the text on the subsequent page. The token adjustment process performs these necessary processes. A detailed description will be given below with reference to the flowchart in FIG. 6.
[0031] <Token Adjustment Process Details> First, in S601, it is determined whether the total number of tokens obtained by decomposing the text of the previous page and the text of the next page is 510 or more, and processing is assigned based on the determination result. Here, the reason for performing threshold processing using 510, which is two less than the upper limit of 512, as the standard is to ensure there are enough tokens for the special tokens to be assigned in S603, which will be described later. If the total number of tokens on both pages is 510 or more, the process proceeds to S602, and if it is less than 510, the process proceeds to S603.
[0032] In S602, a process is performed to discard some of the tokens obtained by breaking down one page of text. This truncation process reduces the total number of tokens to 509 or less. Here, we will explain a method for reducing the total number of tokens to 509 or less by extracting only tokens that are included in the specified ranges of the previous and next pages and discarding the remaining tokens.
[0033] Now, let's say that the token extraction range for both pages is set to extend from the first token to the 256th token. In this case, it would be impossible to input tokens that correspond to the text at the bottom of each page. This would result in a loss of information about the continuity of the text between the previous and next pages, increasing the likelihood of incorrect inference. On the other hand, let's say that the token extraction range for the previous page is set to extend from the last token back to the 256th token. In this case, it would be impossible to input tokens that correspond to the text at the top of the previous page. This would result in a loss of information such as the title or heading, again increasing the likelihood of incorrect inference. Therefore, we use one of the two extraction patterns shown below to extract only tokens that are effective for determining document boundaries from each of the previous and next pages of the page pair.
[0034] <Extraction pattern 1> In the first extraction pattern, as shown in FIG. 7( a), the token extraction range for the previous page is set to the upper region 701 and lower region 702 of the page image, and the token extraction range for the next page is set to the upper region 703 of the page image. This means that for the previous page, only tokens 1 through 127 counting from the top of the page and tokens 1 through 127 counting from the bottom of the page are extracted, and all other tokens are truncated. For the next page, only tokens 1 through 255 counting from the top of the page are extracted, and all other tokens are truncated. In this case, the total number of tokens after truncation is 508. This truncation process allows the input data to include information for determining the similarity of header information, such as titles and headings, commonly written at the top of pages, and the continuity of text, i.e., whether the text flows naturally across pages. In this way, only tokens within the specified range are extracted for each page, and tokens outside the specified range are truncated, so that the total number of tokens is 509 or less. In the above example, the same number of tokens are extracted from the top and bottom areas of the previous page, but they may be different. Also, the same number of tokens are extracted from the previous page and the next page, but they may be different between pages. In other words, tokens are extracted from the top and bottom areas of the previous page and the top area of the next page, and the total number of tokens should be 509 or less.
[0035] <Extraction pattern 2> In the second extraction pattern, as shown in Figure 7(b), the token extraction range for the previous page is set to the top region 711 and bottom region 712 of the page, and the token extraction range for the next page is set to the bottom region 713 and bottom region 714 of the page. This means that for both the previous and next pages, only tokens 1 to 127 counting from the top of the page and tokens 1 to 127 counting from the bottom of the page are extracted, and all other tokens are truncated. In this case, as with extraction pattern 1, the total number of tokens after truncation is 508. In this case, the input data can further include information for determining the continuity of the page numbers commonly written at the bottom of the pages and the similarity of the text written as footer information, making it possible to more accurately determine breaks.
[0036] As described above, only tokens corresponding to text within a predetermined range are extracted, and other tokens are discarded, so that the total number of tokens is kept below the number of tokens that can be input to the neural network model.
[0037] In the next step S603, a predetermined special token is assigned to the sequence of tokens extracted from each of the previous and following pages (hereinafter referred to as the "token array"). In the specific example shown in FIG. 5, the "[cls]" token 505a placed at the beginning of the token array for the previous page and the "[sep]" token 505b placed at the beginning of the token array for the following page correspond to these special tokens. Note that the "##" in the "##△□" token 505c on the previous page is a subtoken indicating that it is connected to the immediately preceding token (here, the "〇×" token).
[0038] In the next step S604, the token array of the previous page and the token array of the next page, each with a special token added, are combined to generate a single token array (hereinafter referred to as the "combined token array") in which tokens corresponding to the text on both pages are consecutive.
[0039] In the next step S605, in order to fix the length of the input vector to the neural network model (here, 512 tokens), it is determined whether the total number of tokens constituting the combined token array is 512. If the result of the determination shows that the number of tokens is less than 512, the process proceeds to S606, and if the total number of tokens is 512, the process ends.
[0040] In S606, padding tokens are added to the end of the combined token array to make up for any tokens that are insufficient to fit the fixed length.
[0041] This completes the token adjustment process, resulting in a token sequence corresponding to the text of the page pair that is suitable for the neural network model.
[0042] Returning to the explanation of the flowchart in FIG.
[0043] In S306, the token sequence obtained by the token adjustment process is converted into a numerical vector using a token conversion dictionary specified by BERT. In the specific example shown in Figure 5, the numerical values 506, such as "2," "27941," and "3999," associated with each token indicate the converted numerical vector. Furthermore, vectors (not shown) representing the start positions of tokens on the next page and vectors (not shown) representing the start positions of padding tokens are assigned to predetermined positions. These vectors assigned to predetermined positions enable BERT to reduce processing time and efficiently update the weights of its internal attention mechanism. In this way, an input vector in a format suitable for a neural network model is obtained.
[0044] In the next step S307, the break position determination unit 224 inputs the input vector generated in S306 into the neural network model to derive a break determination score. This break determination score is a numerical representation of the likelihood that the two page images in the page pair belong to different documents. When deriving this break determination score, the output value of the neural network model may be used as is, or a value obtained by applying an activation function such as a softmax function or a sigmoid function to the output value may be used.
[0045] In the next step S308, it is determined whether the break determination score derived in S307 is equal to or greater than a predetermined threshold. This threshold is used to determine whether the two page images constituting the page pair being processed belong to the same document or to different documents. Therefore, the threshold here may be, for example, an intermediate value among values that can be derived as the break determination score, or a value that provides the highest discrimination accuracy during fine tuning. If the derived score is equal to or greater than the threshold, the process proceeds to S309; if it is less than the threshold, the process proceeds to S310. In S309, the break position determination unit 224 determines that the space between the two page images constituting the page pair being processed corresponds to a break position of the document, and stores the page number of the previous page, which is the page of interest, in RAM 103 as the page number of the break page.
[0046] In the next step S310, it is determined whether the above-described process has been performed on all page pairs in the input scanned image data. If the result of the determination is that there are unprocessed page pairs remaining, the process returns to S302, where a page pair with the current page after it as the next target page (previous page) is acquired, and the process continues. On the other hand, if all page pairs have been processed, the process proceeds to S311.
[0047] Then, in S311, the image dividing unit 230 divides the input scanned image data into document units based on the page number information saved in S309. For example, suppose the scanned image data is made up of page images for 10 pages, and the break determination unit 220 has saved "3" and "7" as the page numbers of the break pages. In this case, the scanned image data is divided into three pieces: image data for pages 1 to 3, image data for pages 4 to 7, and image data for pages 8 to 10. The document-unit image data thus divided is output to the host PC 115 or the like via the network I / F 108.
[0048] <Variation 1> In the token adjustment process (S602) in the flow of Fig. 6 described above, a method of extracting only tokens included in a specified range of each page was explained as a method of discarding some tokens. Instead of this, the number of tokens may be reduced by applying, for example, the following methods 1) or 2).
[0049] 1) Reducing the number of tokens by shortening the text using summarization techniques 2) Perform morphological analysis on the text and extract only tokens corresponding to specific parts of speech Below, we will briefly explain each method.
[0050] When performing token reduction processing using the above method 1), the token adjustment unit 222 has a text summarization function. If the total number of tokens on both pages is 510 or more, the text on both pages generated by the text data generation unit 212 is summarized. Specifically, for each of the previous and following pages, a short text that retains only the main points is created from the entire text. For example, if the target of reading is a report, text data of 1,000 characters or more containing details of the background, issues, and results is converted into text data of approximately 200 characters, each of which summarizes the background, issues, and conclusion in one sentence, allowing for an overview of the page. The summarized text data is then returned to the tokenization unit 221, which again decomposes the summarized text into tokens. The tokens corresponding to the resulting summarized text are subjected to processing from S603 onward and converted into vectors suitable for input to a neural network model. The determination processing of S601 may be performed again on the summarized text to confirm that the total number of tokens on both pages is less than 510 before proceeding to S603. If there are 510 or more, the summarized text may be further summarized, or the above-mentioned method of extracting tokens from a specified range within a page may be applied to the summarized text.
[0051] The token adjustment unit 222, which performs the token reduction process using method 2) above, has a morphological analysis function. If the total number of tokens on both pages is 510 or more, morphological analysis is performed on the tokens on both pages. Morphological analysis technology divides a sentence written in a natural language into the smallest units of meaning in the language (morphemes) and determines the parts of speech and conjugations of each. In the case of Japanese, this is achieved by a morphological analysis engine such as MeCab. This makes it possible, for example, to extract only verbs and nouns from the text and remove words such as particles and conjunctions. Only the tokens corresponding to the nouns and verbs obtained in this way are extracted, and the process from step S603 onward is performed to convert them into vectors that can be input to a neural network model. It is sufficient to keep the total number of tokens below 510, and the parts of speech to be retained can be changed depending on the number of tokens on each page.
[0052] As a variation other than the above, for example, only tokens corresponding to specific character strings such as the title or author name may be extracted from the text. Furthermore, truncation processing using all of the above methods may be performed during fine tuning, and the method with the highest accuracy may be applied during inference. Alternatively, image analysis may be performed on each page image of the scanned image data to automatically determine the truncation method. For example, if a title or header information is detected by image analysis, the token extraction range may be from the first token to a certain number of tokens, and if a page number or footer information is detected, the token extraction range may be from the last token to a certain number of tokens.
[0053] <Variation 2> In the flowchart of FIG. 3 described above, the boundary between two page images constituting a page pair for which the boundary determination score is equal to or greater than a threshold is determined as the boundary position. In this case, if a score equal to or greater than the threshold is not derived, the scanned image data is not divided into document-based image data. This is not a problem if the scanned image data corresponds to a single document, but it can be problematic if, for some reason, the scanned image data of multiple documents does not yield a score equal to or greater than the threshold. Therefore, for example, if the number of scanned documents is known in advance, the boundary determination scores for all page pairs are temporarily stored. Then, the boundary positions may be determined and the page pairs with the highest scores, calculated by subtracting 1 from the number of scanned documents, may be divided accordingly. For example, if it is known that 10 documents have been scanned, it is possible to divide the scanned image data between the pages of the page pairs with the highest scores.
[0054] <Variation 3> In the above-described embodiment, the MFP 100 in FIG. 1 is described as including the image processing unit 106, but the present invention is not limited to this configuration. The image processing unit 106 in FIG. 2 may be implemented on an image processing system such as a server device or a virtual server on cloud computing. In this case, for example, the MFP 100 transmits a document image scanned by a scanner unit to a server via a network. The server then performs OCR processing, segmentation determination processing, and image segmentation processing on the received document image using the image processing unit 106 within the server. The server then transmits the results of the segmentation processing to a host PC via the network (or via the MFP 100).
[0055] Note that instead of implementing all of the processing units 210 to 230 of image processing unit 106 in FIG. 2 on a server or cloud, some of the components may be implemented on a server or cloud. For example, MFP 100 may include text generation unit 210 and image division unit 230, and the server or cloud may include division determination unit 220. In this case, MFP 100 performs OCR processing on each page of the scanned document image to generate text data, and transmits the text data to a division determination unit of the server or cloud while also storing the image data. The division determination unit of the server or cloud that receives the text data then performs the processes of text decomposition unit 221 to division position determination unit 224 and returns information on the determined division positions to the MFP. The MFP then divides the stored image data based on the received division position information and outputs the division results.
[0056] As described above, according to this embodiment and its modified examples, scanned image data obtained by collectively scanning a plurality of documents can be accurately divided into image data for each document.
[0057] (Other Examples) The present invention can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program.The present invention can also be realized by a circuit (e.g., ASIC) that realizes one or more functions.
Claims
1. An image processing device that divides scanned image data consisting of a plurality of page images obtained by collectively scanning a plurality of documents on a page-by-page basis into image data on a document-by-document basis, generating means for performing character recognition processing on the page image to generate text data; a determining means for sequentially obtaining pairs of consecutive page images from the plurality of page images and determining document break positions based on text data of the two page images that make up the pairs; a dividing means for dividing the read image data at the division positions determined by the determining means; and The determining means It has a neural network model, The tokens obtained by decomposing the text of each of the two page images constituting the pair are subjected to an adjustment process to conform to the specifications of the neural network model, and the resulting tokens are converted to generate vectors, which are then input to the neural network model; determining the delimiting position using a score that is a numerical representation of the likelihood that the two page images constituting the pair output from the neural network belong to different documents; The adjustment process includes a reduction process for reducing the number of tokens so that they can be input to the neural network model; The reduction process is a process of discarding some of the tokens obtained by decomposing the text of each page, and is a process of extracting only tokens corresponding to the text in the upper and lower regions of the page image for the previous page of the two page images constituting the pair, and extracting only tokens corresponding to the text in the upper region of the page image for the next page.
1. An image processing device comprising:
2. An image processing device that divides scanned image data consisting of a plurality of page images obtained by collectively scanning a plurality of documents on a page-by-page basis into image data on a document-by-document basis, generating means for performing character recognition processing on the page image to generate text data; a determining means for sequentially obtaining pairs of consecutive page images from the plurality of page images and determining document break positions based on text data of the two page images that make up the pairs; a dividing means for dividing the read image data at the division positions determined by the determining means; and The determining means It has a neural network model, The tokens obtained by decomposing the text of each of the two page images constituting the pair are subjected to an adjustment process to conform to the specifications of the neural network model, and the resulting tokens are converted to generate vectors, which are then input to the neural network model; determining the delimiting position using a score that is a numerical representation of the likelihood that the two page images constituting the pair output from the neural network belong to different documents; The adjustment process includes a reduction process for reducing the number of tokens so that they can be input to the neural network model; The reduction process is a process of summarizing the text of each of the two page images constituting the pair and shortening the text to be decomposed, or a process of performing morphological analysis on the text of each of the two page images constituting the pair and extracting only tokens corresponding to specific parts of speech.
1. An image processing device comprising:
3. 3. The image processing apparatus according to claim 1, wherein the determining unit determines, as the delimiter position, a position between two page images constituting the pair whose scores are equal to or greater than a threshold value.
4. The image processing device according to claim 1 or 2, characterized in that, when the number of the plurality of documents is known in advance, the determination means determines the boundary position between two page images constituting a pair of pairs in a number obtained by subtracting one from the number of the plurality of documents, in descending order of the output score.
5. 5. The image processing device according to claim 1, wherein the neural network model has a unique discriminant layer added to a pre-trained natural language processing model, and is fine-tuned for the purpose of determining the delimiter positions.
6. 6. The image processing device according to claim 5, wherein the pre-trained natural language processing model is BERT (Bidirectional Encoder Representations from Transformers).
7. 7. The image processing apparatus according to claim 1, further comprising a scanner unit for reading the plurality of documents to obtain the plurality of page images.
8. 8. The image processing device according to claim 1, wherein the image processing device is a server device.
9. The image processing device according to claim 1 , wherein the image processing device is a virtual server provided by cloud computing.
Citation Information
Patent Citations
Image acquisition device and method, and computer- readable recording medium recorded with image acquisition processing program
JP2002024258A
Document automated dividing device
JP2002312385A
Method and program for generating genre model for identifying document genre, method and program for identifying document genre, and image processing system
JP2011018316A
Image reading apparatus, learning apparatus, method, and program
JP2021057710A
Information processing system, program, and information processing method
JP2021086480A