Image processing device, image processing method, and non-transitory computer-readable medium

By using the learning model of the generative adversarial network, the image processing device can effectively remove grid lines and unwanted marks, solve the problems of grid line interference and mark influence, and achieve more accurate text recognition.

CN111797667BActive Publication Date: 2025-09-16FUJIFILM BUSINESS INNOVATION CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202010081648.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-04-01
Filing Date
2020-02-06
Publication Date
2025-09-16
Estimated Expiration
2040-02-06

AI Technical Summary

Technical Problem

In the prior art, when an image processing device recognizes information formed on a form, the grid frame and grid lines interfere with the text recognition process, and unwanted marks may be superimposed on the sheet under the transfer member, resulting in degraded image quality and affecting text recognition results.

Method used

An image processing device comprising first and second image generators uses a learning model (such as a generative adversarial network) to remove grid lines and unwanted markings, generating a clear recording image. The first image generator generates a first image comprising the grid lines and the recording image, while the second image generator uses the learning model to remove the remaining image and output accurate recording information.

Benefits of technology

Even when there are grid lines and unexpected marks in the image, the information entered by the user can be accurately extracted, improving the precision and accuracy of text recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111797667B_ABST
    Figure CN111797667B_ABST
Patent Text Reader

Abstract

Image processing apparatus, image processing method, and non-transitory computer-readable medium. An image processing apparatus includes a first image generator and a second image generator. The first image generator generates a first image including a predetermined grid image and a recording image based on a second sheet in a sheet group. The sheet group is obtained by stacking a plurality of sheets including a single first sheet and a second sheet. The first sheet has recording information recorded on the first sheet. The second sheet has a recording image corresponding to the recording information transferred to the second sheet and includes a grid image. The second image generator generates a second image by removing a residual image from the first image generated by the first image generator based on a learning model, the learning model having learned to remove the residual image that is different from the grid image and the recording image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to an image processing apparatus, an image processing method, and a non-transitory computer-readable medium. Background Art

[0002] An image processing device known in the prior art recognizes information formed on a form (e.g., text entered on a form based on a form image) from a form image. The form has a grid frame and grid lines in advance according to the entered content, and these grid frames and grid lines interfere with the text recognition process. As a technology that takes grid frames and grid lines into consideration, for example, Japanese Unexamined Patent Application Publication No. 2000-172780 discloses a form registration device. Specifically, when the format of a form is to be converted into data and registered, the form registration device removes the grid frame and grid lines from the read image of the form after recognizing the grid frame and grid lines, and recognizes the pre-printed text from the image from which the grid frame and grid lines have been removed.

[0003] In order to form the text entered by the user onto multiple sheets of a form, a transfer member (e.g., a carbon sheet) is sometimes inserted between the multiple sheets so that the text entered on the top sheet is transferred to the sheet below. However, one or more sheets (e.g., carbon sheets) located below the transfer member may sometimes have information other than the text entered by the user (e.g., undesirable marks, including scratches or wear marks) superimposed thereon. When obtaining text information by recognizing text from an image obtained as a result of scanning the sheet, undesirable marks (e.g., scratches or wear marks) become an obstacle. Therefore, there is room for improvement in recognizing text from an image on which undesirable marks (e.g., information other than the text entered by the user) are superimposed. Summary of the Invention

[0004] Therefore, an object of the present disclosure is to provide an image processing device, an image processing method, and a non-transitory computer-readable medium, which can extract an entry image corresponding to entry information entered by a user even when the image includes a residual image.

[0005] According to a first aspect of the present disclosure, there is provided an image processing device comprising a first image generator and a second image generator. The first image generator generates a first image comprising a predetermined grid image and a recording image based on a second sheet in a sheet group. The sheet group is obtained by stacking a plurality of sheets comprising a single first sheet and a second sheet. The first sheet has recording information recorded on the first sheet. The second sheet has a recording image corresponding to the recording information transferred to the second sheet and includes a grid image. The second image generator generates a second image by removing a residual image from the first image generated by the first image generator based on a learning model, wherein the learning model has learned to remove a residual image different from the grid image and the recording image.

[0006] According to a second aspect of the present disclosure, in an image processing device according to the first aspect, the learning model may be a model that has learned to generate an original image based on an input image based on a combination of an input image including a residual image and an original image not including a residual image corresponding to the input image.

[0007] According to a third aspect of the present disclosure, in the image processing apparatus according to the first aspect or the second aspect, the learning model may be a model generated as a result of learning by using a generative adversarial network.

[0008] According to a fourth aspect of the present disclosure, the image processing device according to any one of the first to third aspects may further include an output unit that outputs information representing a recognition result obtained as a result of recognizing the second image generated by the second image generator as entry information.

[0009] According to the fifth aspect of the present disclosure, in the image processing device according to the fourth aspect, the output unit can adjust the position of the second image based on the grid position information representing the predetermined position of the grid image on the sheet, so that the position of the grid image in the second image matches the position according to the grid position information.

[0010] According to the sixth aspect of the present disclosure, in the image processing device according to the fifth aspect, attribute information related to the grid frame representing the entry item to be entered in the area formed by the grid frame of the grid image can be pre-set, and wherein the output unit can identify the second image relative to the area formed by the grid frame, and output the recognition result of the area and the attribute information related to the grid frame corresponding to each other.

[0011] According to a seventh aspect of the present disclosure, in the image processing apparatus according to any one of the first to fifth aspects, the entry image may be a written text image.

[0012] According to an eighth aspect of the present disclosure, in the image processing apparatus according to any one of the first to seventh aspects, the sheet set may include a sheet coated with carbon on one side.

[0013] According to a ninth aspect of the present disclosure, in the image processing apparatus according to any one of the first to eighth aspects, the remaining image may be a mark image corresponding to mark information about at least one of a scratch and a wear mark.

[0014] According to a tenth aspect of the present disclosure, a non-transitory computer-readable medium storing a program causing a computer to execute a process is provided. The process includes: generating a first image including a predetermined ruled line image and a recorded image based on a second sheet in a sheet set, the sheet set being obtained by stacking a plurality of sheets including a single first sheet and a second sheet, the first sheet having recorded information recorded thereon, and the second sheet having a recorded image corresponding to the recorded information transferred thereto and including a ruled line image; and generating a second image by removing a residual image from the generated first image based on a learning model, the learning model having learned to remove the residual image that is different from the ruled line image and the recorded image.

[0015] According to the eleventh aspect of the present disclosure, there is provided an image processing method, which includes: generating a first image including a predetermined grid image and a recording image based on a second sheet in a sheet group, the sheet group being obtained by stacking a plurality of sheets including a single first sheet and a second sheet, the first sheet having recording information recorded thereon, the second sheet having a recording image corresponding to the recording information transferred thereto and including a grid image; and generating a second image by removing a residual image from the generated first image based on a learning model, the learning model having learned to remove a residual image different from the grid image and the recording image.

[0016] According to the first aspect, the tenth aspect, and the eleventh aspect of the present disclosure, even when an image includes a residual image, an entry image corresponding to entry information entered by a user can be extracted.

[0017] According to the second aspect of the present disclosure, a taken-in image can be extracted more accurately compared to a case where a model that has undergone a learning process is not used.

[0018] According to the third aspect of the present disclosure, an input image can be extracted more accurately compared to a case where a learning process using a generative adversarial network is not used.

[0019] According to the fourth aspect of the present disclosure, compared with a case where the output unit is not provided, entry information corresponding to the entry image can be output more accurately.

[0020] According to the fifth aspect of the present disclosure, compared with the case where the second image is generated without using the ruled line position information, the entry image corresponding to the entry information can be identified more accurately.

[0021] According to the sixth aspect of the present disclosure, compared with the case where the recognition result is output without considering the entry items in the ruled frame, the entry information can be output corresponding to the entry items in the ruled frame.

[0022] According to the seventh aspect of the present disclosure, even if the text is written by a user, the written text can be extracted.

[0023] According to the eighth aspect of the present disclosure, even in the case where the sheet set includes a sheet coated with carbon on one side, it is possible to extract a registration image corresponding to registration information.

[0024] According to the ninth aspect of the present disclosure, even in the case where the sheet has an undesired mark, it is possible to extract a registered image corresponding to the registered information. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Exemplary embodiments of the present disclosure will be described in detail based on the following drawings, in which:

[0026] Figure 1 is a block diagram illustrating a functional configuration of an image processing apparatus according to a first exemplary embodiment;

[0027] Figure 2 A learning process for learning a written text extraction learning model is shown;

[0028] Figure 3 A learning process of learning a ruled line extraction learning model is shown;

[0029] Figure 4 is a block diagram illustrating an example in which a learning processor is configured as a generative adversarial network (GAN);

[0030] Figure 5 FIG is an image diagram showing an example of a slip image;

[0031] Figure 6 is an image diagram showing an example of a written text image;

[0032] Figure 7 is an image diagram showing an example of a ruled line image;

[0033] Figure 8 is an image diagram showing an example of an image obtained by combining a written text image and a ruled line image;

[0034] Figure 9 A diagram related to recognition of an area within a grid frame in a text image;

[0035] Figure 10 is a block diagram illustrating an example in which an image processing apparatus includes a computer;

[0036] Figure 11 is a flowchart illustrating an example of an image processing flow according to the first exemplary embodiment;

[0037] Figure 12 is a block diagram illustrating a functional configuration of an image processing apparatus according to a second exemplary embodiment;

[0038] Figure 13 is a flowchart illustrating an example of an image processing flow according to the second exemplary embodiment; and

[0039] Figure 14 is a block diagram illustrating a functional configuration of an image processing apparatus according to a third exemplary embodiment. DETAILED DESCRIPTION

[0040] Hereinafter, exemplary embodiments of the present invention will be described in detail with reference to the accompanying drawings. In all drawings, components and processes having the same effects and functions will be given the same reference numerals, and redundant descriptions may be omitted where appropriate.

[0041] First Exemplary Embodiment

[0042] Figure 1 An example of the configuration of the image processing apparatus 1 according to the first exemplary embodiment of the present disclosure is shown.

[0043] In this exemplary embodiment, the disclosed technology is applied to image processing for recognizing text from an input image.

[0044] This exemplary embodiment relates to an example in which text recognition is performed by using an image reader (e.g., a scanner) to read any one of a plurality of sheets, the plurality of sheets including an upper sheet on which text or a mark is recorded and a lower sheet onto which the text or a mark is transferred by a transfer member (e.g., carbon). Hereinafter, a plurality of stacked sheets (sheet group) will be referred to as a "form". Hereinafter, each sheet included in the form will be referred to as a "document". In particular, a sheet on which someone has written can be referred to as an "original", and a sheet onto which the written content is transferred can be referred to as a "copy".

[0045] The image processing apparatus 1 according to this exemplary embodiment outputs a text code by recognizing a text image (eg, written text) included in a document image obtained as a result of an image reader (eg, a scanner) reading the document.

[0046] Although this exemplary embodiment relates to the case of using a sheet that has undergone a transfer process using a carbon sheet, the present disclosure is not limited to carbon copying. The present disclosure is also applicable to other copying technologies in which text or marks written on an original are transferred to a dark copy that reflects the level of pressure applied when writing the text or marks on the original.

[0047] Furthermore, the terms "transfer," "copy," and the like are not intended to be limiting. For example, the line thickness of the text or marking on the original may differ from the line thickness of the text or marking transferred to the copy. For example, when using carbon copying technology, the line thickness of the text or marking on the original depends on the line thickness of the writing tool itself, but the line thickness of the text or marking transferred to the copy depends on the force (e.g., writing pressure).

[0048] A copy obtained using a method such as carbon copying may have undesirable marks (such as stains, scratches, or abrasions) in addition to the written content on the original. A document image with undesirable marks (such as stains, scratches, or abrasions) obtained as a result of reading a document using an image reader (such as a scanner) is a degraded image compared to a document image that does not have such marks. Due to the degraded image quality, a standard recognizer may sometimes be unable to recognize a degraded image with undesirable marks (such as stains, scratches, or abrasions) as noise. The image processing device 1 according to this exemplary embodiment has a function of performing text recognition by removing noise from such an image degraded by undesirable marks (such as stains, scratches, or abrasions).

[0049] In this exemplary embodiment, it is conceivable that the noise caused by undesirable marks (such as stains, scratches or wear marks) (which is the cause of image degradation) is an image (residual image) that is different from the pre-formed grid lines and different from the text and marks (entry information) entered by the user. In addition to the text and marks (entry information) being entered by the user, the residual image is formed by the force acting on the transfer member (such as a carbon sheet). Therefore, the noise in this exemplary embodiment refers to an image that is different from the pre-formed grid lines and different from the text and marks (entry information) entered by the user. The term "straight lines" is not intended to be limited to the use of straightness or a ruler. For example, lines representing areas for providing different numbers of information in a certain form are within the range of "grid lines". Entry information is information entered by the user, and also represents an entry image (such as a text image) corresponding to information such as text information.

[0050] For example, if a first sheet and a second sheet are stacked, the input information is formed on both the first and second sheets, while the residual image can be considered to be information formed on the second sheet (an image based on that information) and facing forward, not on the first sheet. Specifically, the input information refers to information that a user wants to enter (such as an address and name). Therefore, for example, in the case of stacked sheets, the input information is present on the top sheet and is recorded on the underlying sheets via a transfer member. On the other hand, since scratches or wear marks do not involve the use of a writing instrument (such as a pen) but are recorded on one or more sheets below the top sheet via a transfer member, the residual image caused by the scratches or wear marks does not exist on the top sheet. Specifically, the input information is present on each sheet of the stack, but the residual image is not present on the top sheet of the stack, but on the second sheet and located forward, below the top sheet. In other words, the residual image remains as a trace based on the specific type of information, but does not exist as a color image on the first sheet, but as a color image on the second sheet and facing forward.

[0051] For example, if the writing tool used to enter data on a form is, for example, a pen, the entered information can be considered to be information printed with the pen, and the residual image can be considered to be information about content different from the information printed with the pen (an image based on the information). Specifically, since the entered information is information (such as an address or a name) entered by using a writing tool or a printer, it can be considered that the entered information is entered with a substantially uniform pressure. Therefore, the thickness and shape of the text as the entered information are substantially uniform in a single document. The same applies to the case where a printer is used. On the other hand, in the residual image formed by scratches or wear marks, the pressure changes, and the corresponding thickness and shape also change. A learning device (such as a neural network) is used to perform a learning process to determine whether the thickness and shape are uniform, and to distinguish the residual image and the entered information from each other. In another expression, the residual image can also be considered to be an image based on randomly entered information.

[0052] In addition, for example, considering the grid lines pre-formed on the sheet, the residual image can be considered as information (an image corresponding to the information) formed across the entry information and the grid lines. Specifically, the residual image formed by scratches or wear marks is different from the entry information entered by the user in that the residual image may sometimes exist across another entry information or already existing grid line information. In other words, information that overlaps with or covers the grid line information or entry information can be considered as a residual image. Only information that overlaps with the grid line information or entry information can be considered as a residual image. The thickness and shape of the entry image can be identified from the overlapping residual image, and an image with a thickness and shape similar to that of the image can be identified as a residual image. Even if the image does not cross the grid line information or entry information or does not overlap with the grid line information or entry information, the image can be identified as a residual image.

[0053] Furthermore, since input information is information entered by the user, a specific input image (e.g., a circular image or a check image) used to identify a product number or correction lines (elimination lines) used to correct errors in addresses or names can also be considered input information. Since a specific input image is also information entered on the first sheet of a stack of sheets in a form, and the thickness and shape entered using a writing tool such as a pen are the same type of information, the specific input image can be distinguished as input information.

[0054] In this exemplary embodiment, the residual image is removed from an image (first image) including a grid image and a recorded image, thereby generating an image (second image) including a separate grid image and a separate recorded image corresponding to the grid image and the recorded information. The recorded image can be considered to be composed of information including an image corresponding to the recorded information and the residual image. Then, it can be considered that an image (second image) including a separate grid image and a separate recorded image corresponding to the grid image and the recorded information is generated by removing the residual image from the image (first image) including the grid image and the recorded image including the residual image.

[0055] like Figure 1 As shown, the image processing device 1 includes an image input unit 2, a text and ruled line extractor 3, and a text recognizer 4. The text and ruled line extractor 3 includes a written text extractor 31, a written text extraction learning model 32, a ruled line extractor 33, and a ruled line extraction learning model 34. The text recognizer 4 includes a written text positioning unit 41, a registration form frame position information storage unit 42, a written frame position detector 43, a written text recognizer 44, and a written text recognition dictionary 45.

[0056] The image input unit 2 receives at least one document image and outputs the document image to the text and ruled line extractor 3. The text and ruled line extractor 3 extracts written text and ruled lines using the document image from the image input unit 2 and outputs the written text and ruled lines. In other words, the written text image and ruled line image are extracted as intermediate products from the document image. The text recognizer 4 recognizes the written text image formed on the document using the intermediate products (i.e., the written text image and ruled line image) generated by the text and ruled line extractor 3 and outputs the recognition result.

[0057] In detail, the image input unit 2 receives at least one document image. For example, the image input unit 2 receives a scanned image generated as a result of an image reader (e.g., a scanner) performing an image scan on each of a plurality of documents included in a form as a document image (e.g., Figure 5 document image shown).

[0058] exist Figure 5 In the example shown, a form used when sending an item is provided with a plurality of entry fields, in which the user enters information. Specifically, fields are provided for entering a postal code, a telephone number, an address, and a name as information regarding the shipping destination and information regarding the sender, fields are provided for entering a desired receipt date and a desired receipt time range as information related to the receipt of the item, and fields are provided for entering information indicating the contents of the item.

[0059] The text and ruled line extractor 3 extracts written text and ruled lines using the document image from the image input unit 2. Specifically, the written text extractor 31 extracts a written text image from the document image using a written text extraction learning model 32, and outputs the extracted written text image as an intermediate product (e.g., Figure 6 ). In addition, the ruled line extractor 33 extracts a ruled line image from the document image using the ruled line extraction learning model 34, and outputs the extracted ruled line image as an intermediate product (e.g., Figure 7 image shown).

[0060] The written text image extracted by the written text extractor 31 is an intermediate product generated (or estimated) based on the document image. The ruled line image extracted by the ruled line extractor 33 is also an intermediate product generated (or estimated) based on the document image. In other words, the text and ruled line extractor 33 generates (or estimates) a written text image and a ruled line image in the document image that does not include noise caused by undesirable marks (such as scratches or wear marks).

[0061] Next, the text and ruled line extractor 3 will be described in detail.

[0062] In the text and ruled line extractor 3, the written text extraction learning model 32 is a learning model that has undergone a learning process and is used to generate an entry image corresponding to the entry information including the written text formed on the document from the document image (i.e., an image obtained by reading the copies included in the multiple sheets constituting the form). The written text extraction learning model 32 is, for example, a model that defines a learned neural network and is represented as a set of information about the weights (strengths) of the connections between the nodes (neurons) constituting the neural network.

[0063] According to the learning processor 35 (see Figure 2 ) to generate the written text extraction learning model 32. The learning processor 35 performs the learning process by using a large number of pairs of input images and real images. The input images represent written information that has been degraded due to noise caused by undesirable marks (such as scratches and wear marks), while the real images represent written information that has not been degraded and corresponds to the input images. The learning process performed by the learning processor 35 will be described later.

[0064] As an alternative to this exemplary embodiment of performing the learning process by using an input image representing entry information degraded by noise caused by an undesired handprint and a real image, a degraded image may be learned by including therein an image degraded by stains and skew.

[0065] The ruled line extraction learning model 34 is a learning model that has undergone a learning process for generating a ruled line image representing the ruled lines formed on a document based on a document image. The ruled line extraction learning model 34 is, for example, a model that defines a learned neural network and is represented as a set of information regarding the weights (strengths) of the connections between the nodes (neurons) that constitute the neural network.

[0066] According to the learning processor 36 (see Figure 3 ) to generate a ruled line extraction learning model 34. The learning processor 36 performs the learning process by using a large number of pairs of input images and real images, the paired input images including ruled lines degraded due to noise, and the real images representing ruled line images corresponding to the input images. The learning process performed by the learning processor 36 will be described later.

[0067] Next, we will refer to Figure 4 The learning processors 35 and 36 are described. The learning processor 35 includes a generator 350 and a discriminator 352, which together constitute a generative adversarial network (GAN).

[0068] The learning processor 35 retains a large number of pairs of input images 200 and real images 202 as learning data. Figure 5As shown, the input image 200 is a document image containing noise caused by undesirable marks (e.g., scratches or wear marks). Figure 5 In the input image 200 shown, noise caused by undesirable marks (e.g., scratches or wear marks) appears in the document image. The noise caused by undesirable marks (e.g., scratches or wear marks) interferes with the process of recognizing text from the image. Figure 6 The real image 202 shown is an image of a single text written therein. The text can be recognized from the real image 202.

[0069] Figure 4 The generator 350 shown is a neural network that generates a generated image 204 based on the input image 200. The generated image 204 is an estimated image of the real image 202 corresponding to the input image 200. Specifically, the generator 350 generates a generated image 204 similar to the real image 202 based on the input image 200 containing noise caused by undesirable marks (e.g., scratches or wear marks). The generator 350 performs a learning process by using a large number of input images 200 to generate a generated image 204 that is more similar to the real image 202.

[0070] The discriminator 352 is a neural network that discriminates whether the input image is either the real image 202 corresponding to the input image 200 or the generated image 204 generated by the generator 350 based on the input image 200. The learning processor 35 inputs the real image 202 (and its corresponding input image 200) or the generated image 204 (and its corresponding input image 200) to the discriminator 352. Therefore, the discriminator 352 discriminates whether the input image is the real image 202 (true) or the generated image 204 (false), and outputs a signal indicating the discrimination result.

[0071] The learning processor 35 compares the result indicating whether the image input to the discriminator 352 is real or fake with the output signal from the discriminator 352, and feeds back a loss signal based on the comparison result to the weight parameters of the connections between the nodes of the respective neural networks of the generator 350 and the discriminator 352. Accordingly, the generator 350 and the discriminator 352 perform a learning process.

[0072] The generator 350 and discriminator 352 constituting the GAN perform a learning process while working and improving on each other, such that the generator 350 attempts to generate a virtual image (i.e., generated image 204) that is as similar as possible to the teacher data (i.e., real image 202), while the discriminator 352 attempts to correctly identify the fake image.

[0073] For example, the learning processor 35 may use an approach similar to the algorithm known as “pix2pix” (see “Image-to-Image Translation with Conditional Adversarial Networks,” Phillip Isola et al., Berkeley Artificial Intelligence Research (BAIR) Lab, University of California, Berkeley). In this case, the difference between the real image 202 and the generated image 204 is fed back to the learning process of the generator 350 in addition to the loss signal of the discriminator 352.

[0074] As another example, a GAN called a cycle GAN may be used in the learning processor 35. If the cycle GAN is used, the learning process can be performed even if real images are not prepared for all input images.

[0075] In the image processing apparatus 1 according to the present exemplary embodiment, the generator 350 that has undergone a learning process and is generated according to the above-described technique is used as the learned written text extraction learning model 32. The written text extractor 31 generates (or estimates) an image representing written text from a document image by using the learned written text extraction learning model 32 to extract the written text image.

[0076] By using the well-learned written text extraction learning model 32, it is not impossible to extract a recognizable written text image from a document image containing noise caused by undesired marks such as scratches or abrasions.

[0077] Next, the learning processor 36 will be described. The learning processor 36 includes a GAN (see Figure 4 ) generator 350 and discriminator 352. Since the learning processor 36 is similar to the above-mentioned learning processor 35, a detailed description will be omitted. The learning processor 36 is different from the learning processor 35 in that it uses Figure 7 An image consisting only of grid lines is shown as the real image 202 .

[0078] In the image processing apparatus 1 according to the present exemplary embodiment, the generator 350 that has undergone the learning process is used as the learned ruled line extraction learning model 34. The ruled line extractor 33 generates (or estimates) an image representing ruled lines from the document image by using the learned ruled line extraction learning model 34 to extract the ruled line image.

[0079] By using the well-learned written text extraction learning model 32, it is not impossible to extract a ruled line image from a document image containing noise caused by undesirable marks such as scratches or wear marks.

[0080] Next, the text recognizer 4 will be described.

[0081] The text recognizer 4 recognizes the written text formed on the document and outputs the recognition result. In detail, the written text positioning unit 41 positions the written text image by using the registration form frame position information stored in the registration form frame position information storage unit 42. The registration form frame position information storage unit 42 stores therein information related to the grid, such as the position, shape, and size of the writing frame in the grid image detected by the writing frame position detector 43, as the registration form frame position information. The writing frame position detector 43 detects the area within the frame in the grid image extracted by the grid extractor 33 as the entry area, and detects the frame of the entry area as the writing frame. Therefore, the written text positioning unit 41 positions the written text image corresponding to the writing frame by using the registration form frame position information.

[0082] In detail, in the positioning process of the written text image performed in the written text positioning unit 41, the writing frame position detector 43 detects the position, shape and size of the frame image formed by the plurality of grid line images by using the grid line image extracted by the grid line extractor 33. The area within the frame represented by the frame image corresponds to the entry area in which the user enters information. The writing frame position detector 43 stores the writing frame position information representing the writing frame in the registration form frame position information storage unit 42 based on the position, shape and size of the frame image representing the entry area. In the registration form frame position information storage unit 42, information representing the grid lines formed in the table (i.e., grid frame position information representing the grid frame based on the position, shape and size of the grid frame image) is pre-registered as the registration form frame position information.

[0083] The written text positioning unit 41 positions the written text image using the written text frame position information stored in the registration form frame position information storage unit 42 and the registered form frame position information. Specifically, the unit compares the registered registration form frame position information with the detected written text frame position information to calculate a difference, thereby calculating a positional offset. The written text positioning unit 41 performs correction processing by shifting either the written text image 204M or the ruled line image 204K by the calculated positional offset, so that the written text image is positioned within the ruled line frame.

[0084] For example, Figure 8 As shown, when the written text image extracted by the ruled line extractor 33 is superimposed on the written text image extracted from the document image 200 by the written text extractor 31, the text image is set within the ruled line frame. The written text positioning unit 41 associates the written frame detected by the ruled line extractor 33 with the written text image extracted by the written text extractor 31. In the superimposed image of the extracted written text image, the written text image 204M1 is located in the area 204A within the ruled line frame 204K1 representing the item content.

[0085] The written text recognizer 44 uses the written text recognition dictionary 45 to recognize the written text image from the superimposed image obtained by the written text positioning unit 41. The written text recognition dictionary 45 has stored therein a database representing the correspondence between the written text image and the text code corresponding to the plain text of the written text image. Specifically, the text recognizer 4 generates a text code corresponding to the input information based on the written text image generated as a result of the text and the ruled line extractor 3 from which noise is removed (or suppressed).

[0086] In this process of recognizing a written text image, the written text recognizer 44 recognizes a text image of each area within the ruled frame located by the written text location unit 41. In detail, for example, Figure 9 As shown, the written text recognizer 44 recognizes a portion of the written text image 204A1 within the ruled frame 204K1 in the area 204A by using the written text recognition dictionary 45. Figure 9 In the example shown, a text code representing "GOLF CLUBS" is generated as a recognition result.

[0087] In document image 200, multiple entry fields are provided corresponding to the grid frame, where the user enters information. Each entry field is provided for each piece of information to be entered by the user. Therefore, the recognition result of the written text image in each area of ​​the grid frame corresponds to each piece of information.

[0088] By associating information representing each piece of entry information with each area of ​​the ruled frame in the writing frame position information or the registration form frame position information, the written text recognizer 44 becomes able to associate the recognition result with the information representing each piece of entry information. Figure 9 In the example shown, item information indicating the content of the item is added as attribute information to the text code indicating "GOLF CLUBS" as the recognition result. Therefore, the item indicated by the text of the recognition result is identifiable.

[0089] The image input unit 2 is an example of a first image generator according to an exemplary embodiment of the present disclosure. The written text extractor 31 and the text and ruled line extractor 3 are examples of a second image generator according to an exemplary embodiment of the present disclosure. The text recognizer 4 is an example of an output unit according to an exemplary embodiment of the present disclosure.

[0090] For example, the above-described image processing apparatus 1 can be realized by causing a computer to execute a program representing the above-described functions.

[0091] Figure 10 An example is shown in which a computer is included as an apparatus that executes processing to realize various functions of the image processing apparatus 1 .

[0092] Used as Figure 10 The computer of the image processing apparatus 1 shown includes a computer unit 100. The computer unit 100 includes a central processing unit (CPU) 102, a random access memory (RAM) 104 such as a volatile memory, a read-only memory (ROM) 106, an auxiliary storage device 108 such as a hard disk drive (HDD), and an input / output (I / O) interface 110. The CPU 102, RAM 104, ROM 106, auxiliary storage device 108, and I / O interface 110 are connected via a bus 112 so that data and commands can be exchanged with each other. The I / O interface 110 is connected to the image input unit 2, a communication interface (I / F) 114, and an operation display unit 116 such as a display and a keyboard.

[0093] The auxiliary storage device 108 has a control program 108P stored therein for causing the computer unit 100 to function as the image processing apparatus 1 according to this exemplary embodiment. The CPU 102 reads the control program 108P from the auxiliary storage device 108 and develops the control program 108P in the RAM 104 to execute processing. Thus, the computer unit 100 executing the control program 108P operates as an information processing apparatus according to the exemplary embodiment of the present disclosure.

[0094] The auxiliary storage device 108 has stored therein a learning model 108M including the written text extraction learning model 32 and the ruled line extraction learning model 34, and data 108D including the registration form frame position information storage unit 42 and the written text recognition dictionary 45. The control program 108P may be provided by a recording medium such as a CD-ROM.

[0095] Next, image processing in the image processing apparatus 1 realized by a computer will be described.

[0096] Figure 11 An example of the flow of image processing executed according to the control program 108P in the computer unit 100 is shown.

[0097] When the power of the computer unit 100 is turned on, the CPU 102 executes Figure 11 The image processing shown.

[0098] First, the CPU 102 acquires the document image 200 from the image input unit 2 in step S100 and extracts the written text image in step S104 . Specifically, the written text image 204M as an intermediate product is extracted from the document image 200 by using the written text extraction learning model 32 .

[0099] In step S106 , a ruled line image is extracted. Specifically, a ruled line image 204K as an intermediate product is extracted from the document image 200 using the ruled line extraction learning model 34 .

[0100] In step S108, the positional offset within the ruled frame in document image 200 is detected. Specifically, the position, shape, and size of the frame image formed by the plurality of ruled line images are detected using the ruled line images extracted in step S106. Then, writing frame position information representing the writing frame based on the position, shape, and size of the frame images is stored in data 108D in auxiliary storage device 108. In data 108D, information representing the ruled lines formed in the form (i.e., ruled frame position information representing the ruled frame based on the position, shape, and size of the ruled frame images) is pre-registered as registered form frame position information.

[0101] In step S110, the position of the written text is corrected for each ruled frame. Specifically, the written frame position information stored in the data 108D and the registered form frame position information are compared with each other to calculate a difference, thereby calculating the offset of the frame position. Then, correction processing is performed by moving either the written text image 204M or the ruled line image 204K by the calculated offset of the frame position so that the written text image is positioned within the ruled line frame (see Figure 8 ).

[0102] In step S112, the written text image is recognized. Specifically, the written text image in each area of ​​the grid frame corrected in step S110 is recognized by using the written text recognition dictionary 45. In step S114, the recognition result (e.g., text code) in step S112 is output, and the process flow ends.

[0103] Figure 11 The illustrated image processing is an example of processing performed by the image processing apparatus 1 according to this exemplary embodiment.

[0104] Second Exemplary Embodiment

[0105] Next, a second exemplary embodiment will be described. In this second exemplary embodiment, the disclosed technology is applied to a case where image processing for recognizing text is performed after predetermined pre-processing has been performed on a document image. Since the components in the second exemplary embodiment are substantially similar to those in the first exemplary embodiment, the same reference numerals will be assigned to the same components, and detailed descriptions thereof will be omitted.

[0106] Figure 12 An example of the configuration of the image processing apparatus 12 according to the second exemplary embodiment is shown.

[0107] like Figure 12 As shown, the image processing device 12 according to the second exemplary embodiment includes an image input unit 2, a preprocessor 5, a text and ruled line extractor 3, and a text recognizer 4. The second exemplary embodiment is different from the first exemplary embodiment in that the document image received by the image input unit 2 is output to the text and ruled line extractor 3 after being preprocessed by the preprocessor 5.

[0108] The preprocessor 5 includes a preprocessing executor 50. The preprocessing executor 50 included in the preprocessor 5 performs predetermined preprocessing on the document image from the image input unit 2 and outputs the document image. The preprocessing executor 50 performs simple image processing on the document image. Examples of simple image processing include color processing, grayscale correction, fixed noise processing, and sharpening.

[0109] An example of color processing includes removing background color from a document image. For example, if information such as text or graphics is to be entered using black ink with a writing tool, the document included in the form may sometimes have fixed text against a blue background. In this case, the pre-processing executor 50 pre-removes the text against the blue background, thereby pre-removing any fixed text that differs from the user-entered text from the document image to be input to the text and ruled line extractor 3. This improves the accuracy of the output recognition result.

[0110] An example of grayscale correction includes increasing the density of a written text image. For example, due to a lack of writing pressure from the user or insufficient carbon on the carbon sheet, the recorded image corresponding to the recorded information entered by the user may be formed to have a density lower than a pre-assumed density. In this case, grayscale correction is performed to increase the density of the recorded image, which has a density lower than the pre-assumed density, by a predetermined density. As a result of the grayscale correction, the recognition rate of the text image (i.e., the accuracy of the recognition result to be output) can be improved.

[0111] Examples of fixed noise processing include simple noise removal, which can be performed without using a pre-learned learning model. This simple noise removal process, for example, removes noise caused by scattering of point images within predetermined pixels (so-called salt and pepper noise). By performing this simple noise removal process, simple noise images with little correlation to the input information can be pre-removed from the document image. This improves the accuracy of the output recognition results.

[0112] Examples of sharpening include simple image processing to sharpen an image with a density gradient, such as in a so-called blurred image. By performing this sharpening on the document image, the input image in the document image can be pre-processed to have an improved recognition rate. This improves the accuracy of the output recognition result.

[0113] Figure 13 An example of the flow of image processing according to this exemplary embodiment is shown.

[0114] Figure 13 The image processing flow shown in Figure 11 The image processing flow shown has an additional step S102 between step S100 and step S104. The CPU 102 performs the above-described pre-processing on the document image 200 acquired from the image input unit 2 in step S102, and then proceeds to step S104.

[0115] In this exemplary embodiment, image processing for recognizing text is performed after predetermined preprocessing is performed on the document image. Therefore, it is desirable that the written text extraction learning model 32 and the ruled line extraction learning model 34 in the text and ruled line extractor 3 perform learning processing by using the document image that has undergone predetermined preprocessing.

[0116] Third exemplary embodiment

[0117] Next, a third exemplary embodiment will be described. In the third exemplary embodiment, the disclosed technology is applied to a case where a process of correcting a recognition result of written text is performed. Since the components in the third exemplary embodiment are substantially similar to those in the first exemplary embodiment, the same reference numerals will be given to the same components, and a detailed description thereof will be omitted.

[0118] Figure 14 An example of the configuration of the image processing apparatus 13 according to the third exemplary embodiment is shown.

[0119] like Figure 14 As shown, the image processing apparatus 13 according to the third exemplary embodiment further includes a recognition corrector 4A in the text recognizer 4. The third exemplary embodiment differs from the first exemplary embodiment in that a recognition result obtained by the write text recognizer 44 of the text recognizer 4 is corrected by the recognition corrector 4A, and the corrected recognition result is output.

[0120] The recognition corrector 4A includes a recognition result corrector 46 and a database (DB) 47. An example of the DB 47 is an address DB. The address DB contains county names and city names. Another example of the DB 47 is a postal code DB. The postal code DB is a database in which postal codes and address DBs are linked to each other. The recognition result corrector 46 of the recognition corrector 4A uses the DB 47 to correct the recognition result obtained by the written text recognizer 44 and outputs the corrected recognition result.

[0121] Specifically, the recognition result corrector 46 extracts the recognition result obtained by the written text recognizer 44 and data similar to the recognition result from the data registered in the DB 47. For example, if the address DB is registered as the DB 47, a text string matching or similar to the text string of the address of the recognition result obtained by the written text recognizer 44 is extracted. If the text string of the address of the recognition result obtained by the written text recognizer 44 matches the text string extracted from the address DB, the recognition result obtained by the written text recognizer 44 is output.

[0122] In contrast, if the text string of the address of the recognition result obtained by the written text recognizer 44 does not match the text string extracted from the address DB, that is, if the text string of the recognition result obtained by the written text recognizer 44 is not registered in the address DB, then there is a high probability that an erroneously recognized text string is included in the text string of the recognition result obtained by the written text recognizer 44. Therefore, the recognition result corrector 46 compares the text string of the address of the recognition result with the text string extracted from the address DB, and corrects the erroneously recognized text.

[0123] For example, a text string with a high matching rate with the text string of the address of the recognition result is extracted from the address DB and replaced with the recognition result. Regarding the matching rate, the percentage of the number of matching characters between the text string of the address of the recognition result and the text string extracted from the address DB can be used. Multiple (for example, three) text strings can be extracted from the address DB starting from the highest matching rate, and one of the multiple (for example, three) text strings can be selected. In this case, priorities can be given to the multiple extracted (for example, three) text strings respectively. For example, the text string with the highest matching rate can be automatically set according to the priority, or can be selected by the user.

[0124] If a postal code and an address are obtained as the recognition result, the recognition result can be corrected by using the postal code and the address. For example, the address corresponding to the postal code in the recognition result is extracted from the postal code in the recognition result by using the postal code DB, and compared with the address in the recognition result. The address in the recognition result is corrected based on the matching rate of the comparison result. In addition, the postal code corresponding to the address in the recognition result is extracted from the address in the recognition result by using the postal code DB, and compared with the postal code in the recognition result. The postal code in the recognition result is corrected based on the matching rate of the comparison result. Alternatively, a plurality of candidates can be extracted by determining the matching rate based on a combination of the postal code and the address, and one of the extracted candidates can be selected.

[0125] Although the exemplary embodiments have been described above, the technical scope of the present disclosure is not limited to the scope defined in the above exemplary embodiments. Various modifications or changes may be made to the above exemplary embodiments without departing from the scope of the present disclosure, and the exemplary embodiments obtained by adding these modifications or changes thereto are included in the technical scope of the present disclosure.

[0126] As an alternative to realizing each of the above-described exemplary embodiments of the inspection processing by a software configuration based on the processing using the flowchart, each processing may be realized by a hardware configuration, for example.

[0127] Furthermore, a part of the image processing apparatus 1 (for example, a neural network such as a learning model) may be configured as a hardware circuit.

[0128] The foregoing description of exemplary embodiments of the present disclosure has been provided for the purposes of illustration and description, but is not intended to be exhaustive or to limit the disclosure to the precise forms disclosed. Obviously, many modifications and variations will be apparent to those skilled in the art. These embodiments have been selected and described in order to best explain the principles of the present disclosure and its practical application, thereby enabling others skilled in the art to understand the various embodiments of the present disclosure and the various modifications suitable for the intended specific use. The scope of the present disclosure is intended to be defined by the appended claims and their equivalents.

Claims

1. An image processing device, comprising: a first image generator configured to generate a first image including a predetermined ruled line image and a registered image based on a second sheet in a sheet group, the sheet group being obtained by stacking a plurality of sheets including a single first sheet having registered information registered thereon and a second sheet having the registered image corresponding to the registered information transferred thereto and including the ruled line image; a second image generator for generating a second image by removing a residual image from the first image generated by the first image generator according to a learning model, the learning model having learned to remove the residual image that is different from the ruled line image and the input image; wherein the learning model includes a first learning model and a second learning model, the first learning model is used to extract the input image with the residual image and the grid image removed, and the second learning model is used to extract the grid image with the residual image removed; and an output unit that outputs information indicating a recognition result obtained as a result of recognizing the second image generated by the second image generator as the entry information, The output unit adjusts the position of the second image based on ruled line position information indicating a predetermined position of the ruled line image on the sheet so that the position of the ruled line image in the second image matches the position according to the ruled line position information.

2. The image processing apparatus according to claim 1, in, The learning model is a model that has learned to generate the original image from the input image according to a combination of the input image including the residual image and the original image not including the residual image corresponding to the input image.

3. The image processing apparatus according to claim 1 or 2, in, The learning model is a model generated as a result of learning by using a generative adversarial network.

4. The image processing apparatus according to claim 1 or 2, in, Attribute information related to the grid frame representing the entry item to be entered in the area formed by the grid frame of the grid image is pre-set, and wherein the output unit identifies the second image relative to the area formed by the grid frame, and outputs the recognition result of the area and the attribute information related to the grid frame corresponding to each other.

5. The image processing apparatus according to claim 1 or 2, in, The input image is a written text image.

6. The image processing apparatus according to claim 1 or 2, in, The sheet set includes sheets coated with carbon on one side.

7. The image processing apparatus according to claim 1 or 2, in, The remaining image is a mark image corresponding to mark information about at least one of a scratch and a wear mark.

8. A non-transitory computer-readable medium storing a program causing a computer to perform a process, the process comprising: generating a first image including a predetermined ruled line image and a recorded image based on a second sheet in a sheet set obtained by stacking a plurality of sheets including a single first sheet having recorded information recorded thereon and a second sheet having the recorded image corresponding to the recorded information transferred thereto and including the ruled line image; generating a second image by removing a residual image from the generated first image according to a learning model, the learning model having learned to remove the residual image that is different from the ruled line image and the input image, wherein the learning model includes a first learning model and a second learning model, the first learning model is used to extract the input image with the residual image and the grid image removed, and the second learning model is used to extract the grid image with the residual image removed; and outputting, as the entry information, information indicating a recognition result obtained as a result of recognizing the second image generated by the second image generator, The position of the second image is adjusted based on ruled line position information indicating a predetermined position of the ruled line image on the sheet so that the position of the ruled line image in the second image matches the position according to the ruled line position information.

9. An image processing method, comprising: generating a first image including a predetermined ruled line image and a recorded image based on a second sheet in a sheet set obtained by stacking a plurality of sheets including a single first sheet having recorded information recorded thereon and a second sheet having the recorded image corresponding to the recorded information transferred thereto and including the ruled line image; as well as generating a second image by removing a residual image from the generated first image according to a learning model, the learning model having learned to remove the residual image that is different from the ruled line image and the input image, wherein the learning model includes a first learning model and a second learning model, the first learning model is used to extract the input image with the residual image and the grid image removed, and the second learning model is used to extract the grid image with the residual image removed; and outputting, as the entry information, information indicating a recognition result obtained as a result of recognizing the second image generated by the second image generator, The position of the second image is adjusted based on ruled line position information indicating a predetermined position of the ruled line image on the sheet so that the position of the ruled line image in the second image matches the position according to the ruled line position information.

Citation Information

Patent Citations

  • Character recognizing device

    JP1994060222A

  • Compression of missing form document image

    JP1995200720A

  • Image processor, image processing method and program for making computer perform the methods

    JP2003030644A