Image processing apparatus

The image processing device addresses the issue of misrecognizing characters in large-spaced words by combining adjacent regions into a single word, enhancing the accuracy of information extraction.

JP2025175390APending Publication Date: 2025-12-03KYOCERA DOCUMENT SOLUTIONS INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024081467
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-05-20
Publication Date
2025-12-03

AI Technical Summary

Technical Problem

Existing image processing devices struggle to accurately recognize characters in a word when the spacing between them is large, leading to individual characters being treated as separate words, which hinders the extraction of predetermined information.

Method used

An image processing device that detects character areas by comparing the width differences in two directions and uses reference areas to combine adjacent character regions aligned in the writing direction into a single word, ensuring accurate extraction of information.

Benefits of technology

Prevents multiple characters in a word from being recognized as separate words, enabling precise extraction of predetermined information from documents.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025175390000001_ABST
    Figure 2025175390000001_ABST
Patent Text Reader

Abstract

To prevent a plurality of characters constituting one word from being recognized each as an independent word.SOLUTION: A control unit detects a character region from image data of a document, recognizes, as an independent character region, a region of which an absolute value of a difference between the width in a first direction, which is the writing direction in the image data, and the width in a second direction, which is orthogonal to the first direction, is smaller than a first threshold value, sets a character region, as a reference region, which is adjacent in the second direction to a plurality of independent character regions aligned in the first direction but which is not an independent character region, and determines, based on the width in the first direction of the reference region, whether a character string composed of characters in the plurality of independent character regions is one word. When the character string is determined to be one word, the control unit combines the characters in the independent character regions into one character string, and treats the resultant character string as one word.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an image processing device. [Background technology]

[0002] 2. Description of the Related Art Conventionally, there has been known an image processing device that reads an original document and detects a character area from image data obtained by reading the original document. Such an image processing device is disclosed in, for example, Patent Document 1. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2020-177521 Summary of the Invention [Problem to be solved by the invention]

[0004] For example, if the spacing between characters in a word is large, the word may be divided into multiple character regions for detection. In this case, each character in the divided character regions is treated as an independent word. As a result, when extracting predetermined information based on a character string (word) composed of each character in the divided character regions, the predetermined information cannot be extracted.

[0005] The present invention has been made to solve the above-mentioned problems, and aims to provide an image processing device that can prevent multiple characters that make up a word from being recognized as independent words in a configuration that extracts specified information from image data obtained by reading a document. [Means for solving the problem]

[0006] To achieve the above object, an image processing device according to one aspect of the present invention includes an image reading unit that reads a document containing predetermined information, and a control unit that extracts the predetermined information by performing OCR processing on image data of the document obtained by reading the image reading unit. When extracting the predetermined information, the control unit detects character areas from the image data and recognizes character areas in which the absolute value of the difference between the width in a first direction (the writing direction in the image data) and the width in a second direction perpendicular to the first direction is smaller than a predetermined first threshold as independent character areas. Character areas that are adjacent to the multiple independent character areas aligned in the first direction but are not independent character areas in the second direction are set as reference areas. Based on the width of the reference area in the first direction, the control unit determines whether a character string formed by the characters of the multiple independent character areas aligned in the first direction is a single word. When the control unit determines that a character string formed by the characters of the multiple independent character areas aligned in the first direction is a single word, the control unit combines the characters of the multiple independent character areas aligned in the first direction into a single character string and treats the combined character string as a single word. [Effects of the Invention]

[0007] In the present invention, in a configuration in which predetermined information is extracted from image data obtained by scanning a document, it is possible to prevent multiple characters that make up one word from being recognized as independent words. [Brief explanation of the drawings]

[0008] [Figure 1] FIG. 1 is a schematic diagram of a multifunction peripheral according to an embodiment. [Figure 2] FIG. 1 is a block diagram of a multifunction peripheral according to an embodiment. [Figure 3] FIG. 1 is a diagram schematically illustrating an original that can be read by a multifunction peripheral according to an embodiment. [Figure 4] 10 is a flowchart illustrating a flow of an information extraction job executed by the multifunction peripheral according to an embodiment. [Figure 5] 1A and 1B are diagrams illustrating image data generated by scanning in a multifunction peripheral according to an embodiment. [Figure 6]10A and 10B are diagrams illustrating the positions and sizes of each character area in image data generated by scanning with a multifunction peripheral according to an embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0009] <Configuration of the multifunction device> An image processing apparatus according to an embodiment of the present invention will be described below with reference to FIGS. 1 to 6, taking as an example a multifunction peripheral 100 having multiple functions such as a scanning function, a printing function, and a transmission function.

[0010] As shown in FIG. 1, the multifunction device 100 (corresponding to an "image processing device") includes a printing unit 1. The printing unit 1 constitutes the main body of the multifunction device 100. The printing unit 1 prints an image on a sheet S. The printing method of the printing unit 1 is an electrophotographic method. However, this is not limited to this. The printing method of the printing unit 1 may also be an inkjet method.

[0011] The printing unit 1 forms an image based on image data input to the multifunction device 100. The printing unit 1 also transports a sheet S along a sheet transport path. The printing unit 1 prints an image on the sheet S while it is being transported. In FIG. 1, the sheet transport path is indicated by a dashed line.

[0012] The printing unit 1 includes a paper feed roller 11. The paper feed roller 11 comes into contact with a sheet S stored in a sheet cassette CA and rotates in this state, thereby feeding the sheet S from the sheet cassette CA to a sheet transport path.

[0013] The printing unit 1 includes an image forming unit 12. The image forming unit 12 includes a photosensitive drum 12a and a transfer roller 12b. The photosensitive drum 12a carries a toner image on its peripheral surface. The transfer roller 12b is in pressure contact with the photosensitive drum 12a, forming a transfer nip between the photosensitive drum 12a and the transfer roller 12b. The transfer roller 12b rotates together with the photosensitive drum 12a. The image forming unit 12 transfers the toner image onto the sheet S while transporting the sheet S that has entered the transfer nip.

[0014] Although not shown, the image forming unit 12 further includes a charging device, an exposure device, and a developing device. The charging device charges the circumferential surface of the photosensitive drum 12a. The exposure device forms an electrostatic latent image on the circumferential surface of the photosensitive drum 12a. The developing device develops the electrostatic latent image on the circumferential surface of the photosensitive drum 12a into a toner image.

[0015] The printing unit 1 includes a fixing unit 13. The fixing unit 13 includes a heating roller 13a and a pressure roller 13b. The heating roller 13a has a built-in heater (not shown). The pressure roller 13b is pressed against the heating roller 13a, forming a fixing nip between the heating roller 13a and the pressure roller 13b. The pressure roller 13b rotates together with the heating roller 13a. The fixing unit 13 fixes the toner image transferred onto the sheet S onto the sheet S while transporting the sheet S that has entered the fixing nip. The sheet S that has exited the fixing nip is discharged onto an ejection tray ET.

[0016] The multifunction device 100 also includes an image reading unit 2. The image reading unit 2 is disposed on top of the main body of the multifunction device 100. In a job involving reading an original D, the original D is set in the image reading unit 2. The image reading unit 2 reads the original D set in the image reading unit 2 and generates image data of the read original D.

[0017] The image reading unit 2 includes contact glasses G1 and G2. The contact glasses G1 and G2 are installed in a housing RH of the image reading unit 2. The housing RH has an opening on the top surface. The contact glasses G1 and G2 are attached to the opening on the top surface of the housing RH.

[0018] The image reading unit 2 includes a document transport device DP. The document transport device DP is attached to the housing RH. When viewed from the front of the multifunction device 100, the document transport device DP pivots around its rear portion as if swinging its front portion up and down. The document transport device DP opens and closes relative to the top surface of the housing RH.

[0019] The document transport device DP has a set tray ST on which documents D are set. The document transport device DP transports the documents D set on the set tray ST onto the contact glass G1.

[0020] In the conveying reading mode, the user sets the document D on the set tray ST. Then, the document D is automatically conveyed onto the contact glass G1 by the document conveying device DP (in other words, the document D passing over the contact glass G1) and is read. On the other hand, in the placed reading mode, the user sets the document D on the contact glass G2, and the document D on the contact glass G2 is read.

[0021] The image reading unit 2 includes a light source 21, an image sensor 22, a mirror 23, and a lens 24. The light source 21, the image sensor 22, the mirror 23, and the lens 24 are installed inside the housing RH. The image reading unit 2 performs a scanning operation in which light is irradiated from the light source 21 toward the contact glass G1 or G2 and photoelectrically converted by the image sensor 22.

[0022] The light source 21 has a plurality of LED elements. The plurality of LED elements are arranged in a line in the main scanning direction (a direction perpendicular to the paper surface of FIG. 1). The image sensor 22 has a plurality of photoelectric conversion elements arranged in the main scanning direction. The mirror 23 reflects light toward the lens 24. The lens 24 collects the light reflected by the mirror 23 and guides it to the image sensor 22.

[0023] The light source 21 and the mirror 23 are mounted on a carriage 25 that is movable in a sub-scanning direction (left and right direction in FIG. 1) that is perpendicular to the main scanning direction. As the carriage 25 moves in the sub-scanning direction, the reading line of the image reading unit 2 moves in the sub-scanning direction.

[0024] As shown in FIG. 2, the multifunction device 100 also includes an operation display unit 3. The operation display unit 3 is an operation panel having a touch screen. The operation display unit 3 displays software buttons, messages, and the like on the touch screen. The operation display unit 3 is also provided with a plurality of hardware buttons. The operation display unit 3 accepts operations from the user. The user can configure the multifunction device 100 via the operation display unit 3 regarding various jobs, such as an information extraction job, which will be described later.

[0025] The multifunction peripheral 100 includes a control unit 10. The control unit 10 includes a CPU, an ASIC, a memory, and the like. The control unit 10 also includes an image processing circuit. The control unit 10 performs various image processing operations on image data. The control unit 10 also controls the printing of an image onto a sheet S by the printing unit 1, and the reading of an original D by the image reading unit 2.

[0026] The control unit 10 also controls the operation and display unit 3. Specifically, the control unit 10 controls the display operation of the touch screen. The control unit 10 detects operations on software buttons and hardware buttons. The control unit 10 sets jobs based on operations received by the operation and display unit 3 from the user.

[0027] The multifunction peripheral 100 includes a storage unit 101. The storage unit 101 is a non-volatile storage device. An HDD, an SSD, or the like may be used as the storage unit 101. The storage unit 101 is connected to the control unit 10. The control unit 10 writes information to the storage unit 101 and reads information from the storage unit 101.

[0028] The storage unit 101 stores a character recognition program in advance. The control unit 10 performs character recognition processing such as OCR (Optical Character Recognition) processing based on the character recognition program. The control unit 10 subjects the image data obtained by the image reading unit 2 reading the document D to the character recognition processing.

[0029] The multifunction peripheral 100 includes a communication unit 102. The communication unit 102 is an interface for communicatively connecting an external device to the multifunction peripheral 100. The communication unit 102 includes a communication circuit, a communication memory, a communication connector, and the like. The communication unit 102 is connected to the control unit 10. The control unit 10 uses the communication unit 102 to send and receive data to and from the external device.

[0030] The communication unit 102 is communicably connected to an external device via a network NT such as a LAN or the Internet. Although not shown, the communication unit 102 may be directly connected to an external device via a communication cable. An external device connected to the communication unit 102 is, for example, a personal computer 1000 (hereinafter referred to as PC 1000) used by a user of the multifunction peripheral 100. An external device other than the PC 1000 may be communicably connected to the multifunction peripheral 100. By connecting the PC 1000 to the multifunction peripheral 100, image data of the original D obtained by reading the original D with the image reading unit 2 can be transmitted to the PC 1000. This allows the image data of the original D to be saved in the PC 1000.

[0031] <Extraction of personal information> The multifunction peripheral 100 has an information extraction function. In other words, the multifunction peripheral 100 is capable of executing a job related to the information extraction function (hereinafter referred to as an information extraction job). In an information extraction job, the image reading unit 2 reads an original D on which various information such as personal information is written. The control unit 10 performs OCR processing on the image data of the original D obtained by the image reading unit 2 reading the original D. As a result, the control unit 10 recognizes information such as personal information written on the original D.

[0032] By using the information extraction function, it is possible to extract only the predetermined information from the information written on the document D. In other words, in the information extraction job, it is possible to extract only the text data as the predetermined information from the image data of the document D.

[0033] In the information extraction job, for example, predetermined information extracted from the image data of the document D can be sent to the PC 1000 and displayed on the PC 1000 or saved in the PC.

[0034] Furthermore, in an information extraction job, for example, image processing is performed on the original image data obtained by reading the document D, thereby generating output image data in which the area corresponding to the predetermined information is anonymized. Then, an image based on the output image data (i.e., an image in which the predetermined information has been anonymized) can be printed on a sheet S. The output image data can also be sent to the PC 1000 and saved in the PC 1000. The output image data is image data generated from the original image data, and is image data in which a portion of the original image data has been altered.

[0035] There are various types of documents D that can be the target of an information extraction job. For example, a document D as shown in FIG. 3 is the target of an information extraction job. As an example, FIG. 3 schematically shows a document D that includes a field for entering personal information. Hereinafter, this document will be simply referred to as document D.

[0036] In manuscript D, generally, multiple sets of item names related to personal information and item values ​​corresponding to the item names are written. In the example shown in Figure 3, "financial institution name," "branch name," "account type," "account number," "furigana," and "account holder" correspond to the item names. The descriptions following these item names correspond to the item values. Item values ​​are not shown in Figure 3.

[0037] When executing an information extraction job, the user places the document D on the image reading unit 2. In this state, the user performs a start operation to instruct execution of the information extraction job on the operation display unit 3. When the start operation is performed on the operation display unit 3, the control unit 10 starts the information extraction job.

[0038] The information extraction job will be specifically described with reference to the flowchart shown in Fig. 4. The flow shown in Fig. 4 starts when a start operation for the information extraction job is performed on the operation display unit 3. That is, when extracting predetermined information written on the document D, the control unit 10 performs processing in accordance with the flow shown in Fig. 4.

[0039] In step #1, the control unit 10 causes the image reading unit 2 to read the original document D. The image reading unit 2 reads the original document D and generates image data of the read original document D. The image data generated at this time becomes the original image data. In the following description, the image data of the original document D obtained by reading by the image reading unit 2 will be referred to as the original image data.

[0040] The control unit 10 acquires the original image data. Then, the control unit 10 performs OCR processing on the original image data. As part of the OCR processing, the control unit 10 performs layout analysis, line segmentation, character segmentation, etc. Furthermore, the control unit 10 recognizes the writing direction in the original image data (i.e., the direction in which the characters are written) based on the orientation of the characters present in the original image data.

[0041] The following description will be given assuming that the document D shown in Fig. 3 is the target of the information extraction job. When the document D shown in Fig. 3 is the target of the information extraction job, original image data such as that shown in Fig. 5 is generated by reading the document D using the image reading unit 2. The control unit 10 recognizes the position (coordinate value) of each character area in the original image data in a coordinate system (screen coordinate system) whose origin is (0,0) at the upper left corner of the original image data.

[0042] In the example shown in Fig. 5, the writing direction is the X direction. That is, the X direction corresponds to the "first direction," and the Y direction perpendicular to the X direction corresponds to the "second direction." Although not shown, when the writing direction is the Y direction, the Y direction corresponds to the "first direction," and the X direction corresponds to the "second direction."

[0043] In the following description, the X coordinate value of a character area (including an independent character area) corresponds to the coordinate value of one end of the character area in the X direction, for example, the X coordinate value of the upper left corner of the character area. The Y coordinate value of a character area (including an independent character area) corresponds to the coordinate value of one end of the character area in the Y direction, for example, the Y coordinate value of the upper left corner of the character area.

[0044] In step #2, the control unit 10 detects a character area from the original image data. In the example shown in FIG. 5, the area surrounded by the dashed line in the figure is detected as a character area. Typically, multiple character areas are detected from the original image data. In the following description, the multiple character areas may be distinguished by assigning the symbols A1, A21, A22, A23, A3, A4, A5, A6, A71, A72, and A73 to each of them.

[0045] The control unit 10 recognizes the position (coordinate value) of each of the multiple character regions within the original image data. The control unit 10 also recognizes the width in the X direction and the width in the Y direction (character height) for each of the multiple character regions. The results are shown in Figure 6.

[0046] In step #3, the control unit 10 determines whether each of the multiple character regions in the original image data is an independent character region. In other words, the control unit 10 recognizes the independent character regions that exist in the original image data.

[0047] Specifically, the control unit 10 calculates the difference between the width in the X direction and the width in the Y direction for each of the multiple character regions in the original image data. Then, the control unit 10 recognizes a character region where the absolute value of the difference between the width in the X direction and the width in the Y direction is smaller than a predetermined first threshold as an independent character region. A character region whose width in the X direction and the width in the Y direction are approximately the same is recognized as an independent character region. Since the widths in the X direction and the Y direction of a character region with one character are approximately the same, the character region with one character is recognized as an independent character region.

[0048] For example, in manuscript D shown in Figure 3, both ends of the writing direction of each of the multiple item names are aligned. The writing direction position of the first character of each item name is aligned, and the writing direction position of the last character of each item name is aligned. As a result, the character spacing of the writing direction of each item name is set to a distance according to the number of characters.

[0049] The item names "Financial institution name" and "Account holder name" have the most characters, five, so the character spacing is set to the smallest. The item names "Account type", "Account number", and "Furigana" have four characters, so the character spacing is set to be slightly larger than if the item had five characters. The item name "Branch name" has only three characters, so the character spacing is set to the largest. This aligns both ends of the writing direction for each item name.

[0050] In addition, the "Normal / Current" associated with the item name "Account Type" is made larger to make it easier for the person filling out the form to check (circle) the box, and the space between "Normal" and "· (midpoint)" and between "· (midpoint)" and "Current" is made larger.

[0051] In this example, "Financial institution name" and "Account holder name" are each detected as a single character region consisting of five characters. "Account type," "Account number," and "Furigana" are each detected as a single character region consisting of four characters. That is, character regions A1, A3, A4, A5, and A6 are detected from the original image data.

[0052] On the other hand, in "branch name," "branch," "store," and "name" are each detected as separate character regions. Each of "branch," "store," and "name" is recognized as an independent character region. That is, character regions A21, A22, and A23 are each recognized as an independent character region.

[0053] In addition, in "Normal·Temporary", "Normal", "· (middle dot)", and "Temporary" are each detected as separate character regions. "Normal", "· (middle dot)", and "Temporary" are each recognized as independent character regions. That is, character regions A71, 72, and 73 are each recognized as independent character regions.

[0054] When recognizing the item value (actual name of the branch) corresponding to the item name "branch name" as personal information, for example, "branch name" is recognized as the item name, and the character string following the item name "branch name" is recognized as the item value corresponding to the item name "branch name." However, if "branch," "store," and "name" are each detected as separate character regions, and each character in these three character regions is treated as a separate word, "branch name" cannot be recognized as an item name.

[0055] Therefore, in this embodiment, a process for determining whether or not the characters in the multiple independent character regions should be combined into a single character string and treated as a single word is performed. Specifically, the process proceeds from step #3 to step #4.

[0056] In step #4, the control unit 10 groups the multiple independent character regions in the original image data. The control unit 10 classifies multiple independent character regions that are aligned in a line in the X direction into the same group. To perform this grouping, the control unit 10 recognizes the Y coordinate value of each of the multiple independent character regions.

[0057] The control unit 10 compares the Y coordinate value of each of the multiple independent character regions with the Y coordinate values ​​of other independent character regions and determines whether the absolute value of the difference between the two Y coordinate values ​​is smaller than a predetermined second threshold. In other words, the control unit 10 determines whether the absolute value of the difference between the Y coordinate value of a first independent character region and the Y coordinate value of a second independent character region is smaller than the second threshold. If the absolute value of the difference between the Y coordinate value of the first independent character region and the Y coordinate value of the second independent character region is smaller than the second threshold, the control unit 10 determines that the first independent character region and the second independent character region are aligned in a line in the X direction and classifies the first independent character region and the second independent character region into the same group.

[0058] 5, the independent character regions A21, A22, and A23 are classified into the same group because their Y coordinate values ​​are all 130. The independent character regions A71, A72, and A73 are classified into the same group because their Y coordinate values ​​are all 140. In the following description, the group to which the independent character regions A21, A22, and A23 belong is designated by the symbol G2, and the group to which the independent character regions A71, A72, and A73 belong is designated by the symbol G7.

[0059] In step #5, the control unit 10 sets, for each group, a reference area to be used in determining whether or not each character in the multiple independent character areas in the group (i.e., multiple independent character areas aligned in a line in the X direction) should be combined into one character string. Specifically, the control unit 10 sets, as the reference area, a character area that is adjacent in the Y direction to the multiple independent character areas aligned in a line in the X direction but is not an independent character area.

[0060] When setting a reference area for a certain group, the control unit 10 sets the Y coordinate value of the independent character area with the smallest X coordinate value among the multiple independent character areas belonging to the target group as a first reference value. The control unit 10 also calculates a first difference value, which is the absolute value of the difference between the Y coordinate value and the first reference value, for each of the multiple character areas that are not independent character areas. The control unit 10 then sets the character area with the first difference value equal to or less than a predetermined value as a candidate for the reference area. If the candidate character area is adjacent in the Y direction to the multiple independent character areas belonging to the target group, the control unit 10 sets the candidate character area as the reference area. The predetermined value is, for example, a predetermined multiple (e.g., double) of the Y width of the independent character area with the smallest X coordinate value among the multiple independent character areas belonging to the target group.

[0061] In addition, when character areas that are not independent character areas exist adjacent to multiple independent character areas belonging to the target group on both one side and the other side in the Y direction, the control unit 10 sets each of the multiple character areas adjacent to the multiple independent character areas belonging to the target group in the Y direction as a reference area.

[0062] 5, when the reference area is set for group G2, the Y-direction widths of the independent character areas A21, A22, and A23 are all 5, so if the predetermined magnification is 2, the predetermined value is 10 (=5×2). Furthermore, among the independent character areas A21, A22, and A23, the independent character area A21 has the smallest X-coordinate value, so the first reference value is 130.

[0063] Focusing on character area A1, the coordinate value is (15,120) and the first difference value is "10 (=130-120)", so the first difference value is less than the predetermined value. Focusing on character area A3, the coordinate value is (15,140) and the first difference value is "10 (=140-130)", so the first difference value is less than the predetermined value. For the other character areas, the first difference value is greater than the predetermined value. As a result, character areas A1 and A3 are set as reference areas for group G2.

[0064] Here, in some cases, a reference area is added. Specifically, when setting a reference area for a certain group, the control unit 10 sets the absolute value of the difference between the Y coordinate value of the set reference area and the first reference value as the second reference value. The control unit 10 also calculates a second difference value, which is the absolute value of the difference between the Y coordinate value of a character area that is adjacent to the set reference area in the Y direction but is not an independent character area, and the Y coordinate value of the set reference area. The control unit 10 then further sets, as a reference area, a character area whose absolute value of the difference between the second difference value and the second reference value is equal to or less than a predetermined value (for example, the same value as the above-mentioned predetermined value).

[0065] In the example shown in FIG. 5, focusing on character area A3 as the reference area, the coordinate value is (15,140), so the second reference value is "10 (=140-130)." Furthermore, character area A3 as the reference area is adjacent to character area A4 in the Y direction. Character area A4 has coordinate value (15,150), and the second difference value is the same as the second reference value, "10 (=140-130)." In other words, the absolute value of the difference between the second difference value and the second reference value is equal to or less than a predetermined value. As a result, character area A4 is further set as the reference area of ​​group G2.

[0066] When character area A4 is set as the reference area, character area A5, which is adjacent to character area A4 in the Y direction, is also set as the reference area if the conditions for setting it as a reference area are met.When character area A5 is set as the reference area, character area A6, which is adjacent to character area A5 in the Y direction, is also set as the reference area if the conditions for setting it as a reference area are met.Although detailed explanation will be omitted, in the example shown in Figure 5, character areas A5 and A6 are also set as reference areas of group G2.

[0067] For group G7, there are no character areas adjacent to the independent character areas A71, A72, and A73 in the Y direction. Therefore, a reference area is not set for group G7. As a result, the independent character areas A71, A72, and A73 are not combined, and no combining determination is performed.

[0068] In step #6, the control unit 10 performs a join determination process for each group. The control unit 10 determines whether or not to join the characters of multiple independent character regions belonging to the same group (i.e., multiple independent character regions aligned in a row in the X direction) into a single character string. In other words, the control unit 10 determines whether or not to treat a character string made up of the characters of multiple independent character regions belonging to the same group as a single word.

[0069] When performing a join determination process on a certain group, the control unit 10 recognizes the width in the X direction of the reference area assigned to the target group. Then, based on the width in the X direction of the reference area, the control unit 10 determines whether a character string made up of each character in multiple independent character areas belonging to the target group is one word.

[0070] Specifically, the control unit 10 recognizes the coordinate values ​​of one end (for example, the upper left corner) and the other end (for example, the upper right corner) in the X direction of the reference area assigned to the target group. Then, when the X coordinate values ​​of each of the multiple independent character areas belonging to the target group are within the range from the coordinate value of one end to the coordinate value of the other end in the X direction of the reference area, the control unit 10 determines that the character string made up of each character of the multiple independent character areas belonging to the target group is one word.

[0071] When there are multiple reference areas assigned to a target group, the control unit 10 recognizes the coordinate values ​​of one end and the other end in the X direction of the largest reference area that is the longest in the X direction among the multiple reference areas. Then, when the X coordinate values ​​of each of the multiple independent character areas belonging to the target group are within the range from the coordinate value of one end of the largest reference area in the X direction to the coordinate value of the other end, the control unit 10 determines that the character string formed by the characters of the multiple independent character areas belonging to the target group is one word.

[0072] In the example shown in FIG. 5, the independent character regions A21, A22, and A23 belonging to group G2 are subject to the combination determination process. A plurality of reference regions (i.e., character regions A1, A3, A4, A5, and A6) are assigned to group G2. The width in the X direction of each of the character regions A1, A3, A4, A5, and A6 as reference regions is all "15", and their X coordinate values are all "15". Therefore, the coordinate value of one end in the X direction of the maximum reference region is "15", and the coordinate value of the other end is "30 (= 15 + 15)".

[0073] Also, the X coordinate values of the independent character regions A21, A22, and 23 are "15", "21", and "27", respectively. That is, the X coordinate values of the independent character regions A21, A22, and A23 are within the range from the coordinate value of one end in the X direction of the maximum reference region to the coordinate value of the other end. Thereby, it is determined that the character string composed of the characters in each of the independent character regions A21, A22, and 23 is one word.

[0074] In step #7, when it is determined that the character string composed of the characters in each of the plurality of independent character regions belonging to the same group is one word, the control unit 10 combines the characters in each of those plurality of independent character regions into one character string. And the control unit 10 treats the combined character string as one word.

[0075] In the example shown in FIG. 5, the process of combining the characters "branch", "store", and "name" in each of the independent character regions A21, A22, and A23 belonging to group G2 into one character string is performed. And the character string "branch store name" combined by this process is treated as one word.

[0076] In the present embodiment, as described above, the control unit 10 performs the combination determination process for each group. And when it is determined that the character string composed of the characters in each of the plurality of independent character regions belonging to the target group of the combination determination process is one word, the control unit 10 combines the characters in each of the plurality of independent character regions belonging to the target group into one character string and treats the combined character string as one word.

[0077] This prevents multiple characters that make up a single word from being recognized as separate words. As a result, it is possible to extract predetermined information from the original image data with high accuracy. For example, when extracting an item value (the actual name of the branch) corresponding to the item name "branch name" as predetermined information, it is possible to prevent the string "branch name" from being omitted from extraction as the item name. As a result, it is possible to prevent the inconvenience of not being able to extract an item value corresponding to the item name "branch name".

[0078] In this embodiment, the grouping of multiple independent character regions is performed based on the Y coordinate value. This configuration ensures that multiple independent character regions aligned in a line in the X direction are classified into the same group. In other words, it is possible to prevent multiple independent character regions that are unrelated to each other from being classified into the same group.

[0079] In this embodiment, whether a character string consisting of characters in multiple independent character areas belonging to the same group is a single word is determined based on the width of the reference area in the X direction (writing direction). This allows for accurate combination determination processing when a document D in which both ends of the writing direction of multiple item names are aligned is the target of an information extraction job.

[0080] In this embodiment, the reference area is set based on the Y coordinate value. This configuration can prevent a character area that is far away from the independent character area in the Y direction (in other words, an unrelated character area) from being set as the reference area.

[0081] Furthermore, in this embodiment, whether a character string consisting of the characters of multiple independent character regions belonging to the same group is a single word is determined based on the X-direction width of the largest reference region among the multiple reference regions. For example, if the combination determination process is performed based on the range of a reference region with a small X-direction width, an independent character region located at the edge may fall outside the range of the reference region, resulting in a problem in which the character string consisting of the characters of the multiple independent character regions is not determined to be a single word. In other words, the accuracy of the combination determination process will be reduced. For this reason, it is preferable to perform the combination determination process based on the range of the largest reference region among the multiple reference regions.

[0082] The embodiments disclosed herein should be considered to be illustrative in all respects and not restrictive. The scope of the present invention is defined by the claims rather than the description of the above embodiments, and further includes all modifications within the meaning and scope of the claims. [Explanation of symbols]

[0083] 2 Image reading unit 10 Control Unit 100 Multifunction printers (image processing devices) D Manuscript

Claims

1. an image reading unit that reads a document on which predetermined information is written; a control unit that extracts the predetermined information by performing OCR processing on the image data of the document obtained by reading the image by the image reading unit, When extracting the predetermined information, the control unit: Detecting a character region from the image data; recognizes the character region, in which an absolute value of a difference between a width in a first direction, which is a writing direction in the image data, and a width in a second direction perpendicular to the first direction is smaller than a predetermined first threshold, as an independent character region; a character region that is adjacent to the plurality of independent character regions arranged in the first direction in the second direction and is not an independent character region, is set as a reference region; determining whether a character string formed by each character of the plurality of independent character regions aligned in the first direction is one word based on the width of the reference region in the first direction; When the control unit determines that a character string consisting of each character of the multiple independent character areas aligned in the first direction is a single word, it combines each character of the multiple independent character areas aligned in the first direction into a single character string and treats the combined character string as a single word.

2. when an absolute value of a difference between a coordinate value of one end of the first independent character region in the second direction and a coordinate value of one end of the second independent character region in the second direction is smaller than a predetermined second threshold, the control unit determines that the first independent character region and the second independent character region are aligned in the first direction, and classifies the first independent character region and the second independent character region into the same group; 2 . The image processing device according to claim 1 , wherein the control unit sets the reference area for each of the groups and determines whether a character string composed of characters in the plurality of independent character areas aligned in the first direction is one word.

3. the control unit recognizes coordinate values ​​of one end and the other end of the reference area in the first direction; 2. The image processing device of claim 1, wherein when the coordinate values ​​of one end of each of the plurality of independent character areas aligned in the first direction are within a range from the coordinate values ​​of one end of the reference area in the first direction to the coordinate values ​​of the other end, the control unit determines that a character string composed of each character of the plurality of independent character areas aligned in the first direction is a single word.

4. The control unit a coordinate value of one end in the second direction of the independent character region having the smallest coordinate value of one end in the first direction among the plurality of independent character regions arranged in the first direction is set as a first reference value; a first difference value, which is an absolute value of a difference between a coordinate value of one end in the second direction and the first reference value, for each of the plurality of character regions that are not the independent character region; The image processing apparatus according to claim 1 , wherein the character area where the first difference value is equal to or smaller than a predetermined value is set as the reference area.

5. 5. The image processing device according to claim 4, wherein the predetermined value is a value obtained by multiplying the width in the second direction of the independent character region having the smallest coordinate value of one end in the first direction among the plurality of independent character regions aligned in the first direction by a predetermined number.

6. The control unit an absolute value of a difference between a coordinate value of one end of the reference area in the second direction and the first reference value is set as a second reference value; calculating a second difference value which is the absolute value of the difference between the coordinate value of one end in the second direction of the character region that is adjacent to the reference region in the two directions but is not the independent character region and the coordinate value of one end in the second direction of the reference region; further setting the character region where the absolute value of the difference between the second difference value and the second reference value is equal to or less than a predetermined value as the reference region; When there are a plurality of reference areas, the control unit recognizes coordinate values ​​of one end and the other end in the first direction of a maximum reference area that is the longest in the first direction among the plurality of reference areas, and 5. The image processing device of claim 4, wherein when the first-direction position of each of the plurality of independent character areas aligned in the first direction is within a range from one end position to the other end position in the first direction of the maximum reference area, the control unit determines that a character string composed of each character of the plurality of independent character areas aligned in the first direction is a single word.

7. When the character areas that are not the independent character areas are adjacent to the plurality of independent character areas aligned in the first direction on both one side and the other side in the second direction, the control unit sets the plurality of character areas that are adjacent to the plurality of independent character areas aligned in the first direction in the second direction as the reference areas, respectively; When there are a plurality of reference areas, the control unit recognizes coordinate values ​​of one end and the other end in the first direction of a maximum reference area that is the longest in the first direction among the plurality of reference areas, and 2. The image processing device of claim 1, wherein when the first-direction position of each of the plurality of independent character areas aligned in the first direction is within a range from one end position to the other end position in the first direction of the maximum reference area, the control unit determines that a character string composed of each character of the plurality of independent character areas aligned in the first direction is a single word.

Citation Information

Patent Citations

  • Image processing device for character input via touch panel, control method therefor, and program

    JP2020177521A