Character Recognition System and Character Recognition Method

The character recognition system effectively addresses the challenge of recognizing handwritten symbols and characters by grouping and coding them, facilitating efficient search and classification within document images.

JP7707705B2Active Publication Date: 2025-07-15KONICA MINOLTA INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2021115057
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-07-12
Publication Date
2025-07-15
Estimated Expiration
2041-07-12

AI Technical Summary

Technical Problem

Conventional character recognition systems struggle to recognize and digitize handwritten special symbols and characters, as they are not typically registered in advance, leading to inefficiencies in search and classification processes.

Method used

A character recognition system and method that includes an image reading unit, character recognition unit, and storage unit, capable of extracting and grouping unrecognized characters, assigning a predetermined character code to similar characters, and storing the data for efficient search and classification.

Benefits of technology

Enables the extraction and digitization of handwritten special symbols and characters, allowing for efficient search and classification, even when they are included in document images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007707705000001
    Figure 0007707705000001
  • Figure 0007707705000002
    Figure 0007707705000002
  • Figure 0007707705000003
    Figure 0007707705000003
Patent Text Reader

Abstract

To provide a character recognition system and a character recognition method that enable extraction and data conversion of symbols and characters even if hand-written special symbols and special characters are included in an original image, and efficiently perform retrieval, classification, and the like even for the symbols and characters.SOLUTION: A character recognition system 1 includes: an image formation device 10 having an image reading unit 13 for reading an image of a manuscript; a character recognition processing unit 223 for extracting characters included in the image of the manuscript read by the image reading unit 13 to perform character recognition of the extracted character; and an information processing device having a storage unit 26 for storing results of character recognition relative to the image of the manuscript. The character recognition processing unit 223 extracts unrecognized characters that cannot be recognized by character recognition, allocates prescribed character codes to the unrecognized characters, and outputs data associating the unrecognized characters and the prescribed character codes to the storage unit 26 to store them.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a character recognition system and a character recognition method.

Background Art

[0002] Conventionally, in a character recognition system, by converting a character image included in a manuscript image into character data and storing it, search, classification, etc. by characters can be performed, and the storage and reusability of the manuscript are improved. Therefore, various character recognition technologies have been proposed conventionally (for example, see Patent Document 1).

[0003] Patent Document 1 discloses an information processing apparatus that performs optical character recognition (OCR) processing on image data and converts it into document data. In the information processing apparatus disclosed in Patent Document 1, for example, a special-shaped character string such as a company logo is registered in advance as a specific pattern. Then, the information processing apparatus performs matching between the registered specific pattern and an image including the specific pattern, thereby eliminating misdetection of a special-shaped character string such as a company logo and improving the accuracy of the OCR analysis result.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] Incidentally, when the document is, for example, a memo or the like, an image of a special symbol or special character handwritten in the memo may be included in the document image. Since such handwritten special symbols and special characters are usually not registered in the character recognition system, they cannot be recognized by the character recognition system or are recognized as another character. In this case, problems occur such that it is impossible to perform character-based search and classification on unrecognizable characters such as handwritten special symbols and special characters, and they cannot be utilized as digital data. In addition, since the writing style and shape of handwritten special symbols and special characters vary depending on the writer, it is difficult to register these handwritten symbols and characters in the character recognition system in advance. Note that the technique disclosed in Patent Document 1 described above is a character recognition technique that performs matching between a special-shaped character string such as a pre-registered company logo and an image including the specific pattern, and thus cannot solve the above problems. Therefore, conventionally, there has been a demand for the development of a character recognition system that enables extraction and data conversion of handwritten special symbols and special characters even when they are included in a document image.

[0006] The present invention has been made to solve the above problems. An object of the present invention is to provide a character recognition system and a character recognition method that enable extraction and data conversion of such symbols and characters and can efficiently perform search, classification, etc. on those symbols and characters even when a document image includes handwritten special symbols and special characters.

Means for Solving the Problems

[0007] The character recognition system of the present invention includes an image reading unit that reads an image of a document, a character recognition unit that extracts characters included in the image of the document read by the image reading unit and performs character recognition on the extracted characters, and a storage unit that stores the result of character recognition for the image of the document. The character recognition unit extracts unrecognized characters that cannot be recognized by character recognition, and outputs and stores the data for the unrecognized characters in the storage unit. Associate similar characters, compare the similar characters associated with each unrecognized character among a plurality of unrecognized characters, group at least a part of the plurality of unrecognized characters based on the comparison result, assign the same predetermined character code to each grouped unrecognized character, and assign the same predetermined character code to each grouped unrecognized character The data is output to the storage unit for storage. In addition, the character recognition system of the present invention includes an image reading unit that reads an image of a document, a character recognition unit that extracts characters included in the image of the document read by the image reading unit and performs character recognition on the extracted characters, and a storage unit that stores the result of character recognition for the image of the document. When there are a plurality of images of documents read by the image reading unit, the character recognition unit performs character recognition for each image of a single document, extracts unrecognized characters that cannot be recognized by character recognition, and groups the unrecognized characters extracted from a predetermined document image among the plurality of document images and the unrecognized characters extracted from a document image different from the predetermined document image among the plurality of document images, assigns the same predetermined character code to each grouped unrecognized character, and outputs and stores the data with the same predetermined character code assigned to each grouped unrecognized character in the storage unit.

[0008] Further, the character recognition method of the present invention is a character recognition method in a character recognition system including an image reading unit that reads an image of a document, a character recognition unit that extracts characters included in the image of the document read by the image reading unit and performs character recognition on the extracted characters, and a storage unit that stores the result of character recognition for the image of the document. The character recognition unit extracts unrecognized characters that cannot be recognized by character recognition, and Associate similar characters, compare the similar characters associated with each unrecognized character among a plurality of unrecognized characters, group at least a part of the plurality of unrecognized characters based on the comparison result, assign the same predetermined character code to each grouped unrecognized character, and assign the same predetermined character code to each grouped unrecognized character outputs and stores the data obtained for the unrecognized characters in the storage unit. Also, the character recognition method of the present invention is a character recognition method in a character recognition system including an image reading unit that reads an image of a document, a character recognition unit that extracts characters included in the image of the document read by the image reading unit and performs character recognition on the extracted characters, and a storage unit that stores the result of character recognition for the image of the document. When there are a plurality of images of the document read by the image reading unit, the character recognition unit performs character recognition for each image of one document, extracts unrecognized characters that cannot be recognized by the character recognition, groups the unrecognized characters extracted from the image of a predetermined document among the plurality of images of the document and the unrecognized characters extracted from an image of a document different from the predetermined document among the plurality of images of the document, assigns the same predetermined character code to each grouped unrecognized character, and outputs and stores in the storage unit the data in which the same predetermined character code is assigned to each grouped unrecognized character.

Advantages of the Invention

[0009] According to the present invention having the above configuration, even if a manuscript image includes handwritten special symbols or special characters, etc., it is possible to extract and digitize those symbols and characters, and it is also possible to efficiently perform searches, classifications, etc. on those symbols and characters.

Brief Description of the Drawings

[0010]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Embodiments for Carrying Out the Invention

[0011] Hereinafter, the configuration of the character recognition system and the character recognition method according to an embodiment of the present invention will be specifically described with reference to the drawings. Note that the present invention is not limited to the following examples.

[0012] [An example of a manuscript image including special symbols to be recognized] First, before explaining the configuration of the character recognition system and the character recognition method of the present embodiment, an example of an image of a manuscript including handwritten special symbols to be recognized in the character recognition system of the present embodiment will be described. FIG. 1 is a diagram showing an example of an image of a manuscript (note) including handwritten special symbols.

[0013] In the example shown in FIG. 1, one manuscript image N includes images of four notes N1 to N4. Also, in the example shown in FIG. 1, in the center of each image of notes N1 to N4, text information TX1 to TX4 indicating the content of the notes is shown respectively, and handwritten special symbols SC1 to SC4 are shown above the text information TX1 to TX4 respectively. The special symbols SC1 to SC3 shown in each of the notes N1 to N3 are special symbols that respectively enclose the handwritten character "heavy" with a circle, and the special symbol SC4 shown in the note N4 is a special symbol imitating the shape of the moon.

[0014] Handwritten special symbols such as the special symbols SC1 to SC4 are symbols written for the purpose of organizing the memory of the person who created the note and for subsequent classification, etc., and it is usually difficult to register them in the character recognition system in advance. However, the character recognition system of the present embodiment extracts (identifies) the special symbols SC1 to SC4 as unrecognized characters in the character recognition process, assigns (associates) a predetermined character code to the unrecognized characters, and thereby digitizes the unrecognized characters.

[0015] Also, for example, although the shapes of the special symbols SC1 to SC3 are different from each other, the meanings of each symbol are the same. Therefore, if such a plurality of special symbols having the same meaning are recognized as different special symbols (character codes) respectively, these plurality of special symbols cannot be managed collectively. The character recognition system of the present embodiment assigns and manages the same character code to a plurality of special symbols having the same meaning.

[0016] [Configuration of Character Recognition System] FIG. 2 is a block diagram showing the configuration of the character recognition system according to the present embodiment. As shown in FIG. 2, the character recognition system 1 includes an image forming apparatus 10 and an information processing apparatus 20. The image forming apparatus 10 and the information processing apparatus 20 can transmit and receive information data to and from each other via a network 30. Note that FIG. 2 shows an example in which the character recognition system 1 includes one image forming apparatus 10, but the present invention is not limited to this, and the character recognition system 1 may include a plurality of image forming apparatuses.

[0017] When performing character recognition by the character recognition system 1, first, the image forming apparatus 10 reads an image of a document to be subjected to character recognition and generates document image data. Next, the image forming apparatus 10 transmits the generated document image data (hereinafter referred to as "document image") to the information processing apparatus 20 via the network 30. Next, the information processing apparatus 20 performs character recognition processing on the document image received via the network 30 and outputs the result of the character recognition processing.

[0018] Note that, in the present embodiment, an example in which a functional unit for character recognition processing (a character recognition apparatus unit 22 described later) is provided in the information processing apparatus 20 will be described, but the present invention is not limited to this. For example, the functional unit for character recognition processing may be provided inside the image forming apparatus 10. In this case, the character recognition system 1 is configured only by the image forming apparatus 10.

[0019] [Configuration of Image Forming Apparatus] The image forming apparatus 10 is a multi-functional peripheral (MFP) equipped with a plurality of functions such as a scanner function, a copy function, a facsimile function, a network function, and a data box function. The image forming apparatus 10 includes an operation display unit 11, an automatic document feeder (ADF) 12, an image reading unit 13, a printer unit 14, a central processing unit (CPU) 15, a read only memory (ROM) 16, a random access memory (RAM) 17, a storage unit 18, a communication unit 19, and a bus 101. The bus 101 electrically connects each component and is a signal path for input and output of signals between each component.

[0020] The operation display unit 11 includes a display unit composed of a display device such as a liquid crystal display (LCD) or an organic electro-luminescence (EL) display, and an operation unit composed of a touch sensor or the like. The display unit and the operation unit are integrally formed as a touch panel, for example. The operation display unit 11 generates an operation signal representing the operation content from the operator input to the operation unit and supplies the operation signal to the CPU 15. Further, the operation display unit 11 displays the operation content and setting information by the operator on the display unit based on the display signal supplied from the CPU 15. Note that the operation unit may be configured by a mouse, a tablet, or the like and may be configured separately from the display unit.

[0021] The automatic document feeder 12 is composed of a placement tray for placing a recording medium, a mechanism for transporting the recording medium, a transport roller, etc., and transports the recording medium to a predetermined transport path.

[0022] The image reading unit 13 optically reads the image of the document fed by the automatic document feeder 12, and performs A / D (Analog to Digital) conversion on the read image to generate image data (document image). The document image read by the image reading unit 13 is transmitted to the information processing apparatus 20 via the communication unit 19 and the network 30. Note that the image reading unit 13 can also read an image from a document on the platen glass.

[0023] The printer unit 14 is composed of components necessary for image formation, and prints a predetermined image on a recording medium based on the designation information of the print job. Specifically, the printer unit 14 forms an electrostatic latent image by irradiating the photosensitive drum charged by the charging device with light corresponding to the image from the exposure device on the recording medium, and develops the image by attaching the charged toner with the developing device. Then, the printer unit 14 performs a process of primarily transferring the developed toner image to the transfer belt, secondarily transferring it from the transfer belt to the recording medium, and further fixing the toner image on the recording medium with the fixing device.

[0024] The CPU 15 controls the operations of each part within the image forming apparatus 10. For example, the CPU 15 controls the document image reading process of the image reading unit 13, the image forming process of the printer unit 14 based on the print instruction of the user performed via the operation display unit 11, the data information transmission and reception process between the information processing apparatus 20 by the communication unit 19, and the like.

[0025] The ROM 16 is composed of a storage medium such as a non-volatile memory, and stores programs, data, etc. that the CPU 15 executes and references.

[0026] The RAM 17 is composed of a storage medium such as a volatile memory, and temporarily stores information (data) necessary for each process performed by the CPU 15.

[0027] The storage unit 18 is composed of a computer-readable non-transitory recording medium storing programs executed by the CPU 15, and is composed of a storage device such as an HDD (Hard Disk Drive), for example. The storage unit 18 stores programs for the CPU 15 to control each unit, an OS (Operating System), programs such as a controller, and data. Note that part of the programs and data stored in the storage unit 18 may be stored in the ROM 16. Also, the computer-readable non-transitory recording medium storing the programs executed by the CPU 15 is not limited to an HDD, and may be, for example, a recording medium such as an SSD (Solid State Drive), a CD (Compact Disc)-ROM, or a DVD (Digital Versatile Disc)-ROM.

[0028] The communication unit 19 transmits and receives various data information with an external device connected via the network 30. The communication unit 19 is, for example, a communication interface having a communication IC (Integrated Circuit) for communication and a communication connector, etc., and can transmit and receive various data information with the information processing device 20 connected via the network 30 using a predetermined communication protocol under the control of the CPU 15. Also, the communication unit 19 has a configuration such as an antenna, a demodulation circuit, a signal processing circuit, etc., and can transmit and receive various data information by wireless communication with the information processing device 20 connected via the network 30 by a wireless communication method such as Wi-Fi (registered trademark).

[0029] [Configuration of Information Processing Device] The information processing device 20 receives the document image read by the image reading unit 13 of the image forming device 10 via the network 30, and performs character recognition processing on the received document image. As shown in FIG. 2, the information processing device 20 has an operation display unit 21, a character recognition device unit 22, a CPU 23, a ROM 24, a RAM 25, a storage unit 26, a communication unit 27, and a bus 201. The bus 201 electrically connects each component, and is a signal path through which signals are input and output between each component.

[0030] The operation display unit 21 is composed of a display unit including a display device such as an LCD or an organic EL display, and an operation unit including a touch sensor or the like. The display unit and the operation unit are integrally formed as a touch panel, for example. The operation display unit 21 generates an operation signal representing the operation content from the operator input to the operation unit, and supplies the operation signal to the CPU 23. Further, the operation display unit 21 displays on the display unit the operation content, setting information, etc. by the operator based on the display signal supplied from the CPU 23. Further, the display unit can display the recognition result of the character recognition process (text information etc. shown in FIG. 11 described later) for the manuscript image output from the character recognition device unit 22. Note that the operation unit may be configured by a mouse, a tablet, etc. and may be configured separately from the display unit.

[0031] The character recognition device unit 22 (character recognition unit) performs a predetermined character recognition process on the manuscript image to be recognized for character recognition received from the image forming apparatus 10 via the network 30 under the control of the CPU 23. The character recognition process of the present embodiment adopts a processing method used in OCR, that is, a method of performing layout analysis, line cutting, character cutting, character recognition, and result output on the manuscript image in this order. However, the present invention is not limited to this, and any method can be adopted as long as the method of the character recognition process can extract various special symbols of handwritten characters etc. described in FIG. 1 from the manuscript image. In the present embodiment, the function of the character recognition device unit 22 is configured by software, but the character recognition device unit 22 may be configured by an arithmetic processing device dedicated to character recognition (hardware). In this case, the character recognition process is controlled within the arithmetic processing device dedicated to character recognition. The configuration and function of the character recognition device unit 22 will be described in detail later.

[0032] The CPU 23 controls the operations of each part within the information processing apparatus 20. Therefore, the CPU 23 controls the operation of the character recognition process performed by the character recognition apparatus unit 22. Specifically, the CPU 23 controls the process of analyzing the layout of the manuscript image by the layout analysis unit 221 described later, the character extraction process by the character extraction unit 222 described later, and the grouping and character code assignment process by the character recognition processing unit 223 described later, etc.

[0033] The ROM 24 is composed of a storage medium such as a non-volatile memory, and stores programs, data, etc. that the CPU 23 executes and refers to.

[0034] The RAM 25 is composed of a storage medium such as a volatile memory, and temporarily stores information (data) necessary for each process performed by the CPU 23. Also, the RAM 25 temporarily stores, for example, unrecognized characters extracted by the character recognition processing unit 223 (refer to FIG. 8 described later), etc.

[0035] The storage unit 26 is composed of a computer-readable non-transitory recording medium that stores the programs executed by the CPU 23, and is composed of a storage device such as an HDD, for example. The storage unit 26 stores programs for the CPU 23 to control each part, an OS (Operating System), programs such as a controller, and data. Also, the storage unit 26 stores various data generated in various processes related to character recognition performed by the character recognition apparatus unit 22. Specifically, the storage unit 26 stores text data (refer to FIG. 11 described later) and output data related to special symbols (refer to FIGS. 9, 10, 12, and 13 described later) obtained as the recognition result of the character recognition process performed by the character recognition apparatus unit 22. Note that a part of the programs and data stored in the storage unit 26 may be stored in the ROM 24. Also, the computer-readable non-transitory recording medium that stores the programs executed by the CPU 23 is not limited to an HDD, and may be, for example, a recording medium such as an SSD, a CD-ROM, or a DVD-ROM.

[0036] The communication unit 27 transmits and receives various data information with an external device connected via the network 30. The communication unit 27 is, for example, a communication interface having a communication IC, a communication connector, etc., and can transmit and receive various data information with the image forming apparatus 10 connected via the network 30 using a predetermined communication protocol under the control of the CPU 23. Further, the communication unit 27 has a configuration such as an antenna, a demodulation circuit, a signal processing circuit, etc., and can also transmit and receive various data information by wireless communication with the image forming apparatus 10 connected via the network 30 by a wireless communication method such as Wi-Fi (registered trademark).

[0037] [Configuration and Function of Character Recognition Device Unit] As shown in FIG. 2, the character recognition device unit 22 (character recognition unit) includes a layout analysis unit 221, a character extraction unit 222, and a character recognition processing unit 223.

[0038] (1) Layout Analysis Unit The layout analysis unit 221 analyzes the layout (arrangement) of regions such as an image group and a character group in the document image based on the document image input from the image forming apparatus 10. Specifically, the layout analysis unit 221 first separates each component in the document image, classifies the separated components according to the type of their content (character, drawing, table, photograph, line, etc.), and analyzes the layout of each region such as the character region and the figure region in the document image.

[0039] FIG. 3 is a diagram showing an example of a document image 100 read by the image forming apparatus 10 of the character recognition system 1. In the example shown in FIG. 3, the document image 100 has, as components, a drawing portion 110 of a graph arranged in the left half region in the document image 100, and text portions 120 and character portions 130 arranged in the upper right and lower right of the document image 100, respectively. Further, the document image 100 also has a ruled line 140 between the text portion 120 and the character portion 130 as a component. Note that the components of the document image 100 are not limited to these, and may include other components such as a table and a photograph.

[0040] FIG. 4 is a diagram showing an example of the analysis result of the layout of the manuscript image 100 performed by the layout analysis unit 221. In FIG. 4, for convenience of explanation, only the arrangement areas of the drawing part 110, the text part 120, the character part 130, and the ruled line 140 constituting the manuscript image 100 are indicated by a one-dot chain line, and the description of the contents of each constituent part shown in FIG. 3 is omitted. In FIG. 4, the arrangement areas (drawing area 210, character area 220, and character area 230) of the drawing part 110, the text part 120, and the character part 130 analyzed by the layout analysis unit 221 are each shown by a rectangular frame. Also, in FIG. 4, the arrangement area (line area 240) of the ruled line 140 analyzed by the layout analysis unit 221 is shown by a straight line.

[0041] Next, the layout analysis unit 221 extracts the character areas from the arrangement areas (drawing area 210, character area 220, character area 230, and line area 240) of each component in the analyzed (classified) manuscript image 100. In the example shown in FIG. 4, the character area 220 and the character area 230 are extracted. Then, the data of the character areas (character area 220 and character area 230) extracted by the layout analysis unit 221 is output to the character extraction unit 222.

[0042] (2) Character Extraction Unit The character extraction unit 222 decomposes the character area in the manuscript image 100 input from the layout analysis unit 221 into each line (in the case of horizontal writing) or each column (in the case of vertical writing). That is, the character extraction unit 222 extracts the data of the character strings of each row or each column constituting the character area from the character area.

[0043] Fig. 5 is a diagram showing an example of a character string extraction operation (decomposition operation) performed by character extraction unit 222. The example shown in Fig. 5 is a result when each character string constituting character portion 130 is extracted from the contents of character area 230 shown in Fig. 4, i.e., character portion 130 shown in Fig. 3. The area surrounded by a dashed line in Fig. 5 indicates the area of the character string of each line to be extracted. For example, in the third line from the right of character portion 130, character string 231 "business partner company on a fixed day" is extracted by character extraction unit 222.

[0044] Furthermore, the character segmentation unit 222 breaks down the character strings of each row or column into individual characters. That is, the character segmentation unit 222 segments data of each character constituting the character string from the character strings of each row or column.

[0045] Fig. 6 is a diagram showing an example of a method of cutting out (decomposing) each character and a result of the cutting out performed by the character cutting unit 222. Note that the example shown in Fig. 6 shows the cutting out process of each character performed on the character string 231 "business partner company on a fixed day" cut out from the character portion 130 shown in Fig. 5.

[0046] In the process of extracting each character, the character extraction unit 222 first moves the scanning line L1 in a direction from the beginning to the end of the character string 231 to be processed (the direction of the arrow L2 in FIG. 6). At this time, the character extraction unit 222 counts the number of intersections (number of intersection dots) between the scanning line L1 and the constituent parts of each character. FIG. 6(a) is a diagram showing the change in the number of intersections with respect to the scanning direction of the scanning line L1. As shown in FIG. 6(a), in the section between adjacent characters, the number of intersections is "0" consecutively.

[0047] When the interval where the intersection number is "0" consecutively exceeds a predetermined threshold, the character segmentation unit 222 determines that the interval where the intersection number is "0" consecutively is a character interval, breaks down the character string into each character interval, and segments each character. As a result, as shown in Fig. 6(b), the character string 231 "te kette wa hi tou koto kaisha" is broken down into 12 characters (area surrounded by dashed lines in the figure) "te", "kette", "ma", "tsu", "ta", "hi", "ni", "tori", "biki", "saki", "ki" and "gyo".

[0048] The character segmentation unit 222 performs the above-mentioned process of segmenting a character string of one row or one column and the process of segmenting each character from the character string for all character regions constituting the original image 100. Then, each character image data segmented by the character segmentation unit 222 is output to the character recognition processing unit 223.

[0049] (3) Character recognition processing section First, the character recognition processing unit 223 normalizes the size of each character image input from the character segmentation unit 222 to the size (predetermined size) of a character for matching registered in the character recognition system 1. Then, the character recognition processing unit 223 performs character recognition by matching each normalized character image with a character already registered in the character recognition system 1 (hereinafter referred to as a "registered character").

[0050] Fig. 7 is a diagram for explaining the character normalization process performed in the character recognition processing unit 223. Character image T1 shown in Fig. 7 is, for example, an image of the character "業" before normalization input from the character segmentation unit 222, and character image T2 is an image of the character "業" after normalization. Fig. 7 shows an example in which character image T1 input from the character segmentation unit 222 is normalized to character image T2 with a height of "a" and a width of "a". Note that the size (predetermined size) of character image T2 after normalization is the size of a character registered in the character recognition system 1, and therefore changes depending on, for example, the specifications of the character recognition system 1.

[0051] After the text normalization process, the character recognition processing unit 223 collates the normalized character image with registered characters. In the result of the collation, when the degree of match between the normalized character image and a specific registered character exceeds a predetermined threshold, the character recognition processing unit 223 determines that the character recognition is successful and converts the normalized character image into the specific registered character. That is, the character recognition processing unit 223 digitizes the character image cut out from the manuscript image 100. Note that, as the collation method between the normalized character image and the registered character, that is, the calculation method of the degree of match between the two, for example, a matching method used in conventional OCR can be adopted.

[0052] On the other hand, in the result of the character collation, when the degree of match of the normalized character image is equal to or less than a predetermined threshold for all registered characters, or when the degree of match exceeds a predetermined threshold for a plurality of registered characters, the character recognition processing unit 223 determines the normalized character image as an unrecognized character. Note that, in the character recognition system 1 of the present embodiment, the handwritten special symbols SC1 to SC4 in the manuscript image N shown in FIG. 1 usually do not have registered characters whose degree of match with them exceeds a predetermined threshold, that is, registered characters recognized as the same. Therefore, in the character recognition by the character recognition processing unit 223, they are recognized and determined as unrecognized characters. Then, the unrecognized characters are temporarily stored in the RAM 25.

[0053] After temporarily storing the unrecognized characters in the RAM 25, the character recognition processing unit 223 searches for registered characters (hereinafter referred to as "similar characters") similar to the unrecognized characters based on the degree of match obtained in the collation process between the unrecognized characters and the registered characters. Then, the character recognition processing unit 223 associates, for example, similar characters having a degree of match exceeding the above-described predetermined threshold with the unrecognized characters. At this time, only one similar character may be associated with the unrecognized character, or a plurality of similar characters may be associated with the unrecognized character. In the former case, only the similar character having the highest degree of match is associated with the unrecognized character. On the other hand, in the latter case, a specific number of similar characters may be associated with the unrecognized character in descending order of the degree of match, or all similar characters exceeding a specific threshold may be associated with the unrecognized character. Further, characters having a degree of match not exceeding a specific threshold may be associated with the unrecognized character.

[0054] Furthermore, the character recognition processing unit 223 generates a data set (hereinafter referred to as a "similar character data set") that combines unrecognized characters, similar characters or similar character groups associated with the unrecognized characters, and the degree of agreement of each similar character. The generated similar character data set is stored in the storage unit 26. FIG. 8 is a diagram showing an example of a similar character data set generated for each of the handwritten special symbols SC1 to SC4 (unrecognized characters) in the document image N shown in FIG. 1. FIG. 8 shows an example of the configuration of a similar character data set when three similar characters are associated with each special symbol that is an unrecognized character. For example, in the similar character data set for the special symbol SC1, three similar characters, "重", "動" and "動", associated with the special symbol SC1, are associated with the degrees of agreement of the similar characters ("0.75", "0.61" and "0.55"). Since similar character data sets are configured in a similar manner for the special symbols SC2 to SC4, the configuration of the special symbols SC2 to SC4 will not be described here. In addition, in the configuration of the similar character data set shown in FIG. 8, an example is shown in which similar characters are arranged in descending order of the degree of match, but the order in which similar characters are arranged is arbitrary.

[0055] When an unrecognized character is extracted in the character recognition process described above, the character recognition processing unit 223 performs a process of assigning a specific character code to the extracted unrecognized character. Then, in a process of outputting the result of character recognition described below, the character recognition processing unit 223 outputs a data set in which the extracted unrecognized character is associated with the assigned character code to the storage unit 26. That is, the data set in which the unrecognized character is associated with the assigned character code is stored and registered in the storage unit 26. As a result, the unrecognized character extracted by character recognition is converted into digital data.

[0056] In the character code assignment process, the character recognition processing unit 223 first determines whether a character code corresponding to an unrecognized character (unrecognized character group) having the same meaning as the unrecognized characters (unrecognized character group) extracted this time has already been registered in the character recognition system 1. And when a character code corresponding to an unrecognized character (unrecognized character group) having the same meaning as the unrecognized characters (unrecognized character group) extracted this time has already been registered in the character recognition system 1, the character recognition processing unit 223 assigns the registered specific character code to the unrecognized characters (unrecognized character group) extracted this time. That is, the unrecognized characters (unrecognized character group) extracted this time and the previously registered unrecognized characters are grouped together.

[0057] On the other hand, when there is one unrecognized character in the unrecognized character group extracted this time that does not correspond to the registered unrecognized characters, the character recognition processing unit 223 assigns a new character code to the unrecognized character extracted this time. Also, when there are a plurality of unrecognized characters in the unrecognized character group extracted this time that do not correspond to the registered unrecognized characters, the character recognition processing unit 223 performs a grouping process on the plurality of unrecognized characters extracted this time. And when at least some of the plurality of unrecognized characters are grouped, the character recognition processing unit 223 assigns the same new character code to the grouped unrecognized characters. Also, when there are unrecognized characters that are not grouped by this grouping process, the character recognition processing unit 223 assigns another new character code to the unrecognized characters.

[0058] In this embodiment, in the process of assigning character codes to unrecognized characters, as described above, the same character code is assigned to a plurality of unrecognized characters having the same meaning (content), the plurality of unrecognized characters are grouped, and an unrecognized character group is generated. And in the grouping process for a plurality of unrecognized characters, the similar characters associated with each unrecognized character are compared, and the unrecognized characters having overlapping similar characters with each other are put into one group.

[0059] Therefore, for example, in the example shown in FIG. 8, the special symbols SC1 to SC4 included in the document image N are extracted as unrecognized characters, but in the process of assigning character codes to each unrecognized character, the special symbols SC1 to SC3, which are overlapping similar characters "重" and "動", are grouped as unrecognized characters having the same meaning, and a specific character code is assigned to the grouped unrecognized characters. On the other hand, the three similar characters associated with the special symbol SC4 do not include any characters that overlap with the similar characters associated with the special symbols SC1 to SC3, so a new character code different from the specific character code assigned to the special symbols SC1 to SC3 is assigned to the special symbol SC4. When multiple unrecognized characters are grouped, a data set that associates the grouped unrecognized characters with the character codes assigned thereto is output to and stored in the storage unit 26 in the process of outputting the result of character recognition described later. Note that the method of grouping multiple unrecognized characters is arbitrary.

[0060] Character recognition processing unit 223 converts all character images in the character area of the document image, including special symbols, into character codes by the character recognition processing described above, and then outputs the result of the character recognition processing. In this output processing, character recognition processing unit 223 outputs data on the result of the character recognition of the document image to storage unit 26 and stores it. In addition, in this output processing, character recognition processing unit 223 may output data on the result of the character recognition of the document image to storage unit 26 and to operation display unit 21, so that the result of the character recognition is displayed on the display unit of operation display unit 21. In addition, character recognition processing unit 223 may output data on the result of the character recognition of the document image to storage unit 26 and to a printing unit (not shown) so that the result of the character recognition is printed on a printing sheet.

[0061] FIG. 9 is a diagram showing an example of text data which is the result data of character recognition output from the character recognition processing unit 223. FIG. 9 shows the result data of character recognition output when character recognition is performed on the original manuscript image N shown in FIG. 1. The result data of character recognition output from the character recognition processing unit 223 is text data as shown in FIG. 9. Also, in the example shown in FIG. 9, it is assumed that character recognition is performed in the order of memo N1, memo N2, memo N3, and memo N4 within the original manuscript image N.

[0062] The text group starting with the name "<Block 1>" in FIG. 9 corresponds to the result of character recognition of memo N1 in the original manuscript image N shown in FIG. 1. And in "<Paragraph 1>" within <Block 1>, the recognition result of the special symbol SC1 within memo N1 is shown, and in "<Paragraph 2>", the recognition result of the text information TX1 within memo N1 is shown. In this example, it is assumed that the character code "code a" is assigned to the special symbol SC1 within memo N1 extracted as an unrecognized character. Therefore, in "<Paragraph 1>" within <Block 1>, the character code "code a" assigned to the special symbol SC1 is shown as the recognition result. Also, in "<Paragraph 2>" within <Block 1>, the same sentence (character string) as the content of the text information TX1 (refer to FIG. 1) within memo N1 is shown as the recognition result.

[0063] The text group starting with the name <Block 2> in FIG. 9 corresponds to the result of character recognition of the memo N2 in the manuscript image N shown in FIG. 1. And in <Paragraph 1> within <Block 2>, the recognition result of the special symbol SC2 in the memo N2 is shown, and in <Paragraph 2>, the recognition result of the text information TX2 in the memo N2 is shown. In this example, since the special symbol SC2 in the memo N2 extracted as an unrecognized character has the same meaning as the special symbol SC1 in the memo N1, the two are grouped and the same character code is assigned. Therefore, in <Paragraph 1> within <Block 2>, the character code "code a" is shown as the recognition result. Also, in <Paragraph 2> within <Block 2>, the same sentence (character string) as the content of the text information TX2 (see FIG. 1) in the memo N2 is shown as the recognition result.

[0064] The text group starting with the name <Block 3> in FIG. 9 corresponds to the result of character recognition of the memo N3 in the manuscript image N shown in FIG. 1. And in <Paragraph 1> within <Block 3>, the recognition result of the special symbol SC3 in the memo N3 is shown, and in <Paragraph 2>, the recognition result of the text information TX3 in the memo N3 is shown. In this example, since the special symbol SC3 in the memo N3 extracted as an unrecognized character has the same meaning as the special symbol SC1 in the memo N1 and the special symbol SC2 in the memo N2, these special symbols are grouped and the same character code is assigned. Therefore, in <Paragraph 1> within <Block 3>, the character code "code a" is shown as the recognition result. Also, in <Paragraph 2> within <Block 3>, the character string that is the same as the content of the text information TX3 (see FIG. 1) in the memo N3 is shown as the recognition result.

[0065] The text group starting with the name <Block 4> in FIG. 9 corresponds to the result of character recognition of the memo N4 in the manuscript image N shown in FIG. 1. And in <Paragraph 1> within <Block 4>, the recognition result of the special symbol SC4 in the memo N4 is shown, and in <Paragraph 2>, the recognition result of the text information TX4 in the memo N4 is shown. In this example, the special symbol SC4 in the memo N4 extracted as an unrecognized character is recognized as a special symbol different from the special symbols SC1 to SC3 in the memos N1 to N3. Therefore, in <Paragraph 1> within <Block 4>, a character code "code b", which is different from the "code a" assigned to the special symbols SC1 to SC3, is shown as the recognition result. Also, in <Paragraph 2> within <Block 4>, the same sentence (character string) as the content of the text information TX4 (see FIG. 1) in the memo N4 is shown as the recognition result.

[0066] Also, when unrecognized characters are extracted in the character recognition process, the character recognition processing unit 223 outputs and stores in the storage unit 26 the data associating the image data of the unrecognized characters with the character codes assigned thereto, that is, the correspondence data between the unrecognized characters and the character codes. Thereby, in the character recognition system 1, unrecognized characters can also be managed. Also, in this output process, the character recognition processing unit 223 may output the correspondence data between the unrecognized characters and the character codes to the storage unit 26 and also to the operation display unit 21 so that the correspondence between the unrecognized characters and the character codes is displayed on the display unit of the operation display unit 21. Also, the character recognition processing unit 223 may output the correspondence data between the unrecognized characters and the character codes to the storage unit 26 and also to a printing unit (not shown) so that the correspondence data is printed on a printing sheet.

[0067] The configuration (mode) of the correspondence data between the unrecognized characters output from the character recognition processing unit 223 and the character codes is arbitrary. Here, with reference to FIGS. 10 to 12, various modes (output examples 1 to 2) of the correspondence data between the unrecognized characters and the character codes will be described. Note that FIGS. 10 to 12 are diagrams showing various modes (output examples) of the correspondence data between the unrecognized characters and the character codes, and are diagrams showing a configuration example of the correspondence data between the unrecognized characters and the character codes output when character recognition is performed on the original manuscript image N shown in FIG. 1.

[0068] The correspondence data of output example 1 shown in FIG. 10 has a configuration in which character codes are separately associated with each of the special symbols SC1 to SC4 in the original manuscript image N. That is, the correspondence data of output example 1 shown in FIG. 10 is a table-like correspondence data that enumerates data sets of each special symbol and the character code associated therewith for each special symbol. Therefore, in the example shown in FIG. 10, the correspondence data between the unrecognized characters output from the character recognition processing unit 223 and the character codes includes a data set of the special symbol SC1 and the character code ("code a") associated therewith, a data set of the special symbol SC2 and the character code ("code a") associated therewith, a data set of the special symbol SC3 and the character code ("code a") associated therewith, and a data set of the special symbol SC4 and the character code ("code b") associated therewith. Note that in the example shown in FIG. 10, in the original manuscript image N, the special symbols SC1 to SC4 are extracted (determined) as unrecognized characters in this order, so in the correspondence table, the data of the special symbol SC1 is arranged at the top, and the data of the special symbol SC4 is arranged at the bottom.

[0069] The corresponding data of output example 2 shown in FIG. 11 is the corresponding data in the corresponding data of output example 1 shown in FIG. 10, taking into account the grouped special symbol group (Group I). In the corresponding data of output example 2 shown in FIG. 11, for a plurality of grouped special symbols, one special symbol is selected from among the plurality of special symbols, and only the data set of the selected special symbol and the character code associated therewith is included in the corresponding data. The method of selecting one special symbol from among a plurality of special symbols is arbitrary. For example, the special character first recognized as an unrecognized character among the plurality of special symbols may be selected, or the special character last recognized as an unrecognized character may be selected. In the example shown in FIG. 11, an example is shown in which the special symbol SC1 first recognized as an unrecognized character is selected from among the grouped special symbols SC1 to SC3 (Group I). Therefore, in the example shown in FIG. 11, the corresponding data of the unrecognized character and the character code output from the character recognition processing unit 223 includes the data set of the special symbol SC1 and the character code (``Code a'') associated therewith, and the data set of the special symbol SC4 and the character code (``Code b'') associated therewith.

[0070] As an output example of the corresponding data shown in FIGS. 10 and 11, information for identifying an area recognized as an unrecognized character in the manuscript image (for example, the placement area, etc.) and the character code associated therewith may be included in the corresponding data as a data set. FIG. 12 is a diagram showing the correspondence (output example 3) between the placement areas of special symbols SC1 to SC4 extracted as unrecognized characters in the manuscript image 100 and the character codes associated with each special symbol. In FIG. 12, the placement area of the special symbol and the character code associated therewith are shown connected by a dashed line. Although not shown in FIG. 12, the configuration of the corresponding data of output example 3 actually stored in the storage unit 26 is a table-like corresponding data in which a data set of an unrecognized character, information for identifying its placement area, and the character code associated therewith is listed for each special symbol as shown in FIG. 10. However, for example, when this corresponding data of output example 3 is to be displayed on the display unit of the operation display unit 21 or printed on a printing sheet, it is controlled to be displayed or printed in the display mode shown in FIG. 12. Note that the configuration of the corresponding data of output example 3 actually stored in the storage unit 26 may be a configuration in which a data set of an unrecognized character, information for identifying its placement area, the character code associated therewith, and the image data of the unrecognized character is listed in a table form.

[0071] Note that the output mode of the corresponding data indicating the correspondence between the special symbol and the character code assigned thereto is not limited to the output modes described in FIGS. 10 to 12, and any mode can be applied as long as it can clearly represent the correspondence between the special symbol and the character code assigned thereto.

[0072] [Procedure of Character Recognition Processing for One Manuscript Image] Next, the character recognition processing for one manuscript image performed in the character recognition system 1 will be described. FIG. 13 is a flowchart showing the procedure of the character recognition processing for one manuscript image performed in the character recognition system 1 of the present embodiment. The processing described below starts when one manuscript to be subjected to character recognition is fed to the image reading unit 13 of the image forming apparatus 10.

[0073] First, the image reading unit 13 of the image forming apparatus 10 reads an image of a single document to be subjected to character recognition (step S101). In this process, the image reading unit 13 reads the image of a single document and generates image data (document image) of the read document. Further, the generated document image is transmitted to the information processing apparatus 20 via the network 30.

[0074] Next, the layout analysis unit 221 of the character recognition apparatus unit 22 of the information processing apparatus 20 performs layout analysis processing on the document image received from the image forming apparatus 10 (step S102: refer to FIGS. 3 and 4). In this process, the layout analysis unit 221 extracts data of character regions included in the document image. Further, in this process, the layout analysis unit 221 outputs the extracted character regions to the character cutting unit 222 of the character recognition apparatus unit 22.

[0075] Next, the character cutting unit 222 cuts out character groups included in the character regions, row by row or column by column, for the character regions input from the layout analysis unit 221 (step S103: refer to FIG. 5).

[0076] Next, the character cutting unit 222 performs cutting processing for each character on each row or column of character strings cut out from the character regions (step S104: refer to FIG. 6). Further, in this process, the character cutting unit 222 outputs images of the cut-out characters to the character recognition processing unit 223.

[0077] Next, the character recognition processing unit 223 normalizes the size of the character image input from the character cutting unit 222 to a predetermined size (step S105: refer to FIG. 7). Note that the processing of steps S105 to step S111 described later is performed for each character.

[0078] Next, the character recognition processing unit 223 collates the normalized character image with registered characters stored in the storage unit 26 (step S106). In this process, the character recognition processing unit 223 calculates the degree of coincidence between the normalized character image and the registered characters.

[0079] Next, based on the collation result (degree of match) between the normalized character image and the registered characters, the character recognition processing unit 223 determines whether the degree of match between the normalized character image and a specific (one) registered character exceeds a predetermined threshold (step S107).

[0080] In the process of step S107, when the character recognition processing unit 223 determines that the degree of match between the normalized character image and a specific character exceeds a predetermined threshold (when step S107 is a YES determination), the character recognition processing unit 223 converts the normalized character image into the specific character and converts it into data (step S108). After the process of step S108, the character recognition processing unit 223 performs the process of step S111 described later.

[0081] On the other hand, in the process of step S107, when the character recognition processing unit 223 determines that the degree of match between the normalized character image and a specific character does not exceed a predetermined threshold (when step S107 is a NO determination), the character recognition processing unit 223 determines the normalized character image as an unrecognized character (step S109). Also, in this process, the character recognition processing unit 223 stores the determined unrecognized character in the RAM 25.

[0082] Next, the character recognition processing unit 223 associates the determined unrecognized character with its similar characters (step S110). In this process, the character recognition processing unit 223 associates the registered characters obtained in the collation process for the unrecognized characters performed in step S106 based on the degree of match, and associates the registered characters similar to the unrecognized character as similar characters with the unrecognized character. In the example shown in FIG. 8, three similar characters with a high degree of match are associated with one unrecognized character. Also, in this process, the character recognition processing unit 223 outputs and stores the dataset (see FIG. 8) in which the unrecognized character and its similar characters are associated in the storage unit 26.

[0083] After the process of step S108 or after the process of step S110, the character recognition processing unit 223 determines whether character recognition for all the character images included in the character area of the manuscript image has been completed (step S111).

[0084] In the process of step S111, when the character recognition processing unit 223 determines that the character recognition for all the character images included in the character area of the manuscript image has not been completed (when step S111 is a NO determination), the character recognition processing unit 223 returns the process to the process of step S105 and repeats the processes of steps S105 to S111.

[0085] On the other hand, in the process of step S111, when the character recognition processing unit 223 determines that the character recognition for all the character images included in the character area of the manuscript image has been completed (when step S111 is a YES determination), the character recognition processing unit 223 determines whether there is an unrecognized character in the group of characters for which character recognition has been performed (step S112). If there is an unrecognized character stored in the RAM 25 at this stage of the process, step S112 is a YES determination; otherwise, step S112 is a NO determination.

[0086] In the process of step S112, when the character recognition processing unit 223 determines that there is no unrecognized character in the group of characters for which character recognition has been performed (when step S112 is a NO determination), the character recognition processing unit 223 performs the process of step S117 described below.

[0087] On the other hand, in the process of step S112, when the character recognition processing unit 223 determines that there is an unrecognized character in the group of characters for which character recognition has been performed (when step S112 is a YES determination), the character recognition processing unit 223 performs the process of acquiring and collating the registered special symbols stored in the storage unit 26 (step S113).

[0088] In this process, first, the character recognition processing unit 223 acquires unrecognized characters for which character codes have already been assigned and which are registered special symbols stored in the storage unit 26. Next, the character recognition processing unit 223 collates the unrecognized characters stored in the RAM 25 with the registered special symbols, and determines whether there is a special symbol having the same meaning as the unrecognized character among the registered special symbol group. Note that any method can be adopted for this determination process. For example, similar to the process of step S106 described above, the degree of coincidence between the similar characters associated with the unrecognized character and the registered special symbols may be calculated for determination, or the similar characters associated with the unrecognized character and the similar characters associated with the registered special symbols may be compared. In the latter case, for example, the most highly coincident similar character associated with the unrecognized character and the most highly coincident similar character associated with the registered special symbol are compared, and if they match, it may be determined that the registered special symbol is a special symbol having the same meaning as the unrecognized character.

[0089] Also, in the process of step S113, if it is determined by the above collation process that there is a special symbol having the same meaning as the unrecognized character among the registered special symbol group, the character recognition processing unit 223 determines the unrecognized character as a special symbol similar to the registered special symbol determined to have the same meaning.

[0090] Next, the character recognition processing unit 223 performs grouping and special symbol determination processing on the unrecognized characters for which it is determined by the above collation process in step S113 that there is no special symbol having the same meaning among the registered special symbol group (step S114). Note that if there are no unrecognized characters for which it is determined by the above collation process in step S113 that there is no special symbol having the same meaning among the registered special symbol group, the process of step S114 is omitted.

[0091] In the stage of the process of step S114, when there is only one unrecognized character determined to have no special symbol with the same meaning among the registered special symbol groups, the character recognition processing unit 223 determines the unrecognized character as a new special symbol. Further, in the stage of the process of step S114, when there are a plurality of unrecognized characters determined to have no special symbol with the same meaning among the registered special symbol groups, the character recognition processing unit 223 first performs a grouping process on the plurality of unrecognized characters. In this grouping process, unrecognized characters having the same meaning among the plurality of unrecognized characters are determined as the same special symbol.

[0092] Note that the grouping process for a plurality of unrecognized characters is performed based on the similar characters associated with each unrecognized character and their degree of match. Specifically, the character recognition processing unit 223 first compares the most highly matching similar characters associated with each unrecognized character with each other, and groups the plurality of unrecognized characters with the same most highly matching similar characters as the same special symbol. Next, the character recognition processing unit 223 compares the second most highly matching similar character associated with the ungrouped unrecognized character with the similar characters associated with the special symbol that is each ungrouped unrecognized character. Next, in the comparison process, if there is a grouped special symbol associated with the same similar character as the second most highly matching similar character associated with the ungrouped unrecognized character, the character recognition processing unit 223 determines the unrecognized character as a special symbol similar to the grouped special symbol and groups it. Thereafter, the character recognition processing unit 223 performs the above comparison process until there are no more similar characters associated with the ungrouped unrecognized characters to be compared, and finally, if there are unrecognized characters that remain ungrouped, the character recognition processing unit 223 determines the unrecognized character as a new special symbol.

[0093] For example, when the grouping process is applied to the manuscript image N shown in FIG. 1, in the first similar character comparison process by the character recognition processing unit 223, the special symbol SC1 and the special symbol SC2, whose most highly matching similar characters are the same character "heavy", are grouped. Next, in the second similar character comparison process by the character recognition processing unit 223, since the similar character "work" with a high degree of match after the special symbol SC3 is included in the similar characters associated with the special symbols SC1 and SC2, the special symbol SC3 is grouped with the special symbols SC1 and SC2. And finally, the similar characters that do not match the similar characters associated with each of the special symbols SC1 to SC3 remain without being grouped with the associated special symbol SC4.

[0094] Here, returning to the description of the process of step S115, the character recognition processing unit 223 assigns a character code to each special symbol determined in step S113 and / or step S114 (step S115). Note that since the special symbol determined in the process of step S113 is a registered special symbol, the character recognition processing unit 223 assigns the character code associated with the registered special symbol to the special symbol determined in the process of step S113. Thereby, the special symbol determined in the process of S113 and the corresponding registered special symbol are grouped. Also, since the special symbol determined in the process of step S114 is a special symbol not registered in the character recognition system 1, the character recognition processing unit 223 assigns a character code different from the character code associated with the registered special symbol to the special symbol determined in the process of step S114.

[0095] In this embodiment, as the character code assigned to the special symbol, a combination of characters determined in advance for the special symbol, for example, character strings such as "Code a" and "Code b" in FIG. 10 are used, but the present invention is not limited thereto. For example, characters other than the registered characters used in the collation process in step S106 may be used as the character code for the special symbol. Further, for example, characters not included in the character strings (text information TX1 to TX4 in FIG. 1 etc.) normally recognized in the collation process in step S106 may be used as the character code for the special symbol. Further, for example, a character string combining characters not included in the character strings (text information TX1 to TX4 in FIG. 1 etc.) normally recognized in the collation process in step S106 may be used as the character code for the special symbol.

[0096] Next, the character recognition processing unit 223 outputs and stores in the storage unit 26 a data set in which the determined special symbol and the character code assigned thereto are associated (step S116).

[0097] If the determination in the process of step S112 is NO (when there are no unrecognized characters), or after the process of step S116, the character recognition processing unit 223 outputs and stores in the storage unit 26 text data (see FIG. 9), which is the result data of the character recognition processing for the document image (step S117). Then, after the process of step S117, the character recognition processing for one document image ends.

[0098] [Procedure for Character Recognition Processing for Multiple Document Images] Next, the character recognition processing for multiple document images in the character recognition system 1 will be described. FIG. 14 is a flowchart showing the procedure of the character recognition processing for multiple document images in the character recognition system 1 of the present embodiment. The processing described below starts when multiple documents to be character-recognized are fed to the image reading unit 13 of the image forming apparatus 10.

[0099] In the character recognition process for multiple manuscript images, the character recognition process for a single manuscript image described in FIG. 13 is repeated the number of times corresponding to the number of manuscript images. Therefore, the processing contents of steps S201 to S217 in the character recognition process for multiple manuscript images shown in FIG. 14 are the same as the processing contents of steps S101 to S117 in the character recognition process for a single manuscript image described in FIG. 13. Therefore, here, the description of the processing of steps S201 to S217 in the character recognition process for multiple manuscript images shown in FIG. 14 is omitted, and the description will start from the processing after step S217.

[0100] After the processing of step S217, the character recognition processing unit 223 determines whether the reading of all manuscript images is completed (step S218).

[0101] In the processing of step S218, when the character recognition processing unit 223 determines that the reading of all manuscript images is not completed (when step S218 is a NO determination), the character recognition processing unit 223 returns the processing to the processing of step S201 and repeats the processing after step S201. On the other hand, in the processing of step S218, when the character recognition processing unit 223 determines that the reading of all manuscript images is completed (when step S218 is a YES determination), the character recognition processing unit 223 ends the character recognition process for multiple manuscript images.

[0102] In the above-described character recognition process for multiple manuscript images, the unrecognized characters extracted from a predetermined manuscript image among the multiple manuscript images and the unrecognized characters extracted from a manuscript image different from the predetermined manuscript image among the multiple manuscript images are grouped, and the same character code can be assigned to each grouped unrecognized character.

[0103] [Effect] In the character recognition system 1 of the present embodiment, as described above, even if a manuscript image includes special symbols such as handwritten special symbols that cannot be normally recognized by character recognition processing, a character code is assigned to the special symbols, data is created, and the data-converted special symbols are stored in the storage unit 26. Further, in the character recognition system 1 of the present embodiment, similar characters are associated with the extracted special characters, and based on the similar characters, a plurality of special symbols having the same meaning are grouped. Therefore, in the character recognition system 1 of the present embodiment, even if a manuscript image includes unrecognized characters such as handwritten special symbols, it is possible to efficiently perform searches, classifications, etc. of the special symbols.

[0104] Also conventionally, there has been proposed a character recognition system that gives users clues for guessing the original characters even when there is misrecognition by simultaneously displaying the candidate character with the highest degree of matching (recognition degree) and other candidate characters for each character that has been character-recognized. However, in such a character recognition system, candidate characters are displayed individually for each character even when there are multiple identical characters in the text to be character-recognized. Therefore, in this case, a plurality of identical characters cannot be managed collectively. On the other hand, in the present embodiment, since unrecognized characters such as handwritten special symbols can be grouped and managed with other unrecognized characters having the same meaning, the above problem can be solved.

[0105] <Various Modification Examples> As described above, the configuration of the character recognition system 1 according to the present embodiment and the character recognition method have been described. However, the present invention is not limited to these, and various other modification example aspects can be adopted without departing from the gist of the present invention described in the claims.

[0106] In the character recognition system 1 of the above-described embodiment, an example in which the result of character recognition of the document image is output to and stored in the storage unit 26 of the information processing apparatus 20 has been described, but the present invention is not limited thereto. For example, the result of character recognition of the document image may be transmitted to the image forming apparatus 10 via the communication unit 27 of the information processing apparatus 20 and the network 30 and also stored in the storage unit 18 of the image forming apparatus 10. In this case, it is also possible to display the result of character recognition of the document image on the display unit of the operation display unit 11 of the image forming apparatus 10.

[0107] In the character recognition system 1 of the above-described embodiment, an example in which character recognition processing is performed on the document image read by the image forming apparatus 10 has been described, but the present invention is not limited thereto. For example, it is also possible to capture an image taken by a digital camera, a smartphone, etc. as a document image and perform character recognition processing by taking it into the character recognition apparatus unit 22 of the information processing apparatus 20.

Description of Reference Numerals

[0108] 1…Character recognition system, 10…Image forming apparatus, 11, 21…Operation display unit, 12…Automatic document feeder, 13…Image reading unit, 14…Printer unit, 15, 23…CPU, 16, 24…ROM, 17, 25…RAM, 18, 26…Storage unit, 19, 27…Communication unit, 20…Information processing apparatus, 22…Character recognition apparatus unit, 30…Network, 101, 201…Bus, 221…Layout analysis unit, 222…Character cutting-out unit, 223…Character recognition processing unit

Claims

1. An image reading unit that reads an image of a document, A character recognition unit that extracts characters included in the image of the document read by the image reading unit and performs character recognition on the extracted characters, A storage unit that stores the result of character recognition for the image of the document, and is provided with, The character recognition unit extracts unrecognized characters that cannot be recognized by character recognition, associates similar characters with the unrecognized characters, compares the similar characters associated with each unrecognized character among the plurality of unrecognized characters, and based on the comparison result, groups at least a part of the plurality of unrecognized characters, assigns the same predetermined character code to each grouped unrecognized character, and outputs and stores in the storage unit data in which the same predetermined character code is assigned to each grouped unrecognized character A character recognition system characterized by this.

2. The character recognition unit selects a predetermined unrecognized character from among the plurality of grouped unrecognized characters, and outputs and stores in the storage unit data associating the selected predetermined unrecognized character with the predetermined character code The character recognition system according to claim 1.

3. The character recognition unit outputs and stores in the storage unit data associating the unrecognized character, a plurality of similar characters associated with the unrecognized character, and the degree of match between the unrecognized character and each similar character, When a plurality of unrecognized characters are extracted by character recognition by the character recognition unit, the character recognition unit compares the plurality of similar characters associated with each unrecognized character among the plurality of unrecognized characters, and assigns the same predetermined character code to the unrecognized characters having the same similar characters among the plurality of unrecognized characters The character recognition system according to claim 1 or 2, characterized by this.

4. When a plurality of unrecognized characters are extracted by character recognition by the character recognition unit, The character recognition unit, Among the plurality of unrecognized characters, compares the plurality of similar characters associated with each unrecognized character, and assigns the same predetermined character code to the unrecognized characters among the plurality of unrecognized characters for which the most similar characters with the highest degree of match are the same Of the plurality of unrecognized characters, compare the plurality of similar characters associated with each of the first unrecognized characters to which the predetermined character code is assigned, and the second most similar characters associated with each of the remaining second unrecognized characters to which the predetermined character code is not assigned. If the second most similar characters are included in the plurality of similar characters associated with each of the first unrecognized characters, assign the predetermined code to the second unrecognized characters. Repeat the comparison process between the plurality of similar characters associated with each of the first unrecognized characters and the similar characters associated with each of the second unrecognized characters until there are no more similar characters associated with the second unrecognized characters to be compared. The character recognition system according to claim 3, characterized in that.

5. When there are a plurality of images of the manuscript read by the image reading unit, The character recognition unit Performs character recognition for each image of the manuscript, Group the unrecognized characters extracted from a predetermined manuscript image among the plurality of manuscript images and the unrecognized characters extracted from a manuscript image different from the predetermined manuscript image among the plurality of manuscript images, and it is possible to assign the same predetermined character code to each grouped unrecognized character. The character recognition system according to any one of claims 1 to 4, characterized in that.

6. An image reading unit that reads an image of a manuscript, A character recognition unit that extracts characters included in the image of the manuscript read by the image reading unit and performs character recognition on the extracted characters, A storage unit that stores the result of character recognition for the image of the manuscript, and is provided with When there are a plurality of images of the manuscript read by the image reading unit, the character recognition unit performs character recognition for each image of the manuscript, extracts unrecognized characters that cannot be recognized by character recognition, and extracts the unrecognized characters from a predetermined manuscript image among the plurality of manuscript images and the unrecognized characters extracted from a manuscript image different from the predetermined manuscript image among the plurality of manuscript images are grouped, and the same predetermined character code is assigned to each grouped unrecognized character, and the data with the same predetermined character code assigned to each grouped unrecognized character is output to the storage unit and saved. A character recognition system, characterized in that.

7. When the character recognition unit outputs and stores the result of the character recognition for the image of the original document in the storage unit, the data associating the unrecognized characters with the predetermined character codes is output and stored in the storage unit. The character recognition system according to any one of claims 1 to 6, characterized in that.

8. In the data associating the unrecognized characters with the predetermined character codes, information specifying the arrangement area of the unrecognized characters in the image of the original document is associated with the predetermined character codes. The character recognition system according to any one of claims 1 to 7, characterized in that.

9. The data associating the unrecognized characters with the predetermined character codes is table data. The character recognition system according to any one of claims 1 to 8, characterized in that.

10. The predetermined character codes are characters other than the characters that can be recognized by the character recognition unit. The character recognition system according to any one of claims 1 to 9, characterized in that.

11. The predetermined character codes are character strings formed by combining characters other than the characters that can be recognized by the character recognition unit. The character recognition system according to any one of claims 1 to 9, characterized in that.

12. A character recognition method in a character recognition system including an image reading unit that reads an image of an original document, a character recognition unit that extracts characters included in the image of the original document read by the image reading unit and performs character recognition on the extracted characters, and a storage unit that stores the result of the character recognition for the image of the original document, wherein the character recognition unit extracts unrecognized characters that cannot be recognized by character recognition, associates similar characters with the unrecognized characters, compares the similar characters associated with each unrecognized character among the plurality of unrecognized characters, and based on the comparison result, groups at least a part of the plurality of unrecognized characters, assigns the same predetermined character code to each grouped unrecognized character, and outputs and stores in the storage unit the data in which the same predetermined character code is assigned to each grouped unrecognized character. The character recognition method is characterized by that. A method for character recognition in a character recognition system comprising an image reading unit that reads an image of a document, a character recognition unit that extracts characters included in the image of the document read by the image reading unit and performs character recognition on the extracted characters, and a storage unit that stores the result of character recognition for the image of the document, wherein: When there are a plurality of images of the document read by the image reading unit, the character recognition unit performs character recognition for each image of the document, extracts unrecognized characters that cannot be recognized by character recognition, groups the unrecognized characters extracted from a predetermined image of the document among the plurality of images of the document and the unrecognized characters extracted from an image of the document different from the predetermined image of the document among the plurality of images of the document, assigns the same predetermined character code to each grouped unrecognized character, and outputs and stores data in which the same predetermined character code is assigned to each grouped unrecognized character in the storage unit The character recognition method is characterized by the above.

Citation Information

Patent Citations

  • Automatic character recognizing device

    JP1993217018A

  • Character recognizing device

    JP1994012522A

  • Display method for non-recognized character

    JP1995334611A

  • Information processor

    JP2019149073A