Document processing device, document processing method and program

The document processing device efficiently extracts annotations in standard documents by detecting and specifying annotation regions outside the standard area, reducing processing time and maintaining accuracy by focusing on predefined areas and adjacent regions.

JP7767764B2Active Publication Date: 2025-11-12KONICA MINOLTA INC

Patent Information

Application Number
JP2021133147
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-08-18
Publication Date
2025-11-12
Estimated Expiration
2041-08-18

AI Technical Summary

Technical Problem

Existing OCR technologies require layout analysis on all areas outside the standard area to extract annotations, negating the advantage of faster processing for standard documents.

Method used

A document processing device and method that detects objects indicating annotations outside the standard area, specifies the annotation region, and performs character recognition only on the identified area, expanding detection until no object is detected, and removes guide lines within the standard area before recognition.

Benefits of technology

Efficiently extracts annotations without overlooking them, reducing processing time and maintaining character recognition accuracy by focusing on predefined standard areas and adjacent regions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007767764000001
    Figure 0007767764000001
  • Figure 0007767764000002
    Figure 0007767764000002
  • Figure 0007767764000003
    Figure 0007767764000003
Patent Text Reader

Abstract

To provide a document processing device, a document processing method, and a program capable of shortening the time required for processing as much as possible when performing character recognition processing on standard documents, even if an annotation such as a memo is written in an area different from the predefined fixed area, while preventing oversight of the annotation.SOLUTION: In a document processing system in which an image reader, a server, and a terminal device are communicatively connected to each other via a network, a control unit 24 of the server, which is a document processing device, includes: an object detection unit 242a for detecting an object indicating the presence of the annotation in a neighborhood area of a given range in a fixed area; an annotation area specifying unit 242c for specifying an annotation area in which an annotation is written when an object is detected by the object detection unit 242a; and a fixed area recognition unit 242 that extracts an annotation from the annotation area identified by the annotation area specifying unit 242c.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a document processing apparatus, a document processing method, and a program capable of character recognition processing of documents such as forms. [Background technology]

[0002] In recent years, a technology called character recognition processing (hereinafter referred to as OCR (Optical Character Recognition) processing) has become popular. OCR processing is compatible with both unstructured and structured documents, and each has its own processing characteristics.

[0003] For non-standard documents such as reports and manuals, which have no set layout, OCR processing is performed on the entire document. This is useful for data with a large amount of text, as it allows for searches within the document.

[0004] On the other hand, because the layout of standard documents is predetermined, OCR processing can be performed only on predefined standard areas. This makes processing fast and efficient when there are a large number of documents in the same format, such as forms, application forms, and purchase orders.

[0005] The general procedure for OCR processing is as follows: (1) Reading (scanning) data: Data on paper is captured using a personal computer or scanner and converted into image data. (2) Layout analysis: Since the layout varies depending on the document, the document is analyzed to determine where the text areas, ruled lines, and image areas are located, and then segmented. The order in which each block of text should be recognized is determined based on the text structure. (3) Line extraction: The character area detected by layout analysis is decomposed line by line. (4) Extracting characters: The extracted line is further broken down into individual characters. (5) Character Recognition: Character feature values ​​are detected, and similar candidates are selected from a pre-registered dictionary. From among the candidates, a candidate is identified based on natural language knowledge to determine whether it can be connected to the preceding and following characters to form a correct Japanese sentence.

[0006] In the case of standard documents, the layout analysis process can be omitted by inputting data in the same format each time and by defining the area to be read in advance as a standard area.

[0007] Incidentally, annotations such as memos may be added to areas outside the preset standard area of ​​a standard document. As a technology for accurately recognizing such annotations outside the standard area together with the standard area, Patent Document 1 proposes an information processing device that includes a display control unit that controls the display of an image area extracted from an area outside the standard area that has been defined in advance as a recognition target on a confirmation screen that displays the recognition results of characters entered in the document.

[0008] In this information processing device, in order to extract annotations outside the standard area, layout analysis is performed on the image data of the form image, and character strings are obtained from the definition data as recognition results within the standard area. After that, unspecified items outside the standard area are extracted to prevent users from overlooking them. [Prior art documents] [Patent documents]

[0009] [Patent Document 1] Japanese Patent Application Publication No. 2020-160629 Summary of the Invention [Problem to be solved by the invention]

[0010] However, when only standard documents are read, it would normally be possible to start character recognition from the coordinates of the set standard area without performing layout analysis, but in Patent Document 1, in order to extract annotations outside the standard area, it was necessary to perform layout analysis on all areas outside the standard area. For this reason, Patent Document 1 has a problem in that by performing layout analysis on all areas outside the standard area, the advantage of not needing the time required for layout analysis in OCR processing on standard documents is negated.

[0011] This invention has been made in consideration of this technical background, and aims to provide a document processing device, document processing method, and program that, when processing a standard document with character recognition, can minimize the time required for processing while preventing annotations such as notes from being overlooked, even if the annotations are written in an area other than a predefined standard area. [Means for solving the problem]

[0012] The above object can be achieved by the following means: (1) A document processing device that performs character recognition processing on a predefined standard area, fixed Katarayo Outside the region a detection means for detecting an object indicative of the presence of an annotation; a specifying means for specifying an annotation area in which the annotation is written when the object is detected by the detecting means; an extraction means for extracting an annotation from the annotation region identified by the identification means; Equipped with 、 When the detection means detects the object, the area in which the object is detected is set as an attention area, and the detection means expands the detection area around the attention area to detect the object, and repeats the expansion and detection of the detection area until no object is detected in the entire expanded detection area; The specifying means specifies the annotation area based on the detection result of the detecting means. A document processing device characterized by: (2) A document processing device that performs character recognition processing on a predefined standard area, a detection means for detecting an object outside the fixed region that indicates the presence of an annotation; a specifying means for specifying an annotation area in which the annotation is written when the object is detected by the detecting means; an extraction means for extracting an annotation from the annotation region identified by the identification means; Equipped with A document processing device characterized in that, when an object detection process outside the standard area is performed before a character recognition process for the standard area, and as a result an indication line is detected outside the standard area and that the indication line has entered the standard area, the character recognition process for the standard area is performed with the indication line within the standard area removed. ( 3 ) The annotation area specified by the specifying means is a predetermined area. 2 The document processing device according to claim 1. ( 4 ) The extraction means extracts annotations by performing character recognition processing on the identified annotation area. ~One of 3 The document processing device according to claim 1. ( 5 ) The extraction means extracts the annotation as an image. ~One of 3 The document processing device according to claim 1. ( 6 ) The detecting means is Note O If a pointer line is detected as one of the objects, the direction of the pointer line is Detection Area and detecting an end of the indication line by enlarging the indication line, and the specifying means specifies an annotation area in the vicinity of the end of the indication line. 5 10. The document processing device according to claim 9, wherein ( 7 The detecting means determines that the end of the indicator line is an arrow even if the end of the indicator line is an arrow. 6 The document processing device according to claim 1. ( 8 The annotation extracted by the extraction means is associated with the most relevant item in the standard domain. 7 10. The document processing device according to claim 9, wherein ( 9 ) The annotations are handwritten and / or printed characters. 8 10. The document processing device according to claim 9, wherein ( 10 ) A document processing device that performs character recognition processing on a predefined standard area, fixed Katarayo Outside the region a detection step of detecting an object indicative of the presence of an annotation; a specifying step of specifying an annotation area in which the annotation is written, when the object is detected by the detecting step; an extraction step of extracting an annotation from the annotation region identified in the identification step; Run death, In the detection step, when the object is detected, the area in which the object is detected is set as an attention area, and the detection area is expanded around the attention area to detect the object, and the expansion and detection of the detection area are repeated until no object is detected in any of the expanded detection areas; In the identifying step, an annotation region is identified based on the detection result of the detecting step. A document processing method comprising: (11) A document processing device that performs character recognition processing on a predefined standard area, a detection step for detecting objects outside the standard region that indicate the presence of an annotation; a specifying step of specifying an annotation area in which the annotation is written, when the object is detected by the detecting step; an extraction step of extracting an annotation from the annotation region identified in the identification step; Run A document processing method characterized in that, when a step of detecting objects outside the standard area is performed before character recognition processing for the standard area, and as a result, an indicator line is detected outside the standard area and it is detected that the indicator line has entered the standard area, character recognition processing for the standard area is performed with the indicator line within the standard area removed. ( 12 The annotation area identified by the identifying step is a predetermined area. It is an area The preceding paragraph 11 A document processing method according to claim 1. ( 13 In the extraction step, the annotation is extracted by performing character recognition processing on the identified annotation region. Any of 10 to 12 A document processing method according to claim 1. ( 14 ) In the extraction step, the annotation is extracted as an image. Any of 10 to 12 A document processing method according to claim 1. ( 15 In the detection step, Note O If a pointer line is detected as one of the objects, the direction of the pointer line is Detection Area the step of identifying an annotation area in the vicinity of the end of the indication line by enlarging the image data. 10 ~ 14 10. The document processing method according to claim 9, wherein ( 16 The annotations extracted by the extraction step are the items associated with the most relevant items in the standard domain. 10 ~ 15 10. The document processing method according to claim 9, wherein ( 17 )Previous paragraph 10 ~ 16 2. A program for causing a computer to execute the document processing method according to claim 1. [Effects of the Invention]

[0013] The preceding paragraph (1) and ( 10According to the invention described in ), character recognition processing is performed on a predefined standard area. outside the typical area An object indicating the presence of an annotation is detected in the annotation region. If the object is detected, an annotation region in which the annotation is written is identified, and the annotation is extracted from the identified annotation region. Furthermore, when an object is detected, the annotation area is identified by repeating the enlargement and detection of the detection area until no object is detected in any of the enlarged detection areas, so that the annotation area can be identified reliably.

[0014] Previous section ( 4 According to the inventions described in (1) and (13), annotations can be extracted as characters.

[0015] Previous section ( 5 According to the inventions described in (1) and (14), annotations can be extracted as images.

[0017] Previous section ( 6 ) and ( 15 According to the invention described in ), when a pointer line is detected as one of the objects in the neighboring area, the direction in which the pointer line extends is Detection Area is enlarged to detect the end of the indicator line, and the annotation area is identified near the end of the indicator line. Therefore, even if the annotation position is far from the standard area, the annotation can be extracted by tracing the indicator line.

[0018] Previous section ( 7 According to the invention described in (1), even if the end of the indicator line is an arrow, the end of the indicator line can be determined.

[0019] Previous section ( 8 ) and ( 16 According to the invention described in ), the extracted annotation can be associated with the most relevant item in the standard area and displayed.

[0020] Previous section ( 3 ) and ( 12 According to the invention described in (1), the annotation area to be identified is set in advance, so that the process of identifying the annotation area can be simplified.

[0022] Previous section ( 9 ) allows annotations to be extracted whether they are handwritten or printed.

[0023] Previous section ( 2 ) and ( 11 According to the invention described in (1), if the object detection process in the nearby area is performed before the character recognition process for the standard area, and as a result a guide line is detected in the nearby area and it is detected that the guide line has entered the standard area, the character recognition process for the standard area is performed with the guide line within the standard area removed, so that the character recognition process for the standard area can be performed in a normal state where no support line exists.

[0024] Previous section ( 17 According to the invention described in Item 1 The document processing method according to any one of 2 to 20 can be executed by a computer. [Brief explanation of the drawings]

[0025] [Figure 1] 1 is a block diagram showing a configuration of a document processing system using a document processing device according to an embodiment of the present invention; [Figure 2] FIG. 2 is a block diagram showing the configuration of a server. [Figure 3] FIG. 2 is a block diagram showing the functional configuration of a control unit of the server. [Figure 4] FIG. 1 is a diagram showing a form, which is an order form, as an example of a fixed form used in this embodiment. [Figure 5] 1 is a flowchart illustrating document processing according to an embodiment of the present invention. [Figure 6] FIG. 1A is a diagram showing an annotation extraction range outside a fixed area in an embodiment of the present invention, and FIG. 1B is a diagram showing the range in a conventional example. [Figure 7] 10(A) to 10(C) are explanatory diagrams of the process of identifying an annotation region. [Figure 8] 10A and 10B are explanatory diagrams of another process for identifying an annotation region. [Figure 9] FIG. 10 is an explanatory diagram of yet another process for identifying an annotation region. [Figure 10] FIG. 10 is a diagram for explaining the continuation of the identification process of FIG. 9. [Figure 11] FIG. 10 is a diagram showing an example of display of a character recognition processing result. [Figure 12] FIG. 10 is a diagram showing another example of displaying the results of the character recognition process. [Figure 13] 10 is a flowchart illustrating document processing according to another embodiment of the present invention. [Figure 14] 10 is a flowchart illustrating document processing according to still another embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0026] Hereinafter, an embodiment of the present invention will be described with reference to the drawings.

[0027] FIG. 1 is a block diagram showing the configuration of a document processing system using a document processing device according to an embodiment of the present invention.

[0028] This document processing system includes an image reading device 1, a server 2 as a document processing device, and a terminal device 3.

[0029] The image reading device 1 is a device that reads various standard and non-standard documents, and examples thereof include a multifunction printer, a handheld scanner, a smartphone equipped with a camera function, a personal computer (PC), etc. The image reading device 1 reads the document, converts it into image data, which is electronic data, and transmits it to the server 2.

[0030] The server 2, which is a document processing device, receives the electronic image data transmitted from the image reading device 1 that read the document, and performs character conversion processing (OCR processing), and is configured as a PC, etc. The terminal device 3 displays the results of the OCR processing by the server 2 so that the user can confirm the results, and is configured as a PC, smartphone, etc.

[0031] The image reading device 1, server 2, and terminal device 3 may each be configured independently as in this embodiment and connected to each other so that they can communicate with each other via a network 4. Alternatively, any two or all three may be configured as a single device. Examples of the network 4 for interconnection include the Internet, a WAN (Wide Area Network), and a LAN (Local Area Network).

[0032] 2 is a block diagram showing the configuration of the server 2. The server 2 includes a display unit 21, a storage unit 22, a communication unit 23, a control unit 24, and an operation unit 25.

[0033] The display unit 21 displays the results of the user's operations, and may also display the results of OCR processing.

[0034] For example, a hard disk drive (HDD) or a solid state drive (SSD) is used as the storage unit 22. The storage unit 22 stores a control program, definition data of standard areas, and image data of documents read by the image reading device 1.

[0035] The communication unit 23 is a communication means for communicating with the image reading device 1 and the terminal device 3, respectively.

[0036] The control unit 24 includes a CPU 24a, a RAM 24b, a ROM 24c, etc., and is connected to each unit via a bus 26. The CPU 24a reads a control program from the ROM 24b or the storage unit 22, expands it in the RAM 24c, and executes it to perform overall control.

[0037] The operation unit 25 is configured by, for example, a mouse, a keyboard, etc., and accepts user input for OCR processing, etc. When the results of the OCR processing are displayed on the display unit 21 of the server 2, the operation unit also accepts user input for the display of the results.

[0038] 3 is a block diagram showing the functional configuration of the control unit 24 of the server 2. As described above, the functions of the control unit 24 are realized by the CPU 24a operating in accordance with the control program.

[0039] The control unit 24 includes a standard area recognition unit 241 , a non-standard area recognition unit 242 , and a recognition result output unit 243 .

[0040] The standard area recognition unit 241 performs OCR processing on standard areas, among the scanned data of the standard document, that should undergo OCR processing and are defined in the standard area definition data 301. The standard area definition data 301 is stored in the storage unit 22 in the server 2, but may also be stored in an external device other than the server 2.

[0041] The outside-of-standard-area recognition unit 242 performs annotation extraction processing, OCR processing, etc. on a predetermined amount of neighboring area (also simply referred to as neighboring area) of the standard area, which is set in advance as an area from which annotations should be extracted, within the area outside the standard area. This outside-of-standard-area recognition unit 242 includes an object detection unit 242a, an annotation area identification unit 242c, and a character recognition unit 242b.

[0042] The object detector 242a detects objects within the vicinity area, such as at least some of the characters of an annotation, or an indication line.

[0043] The annotation region identifying unit 242c identifies a region where an annotation exists, that is, an annotation region, as will be described later.

[0044] The character recognition unit 242b performs OCR processing on the identified annotation area and extracts the annotation as characters. Note that instead of extracting characters by OCR processing, the identified annotation area may be extracted as an image.

[0045] The recognition result output unit 243 outputs the OCR processing result for the standard region and the annotation extraction processing result for the nearby region to the terminal device 3 or the like. The results may be output and displayed on the display unit 21 of the server 2 itself. The processing results may also be stored in the storage unit 22 within the server 2, or in a storage unit of an external terminal device 3 or the like.

[0046] FIG. 4 shows a form 5, which is an order form, as an example of a fixed form document used in this embodiment. Orderer information and information about the ordered items are printed on the form 5. In addition, annotations are written as supplemental information. In this embodiment, the annotations shown are "Changed from 4 / 1" 51, "Changed to 200 sheets of copy paper" 52, and "Checked" 53. The annotations may be printed or handwritten.

[0047] Annotation 51, "Changed from 4 / 1," is written outside the email address in the standard area. Annotation 52, "Changed to 200 sheets of copy paper," is written near the end of instruction line 10 drawn from the corresponding item in the standard area, "500 sheets of copy paper." The end of instruction line 10 may be an arrow. Annotation 53, "Checked," is written in the upper left margin. [Example 1] The document processing executed by the server 2 for the form 5 shown in Fig. 4 will be described with reference to the flowchart of Fig. 5. The processing shown in Fig. 5 and the subsequent flowcharts is executed by the CPU 24a of the control unit 24 of the server 2 operating in accordance with the control programs stored in the ROM 24b and the storage unit 22.

[0048] The server 2 receives and stores electronic data (document data) of the document 5 read by the image reading device 1.

[0049] In step S01, OCR processing is performed on the standard area of ​​the form data according to the definition data. As shown in Figure 6(A), the area inside the gray zone on form 5 is standard area 6. The definition data includes information on multiple items in standard area 6 and the reading position, and OCR processing begins from the reading position, obtaining multiple items such as "orderer," "person in charge," "product number," and "product name" and the corresponding character strings.

[0050] After reading the standard area 6, object detection is performed in the nearby area in step S02. The gray zone shown in Figure 6(A) is the nearby area 7. The gray zone is shown for convenience of explanation and is not actually displayed on the form 5. The nearby area 7 is preset to, for example, 50 pixels or 3 centimeters around the standard area 6. The specific numerical value may be determined arbitrarily. For comparison, in the conventional example shown in Figure 6(B), annotation extraction processing was performed on all areas 71 outside the standard area 6.

[0051] In step S03, it is determined whether a character, which is one of the objects, has been detected in the vicinity region 7. If a character has not been detected (NO in step S03), it is determined in step S05 whether an instruction line 10, which is one of the objects, has been detected. If the instruction line 10 has not been detected (NO in step S05), it is checked in step S06 whether a predetermined threshold value defining the vicinity region 7 has been reached, and if not (NO in step S06), the detection region is expanded in step S07, and then the process returns to step S02, and the detection determination of characters and instruction lines is repeated until the threshold value is reached.

[0052] If characters are detected in step S03 (YES in step S03), an annotation may be written outside the vicinity area 7, so in step S04 an annotation area where characters are written is identified. The process of identifying an annotation area will be described later. Then, after extracting the annotation from the identified annotation area, the process proceeds to step S06. The annotation may be extracted by extracting character information by performing OCR processing on the annotation area, or by extracting the annotation area as an image.

[0053] On the other hand, if the instruction line 10 is detected in step S05 (YES in step S05), an annotation is often written near the end of the instruction line 10. Therefore, the instruction line is traced to its end in step S08, and then the annotation area where the annotation is written is identified in step S04. The process of identifying the annotation area when the instruction line 10 is detected will also be described later. Then, the annotation is extracted from the identified annotation area, and the process proceeds to step S06. In this case, too, the annotation may be extracted by extracting character information by performing OCR processing on the annotation area, or by extracting the annotation area as an image.

[0054] If the threshold value is reached in step S06 (YES in step S06), the detection process ends. Then, in step S09, the OCR process result for the standard area 6 and the annotation extraction result for the nearby area 7 are output to the display unit 21 or the terminal device 3, and the process ends. If no object is detected in the nearby area 7, only the OCR process result for the standard area 6 is output.

[0055] If the object cannot be detected even when the threshold value is reached, the vicinity region 7 may be increased by a predetermined amount beyond the threshold value, and the object detection process may be performed again.

[0056] Next, the annotation region specification process in step S04 will be described.

[0057] When a character that is an object is extracted from the neighborhood region 7, an annotation region is identified because an annotation may be written outside the neighborhood region 7.

[0058] First, the area in which an object is detected in the neighborhood area 7 is set as the area of ​​interest, and the detection area is expanded from the coordinates of the area of ​​interest to the surrounding pixels. For example, as shown in FIG. 7(A), the pixel indicated by the thick frame in which the object was first detected is set as the area of ​​interest 8, and as shown in the same figure (B), object detection is performed on the surrounding pixels of the area of ​​interest 8. For the pixel in which an object is detected, that pixel is set as the area of ​​interest 8, and object detection is performed on the surrounding pixels.

[0059] In this way, the expansion of the detection area and the detection of objects are repeated until no objects are detected in the entire detection area. When no objects are detected, the series of expanded areas are identified as the annotation area 9, as shown by the dotted lines in Fig. 7(C), and an annotation extraction process is performed on this identified annotation area 9. By performing this annotation area identification process, the annotation area 9 can be identified with high accuracy.

[0060] As another process for identifying the annotation area 9, a rectangle may be cut out along a predetermined size and direction based on the position where the character is detected, and this cut-out rectangle may be identified as the annotation area 9, and an annotation extraction process may be performed on this identified annotation area 9.

[0061] For example, if part of the characters of annotation 51 is detected in the vicinity region 7 as shown by the circle in Fig. 8(A), a rectangle that is long in the horizontal direction from that position can be cut out and identified as annotation region 9 as shown in Fig. 8(B), and annotation extraction processing can be performed on this identified annotation region 9. This identification processing simplifies the processing because it is not necessary to extend object detection processing to the surrounding pixels.

[0062] Next, a process for specifying the annotation region 9 when the instruction line 10 is detected as an object will be described.

[0063] If an indication line 10 is detected, the detected area is designated as an attention area 8, just as in the case of text, and the detection area is expanded from the coordinates of the attention area 8 to the surrounding pixels as shown in Figure 9. The dotted area in Figure 9 is the expanded detection area. The Hough transform is used to detect lines, and feature pattern matching technology is used to detect arrows. The area where the indication line 10 is no longer detected is designated as the end 11 of the indication line 10, and it is determined that an annotation is written nearby, and text is detected.

[0064] Specifically, as shown in Fig. 10, the detection area is expanded in the circumferential direction starting from the end point 11 of the indication line 10 as the base point, and characters are detected. In Fig. 10, the expanded detection area is shown by dots. Detection and expansion of the detection area are repeated until no characters are detected in the entire detection area. As with characters, when no characters are detected, the expanded series of areas are identified as an annotation area 9, and annotation extraction processing is performed on this identified annotation area 9.

[0065] As in the case of characters, a rectangle may be cut out along a predetermined size and direction with the end point 11 of the instruction line 10 as the base point, and this cut-out rectangle may be identified as the annotation area 9, and annotation extraction processing may be performed on this identified annotation area 9.

[0066] The character recognition results in the standard area 6 and information such as annotations extracted in the neighboring area 7 are stored in the storage unit 22, and the user can check this stored information on the display unit 21 or the terminal device 3. In this case, the annotations may be displayed as individual images as shown in Fig. 11, or as an image of the entire document including the annotations, or as character information as shown in Fig. 12.

[0067] On the display screen for checking the processing results, annotations are generally written near the related items in the standard area 6, or if this is not possible, they are written in a different location by drawing an instruction line 10. For this reason, it is advisable to associate the extracted annotations with the items in the standard area 6, and when displaying them, to make it easier to visually see which items they are related to when checking the results, so that the associated items and annotations are displayed in correspondence with each other.

[0068] As an example, if only the recognition result of annotation 51 "Changed from 4 / 1" is displayed, it will not be clear what the change is for, so by displaying annotation 51 in association with the items in the standard area 6, as shown in Figure 12, the user can easily understand the changes. In Figure 12, annotation 51 "Changed from 4 / 1" is displayed immediately next to email address item 55, making it easy to understand that the email address will be changed from April 1st.

[0069] Next, another embodiment of the present invention will be described with reference to the flowchart of FIG.

[0070] In this embodiment, OCR processing is performed first in the standard region 6, but if an annotation instruction line 10 is mixed in the standard region 6, this may lead to a decrease in character recognition accuracy. For this reason, processing is performed first in the neighboring region 7, and if an instruction line 10 is detected, the instruction line is detected in the standard region 6, and if an instruction line 10 is detected in the standard region 6, the instruction line 10 in the standard region 6 is removed before OCR processing in the standard region 6 is started.

[0071] In step S11, object detection is performed on the vicinity region 7. In step S12, it is determined whether a character, which is one of the objects, has been extracted from the vicinity region 7. If a character has not been extracted (NO in step S12), it is determined in step S14 whether an instruction line 10, which is one of the objects, has been detected. If the instruction line 10 has not been detected (NO in step S14), it is checked in step S15 whether a predetermined threshold value defining the vicinity region 7 has been reached, and if not (NO in step S15), the detection region is expanded in step S16, and then the process returns to step S11, and the detection determination of characters and instruction lines is repeated until the threshold value is reached.

[0072] If characters are detected in step S12 (YES in step S12), the annotation area 9 in which the annotation is written is identified in step S13. Then, the annotation is extracted from the identified annotation area 9, and the process proceeds to step S15. The annotation may be extracted by extracting character information by performing OCR processing on the annotation area, or by extracting the annotation area 9 as an image.

[0073] On the other hand, if the instruction line 10 is detected in step S14 (YES in step S14), then in step S17, the instruction line 10 is traced to check whether it is within the standard region 6. If the instruction line 10 is within the standard region 6 (YES in step S17), then in step S18 the instruction line 10 within the standard region 6 is removed, and the process proceeds to step S19. If the instruction line 10 is not within the standard region 6 (NO in step S17), then the process proceeds directly to step S19.

[0074] In step S19, the instruction line 10 is traced to its end, and then in step S13, the annotation area 9 in which the annotation is written is identified, and the annotation is extracted from the identified annotation area 9. Then, the process proceeds to step S15. In this case, too, the annotation may be extracted by extracting character information by performing OCR processing on the annotation area 9, or by extracting an image of the annotation area.

[0075] In step S15, if the threshold value is reached (YES in step S15), the object detection process is terminated, and in step S20, OCR processing is performed on the standard area 6. Then, in step S21, the OCR processing result for the standard area 6 and the annotation extraction result for the nearby area 7 are output to the display unit 21 or the terminal device 3, etc., and the process is terminated. If no object is detected in the nearby area 7, the OCR processing result for the standard area 6 is output.

[0076] Thus, in this embodiment, if the indicator line 10 is within the standard area 6, the indicator line 10 within the standard area 6 is removed and OCR processing is performed within the standard area 6, thereby preventing a decrease in character recognition accuracy.

[0077] Another embodiment of the present invention will be described with reference to the flowchart shown in Fig. 14. In this embodiment, when an object is detected in the nearby area 7, OCR processing or image extraction processing is performed on everything outside the standard area 6.

[0078] In step S31, OCR processing is performed on the standard area 6 of the form data in accordance with the definition data.

[0079] Next, in step S32, object detection is performed on the vicinity region 7, and then in step S33, it is determined whether or not a character, which is one of the objects, has been extracted from the vicinity region 7. If no character has been extracted (NO in step S33), it is determined in step S36 whether or not an instruction line 10, which is one of the objects, has been detected. If the instruction line 10 has not been detected (NO in step S36), it is checked in step S37 whether or not a predetermined threshold value defining the vicinity region has been reached, and if not (NO in step S37), the detection region is expanded in step S38, and then the process returns to step S32, and the detection determination of characters and instruction lines is repeated until the threshold value is reached.

[0080] If a character is detected in step S33 (YES in step S33), the process proceeds to step S34. If an instruction line 10 is detected in step S36 (YES in step S36), the process also proceeds to step S34.

[0081] In step S34, annotation extraction is performed for everything outside the standard area 6, and then in step S35, the OCR processing results for the standard area 6 and the extraction results of annotations outside the standard area are output to the display unit 21 or terminal device 3, etc., and the processing is terminated.

[0082] In step S37, if the threshold value is reached (YES in step S37), the detection process for the object in the vicinity area 7 is terminated, and the OCR processing result for the standard area 6 is output in step S39.

[0083] In this manner, in this embodiment, when an object is detected in the neighboring region 7, annotation extraction is performed for everything outside the standard region 6.

[0084] As described above, in this embodiment, OCR processing is performed on the predefined standard region 6. Meanwhile, parts of characters and instruction lines 10, which are objects indicating the presence of annotations, are detected in the nearby region 7. When an object is detected, the annotation region 9 in which the annotations 51-53 are written is identified, and the annotation is extracted from the identified annotation region 9. In other words, the processing for extracting the annotations 51-53 is performed on the nearby region 7, and it is not necessary to perform the processing on all regions outside the standard region 6, thereby reducing the time required for processing. Furthermore, since the annotations 51-53 are often written near the standard region 6 or near the end of the instruction line 10 drawn from the standard region 6, detecting the object in the nearby region 7 allows for efficient object detection and therefore efficient extraction of the annotations 51-53, preventing annotations from being overlooked. [Explanation of symbols]

[0085] 1. Image reader 2 Server 3 Terminal Devices 4 Network 5 Reports 6 Typical area 7 Neighborhood 8. Areas of Interest 9 Annotation Area 10 support line 11 End of support line 11 Light receiving section 21 Display section 22 Memory section 23 Communications Department 24 Control Unit 24a CPU 24b ROM 24c RAM 51~53 Comments 55 items 241 Regular area recognition unit 242 Typical outside recognition unit 242a Object detection unit 242b Character recognition section 242c Annotation area identification unit 243 Recognition result output unit

Claims

1. A document processing device that performs character recognition processing on a predefined standard area, a detection means for detecting an object outside the fixed region that indicates the presence of an annotation; a specifying means for specifying an annotation area in which the annotation is written when the object is detected by the detecting means; an extraction means for extracting an annotation from the annotation region identified by the identification means; Equipped with When the detection means detects the object, the area in which the object is detected is set as an attention area, and the detection means expands the detection area around the attention area to detect the object, and repeats the expansion and detection of the detection area until no object is detected in the entire expanded detection area; The document processing device according to claim 1, wherein the specifying means specifies an annotation area based on the detection result of the detecting means.

2. A document processing device that performs character recognition processing on a predefined standard area, a detection means for detecting an object outside the fixed region that indicates the presence of an annotation; a specifying means for specifying an annotation area in which the annotation is written when the object is detected by the detecting means; an extraction means for extracting an annotation from the annotation region identified by the identification means; Equipped with A document processing device characterized in that, when an object detection process outside the standard area is performed before a character recognition process for the standard area, and as a result an indication line is detected outside the standard area and that the indication line has entered the standard area, the character recognition process for the standard area is performed with the indication line within the standard area removed.

3. 3. The document processing device according to claim 2, wherein the annotation area specified by said specifying means is a predetermined area.

4. 4. The document processing device according to claim 1, wherein the extraction means extracts the annotation by performing character recognition processing on the identified annotation area.

5. 4. The document processing device according to claim 1, wherein the extraction means extracts the annotation as an image.

6. A document processing device according to any one of claims 1 to 5, wherein when the detection means detects a pointer line as one of the objects, the detection means expands the detection area in the direction in which the pointer line extends and detects the end of the pointer line, and the identification means identifies an annotation area near the end of the pointer line.

7. 7. The document processing apparatus according to claim 6, wherein said detecting means determines that the end of the indicator line is an arrowhead.

8. 8. The document processing device according to claim 1, wherein the annotation extracted by said extraction means is associated with the most relevant item in said standard area.

9. 9. The document processing device according to claim 1, wherein the annotations are handwritten characters and / or printed characters.

10. A document processing device that performs character recognition processing on a predefined standard area, a detection step for detecting objects outside the standard region that indicate the presence of an annotation; a specifying step of specifying an annotation area in which the annotation is written, when the object is detected by the detecting step; an extraction step of extracting an annotation from the annotation region identified in the identification step; Run In the detection step, when the object is detected, the area in which the object is detected is set as an attention area, and the detection area is expanded around the attention area to detect the object, and the expansion and detection of the detection area are repeated until no object is detected in any of the expanded detection areas; The document processing method, wherein the identifying step identifies an annotation area based on the detection result of the detecting step.

11. A document processing device that performs character recognition processing on a predefined standard area, a detection step for detecting objects outside the standard region that indicate the presence of an annotation; a specifying step of specifying an annotation area in which the annotation is written, when the object is detected by the detecting step; an extraction step of extracting an annotation from the annotation region identified in the identification step; Run A document processing method characterized in that, when a step of detecting objects outside the standard area is performed before character recognition processing for the standard area, and as a result, an indicator line is detected outside the standard area and it is detected that the indicator line has entered the standard area, character recognition processing for the standard area is performed with the indicator line within the standard area removed.

12. 12. The document processing method according to claim 11, wherein the annotation area identified in the identifying step is a predetermined area.

13. 13. The document processing method according to claim 10, wherein in the extracting step, the annotation is extracted by executing character recognition processing on the identified annotation area.

14. 13. The document processing method according to claim 10, wherein the extraction step extracts the annotation as an image.

15. A document processing method according to any one of claims 10 to 14, wherein in the detection step, if a pointer line is detected as one of the objects, the detection area is expanded in the direction in which the pointer line extends to detect the end of the pointer line, and in the identification step, an annotation area is identified near the end of the pointer line.

16. 16. The document processing method according to claim 10, wherein the annotation extracted in the extraction step is associated with the most relevant item in the standard area.

17. A program for causing a computer to execute the document processing method according to any one of claims 10 to 16.

Citation Information

Patent Citations

  • Image forming device, and method and program for forming electronic document data

    JP2008181485A

  • Contour detection device, contour detection method, and contour detection program

    JP2014102711A

  • Information processing apparatus and information processing program

    JP2015072541A

  • Information processing device and program

    JP2020160629A

Cited By

  • Note generating method and related device thereof

    US20250061736A1